Text processing method, electronic equipment and storage medium
By performing text structure detection and multi-level embedding on text generated by large language model, combined with error correction coding and hash tree verification, the shortcomings of existing text watermarking technology in terms of concealment and robustness are solved, and efficient copyright protection for text generated by large language model are achieved.
Patent Information
- Application Number
- CN202510875946.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-06-27
AI Technical Summary
The existing text watermarking technology is difficult to achieve copyright protection for text generated by large language models while ensuring concealment and robustness, especially when format adjustments, multilingual text and text content modifications are not effective.
By detecting text structures for to be processed, using the Hidden Markov model to identify syntax, logic and chapter structures, determining the embedding bits of watermark information, and using multi-level embedding and error correction coding technology to ensure that the watermark is hidden in the text and does not affect semantics, and self-checking is achieved in combination with the hash tree.
It realizes efficient, hidden and robust copyright protection in text generated by large language models, can resist local modifications, ensure the integrity and reliability of watermark information, and is suitable for content traceability and attribution verification of high-value texts.
Smart Images

Figure CN120409463A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical fields of natural language processing and large model technology, and particularly relates to a text processing method, an electronic device, and a storage medium. Background Art
[0002] With the in-depth application of large language models (LLMs) in various fields such as documents, academic papers, and business reports, the copyright attribution and content traceability of generated texts have become problems to be solved urgently. The usual processing method is to add watermarks to the text to protect the copyright.
[0003] However, traditional text watermarking technologies have many limitations and it is difficult to achieve strong robustness while ensuring the concealment of watermarks. For example, some format perturbation-based technologies are prone to exposing watermarks due to format adjustments and have poor effects when dealing with punctuation-free or multilingual texts; statistic feature-based technologies have low embedding density and are vulnerable to text content modification, resulting in attenuation or even loss of watermark signals; semantic perturbation-based technologies may damage the semantic coherence of the text and lack deep protection for the text structure. Summary of the Invention
[0004] In view of the above problems, the present application provides a text processing method, an electronic device, and a storage medium.
[0005] According to a first aspect of the present application, there is provided a text processing method, including:
[0006] Performing text structure detection on the text to be processed to obtain the text structure information of the text to be processed, where the text structure information includes at least one of the following: the grammatical structures to which multiple words in the sentences of the text to be processed belong respectively, the logical structures to which multiple sentences in the paragraphs of the text to be processed belong respectively, the discourse structures to which multiple paragraphs in the chapters of the text to be processed belong respectively; determining at least one embedding position for adding watermark information in the text to be processed based on the relevance between the text structure information and the context semantics in the text to be processed; adding the target watermark information to be added at the at least one embedding position.
[0007] A second aspect of the present application provides a text processing device, including a detection module, a determination module, and an addition module.
[0008] The detection module is configured to perform text structure detection on the text to be processed to obtain the text structure information of the text to be processed, where the text structure information includes at least one of the following: the grammatical structures to which multiple words in the sentences of the text to be processed belong respectively, the logical structures to which multiple sentences in the paragraphs of the text to be processed belong respectively, the discourse structures to which multiple paragraphs in the chapters of the text to be processed belong respectively;
[0009] A determination module, configured to determine at least one embedding bit for adding watermark information in the text to be processed based on the relevance between the text structure information and the context semantics in the text to be processed;
[0010] An adding module, configured to add the target watermark information to be added at the at least one embedding bit.
[0011] The third aspect of the present application provides an electronic device, including one or more processors; a memory, configured to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors are caused to execute the steps of the above method.
[0012] The fourth aspect of the present application further provides a non-volatile computer-readable storage medium, on which a computer program or instruction is stored, and when the computer program or instruction is executed by a processor, the steps of the above method are implemented.
[0013] The fifth aspect of the present application further provides a computer program product, including a computer program or instruction, and when the computer program or instruction is executed by a processor, the steps of the above method are implemented. Description of the Drawings
[0014] Through the following description of the embodiments of the present application with reference to the drawings, the above content and other objects, features and advantages of the present application will become clearer. In the drawings:
[0015] Figure 1 A diagram showing an application scenario of the text processing method according to an embodiment of the present application;
[0016] Figure 2 A flowchart showing the text processing method according to an embodiment of the present application;
[0017] Figure 3 A flowchart showing the training method of the first detection model according to an embodiment of the present application;
[0018] Figure 4 A flowchart showing the sentence-level watermark embedding method according to an embodiment of the present application;
[0019] Figure 5 A flowchart showing the paragraph-level watermark embedding method according to an embodiment of the present application;
[0020] Figure 6 A flowchart showing the chapter-level watermark embedding method according to an embodiment of the present application;
[0021] Figure 7 A flowchart showing the chapter coding generation method according to an embodiment of the present application;
[0022] Figure 8Shows a schematic diagram of the watermark information embedding process according to an embodiment of the present application;
[0023] Figure 9 Shows a flowchart of the watermark information restoration method according to an embodiment of the present application;
[0024] Figure 10 Shows a schematic diagram of the entire process of watermark embedding and watermark detection according to an embodiment of the present application;
[0025] Figure 11 Shows a structural block diagram of a text processing device according to an embodiment of the present application; and
[0026] Figure 12 Shows a block diagram of an electronic device suitable for implementing the text processing method according to an embodiment of the present application. Detailed implementation manners
[0027] Hereinafter, embodiments of the present application will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present application. In the following detailed description, for the sake of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present application. However, obviously, one or more embodiments can also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present application.
[0028] The terms used herein are merely for describing specific embodiments and are not intended to limit the present application. The terms "including", "comprising", etc. used herein indicate the presence of features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0029] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0030] In the case of using expressions such as "at least one of A, B, and C", generally, it should be interpreted according to the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include, but is not limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C).
[0031] In the technical solution of this application, the user information involved (including but not limited to user personal information, user image information, user device information, such as location information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) are all information and data authorized by the user or fully authorized by all parties. Moreover, for the processing of relevant data, such as collection, storage, use, processing, transmission, provision, application, and request, etc., all comply with relevant laws, regulations, and standards, adopt necessary confidentiality measures, do not violate public order and good customs, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0032] An embodiment of this application provides a text processing method, including:
[0033] Perform text structure detection on the text to be processed to obtain the text structure information of the text to be processed. Among them, the text structure information includes at least one of the following: the grammatical structures to which multiple words in the sentences of the text to be processed belong, the logical structures to which multiple sentences in the paragraphs of the text to be processed belong, and the discourse structures to which multiple paragraphs in the chapters of the text to be processed belong; based on the relevance between the text structure information and the context semantics in the text to be processed, determine at least one embedding position in the text to be processed for adding watermark information; add the target watermark information to be added at the at least one embedding position.
[0034] Figure 1 Schematically shows an application scenario diagram of the text processing method according to an embodiment of this application. As Figure 1 shown, the application scenario 100 according to this embodiment may include a large language model 101, a processor 102, and a client 103. Among them, the processor 102 is used to execute the text processing method provided by the embodiment of this application, and the user interacts with the large language model 101 and the processor 102 through the client 103. For example, the user inputs a prompt word to the large language model 101 through the client, and the large language model 101 generates relevant text according to the prompt word input by the user, such as professional documents, business reports, etc. In order to determine the copyright attribution of the generated text, the user can input the target watermark information to the processor 102 through the client 103. The processor 102 executes the text processing method provided by the embodiment of this application, embeds the target watermark information in the text generated by the large language model, and displays the text with the target watermark information added to the user through the client 103. Similarly, the processor 102 can also perform watermark restoration on the text with the watermark information embedded, so as to trace the content source and prevent the text from being tampered with.
[0035] The following will be based on Figure 1 the scenario described, and through Figures 2 - 6 describe the text processing method of the application embodiment in detail.
[0036] Figure 2 The figure schematically shows a flowchart of a text processing method according to an embodiment of the present application.
[0037] As Figure 2 shown, the method of this embodiment includes operations S210 to S230, and this text processing method can be executed by a main controller in a storage device.
[0038] In operation S210, text structure detection is performed on the text to be processed to obtain text structure information of the text to be processed.
[0039] According to an embodiment of the present application, the text structure information includes at least one of the following: the grammatical structure to which each of multiple words in a sentence of the text to be processed belongs, the logical structure to which each of multiple sentences in a paragraph of the text to be processed belongs, and the discourse structure to which each of multiple paragraphs in a chapter of the text to be processed belongs.
[0040] In operation S220, based on the relevance between the text structure information and the context semantics in the text to be processed, at least one embedding position for adding watermark information in the text to be processed is determined.
[0041] In operation S230, the target watermark information to be added is added at the at least one embedding position.
[0042] In one example, since the text generated by the LLM has structural features such as syntactic diversity, logical rigor, and discourse normativity, it is required that the watermark be embedded to achieve hierarchical protection of the self-similar features at the three levels of sentences, paragraphs, and chapters without damaging semantic coherence and structural integrity. However, related technologies are difficult to meet the copyright protection requirements in the above complex scenarios. Therefore, in the embodiments of the present application, the text structure of the text to be processed is first detected. The text structure detection can be performed using a detection model. For example, the text structure detection is performed based on a Hidden Markov Model (HMM) obtained by pre-training to obtain the text structure information of the text to be processed. Specifically, the text structure information includes the syntactic structures to which the multiple words in the sentences of the text to be processed belong. The syntactic structures can be, for example, subject-predicate-object, subject-linking verb-predicative, adverbial-modifier, attributive-modifier, verb-object, and parallel structures, etc.; the text structure information also includes the logical structures to which the multiple sentences in the paragraphs of the text to be processed belong. The logical structures can be, for example, general-specific structure, specific-general structure, progressive structure, parallel structure, causal structure, transitional structure, and sequential structure, etc.; the text structure information also includes the discourse structures to which the multiple paragraphs in the chapters of the text to be processed belong, such as discourse structures like introduction and conclusion. Through the pre-trained model and efficient parsing algorithms, the structure detection and watermark embedding of the text to be processed can be quickly performed, improving the efficiency of the entire watermark embedding process; based on the detection model, the text structure detection of the text to be processed can accurately identify the syntactic structure, logical structure, and discourse structure in the text, ensuring that the watermark is embedded without damaging the semantic coherence and structural integrity of the text, providing a basis for realizing copyright protection without semantic loss.
[0043] After determining the text structure information, according to the relevance between the text structure information and the context semantics in the text to be processed, one or more embedding positions for adding watermark information in the text to be processed can be determined. The embedding positions determined in operation S220 include at least one or more of the sentence level, paragraph level, and chapter level, providing a basis for realizing hierarchical protection of the self-similar features at the three levels of sentences, paragraphs, and chapters. The target watermark information to be added can be, for example, default watermark information or watermark information edited by the user. The target watermark information to be added is added to the foregoing at least one embedding position to complete the hierarchical watermark embedding at the sentence level, paragraph level, and chapter level.
[0044] Through the text processing method provided by the embodiments of the present application, by detecting the text structure of the text to be processed, the grammatical structure, logical structure, and discourse structure in the text can be accurately identified. This provides a basis for determining the watermark embedding positions subsequently, ensuring that the embedding of the watermark does not damage the semantic and structural integrity of the text. Adding the target watermark information to the determined embedding positions realizes the effective embedding of the watermark information in the text. Compared with the watermark embedding methods in the related art, for example, the method of adding watermarks based on statistical feature techniques (such as the deviation of word frequency distribution) is vulnerable to signal attenuation caused by synonym replacement or sentence restructuring; for example, the method of embedding watermarks based on semantic perturbation techniques by replacing synonyms or adjusting word order may damage semantic coherence and lack correlation between levels, and local modification is likely to cause global failure. In the method of the embodiments of the present application, the watermark information is added to the semantic structure positions in the document. Watermarks are added at various semantic structure positions in the text, such as the end of words in the adverbial-middle structure, the end of sentences in the progressive relationship, the end of paragraphs in the chapter, etc. In this way, the watermark information does not damage the original grammatical structure and does not affect semantic understanding and semantic parsing; in addition, because these watermarks are added to the semantic structure positions, they are usually not modified when the text is modified. Therefore, the text modification does not damage the original watermark information, and it has strong robustness.
[0045] Among them, in the process of determining at least one embedding position for adding watermark information in the text to be processed, the factors considered include: text structure information and the relevance of the context semantics in the text to be processed. For example, after detecting that a sentence in the text to be processed contains multiple grammatical structure positions, further according to the relevance of each grammatical structure position to the context semantics, the positions with lower relevance (less than the preset relevance threshold) to the context semantics among the multiple grammatical structure positions are used as the embedding positions for adding watermarks. Since the positions with lower relevance to the context semantics belong to the semantic segmentation positions in the text, they have weak relevance to the context semantics and strong independence, and the text content at these positions is usually not easily modified by users. Therefore, adding watermark information here is not easily damaged, further improving the robustness of the watermark.
[0046] The text processing method provided by the embodiments of this application can be executed by a processor. The processor and the large language model can be deployed on the same server or on different servers, which is not limited here. For example, the processor can be deployed on the first server, and the large language model can be deployed on the second server. After the first server receives the creative text generated by the user using the language model, it performs text structure detection on the creative text to obtain the text structure information of the creative text. For example, it can be the grammatical structure to which each of the multiple words in the sentences of the creative text belongs, the logical structure to which each of the multiple sentences in the paragraphs of the creative text belongs, and the discourse structure to which each of the multiple paragraphs in the chapters of the creative text belongs, etc. text structure information; based on the relevance between the text structure information and the context semantics in the creative text, determine at least one embedding position in the creative text for adding watermark information; the first server receives the target watermark information input by the user through the client and adds the target watermark information to at least one embedding position; use the client to display the target text with the target watermark information added to the user. Since the watermark information is embedded in the embedding position determined based on the text structure information and the context semantics relevance, it ensures that the embedding of the watermark does not damage the semantic and structural integrity of the text, thereby realizing the effective protection of the copyright of the user's creative text.
[0047] According to the embodiments of this application, to perform text structure detection on the text to be processed and obtain the text structure features of the text to be processed, a detection model can be used for text structure detection. It can be to perform text structure prediction at the sentence level, paragraph level, and chapter level respectively, train three separate detection models respectively, and then perform prediction.
[0048] For example, for the grammatical structure prediction at the sentence level, first train a first detection model, and based on the first detection model, perform a first detection on multiple words to obtain multiple first probabilities that each word belongs to multiple candidate grammatical structures, and determine the grammatical structure to which the word belongs based on the multiple first probabilities.
[0049] For example, for the logical structure prediction at the paragraph level, first train a second detection model, and based on the second detection model, perform a second detection on multiple sentences to obtain multiple second probabilities that each sentence belongs to multiple candidate logical structures, and determine the logical structure to which the sentence belongs based on the multiple second probabilities.
[0050] For example, for the discourse structure prediction at the chapter level, first train a third detection model, and based on the third detection model, perform a third detection on multiple paragraphs to obtain multiple third probabilities that each paragraph belongs to multiple candidate discourse structures, and determine the discourse structure to which the paragraph belongs based on the multiple third probabilities.
[0051] The above three detection models can use the hidden Markov model. The construction processes of these three detection models are introduced separately below.
[0052] After training these three detection models, the model parameters include state transition information (state transition matrix) and observation probability information (observation probability matrix).
[0053] Each element in the state transition matrix represents: the probability of transitioning from the first text structure to the second text structure. That is, the probability that the current text object (word, sentence, paragraph) is the first text structure and the next text object (word, sentence, paragraph) following it is the second text structure. The text structure refers to the grammatical structure of the aforementioned words, the logical structure of the sentences, and the discourse structure of the paragraphs.
[0054] Each element in the observation probability matrix represents the probability of the occurrence of multiple text expression forms under the second text structure. That is, the respective probabilities of the multiple text expression forms corresponding to the second text structure. For detection models at different levels, they respectively represent the respective probabilities of the multiple text expression forms of the observed words corresponding to a certain grammatical structure, the respective probabilities of the multiple text expression forms of the observed sentences corresponding to a certain logical structure, and the respective probabilities of the multiple text expression forms of the observed paragraphs corresponding to a certain discourse structure.
[0055] Figure 3 Schematically shows a flowchart of a training method for a first detection model according to an embodiment of the present application. As Figure 3 shown, it includes operation S310 to operation S330.
[0056] [[ID=1X]]In operation S310, training texts are obtained, and multiple sample sentences in the training texts are respectively tokenized to obtain multiple sample words.
[0057] In operation S320, the grammatical structures to which the multiple sample words respectively belong are determined to generate a first label for the training text.
[0058] In operation S330, a basic model is trained using the multiple sample words and the first label to obtain a first detection model.
[0059] According to an embodiment of the present application, the model parameters of the first detection model include first state transition information and first observation probability information. Among them, the first state transition information is used to represent the probability of transitioning from the first candidate grammatical structure to the second candidate grammatical structure, and the first observation probability information is used to represent the respective probabilities of multiple target sample words belonging to the second candidate grammatical structure. The first candidate grammatical structure and the second candidate grammatical structure belong to at least one of the multiple grammatical structures to which the multiple sample words belong. The first candidate grammatical structure and the second candidate grammatical structure may be the same, and both are one of the multiple grammatical structures; the first candidate grammatical structure and the second candidate grammatical structure may also be different, and both are two different grammatical structures among the multiple grammatical structures.
[0060] In one example, the construction process of the sentence-level Hidden Markov Model in the embodiments of the present application is as follows: First, collect a training text containing multiple sample sentences, tokenize each sample sentence in the training text to obtain multiple sample words. Label the grammatical structure to which each sample word belongs, such as subject-predicate-object, adverbial-modifier, parallelism, etc., to obtain a sentence-level dataset: label the grammatical structure of each sample word (subject-predicate-object , adverbial-modifier , parallelism , etc.), generate sample words and their corresponding grammatical structure labels, and use them as training samples to train a basic model to obtain a first detection model, that is, a sentence-level Hidden Markov Model. After training, the parameters of this model include first state transition information and first observation probability information. Among them, the first state transition information is used to represent the probability of transitioning from a first candidate grammatical structure to a second candidate grammatical structure (assuming that the current word is the first candidate grammatical structure and the next word is the second candidate grammatical structure). The first observation probability information is used to represent the probability of each target sample word appearing under the second candidate grammatical structure (the probabilities of various text expressions of the observed word corresponding to the second candidate grammatical structure appearing respectively).
[0061] According to the embodiments of the present application, the construction process of the paragraph-level Hidden Markov Model is as follows:
[0062] Obtain a training text, and perform clause separation processing on multiple sample paragraphs in the training text to obtain multiple sample sentences. Determine the logical structure to which each of the multiple sample sentences belongs to generate a second label for the training text. Use the multiple sample sentences and the second label to train a basic model to obtain a second detection model. Among them, the model parameters of the second detection model include second state transition information and second observation probability information. Among them, the second state transition information is used to characterize the probability of transitioning from a first candidate logical structure to a second candidate logical structure, and the second observation probability information is used to characterize the probabilities of multiple target sample sentences belonging to the second candidate logical structure. The first candidate logical structure and the second candidate logical structure belong to at least one of multiple logical structures to which multiple sample sentences belong.
[0063] In one example, similar to the sentence-level Hidden Markov training process, training text containing multiple sample paragraphs is first collected. Each sample paragraph in the training text is segmented into sentences to obtain multiple sample sentences. The sentence segmentation process can use commas and periods as delimiters, and sentences ending with commas or periods are considered as one sentence. For example, for the text "First, let's give an example...", "First" and "Let's give an example" are segmented as two sentences. Each sample sentence is labeled with its belonging logical structure, such as the overall-part, progressive, causal, etc. logical structures, to generate the second label of the training text and obtain the paragraph-level dataset. The basic model is trained using the sample sentences and their corresponding logical structure labels to obtain the second detection model, that is, the paragraph-level Hidden Markov model. The parameters of the second detection model include the second state transition information and the second observation probability information. The second state transition information represents the probability of transitioning from the first candidate logical structure to the second candidate logical structure (i.e., the probability that the current sentence is the first candidate logical structure and the next sentence is the second candidate logical structure), and the second observation probability information represents the probability of each target sample sentence appearing under the second candidate logical structure (the probabilities of various text expressions of the observation sentences corresponding to the second candidate logical structure).
[0064] According to an embodiment of the present application, the construction process of the chapter-level Hidden Markov model is as follows: Obtain the training text and perform paragraph segmentation on the training text to obtain multiple sample paragraphs. Determine the discourse structure to which each of the multiple sample paragraphs belongs to generate the third label of the training text. The basic model is trained using the multiple sample paragraphs and the third label to obtain the third detection model, where the model parameters of the third detection model include the third state transition information and the third observation probability information. The third state transition information is used to characterize the probability of transitioning from the first candidate discourse structure to the second candidate discourse structure, and the third observation probability information is used to characterize the probabilities of the multiple target sample paragraphs belonging to the second candidate discourse structure. The first candidate discourse structure and the second candidate discourse structure belong to at least one of the multiple discourse structures to which the multiple sample paragraphs belong.
[0065] In one example, a training text containing multiple sample paragraphs is first collected, and the training text is segmented to obtain multiple sample paragraphs. Each sample paragraph is labeled with the chapter structure to which it belongs, such as the introduction, main text, conclusion, etc., and a third label of the training text is generated to obtain a chapter-level dataset. It should be understood that since the choice of observation set does not affect the watermark embedding and detection effect, it is only used as an example here and is not fully listed. The basic model is trained using sample paragraphs and their corresponding chapter structure labels to determine the third state transition information and third observation probability information in the third detection model, which is a chapter-level hidden Markov model. The third state transition information represents the probability of transitioning from the first candidate chapter structure to the second candidate chapter structure (assuming that the current paragraph is the first chapter structure and the next paragraph is the second chapter structure). The third observation probability information represents the probability of each target sample paragraph appearing under the second candidate chapter structure (the probability of each of the multiple text expressions of the observation paragraph corresponding to the second chapter structure appearing).
[0066] When constructing the data set, we define the state set, observation set, state transition matrix and observation probability matrix of each level of the hidden Markov model. Taking the sentence-level hidden Markov model as an example, the state set is , where it is assumed Representing the subject-verb-object structure, Characterize the structure of the Characterizes the parallel structure. The observation set is , where, for example, the observation set can represent the corresponding multiple text expressions under the current grammatical structure, such as, in the adverbial structure, various prepositions usually appear, then It can represent a variety of prepositional expressions, such as It means "in", Means "from", Indicates "to". Define the state transfer information composed of the transition probability between states, that is, the state transfer matrix , as shown in formula (1):
[0067] ----(1);
[0068] in Indicates the slave state Transfer to state The probability, for example Indicates that the current word is the subject, predicate, and object , the next word is transferred to the adverbial structure The probability is 0.2.
[0069] Define the observation probability information composed of the generation probability of the observed words in the current state, that is, the observation probability matrix , as shown in formula (2):
[0070] ---- (2);
[0071] Wherein represents the state to generate the probability of an observed word, for example indicates that the current statement is a modifier-head structure and the probability of generating the next word is as follows : the probability of "from" is 0.7.
[0072] The state sets and observation set dimensions of the paragraph-level hidden Markov model and the chapter-level hidden Markov model are consistent with those of the aforementioned data set. Other construction processes are the same as those of the above sentence-level hidden Markov model and will not be elaborated here.
[0073] The method for training the above three hidden Markov models can be to use a variety of feasible training strategies for training.
[0074] One of the methods is that when the sample data volume is large enough, it can be through data statistics methods to separately count the probability of each semantic structure transferring to another semantic structure to obtain state transition information; and count the probability of various text expression forms corresponding to each semantic structure to obtain observation probability information.
[0075] For example, for the training of the sentence-level hidden Markov model, separately count the probability that the words following each sample word belong to various grammatical structures (subject-predicate-object, modifier-head, parallel...). In this way, the probability of transferring from the first candidate grammatical structure (the grammatical structure of the current word) to the second candidate grammatical structure (the grammatical structure of the following word) can be obtained, and a state transition matrix is generated. Further statistics are carried out on the probability of the text expression form of the observed word corresponding to each grammatical structure. For example, for the modifier-head structure, count the probability of various prepositions appearing in the sample to generate an observation probability matrix.
[0076] Another method can also be to use the maximum likelihood estimation method, or it can be the combined training of the above three models. For example, first train the first detection model, then train the second detection model based on the model parameters of the first detection model, and then train the third detection model based on the model parameters of the second detection model.
[0077] For example, when training the above three hidden Markov models, the training objective is to maximize the joint likelihood probability of the observation sequence , is the log-likelihood function of the model, representing the logarithm of the joint probability of the observation sequence under the given model parameters . Maximizing means finding the model parameters that make the probability of the observation sequence appear the largest As shown in the following formula (3):
[0078] ---- (3);
[0079] Where l is the level, and its values are 1, 2, and 3, corresponding to the sentence level, paragraph level, and chapter level respectively. is the length of the feature sequence at the l-th level. For the sentence level, is the number of words in the sentence; for the paragraph level, is the number of sentences in the paragraph; for the chapter level, is the number of paragraphs in the chapter; represents the observed value at the l-th level at time step (e.g., the text expression of the observed word, observed sentence, observed paragraph); represents the hidden state at the l-th level at time step (such as the grammatical structure of the word, the logical structure of the sentence, the discourse structure of the paragraph); is the model parameter at level ; is the state transition matrix,[[]] is the observation probability matrix. refers to the emission probability of the observed value given the hidden state and the model parameter . In the hidden Markov model, this is usually determined by the observation probability matrix.
[0080] The objective function is the sum of the log-likelihood probabilities of all levels (sentence level, paragraph level, chapter level). Specifically, for each level l, calculate the emission probability of the observed value at each time step t in this level given the hidden state and the model parameter , and accumulate these emission probabilities over time and levels to maximize the objective function. To maximize the objective function , the forward-backward algorithm of the hidden Markov model, such as the Baum-Welch algorithm, is used to iteratively optimize until the change in the likelihood probability is less than the acceptable error, such as , that is, as shown in the following formula (4):
[0081] ---- (4);
[0082] Where n represents the number of iterations, that is, the current is the n-th iteration.
[0083] , represents the model parameters at the nth iteration (including state transfer matrix and observation matrix);
[0084] , represents the model parameters after the n+1th iteration update;
[0085] , which represents the log-likelihood function value of the model at the nth iteration.
[0086] By training the above hidden Markov model on the sentence-level grammatical state, paragraph-level logical state, and chapter-level text state, a state transition matrix and an observation probability matrix at each level are generated. The embodiments of the present application do not limit the training method of the hidden Markov model.
[0087] After training the three detection models, the semantic structure position detection is performed using the detection models. Below, the method of performing a first detection on multiple words based on the first detection model to obtain multiple first probabilities that each word belongs to multiple candidate grammatical structures is used as an example to illustrate the method of using the hidden Markov model for detection.
[0088] According to an embodiment of the present application, a first detection is performed on multiple words based on a first detection model to obtain multiple first probabilities that each word belongs to multiple candidate grammatical structures, including: using first state transition information and first observation probability information to determine multiple first probabilities that each word belongs to multiple candidate grammatical structures.
[0089] In one example, after training and obtaining state transition information (state transition matrix) and observation probability information (observation probability matrix) for each level of detection model, the text to be processed is first segmented to obtain a word sequence. This word sequence is then input into the first detection model. For each word, the model's state transition matrix and observation probability matrix are used to calculate the first probability of each word belonging to each candidate grammatical structure to determine the word's grammatical structure. For example, using a sentence-level hidden Markov model to detect the grammatical structure of the text to be processed, suppose the sentence "He quickly solved the problem" is being tested. After segmentation, the sentence "He quickly solved the problem" is segmented to obtain the word sequence ["he", "quickly", "solve", "solved", "problem"].
[0090] The specific detection process is as follows:
[0091] Assume the initial state probability distribution is: [0.6, 0.3, 0.1];
[0092] Word sequence: [“he”, “quickly”, “solve”, “了”, “problem”];
[0093] Assume that the observation probability (B) of each word is read according to the observation probability matrix of the model as follows:
[0094] "He": in grammatical structure 0-state The observation probability is 0.6, in the grammatical structure 1-state The observation probability is 0.1, in the grammatical structure 2-state The next observation probability is 0.2.
[0095] "Fast": In state The observation probability is 0.2, in state The observation probability is 0.8, in state The next observation probability is 0.1.
[0096] "Solved": In state The observation probability is 0.7, in state The observation probability is 0.1, in state The next observation probability is 0.2.
[0097] "了": in state The observation probability is 0.2, in state The observation probability is 0.5, in state The probability of the next observation is 0.1.
[0098] "Problem": In the state The observation probability is 0.1, in state The observation probability is 0.1, in state The next observation probability is 0.8.
[0099] Assume that the transition probability (A) between states is read as follows based on the state transition matrix of the model: arrive The probability is 0.3, from arrive The probability is 0.5... and so on. For each word and each state, calculate the maximum probability of reaching that state along the path and record the path. Finally, the most likely state sequence is obtained: .Right now:
[0100] he( , 0.6): The subject in the subject-verb-object structure. It indicates the main body of the sentence, performing an action or assuming a state.
[0101] fast( , 0.8): Adverbial in the adverbial-inflected structure. Modifies the verb "solve" and describes the way, time, or conditions under which the action occurs.
[0102] solve( , 0.7): The predicate in the subject-verb-object structure. It indicates the action taken by the subject or the state it is in.
[0103] already ( , 0.5): The adverbial in the adverbial-center structure. Here it indicates the perfective aspect of the action, supplementing the state of the action of "solving".
[0104] problem ( , 0.8): The object in the parallel structure. It represents the object of action of the predicate verb "solve" and is the object component of the sentence.
[0105] By detecting the text to be processed, it provides a basis for subsequent text watermark embedding, ensuring that the watermark is embedded in the text position with correct grammar, thereby improving the concealment and robustness of the watermark.
[0106] It should be understood that the principle and process of obtaining the sentence logical structure based on the second detection of the sentence by the second detection model and determining the discourse structure of the paragraph attribution based on the third detection of the paragraph by the third detection model are similar to the process of the first detection of multiple words based on the first detection model, and will not be elaborated here.
[0107] After determining the text structure features of the text to be processed, at least one embedding position for adding watermark information in the text to be processed is further determined according to the relevance between the text structure information and the context semantics in the text to be processed.
[0108] According to the embodiments of the present application, determining at least one embedding position for adding watermark information in the text to be processed based on the relevance between the text structure information and the context semantics in the text to be processed includes at least one of the following: based on the grammatical structures to which multiple words belong respectively, taking the position where the target word belonging to the predetermined grammatical structure and the context semantic relevance meets the predetermined condition (for example, the relevance is less than the preset relevance threshold) as the first embedding position; based on the logical structures to which multiple sentences belong respectively, taking the position where the target sentence belonging to the predetermined logical structure and the context semantic relevance meets the predetermined condition as the second embedding position; based on the discourse structures to which multiple paragraphs belong respectively, taking the position where the target paragraph belonging to the predetermined discourse structure and the context semantic relevance meets the predetermined condition as the third embedding position.
[0109] In one example, while obtaining the text structure information according to operation S210, the observation probabilities corresponding to each word, sentence, and paragraph in the text to be processed are also determined. The threshold Q is used to screen the positions with higher observation probabilities as candidate embedding positions. Only the positions with observation probabilities greater than or equal to Q will be determined as candidate embedding positions. The value of the threshold Q can be adjusted according to specific requirements and the experimental results of the validation set. A higher Q value indicates that only the positions with very high observation probabilities are selected for embedding, which can improve the invisibility and robustness of the watermark, but may reduce the number of embeddable positions, thereby reducing the watermark capacity. On the contrary, a lower Q value can increase the number of embedding positions and improve the watermark capacity, but may reduce the invisibility and robustness. Suppose Q = 0.6, then only the positions with observation probabilities ≥ 0.6 will be selected as candidate embedding positions.
[0110] For example, taking the detection of a sentence using a sentence-level hidden Markov model as an example, when detecting the syntactic structure of "He solved the problem this month", "He" (subject-verb-object 0.6) → "In" (adverbial-middle 0.8) → "This month" (adverbial-middle 0.7) → "Solve" (subject-verb-object 0.5) → "Problem" (subject-verb-object 0.8), then it is determined that the positions where "He" ( , 0.6), "In" ( , 0.8), "This month" ( , 0.7), "Problem" ( , 0.8) are located are used as candidate embedding positions.
[0111] After determining the candidate embedding positions, it is also necessary to consider whether the context semantic relevance of the words at the candidate embedding positions meets the predetermined conditions. Specifically, for example, in an adverbial-middle structure, there will probably be some prepositions, time words, place nouns, etc. in this structure. These words serve as semantic segmentation positions and have weak semantic relevance to their previous or subsequent texts. These words are used as the first embedding positions, and embedding watermark information at these positions does not damage the original semantics. For example, by detecting the context semantic relevance of the words at each candidate embedding position (for example, using a language model), it is found that the semantic relevance of "In" to the context (the relevance is 0.3) is lower than the preset relevance threshold of 0.5. Then, "In" is used as the embedding position.
[0112] After determining the logical structure to which each of multiple statements belongs, the position where the statement belonging to the predetermined logical structure is located is used as a candidate embedding position, and then it is determined whether the context semantic relevance meets the predetermined condition, and the position where the statement whose context semantic relevance meets the predetermined condition is located is determined as the second embedding position. For example, in a causal structure, there will be conjunctions such as "because", "so", "therefore" (here, these conjunctions are used as independent statements). After performing relevance detection on these statements, they usually belong to statements with relatively low relevance to the context and are usually used as semantic segmentation positions, with weak semantic relevance to their upper or lower context. Therefore, the position of the conjunction is used as the second embedding position, and embedding the watermark information at these positions does not damage the original semantics.
[0113] Similarly, after determining the discourse structure to which each of multiple paragraphs belongs, the one belonging to the predetermined discourse structure is used as a candidate embedding position, and it is determined whether its context semantic relevance meets the predetermined condition, and the position where the target paragraph whose context semantic relevance meets the predetermined condition is located is used as the third embedding position. For example, the paragraph at the end of a chapter has weak semantic relevance to the next chapter. The starting paragraph of a chapter has weak semantic relevance to the previous chapter. These positions are used as the third embedding position, and embedding the watermark information at these positions does not damage the original semantics.
[0114] According to an embodiment of the present application, embedding the watermark at a position with relatively low semantic relevance does not damage the original semantics, and subsequent text detection using a language model will not result in errors. Moreover, since the semantics of these positions are relatively independent, users are usually not likely to modify them when fine-tuning the text, and the watermark is not easily damaged.
[0115] The embodiment of the present application determines the embedding positions for adding watermark information based on the text structure information, which can ensure that the watermark information is embedded in a relatively concealed position in the text and has a relatively small impact on the semantics. It not only improves the concealment of the watermark but also reduces the interference of the watermark on the readability of the text, while ensuring that the watermark information can be stably embedded in the text. Determining the embedding positions at different levels (statements, paragraphs, chapters) enables the watermark to be flexibly embedded in appropriate positions according to the structural characteristics of the text, enhancing the flexibility and adaptability of watermark embedding.
[0116] According to an embodiment of the present application, adding the target watermark information to be added at at least one embedding position includes: generating a J-bit watermark code in a predetermined format (such as a binary code) based on the target watermark information, where J is a positive integer; converting the j-bit target code in the J-bit watermark code into an invisible j-bit predetermined character, and the predetermined character includes a zero-width character; sequentially adding the j-bit predetermined characters to the j target embedding positions among the multiple embedding positions, where j is a positive integer less than or equal to J.
[0117] In one example, the target watermark information is converted into an encoded sequence in a specific format (such as binary, UTF-8 encoding), where J represents the total number of bits of the encoded sequence. For example, after encoding "©2024 AI-generated" using UTF-8, it is converted into a binary bitstream. For example, a total of J bits are generated, such as a 16-bit binary bitstream "0010101101001001" for processing.
[0118] Convert the j-bit target encoding in the J-bit watermark encoding into an invisible j-bit predetermined character. For example, one of the characters "0" or "1" in the binary encoding can be converted into a zero-width character for embedding.
[0119] Specifically, sequentially add the j-bit predetermined characters to the j target embedding bits among the multiple embedding bits.
[0120] For example, when using "1" as the target encoding, the method of embedding the watermark encoding "0010101101001001" into the statement "Zhang San failed to fulfill the payment obligation agreed in the contract within a reasonable period" is as follows:
[0121] First, the embeddable grammar bits selected after segmenting "Zhang San failed to fulfill the payment obligation agreed in the contract within a reasonable period" are "Zhang San", "in", "within a reasonable period",...
[0122] The first bit "0" of the watermark encoding corresponds to the first embedding bit "Zhang San", the second bit "0" of the watermark encoding corresponds to the second embedding bit "in", the third bit "1" of the watermark encoding corresponds to the third embedding bit "within a reasonable period",... Convert the third bit "1" of the watermark encoding into a zero-width character and embed it correspondingly into the third embedding bit "within a reasonable period" (the positions where the watermark encoding is "0" are not processed). That is, sequentially insert zero-width characters into the selected target embedding bits. The zero-width character can be, for example, a zero-width space (ZeroWidth Space, ZWS), whose Unicode encoding is U+200B, to complete the embedding of the watermark information. The embedded text is: Zhang San failed to fulfill the payment obligation agreed in the contract within a reasonable period\u200B.
[0123] By converting the watermark encoding into invisible characters and sequentially adding them to the embedding bits, an efficient and concealed watermark embedding method is achieved, which can embed the watermark information into the text without affecting the readability of the text.
[0124] Next, it will be combined with Figures 4 - 7 Introduce the process of hierarchical watermark embedding provided according to the embodiments of the present application.
[0125] Figure 4 Schematically shows a flowchart of the statement-level watermark embedding method according to the embodiments of the present application. As Figure 4As shown, the method of this embodiment includes operation S410 and operation S420.
[0126] In operation S410, the target watermark information is encoded in a predetermined format to generate an M-bit initial code, where M is a positive integer.
[0127] In operation S420, at least part of the M-bit initial code is added to the first embedding bit.
[0128] In one example, first preprocess the target watermark information "©2024 Text AI Generated", that is, encode the target watermark information in a predetermined format to generate an M-bit initial code. Assume M = 16, and the initial code (B info ) is 0010101101001001. When performing sentence-level watermark information embedding, at least part of the M-bit initial code is added to the first embedding bit. For example, in the sentence "Zhang San failed to fulfill the payment obligation stipulated in the contract within a reasonable period and still did not pay after being repeatedly urged by Li Si.", the first embedding bits (candidate bit 1, candidate bit 2, and candidate bit 3) with an observation probability ≥ 0.6 are detected and screened out through the sentence-level hidden Markov model as follows:
[0129] Candidate bit 1: between "in" and "within a reasonable period" (adverbial-middle structure, probability 0.8);
[0130] Candidate bit 2: between "contractually agreed" and "of" (attributive-middle structure, probability 0.7);
[0131] Candidate bit 3: between "repeatedly" and "urged" (adverbial-middle structure, probability 0.9).
[0132] Since the first 3 bits of the 16-bit initial code 0010101101001001 are 001, the watermark information embedding operation in this sentence is: embed the 1st bit 0 (do not insert ZWS), embed the 2nd bit 0 (do not insert ZWS), embed the 3rd bit 1 (insert \u200B). Finally, the sentence after embedding the watermark information becomes: "Zhang San failed to fulfill the payment obligation stipulated in the contract within a reasonable period and still did not pay after being repeatedly \u200B urged by Li Si." The actually embedded bit stream needs to match all candidate bits in order. This is only a simplified example here.
[0133] Figure 5 Schematically shows a flowchart of the paragraph-level watermark embedding method according to an embodiment of the present application. As Figure 5 shown, it includes operation 510 and operation 520.
[0134] In operation S510, the M-bit initial code generated based on the target watermark information is encoded based on the error correction mechanism to generate an (M+N)-bit error correction code, where the error correction code includes the M-bit initial code and an N-bit redundant check code, where N is a positive integer; the redundant check code is used to: in the case where the initial code is modified, encode the modified initial code to restore the modified initial code.
[0135] In operation S520, at least part of the N-bit redundancy check code is added to the second embedded bits.
[0136] In one example, the 16-bit initial code generated based on the target watermark information is encoded based on the error correction mechanism to generate a (16+4)-bit error correction code. For example, the Reed-Solomon (RS) coding technology can be used based on the 16-bit initial code. Generate 4-bit redundant check code (Assuming the redundant bits are 1100), the redundant check code can restore the modified initial code when the initial code is modified, and finally obtain a 20-bit error correction code consisting of 16-bit initial code and 4-bit redundant check code:
[0137] ;
[0138] When embedding paragraph-level watermark information, at least part of the N-bit redundant check code is added to the second embedding bit. Assume that the target paragraph contains 3 sentences:
[0139] Statement 1: First, Zhang San failed to fulfill his payment obligation.
[0140] Sentence 2: Secondly, Li Si made repeated demands but to no avail.
[0141] Statement 3: Finally, Zhang San committed a breach of contract.
[0142] Based on the paragraph-level hidden Markov model, the logical structure of the paragraph is detected as a progressive structure (connecting words: "first" (0.7), "second" (0.6), "last" (0.8)). The three second embedding positions are determined based on the observation probability:
[0143] Candidate 1: "First," after (probability 0.7).
[0144] Candidate 2: "Secondly," (probability 0.6).
[0145] Candidate position 3: "Finally," after (probability 0.8).
[0146] Due to the redundant check code is 1100, then the watermark information embedding operation in this paragraph is: embed Position 1 1 (insert \u200B), embedded Position 2 1 (insert \u200B), embedded Bit 3 is 0 (not inserted).
[0147] The final paragraph after embedding the watermark information becomes:
[0148] "First, Zhang San failed to fulfill his payment obligations.
[0149] Secondly, Li Si made repeated demands but to no avail.
[0150] Finally, Zhang San breached the contract.”
[0151] By encoding the target watermark information to generate an initial code, the watermark information can be embedded in the text in digital form. The encoding process can compress or transform the watermark information to accommodate the capacity limitations of the embedding bit, thereby improving the watermark's confidentiality and anti-interference capabilities. The completeness of the embedded watermark information can be flexibly selected based on the available embedding bit space and the importance of the watermark information, thereby increasing the flexibility of watermark embedding. By introducing an error correction coding mechanism to generate an error correction code and adding a redundant check code to the second embedded bit, the robustness of the watermark is enhanced. Even if the watermark information is partially damaged during transmission or storage, the redundant check code can be used to correct the error and restore it, ensuring the integrity and accuracy of the watermark information.
[0152] Figure 6 The flowchart of the chapter-level watermark embedding method according to an embodiment of the present application is schematically shown. Figure 6 As shown, it includes operations S610 to S630.
[0153] In operation 610 , a plurality of groups of reference information codes corresponding to a plurality of paragraphs are generated based on historical record information of each of a plurality of paragraphs included in each chapter of the text to be processed being added with target watermark information.
[0154] During the watermarking process, the watermark information added to each paragraph can be recorded, including the historical records of the initial code and redundant check code. After the adding is completed, the initial code and redundant check code added to each paragraph can be determined based on the historical records, that is, multiple groups of reference information codes corresponding to multiple paragraphs can be obtained.
[0155] In one example, the reference information code is intermediate information used to generate chapter codes and is related to the paragraph watermark information. Specifically, the reference information code is the watermark information added to the paragraph and is used in the subsequent generation of chapter codes to ensure the integrity of the text and the consistency of the watermark information. The reference information code can be understood as a watermark feature identifier for the paragraph, carrying the key characteristics of the paragraph watermark information and used in the subsequent generation of chapter-level codes, thereby achieving watermark protection and verification for the entire chapter.
[0156] In operation S620, a chapter code for a chapter is generated based on multiple sets of reference information codes. For example, it may be to perform a hash calculation on multiple sets of reference information codes of multiple paragraphs included in each chapter to obtain the chapter code.
[0157] In operation S630, at least part of the chapter code is added to the third embedding bit. For example, it is added at the end of the chapter.
[0158] According to an embodiment of the present application, by generating chapter watermark information based on the added sentence-level and paragraph-level watermark information in the chapter paragraphs, the association at the sentence level, paragraph level, and chapter level can be established. The multi-level associated watermark is not easily cracked and has higher security, and the robustness is improved through the hierarchical association of the watermark.
[0159] Figure 7 Schematically shows a flowchart of a chapter code generation method according to an embodiment of the present application. As Figure 7 shown, the method of generating a chapter code for a chapter based on multiple sets of reference information codes may further include operations S710 to S740.
[0160] In operation S710, hash operations are respectively performed on multiple sets of reference information codes to generate multiple sets of paragraph codes corresponding to multiple paragraphs.
[0161] In operation S720, the structure of the hash tree used for performing the hash operation based on multiple sets of reference information codes is determined.
[0162] According to an embodiment of the present application, the hash tree includes multiple layers of nodes, and a node represents a coding value to be determined.
[0163] In operation S730, based on multiple sets of reference information codes, the coding values of each layer of nodes are determined layer by layer until the coding value of the root node is determined.
[0164] According to an embodiment of the present application, the hash tree includes K layers; determining the coding values of each layer of nodes layer by layer based on multiple sets of reference information codes includes: using the coding value of the (k - 1)th layer of nodes to calculate the coding value of the kth layer of nodes, where multiple sets of reference information codes are used as the coding values of multiple first-layer nodes; k = 2,..., K.
[0165] In operation S740, a chapter code is generated based on the coding value of the root node.
[0166] According to an embodiment of the present application, generating a chapter code based on the coding value of the root node includes: performing code bit compression on the coding value of the root node to obtain the chapter code.
[0167] In one example, after generating multiple groups of paragraph codes corresponding to multiple paragraphs, a hash tree is constructed, such as a Merkle tree. The construction process is actually an integration of all paragraph watermark information within a chapter. Determine the structure of the Merkle tree, that is, determine the levels and node organization methods of the hash tree. Usually, the Merkle tree is a binary tree structure, but it can also be other types of hash trees. The number of leaf nodes depends on the number of paragraphs. If the number of paragraphs is not a power of 2, the last layer is filled with repeated leaf nodes.
[0168] For a hash tree with K layers, starting from the first layer, use multiple groups of reference information codes as the coding values of the nodes in the first layer. For the k-th layer (k = 2,..., K), use the coding values of the nodes in the k - 1 layer for splicing and hashing operations to obtain the coding values of the nodes in the k-th layer. Calculate layer by layer upwards until the coding value of the root node is obtained. Specifically, starting from the leaf nodes, calculate the hash values of the parent nodes layer by layer upwards. For each parent node, after splicing the hash values of its two child nodes, perform a hashing operation to obtain the hash value of this parent node.
[0169] For a chapter containing 4 paragraphs, assume the reference information codes of each paragraph are H1, H2, H3, and H4 respectively. Calculate the hash values of multiple groups of reference information codes. For example, generate a 128-bit hash value based on multiple groups of reference information codes and take the first 32 bits to generate multiple groups of paragraph codes H(H1), H(H2), H(H3), H(H4) corresponding to multiple paragraphs. These paragraph codes will be used as the leaf nodes of the hash tree.
[0170] Taking a 3-layer hash tree structure as an example (K = 3):
[0171] The first layer: The paragraph code H(Hx) of paragraph 1, the paragraph code H(H2) of paragraph 2, the paragraph code H(H3) of paragraph 3, the paragraph code H(H4) of paragraph 4.
[0172] The second layer: Calculate the hash of the combined nodes in the first layer.
[0173] The third layer (root node): Calculate the hash of the combined nodes in the second layer.
[0174] The encoding values of the nodes in the first layer are used to calculate the encoding values of the nodes in the second layer. The encoding value of node 1 in the second layer is: H(H(H1)+H(H2)); the encoding value of node 2 in the second layer is: H(H(H3)+H(H4)). The encoding value of the nodes in the second layer is used to calculate the encoding value of the third layer (root node): the encoding value of the root node is: H(H(H(H1)+H(H2))+H(H(H3)+H(H4))). Finally, the hash value of the root node is calculated. The hash value of the root node represents a summary of the content of the entire chapter. The root hash value is unique, and any slight change in any paragraph will cause the root hash value to change significantly. After obtaining the encoding value of the root node, the encoding value of the root node can be directly used as the chapter code. As a preference, in order to improve embedding efficiency and reduce the length of the code, the encoding value of the root node can be compressed to obtain a shorter binary bit stream as the chapter code.
[0175] After obtaining the chapter code, the compressed chapter code is converted into invisible characters, such as zero-width characters, and embedded into the third embedding position. The specific embedding process is the same as the aforementioned sentence-level and paragraph-level watermark information embedding process, which will not be repeated here.
[0176] According to the embodiments of this application, by generating chapter codes based on a hash tree, the information of each paragraph can be implicitly stored in the text, eliminating reliance on external storage, ultimately achieving efficient and reliable copyright protection with semantic losslessness and structural integrity. Furthermore, because the information generated by the hash tree is self-verifying, it also facilitates subsequent verification of watermark information.
[0177] By generating reference information codes based on the watermarking history of a paragraph and then generating chapter codes based on these reference information codes, watermark information can be integrated and identified at the chapter level, further enhancing copyright protection and traceability of the text. Adding at least part of the chapter code to the third embedding position allows watermark information to be embedded at the chapter level, achieving identification and protection of the watermark on a wider scale.
[0178] Figure 8 The following is a schematic diagram showing a watermark information embedding process according to an embodiment of the present application.
[0179] like Figure 8As shown, in one example, a processor embeds a watermark into text (which can be generated by an LLM model, for example). The processor performs semantic structure bit detection, which can be to call a trained multi-level Hidden Markov Model (HMM), for example including a first detection model, a second detection model, and a third detection model, to perform semantic structure bit detection, and determine the final embedding bits from multiple candidate embedding bits obtained from the detection, including a first embedding bit at the target word in a sentence, a second embedding bit at the target sentence in a paragraph, and a third embedding bit at the target paragraph in a chapter, respectively.
[0180] In the watermark embedding stage, first obtain the watermark information input by the user, convert the watermark information into an M-bit binary bitstream, generate an M+N-bit stream through RS coding, where the M-bit binary bitstream is the initial code and the N-bit binary bitstream is the redundant check code.
[0181] Perform word segmentation on the text to obtain a word sequence, input the word sequence into the first detection model, determine the grammatical structures to which the multiple words in the sentence of the text to be processed belong, and based on the grammatical structures to which the multiple words belong, take the position where the target word belonging to a predetermined grammatical structure is located as the first embedding bit, and insert zero-width spaces in sequence according to the M-bit binary stream at the first embedding bit, with 1 for insertion and 0 for non-insertion.
[0182] Through the second detection model, determine the logical structures to which the multiple sentences in the paragraph of the text to be processed belong, and based on the logical structures to which the multiple sentences belong, take the position where the target sentence belonging to a predetermined logical structure is located as the second embedding bit; insert the corresponding zero-width spaces at the second embedding bit according to the redundant check code.
[0183] During the embedding process, sequentially record the watermark information inserted at all the first and second embedding bits within the paragraph to form a paragraph binary sequence, perform a hash operation on the paragraph binary bitstream to obtain a paragraph hash value, use the paragraph hash as a leaf node, calculate the parent node hash layer by layer, finally obtain the root hash, and compress the root hash into a low-bit bitstream. Determine the discourse structures to which the multiple paragraphs in the chapter of the text to be processed belong according to the third detection model, and based on the discourse structures to which the multiple paragraphs belong, take the position where the target paragraph belonging to a predetermined discourse structure is located as the third embedding bit, and insert zero-width spaces at the third embedding bit to complete the hierarchical watermark embedding of the text structure features.
[0184] Furthermore, while inserting zero-width spaces at the third embedding bit, a delimiter (such as a carriage return) can be synchronously inserted to separate each code to distinguish the types of the inserted codes, which is convenient for subsequent restoration of the chapter watermark code.
[0185] For example, when inserting a watermark code of 01011 at the end of a certain chapter, the inserted characters are: \ru2008B\r\ru2008B\ru2008B. In this way, when restoring the chapter watermark subsequently, "u2008B" is restored to the code "1". Since the "\r" is a character segmentation position, the uninserted code "0" can be restored.
[0186] The text processing method provided by the embodiments of this application further includes a watermark detection process, which will be described below in combination with Figure 9 introduce the watermark detection process.
[0187] Figure 9 Schematically shows a flowchart of a watermark information restoration method according to an embodiment of this application. As Figure 9 shown, it includes operations S910 to S930.
[0188] In operation S910, the watermark information of multiple paragraphs included in the chapter of the target text with the target watermark information added is restored respectively to obtain multiple groups of restored watermark information codes corresponding to the multiple paragraphs.
[0189] In operation S920, the multiple groups of restored watermark information codes are processed to obtain a restored code corresponding to the chapter. For example, the same method as generating the chapter code can be used to perform a hash operation on the multiple groups of restored watermark information codes respectively to generate multiple groups of restored paragraph codes corresponding to the multiple paragraphs; based on the restored paragraph codes, a restored code corresponding to the chapter is obtained.
[0190] In operation S930, the chapter code extracted from the chapter containing multiple paragraphs is compared and verified with the restored code.
[0191] In one example, the target text in operation S910 is the text with the target watermark information added. The watermark information embedded in each paragraph is restored. Specifically, the hidden Markov model is used to detect the text structure, identify grammar, logic, and discourse structure, and determine the embedding positions (such as the adverbial position in the adverbial-middle structure, the position after the conjunction in the progressive structure, etc.). The watermark information fragments at the embedding positions are extracted. The watermark information fragments of all paragraphs are integrated to generate the restored code of the chapter. The chapter code embedded in the document is extracted. The restored code and the chapter code are compared to determine whether they are consistent to verify the integrity of the document. The hash value of the restored code is calculated and compared with the chapter code pre-embedded and extracted from the embedding positions of the chapter. If the two are consistent, the verification passes, indicating that the chapter has not been tampered with.
[0192] According to an embodiment of the present application, different from the conventional key verification method, in the method of the embodiment of the present application, since the chapter encoding is calculated by a hash tree during watermark embedding and stored, this method can form a self-verification mechanism during the subsequent watermark verification process. The watermark information can be obtained through hash operations without additional storage and transmission of keys, improving the verification efficiency, especially suitable for the scenario of offline verification.
[0193] According to an embodiment of the present application, the method for separately restoring watermark information for multiple paragraphs included in a chapter of a target text to which target watermark information is added includes:
[0194] First, by performing text structure detection on the target text, at least one detected embedding position where watermark information can be added in the target text is determined;
[0195] After that, the initial encoding and / or redundancy check code added to each of the multiple paragraphs are extracted;
[0196] Then, according to the initial encoding and / or redundancy check code added to each of the multiple paragraphs, and the position information of at least one detected embedding position, watermark information restoration is performed on each paragraph to obtain multiple groups of restored watermark information codes corresponding to the multiple paragraphs.
[0197] Among them, in an example of performing text structure detection on the target text, based on the aforementioned multi-level hidden Markov model, text structure detection is performed on the target text to obtain the structure attribution probabilities of each word, sentence, and paragraph, and the detected embedding positions are determined according to the probabilities. That is, the sentence-level hidden Markov model is used to detect the text to identify the grammatical structure of each word; the paragraph-level hidden Markov model is used to detect the text to identify the logical structure of each sentence; the chapter-level hidden Markov model is used to detect the text to identify the discourse structure of each paragraph. At least one detected embedding position where watermark information can be added in the target text is determined. The specific determination process of the detected embedding position is the same as the determination process of the aforementioned watermark embedding position and will not be elaborated here.
[0198] Among them, when extracting the initial encoding and / or redundancy check code added to each of the multiple paragraphs, since the initial encoding and / or redundancy check code are added to the target text in the form of predetermined characters, it can be to use a character recognition tool to extract all the predetermined characters (such as zero-width spaces) added to the text.
[0199] According to an embodiment of the present application, based on the predetermined characters included in each of the multiple paragraphs, and the position information of at least one detected embedding position, the method for restoring watermark information for each paragraph includes the following method:
[0200] First, according to the initial encoding added to the paragraph and the position information of at least one detected embedding position, first watermark information restoration is performed on the paragraph to obtain an initial restored information code corresponding to the paragraph;
[0201] After that, according to the redundancy check code added to the paragraph and the position information of at least one detected embedded bit, the second watermark information of the paragraph is restored to obtain an error correction and restoration information code corresponding to the paragraph.
[0202] Then, the error correction is performed on the initial restoration information code by using the error correction and restoration information code, and a restored watermark information code corresponding to the paragraph is obtained based on the error-corrected initial restoration information code.
[0203] During the process of restoring the watermark information, error correction can be performed based on the extracted check code (the error correction mechanism of the RS coding algorithm). After the watermark information is partially damaged, it can be restored to ensure the normal progress of watermark verification and better information reliability.
[0204] According to the embodiments of the present application, further, the method for restoring the first watermark information of each paragraph and the method for restoring the second watermark information of each paragraph are the same, and the following method can be adopted:
[0205] Extract the predetermined characters included in each of the multiple paragraphs (the initial code and / or the redundancy check code are added to the target text in the form of predetermined characters); based on the predetermined characters included in each of the multiple paragraphs and the position information of at least one detected embedded bit, the watermark information of each paragraph is restored.
[0206] Because when embedding the watermark information, j target codes in the J-bit watermark code are converted into invisible j-bit zero-width characters and added to j target embedding bits among the multiple embedding bits, that is, the zero-width characters inserted in the paragraph only represent part of the codes in the initial code and / or the redundancy check code. For example, when embedding the watermark information in binary code (such as 0110000...), the "1" in the code is converted into a zero-width character and embedded in a specific position, and the "0" in the code is not embedded. Then, after extracting the predetermined characters included in each of the multiple paragraphs, the obtained codes are only partial codes, all of which are the character "1", such as "111111...", rather than the complete watermark "0110000..." originally added. Therefore, it is necessary to combine the position information of at least one detected embedded bit to restore the watermark information of each paragraph.
[0207] For example, it is detected that a certain paragraph "Great progress has been made in scientific and technological innovation" has 3 embedding bits: the first bit "scientific and technological innovation", the second bit "has been made", and the third bit "great progress"; when extracting the watermark, only "1" is extracted at the second bit "has been made"; then the complete watermark of this paragraph should be "010" after restoration, that is, the second bit corresponds to "1", and the first bit and the third bit correspond to "0".
[0208] According to an embodiment of the present application, the predetermined characters included in each of the multiple paragraphs include: a first predetermined character added to at least one detected embedding bit, and a second predetermined character not in at least one detected embedding bit. The first predetermined character and the second predetermined character belong to the same character. For example, they are both zero-width spaces. The difference is that the first predetermined character is added to the detected embedding bit, and the second predetermined character is added to the non-detected embedding bit. This indicates that the first predetermined character added to the detected embedding bit belongs to the watermark information, and the second predetermined character added to the non-detected embedding bit does not belong to the watermark information and may be added by the user through tampering. Therefore, in order to restore the correct watermark information, the second predetermined character added to the non-detected embedding bit needs to be filtered out.
[0209] Specifically, based on the predetermined characters included in each of the multiple paragraphs and the position information of at least one detected embedding bit, restoring the watermark information for each paragraph includes: restoring the watermark information for each paragraph based on the first predetermined character and the position information of at least one detected embedding bit. That is, only the first predetermined characters added to the detected embedding bits are extracted for processing, and the second predetermined characters added to the non-detected embedding bits are not processed.
[0210] In one example, the predetermined character can be, for example, a zero-width character. The initial encoding and / or redundancy check code are added to the target text in the form of zero-width characters. The text is scanned to determine the positions of all zero-width characters. Since the text may be tampered with during transmission or use, there may be a situation where zero-width characters appear in non-embedding bits. For example, in a certain paragraph, zero-width characters are extracted at embedding bit 1, embedding bit 3, and non-embedding bit 2. At this time, the zero-width characters at embedding bit 1 and embedding bit 3 are the first predetermined characters, and the zero-width character at non-embedding bit 2 is the second predetermined character. Therefore, only the first predetermined characters at the embedding bits need to be concerned. Combining the position information of the detected embedding bits, it is possible to determine which characters belong to the initial encoding or the redundancy check code. Based on the zero-width characters corresponding to the initial encoding and the positions of the detected embedding bits, the watermark information code is restored. Restoring the watermark information based on the first predetermined character and the embedding bit position information avoids the interference of non-embedding bit characters, filters out illegal characters, and improves the accuracy and reliability of the watermark restoration information.
[0211] Figure 10 Schematically shows a schematic diagram of the entire process of watermark embedding and watermark detection according to an embodiment of the present application.
[0212] Hereinafter, in combination with Figure 10 , an exemplary description is given of the method for the watermark information embedding and restoration process.
[0213] As Figure 10As shown in the figure, in the watermark embedding stage, after the text is generated, text structure detection is performed on the text, and the text structure information of the text to be processed is obtained. For example, the positions with an observation probability greater than the threshold Q are detected through a multi-level hidden Markov model as candidate embedding positions, and the positions with an observation probability less than the preset threshold Q are filtered out and this position is skipped. It is also possible to further screen from the candidate embedding positions to determine the final target embedding position. After marking the embedding position, the watermark information is embedded, that is, the watermark information is converted into a predetermined character and inserted into the embedding position.
[0214] In the watermark detection stage, first, for the target text embedded with the watermark, the predetermined characters (such as zero-width characters) are extracted from it, and the positions with an observation probability greater than the threshold Q are detected through a multi-level hidden Markov model as the detected embedding positions, and the predetermined characters at this position are retained, and the zero-width characters added at the non-detected embedding positions are filtered out (filtering interference signals) to obtain the correct watermark coding information.
[0215] Hereinafter, taking the watermark "©2024 Text AI Generated" designed by the user and embedded in the professional document generated by the large model as an example, the method of watermark embedding and watermark detection in the embodiments of the present application will be exemplarily described.
[0216] Exemplarily, the method of embedding watermark information is as follows:
[0217] Watermark preprocessing:
[0218] The original watermark "©2024 Text AI Generated" input by the user through the client is encoded in UTF-8 and then converted into a binary bit stream (a total of 88 bits, and the first 16 bits are used as an example: 0010101101001001); RS(20,16) encoding (M = 16 information bits, N = 4 redundant bits) is used to generate a 20-bit bit stream.
[0219] Sentence-level embedding:
[0220] For the sentence in the "Finding of Facts" part of the target paragraph: "Zhang San failed to perform the payment obligation stipulated in the contract within a reasonable period and still did not pay after being urged by Li Si many times." Syntax analysis is performed, and after word segmentation, it is decoded using a sentence-level hidden Markov model. The positions with an observation probability greater than 0.6 are screened out as candidate embedding positions.
[0221] For example, "within a reasonable period" is a modifier-object structure, and the observation probability is 0.8 (≥0.6), then it is marked as candidate position 1. The determination of other legal embedding positions is the same and will not be elaborated here.
[0222] Insert ZWS () at candidate position 1 (between "at" and "within a reasonable time limit") to represent 1, and embed other information bits in sequence. The final statement becomes: "Zhang San failed to fulfill the payment obligation stipulated in the contract within a reasonable time limit and still did not pay after being urged by Li Si many times."
[0223] Paragraph-level embedding:
[0224] For the entire content of the target paragraph "Finding of Facts" (e.g., containing 3 statements), use the Hidden Markov Model to decode and select the positions with an observation probability greater than 0.6 as redundant embedding positions.
[0225] For example, the probability of the word "First" as a sentence connective in "First, Zhang San failed to fulfill the payment obligation" is 0.7 (≥0.6), which is marked as candidate position 1. The determination of other legal embedding positions is the same and will not be elaborated here.
[0226] Insert ZWS at candidate position 1 to represent redundant bit 1, and the final paragraph becomes: "First, Zhang San failed to fulfill the payment obligation..."
[0227] Chapter-level embedding:
[0228] First, construct a Merkle tree. Calculate the 128-bit hash value of each paragraph bitstream (take the first 32 bits as an example) to generate leaf nodes H(P1), H(P2), H(P3), calculate the parent node hash layer by layer, and obtain the root hash such as H_root=a1b2c3d4, and compress it into an 8-bit bitstream (10110011).
[0229] The full text conforms to the structure of "Finding of Facts, Basis for Finding, Result of Finding". According to the chapter-level Hidden Markov Model decoding, insert ZWS and \r delimiters at the legal position (observation probability 0.95≥0.6) at the end of the "Result of Finding" chapter. For example, insert \r at the end of the chapter: "The result of finding is as follows: Pay liquidated damages of 100,000 yuan."
[0230] Exemplarily, the method for watermark detection is as follows:
[0231] Invisible character extraction: Use a character detection tool to scan the text and locate the positions of ZWS () (1 at the sentence level, 1 at the paragraph level, 1 at the chapter level) and the \r delimiter (1).
[0232] Sentence-level detection of embedded bit detection: Input the word sequence into the sentence-level Hidden Markov Model, confirm that "within a reasonable time limit" is an adverbial-modifier structure, and the observation probability is 0.8 (≥0.6), and retain this ZWS.
[0233] Paragraph-level detection and embedded bit detection: The paragraph logical structure is general-to-specific (P1), the connection word probability of "firstly" is 0.7 (≥0.6), and this ZWS is retained.
[0234] Chapter-level detection and embedded bit detection: The chapter order is "Factual Findings", "Basis for Findings", "Results of Findings", and the observation probability is 0.95 (≥0.6), and the root hash bitstream is retained.
[0235] Hierarchical bitstream extraction and error correction: Based on the position information of the detected embedded bits determined by the sentence-level hidden Markov model and the sentence-level ZWS, 16 bits are restored (Example: 0010101101001001); Based on the position information of the detected embedded bits determined by the paragraph-level hidden Markov model and the paragraph-level ZWS, 4 redundant bits are restored (Example 1010). Through RS decoding, 2-bit errors (such as bitstream errors caused by deletion or modification) are corrected and restored to the original .
[0236] Using the restored Reassembled into the bitstream of each paragraph (the splicing method is like the watermark embedding process), calculate the hash value to construct a Merkle tree, and calculate and obtain a new root hash , Compare the new root hash with the root hash bitstream embedded at the end of the chapter and paragraph. If the comparison is consistent, the watermark verification passes, indicating that the watermark content has not been tampered with.
[0237] The text processing method provided by the embodiments of the present application performs text structure detection on the text to be processed through a pre-trained hidden Markov model, and can accurately identify the grammatical structure, logical structure, and discourse structure in the text. Determining the embedding bits for adding watermark information based on the text structure information can ensure that the watermark information is embedded in a relatively concealed position in the text and has less impact on the semantics. It not only improves the concealment of the watermark, but also reduces the interference of the watermark on the readability of the text, while ensuring that the watermark information can be stably embedded in the text. Through multi-level embedding at the sentence level, paragraph level, and chapter level, it is possible to more comprehensively protect the copyright information of the text and enhance the anti-attack ability of the watermark. Through multi-level embedding and verification mechanisms, the concealment and robustness of the watermark are improved, making the watermark more difficult to detect and remove. At the same time, it is also convenient to quickly and accurately verify the integrity of the text and the authenticity of the copyright information during detection. It effectively solves the problem of copyright protection for texts generated by large language models: it not only ensures the concealment of the watermark, but also enhances the robustness through hierarchical association and RS coding to resist local deletion and modification. At the same time, based on the self-checking hash tree, it avoids external storage dependence, and finally realizes efficient and reliable copyright protection with semantic losslessness and structural integrity, providing a highly adaptable technical solution for content traceability and attribution verification of high-value texts.
[0238] Based on the above text processing method, the present application also provides a text processing device. The following will be combined with Figure 11 to describe the device in detail.
[0239] Figure 11 The structural block diagram of the text processing device according to an embodiment of the present application is schematically shown.
[0240] As Figure 11 shown, the text processing device 1100 of this embodiment includes a detection module 1110, a determination module 1120, and an addition module 1130.
[0241] The detection module 111 is configured to perform text structure detection on the text to be processed to obtain the text structure information of the text to be processed, where the text structure information includes at least one of the following: the grammatical structures to which multiple words in the sentences of the text to be processed belong, the logical structures to which multiple sentences in the paragraphs of the text to be processed belong, and the discourse structures to which multiple paragraphs in the chapters of the text to be processed belong. In one embodiment, the detection module 1110 may be used to perform the operation S210 described above, which will not be elaborated here.
[0242] The determination module 1120 is configured to determine at least one embedding position for adding watermark information in the text to be processed based on the relevance between the text structure information and the context semantics in the text to be processed. In one embodiment, the determination module 1120 may be used to perform the operation S220 described above, which will not be elaborated here.
[0243] The addition module 1130 is configured to add the target watermark information to be added at at least one embedding position. In one embodiment, the addition module 1130 may be used to perform the operation S230 described above, which will not be elaborated here.
[0244] According to an embodiment of the present application, the determination module includes a first determination sub-module, a second determination sub-module, and a third determination sub-module.
[0245] The first determination sub-module is configured to use the position where the target word belonging to a predetermined grammatical structure and having a context semantic relevance meeting a predetermined condition is located as the first embedding position based on the grammatical structures to which multiple words belong.
[0246] The second determination sub-module is configured to use the position where the target sentence belonging to a predetermined logical structure and having a context semantic relevance meeting a predetermined condition is located as the second embedding position based on the logical structures to which multiple sentences belong.
[0247] The third determination sub-module is configured to use the position where the target paragraph belonging to a predetermined discourse structure and having a context semantic relevance meeting a predetermined condition is located as the third embedding position based on the discourse structures to which multiple paragraphs belong.
[0248] According to an embodiment of the present application, the adding module includes a first generating sub-module and a first adding sub-module.
[0249] The first generating sub-module is configured to encode the target watermark information in a predetermined format to generate an M-bit initial code, where M is a positive integer.
[0250] The first adding sub-module is configured to add at least a part of the M-bit initial code to the first embedding bit.
[0251] According to an embodiment of the present application, the adding module further includes a second generating sub-module and a second adding sub-module.
[0252] The second generating sub-module is configured to encode the M-bit initial code generated based on the target watermark information based on an error correction mechanism to generate an (M + N)-bit error correction code, where the error correction code includes the M-bit initial code and N-bit redundant check codes, and N is a positive integer; the redundant check codes are used for: encoding and restoring the modified initial code in the case where the initial code is modified;
[0253] The second adding sub-module is configured to add at least a part of the N-bit redundant check codes to the second embedding bit.
[0254] According to an embodiment of the present application, the adding module further includes a third generating sub-module, a fourth generating sub-module, and a third adding sub-module.
[0255] The third generating sub-module is configured to generate multiple groups of reference information codes corresponding to multiple paragraphs according to the historical record information of each paragraph included in each chapter of the text to be processed, where the target watermark information is added to each paragraph;
[0256] The fourth generating sub-module is configured to generate a chapter code for the chapter based on the multiple groups of reference information codes.
[0257] The third adding sub-module is configured to add at least a part of the chapter code to the third embedding bit.
[0258] According to an embodiment of the present application, the fourth generating sub-module includes a first generating unit and a second generating unit.
[0259] The first generating unit is configured to perform a hash operation on each of the multiple groups of reference information codes to generate multiple groups of paragraph codes corresponding to multiple paragraphs.
[0260] The second generating unit is configured to generate a chapter code for the chapter based on the multiple groups of paragraph codes.
[0261] According to an embodiment of the present application, the second generating unit includes a first determining sub-unit, a second determining sub-unit, and a generating sub-unit.
[0262] A first determination subunit, configured to determine the structure of a hash tree used for performing a hash operation based on multiple groups of reference information codes, where the hash tree includes multiple layers of nodes, and a node represents an encoding value to be determined;
[0263] A second determination subunit, configured to layer by layer determine the encoding values of each layer of nodes based on multiple groups of reference information codes until the encoding value of the root node is determined;
[0264] A generation subunit, configured to generate a chapter code based on the encoding value of the root node.
[0265] According to an embodiment of the present application, the second determination subunit is further configured to calculate the encoding value of the kth layer node by using the encoding value of the (k - 1)th layer node, where multiple groups of reference information codes are used as the encoding values of multiple first layer nodes; k = 2,..., K.
[0266] According to an embodiment of the present application, the generation subunit is further configured to perform code bit compression on the encoding value of the root node to obtain a chapter code.
[0267] According to an embodiment of the present application, the apparatus further includes: a restoration module, a restored code generation module, and a comparison and verification module.
[0268] The restoration module is configured to respectively perform watermark information restoration on multiple paragraphs included in a chapter of a target text added with target watermark information to obtain multiple groups of restored watermark information codes corresponding to the multiple paragraphs;
[0269] The restored code generation module is configured to process multiple groups of restored watermark information codes to obtain a restored code corresponding to the chapter;
[0270] The comparison and verification module is configured to perform comparison and verification on the chapter code extracted from a chapter including multiple paragraphs and the restored code.
[0271] According to an embodiment of the present application, the restoration module includes a fourth determination sub-module, an extraction sub-module, and a restoration sub-module.
[0272] The fourth determination sub-module is configured to determine at least one detected embedding position where watermark information can be added in the target text by performing text structure detection on the target text;
[0273] The extraction sub-module is configured to extract the initial code and / or redundancy check code added to each of the multiple paragraphs;
[0274] The restoration sub-module is configured to perform watermark information restoration on each paragraph according to the initial code and / or redundancy check code added to each of the multiple paragraphs and the position information of at least one detected embedding position to obtain multiple groups of restored watermark information codes corresponding to the multiple paragraphs.
[0275] According to an embodiment of the present application, the reduction sub-module includes a first watermark information reduction unit, a second watermark information reduction unit, and an error correction unit.
[0276] The first watermark information reduction unit is configured to perform first watermark information reduction on a paragraph according to the initial encoding added to the paragraph and the position information of at least one detected embedding bit, so as to obtain an initial reduction information code corresponding to the paragraph;
[0277] The second watermark information reduction unit is configured to perform second watermark information reduction on a paragraph according to the redundancy check code added to the paragraph and the position information of at least one detected embedding bit, so as to obtain an error correction reduction information code corresponding to the paragraph;
[0278] The error correction unit is configured to correct the initial reduction information code by using the error correction reduction information code, and obtain a restored watermark information code corresponding to the paragraph based on the corrected initial reduction information code.
[0279] According to an embodiment of the present application, the extraction sub-module is further configured to extract predetermined characters included in each of the multiple paragraphs. The predetermined characters included in each of the multiple paragraphs include: a first predetermined character added at at least one detected embedding bit, and a second predetermined character not at at least one detected embedding bit
[0280] According to an embodiment of the present application, the reduction sub-module is further configured to perform watermark information reduction on each paragraph based on the predetermined characters included in each of the multiple paragraphs and the position information of at least one detected embedding bit.
[0281] According to an embodiment of the present application, the reduction sub-module is further configured to perform watermark information reduction on each paragraph based on the first predetermined character and the position information of at least one detected embedding bit.
[0282] According to an embodiment of the present application, the addition module is further configured to generate a J-bit watermark code in a predetermined format based on the target watermark information, where J is a positive integer; convert the j-bit target code in the J-bit watermark code into an invisible j-bit predetermined character, and the predetermined character includes a zero-width character; sequentially add the j-bit predetermined characters to j target embedding bits among the multiple embedding bits, where j is a positive integer less than or equal to J.
[0283] According to an embodiment of the present application, the detection module includes a first detection sub-module, a second detection sub-module, and a third detection sub-module.
[0284] The first detection sub-module is configured to perform a first detection on multiple words based on a first detection model, obtain multiple first probabilities that each word belongs to multiple candidate grammatical structures, and determine the grammatical structure to which the word belongs based on the multiple first probabilities;
[0285] The second detection sub-module is used to perform a second detection on multiple statements based on a second detection model, obtain multiple second probabilities of each statement belonging to multiple candidate logical structures, and determine the logical structure to which the statement belongs based on the multiple second probabilities;
[0286] The third detection sub-module is used to perform a third detection on multiple paragraphs based on a third detection model, obtain multiple third probabilities of each paragraph belonging to multiple candidate discourse structures, and determine the discourse structure to which the paragraph belongs based on the multiple third probabilities.
[0287] According to an embodiment of the present application, the first detection sub-module is further used to determine multiple first probabilities of each word belonging to multiple candidate syntactic structures by using first state transition information and first observation probability information.
[0288] According to an embodiment of the present application, the apparatus further includes: a first acquisition module, a first determination module, and a first training module.
[0289] The first acquisition module is used to acquire a training text, and perform word segmentation processing on multiple sample statements in the training text respectively to obtain multiple sample words;
[0290] The first determination module is used to determine the syntactic structure to which each of the multiple sample words belongs to generate a first label of the training text;
[0291] The first training module is used to train a basic model by using multiple sample words and the first label to obtain a first detection model, wherein the model parameters of the first detection model include first state transition information and first observation probability information, wherein the first state transition information is used to represent the probability of transitioning from a first candidate syntactic structure to a second candidate syntactic structure, and the first observation probability information is used to represent the probabilities of multiple target sample words belonging to the second candidate syntactic structure, and the first candidate syntactic structure and the second candidate syntactic structure belong to at least one of multiple syntactic structures to which the multiple sample words belong.
[0292] According to an embodiment of the present application, the apparatus further includes: a second acquisition module, a second determination module, and a second training module.
[0293] The second acquisition module is used to acquire a training text, and perform clause segmentation processing on multiple sample paragraphs in the training text respectively to obtain multiple sample statements;
[0294] The second determination module is used to determine the logical structure to which each of the multiple sample statements belongs to generate a second label of the training text;
[0295] A second training module, configured to train a base model by using a plurality of sample statements and second labels to obtain a second detection model, wherein the model parameters of the second detection model include second state transition information and second observation probability information, and the second state transition information is used to represent the probability of transitioning from a first candidate logical structure to a second candidate logical structure, and the second observation probability information is used to represent the probabilities of respective target sample statements belonging to the second candidate logical structure, and the first candidate logical structure and the second candidate logical structure belong to at least one of multiple logical structures to which the plurality of sample statements belong.
[0296] According to an embodiment of the present application, the apparatus further includes: a third acquisition module, a third determination module, and a third training module.
[0297] The third acquisition module is configured to acquire a training text and perform a segmentation process on the training text to obtain a plurality of sample paragraphs;
[0298] The third determination module is configured to determine the discourse structure to which each of the plurality of sample paragraphs belongs to generate a third label for the training text;
[0299] The third training module is configured to train a base model by using the plurality of sample paragraphs and the third label to obtain a third detection model, wherein the model parameters of the third detection model include third state transition information and third observation probability information, and the third state transition information is used to represent the probability of transitioning from a first candidate discourse structure to a second candidate discourse structure, and the third observation probability information is used to represent the probabilities of respective target sample paragraphs belonging to the second candidate discourse structure, and the first candidate discourse structure and the second candidate discourse structure belong to at least one of multiple discourse structures to which the plurality of sample paragraphs belong.
[0300] According to embodiments of the present application, any one or more of the detection module 1110, the determination module 1120, and the addition module 1130 may be combined and implemented in one module, or any one of them may be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules may be combined with at least part of the functions of other modules and implemented in one module. According to embodiments of the present application, at least one of the detection module 1110, the determination module 1120, and the addition module 1130 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on substrate, a system on package, an application specific integrated circuit (ASIC), or any other reasonable manner that can integrate or package circuits, etc., implemented by hardware or firmware, or implemented in any one of the three implementation manners of software, hardware, and firmware, or in any appropriate combination of several of them. Alternatively, at least one of the detection module 1110, the determination module 1120, and the addition module 1130 may be at least partially implemented as a computer program module, and when the computer program module is run, it can execute the corresponding functions.
[0301] Figure 12 Schematically shows a block diagram of an electronic device suitable for implementing a text processing method according to an embodiment of the present application.
[0302] As Figure 12 shown, the electronic device 1200 according to an embodiment of the present application includes a processor 1201, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1202 or a program loaded from a storage section 1208 into a random access memory (RAM) 1203. The processor 1201 may include, for example, a general microprocessor (such as a CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (such as an application specific integrated circuit (ASIC)), etc. The processor 1201 may also include on-board memory for caching purposes. The processor 1201 may include a single processing unit or multiple processing units for performing different actions of the method flow according to embodiments of the present application.
[0303] In the RAM 1203, various programs and data required for the operation of the electronic device 1200 are stored. The processor 1201, the ROM 1202, and the RAM 1203 are connected to each other via a bus 1204. The processor 1201 performs various operations of the method flow according to the embodiments of the present application by executing the programs in the ROM 1202 and / or the RAM 1203. It should be noted that the programs can also be stored in one or more memories other than the ROM 1202 and the RAM 1203. The processor 1201 can also perform various operations of the method flow according to the embodiments of the present application by executing the programs stored in one or more memories.
[0304] According to an embodiment of the present application, the electronic device 1200 may further include an input / output (I / O) interface 1205, and the input / output (I / O) interface 1205 is also connected to the bus 1204. The electronic device 1200 may further include one or more of the following components connected to the input / output (I / O) interface 1205: an input portion 1206 including a keyboard, a mouse, etc.; an output portion 1207 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage portion 1208 including a hard disk, etc.; and a communication portion 1209 including a network interface card such as a LAN card, a modem, etc. The communication portion 1209 performs communication processing via a network such as the Internet. A drive 1210 is also connected to the input / output (I / O) interface 1205 as needed. A removable medium 1211, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1210 as needed so that a computer program read from it can be installed into the storage portion 1208 as needed.
[0305] The present application also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or may exist alone without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed, the method according to the embodiments of the present application is implemented.
[0306] According to an embodiment of the present application, the computer-readable storage medium may be a non-volatile computer-readable storage medium, which may include, for example, but not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present application, the computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present application, the computer-readable storage medium may include the above-described ROM 1202 and / or RAM 1203 and / or one or more memories other than ROM 1202 and RAM 1203.
[0307] An embodiment of the present application also includes a computer program product, which includes a computer program that contains program code for executing the method shown in the flowchart. When the computer program product runs in a computer system, the program code is used to enable the computer system to implement the text processing method provided by the embodiment of the present application.
[0308] When the computer program is executed by a processor, it executes the above functions defined in the system / apparatus of the embodiment of the present application. According to an embodiment of the present application, the above-described systems, apparatuses, modules, units, etc. can be implemented by computer program modules.
[0309] In one embodiment, the computer program can rely on tangible storage media such as optical storage devices and magnetic storage devices. In another embodiment, the computer program can also be transmitted and distributed in the form of a signal on a network medium, and is downloaded and installed through the communication part, and / or installed from a removable medium. The program code contained in the computer program can be transmitted by any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0310] In such an embodiment, the computer program can be downloaded and installed from the network through the communication part, and / or installed from a removable medium. When the computer program is executed by a processor, it executes the above functions defined in the system of the embodiment of the present application. According to an embodiment of the present application, the above-described systems, devices, apparatuses, modules, units, etc. can be implemented by computer program modules.
[0311] In accordance with embodiments of the present application, program code for executing the computer programs provided by the embodiments of the present application may be written in any combination of one or more programming languages. Specifically, these computing programs may be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, such as Java, C++, Python, the "C" language, or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device may be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or alternatively, may be connected to an external computing device (e.g., by connecting through the Internet using an Internet service provider).
[0312] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above-mentioned module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and combinations of blocks in the block diagram or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0313] Those skilled in the art can understand that the features described in the various embodiments of the present application can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present application. In particular, without departing from the spirit and teachings of the present application, the features described in the various embodiments of the present application can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present application.
[0314] The above describes the embodiments of the present application. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present application. Although the embodiments are described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. Without departing from the scope of the present application, those skilled in the art can make various substitutions and modifications, and all such substitutions and modifications should fall within the scope of the present application.
Claims
1. A text processing method, characterized in that, The method includes: Performing text structure detection on the text to be processed to obtain text structure information of the text to be processed, where the text structure information includes at least one of the following: the grammatical structures to which multiple words in the sentences of the text to be processed belong, the logical structures to which multiple sentences in the paragraphs of the text to be processed belong, and the discourse structures to which multiple paragraphs in the chapters of the text to be processed belong; Determining at least one embedding position for adding watermark information in the text to be processed based on the relevance between the text structure information and the context semantics in the text to be processed; Adding the target watermark information to be added at the at least one embedding position.
2. The method according to claim 1, wherein Determining at least one embedding position for adding watermark information in the text to be processed based on the relevance between the text structure information and the context semantics in the text to be processed includes at least one of the following: Based on the grammatical structures to which the multiple words belong, taking the position where the target word belonging to a predetermined grammatical structure and with context semantic relevance meeting a predetermined condition is located as the first embedding position; Based on the logical structures to which the multiple sentences belong, taking the position where the target sentence belonging to a predetermined logical structure and with context semantic relevance meeting a predetermined condition is located as the second embedding position; Based on the discourse structures to which the multiple paragraphs belong, taking the position where the target paragraph belonging to a predetermined discourse structure and with context semantic relevance meeting a predetermined condition is located as the third embedding position.
3. The method according to claim 2, wherein Adding the target watermark information to be added at the at least one embedding position includes: Encoding the target watermark information in a predetermined format to generate an M-bit initial code, where M is a positive integer; Adding at least part of the M-bit initial code at the first embedding position.
4. The method according to claim 3, characterized in that, Adding the target watermark information to be added at the at least one embedding position further includes: Encoding the M-bit initial code generated based on the target watermark information with an error correction mechanism to generate an (M + N)-bit error correction code, where the error correction code includes the M-bit initial code and N-bit redundant check codes, and N is a positive integer; the redundant check codes are used for: encoding and restoring the modified initial code in the case where the initial code is modified; Adding at least part of the N-bit redundant check codes at the second embedding position.
5. The method according to claim 4, characterized in that Adding the target watermark information to be added at the at least one embedding position further includes: Generating multiple groups of reference information codes corresponding to multiple paragraphs according to the historical record information of the target watermark information added to each of the multiple paragraphs included in each chapter of the text to be processed; Generating a chapter code for the chapter based on the multiple groups of reference information codes; Adding at least part of the chapter code at the third embedding position.
6. The method according to claim 5, wherein Generating a chapter code for the chapter based on the multiple groups of reference information codes includes: Performing a hash operation on each of the multiple groups of reference information codes to generate multiple groups of paragraph codes corresponding to multiple paragraphs; Generating a chapter code for the chapter based on the multiple groups of paragraph codes.
7. The method according to claim 5, characterized in that Generating a chapter code for the chapter based on the multiple groups of reference information codes includes: Determine the structure of the hash tree used for hash operation based on the multiple sets of reference information codes, where the hash tree includes multiple layers of nodes, and each node represents a coding value to be determined; Based on the multiple sets of reference information codes, layer by layer determine the coding values of each layer of nodes until the coding value of the root node is determined; Generate the chapter code based on the coding value of the root node.
8. The method according to claim 5, wherein The method further includes: Restore the watermark information for each of the multiple paragraphs included in the chapter of the target text added with the target watermark information, to obtain multiple sets of restored watermark information codes corresponding to the multiple paragraphs; Process the multiple sets of restored watermark information codes to obtain a restored code corresponding to the chapter; Compare and verify the chapter code extracted from the chapter including the multiple paragraphs and the restored code.
9. The method according to claim 8, wherein Restoring the watermark information for each of the multiple paragraphs included in the chapter of the target text added with the target watermark information includes: By performing text structure detection on the target text, determine at least one detected embedding position where watermark information can be added in the target text; Extract the initial code and / or redundancy check code added to each of the multiple paragraphs; According to the initial code and / or the redundancy check code added to each of the multiple paragraphs, and the position information of the at least one detected embedding position, restore the watermark information for each paragraph to obtain multiple sets of restored watermark information codes corresponding to the multiple paragraphs.
10. The method according to claim 9, characterized in that Restoring the watermark information for each paragraph includes: According to the initial code added to the paragraph and the position information of the at least one detected embedding position, perform a first watermark information restoration on the paragraph to obtain an initial restored information code corresponding to the paragraph; According to the redundancy check code added to the paragraph and the position information of the at least one detected embedding position, perform a second watermark information restoration on the paragraph to obtain an error correction restored information code corresponding to the paragraph; Use the error correction restored information code to correct the initial restored information code, and based on the corrected initial restored information code, obtain the restored watermark information code corresponding to the paragraph.
11. The method according to claim 9, wherein The initial code and / or the redundancy check code are added to the target text in the form of a predetermined character; Extracting the initial code and / or the redundancy check code added to each of the multiple paragraphs includes: extracting the predetermined characters included in each of the multiple paragraphs; Restoring the watermark information for each paragraph includes: based on the predetermined characters included in each of the multiple paragraphs, and the position information of the at least one detected embedding position, restore the watermark information for each paragraph.
12. The method according to claim 11, wherein The predetermined characters included in each of the multiple paragraphs include: a first predetermined character added at the at least one detected embedding position, and a second predetermined character not at the at least one detected embedding position; Based on the predetermined characters included in each of the multiple paragraphs, and the position information of the at least one detected embedding position, restoring the watermark information for each paragraph includes: Restore the watermark information for each of the paragraphs based on the first predetermined character and the position information of the at least one detected embedded bit.
13. The method according to claim 1, characterized in that, The adding the target watermark information to be added to the at least one embedded bit includes: Generating a J-bit watermark code in a predetermined format based on the target watermark information, where J is a positive integer; Converting j target codes in the J-bit watermark code into invisible j predetermined characters, where the predetermined characters include zero-width characters; Sequentially adding the j predetermined characters to j target embedded bits among the multiple embedded bits, where j is a positive integer less than or equal to J.
14. The method according to claim 1, wherein Performing text structure detection on the text to be processed, and obtaining the text structure features of the text to be processed including at least one of the following: Performing a first detection on the multiple words based on a first detection model to obtain multiple first probabilities that each of the words belongs to multiple candidate grammatical structures, and determining the grammatical structure to which the word belongs based on the multiple first probabilities; Performing a second detection on the multiple sentences based on a second detection model to obtain multiple second probabilities that each of the sentences belongs to multiple candidate logical structures, and determining the logical structure to which the sentence belongs based on the multiple second probabilities; Performing a third detection on the multiple paragraphs based on a third detection model to obtain multiple third probabilities that each of the paragraphs belongs to multiple candidate discourse structures, and determining the discourse structure to which the paragraph belongs based on the multiple third probabilities.
15. The method according to claim 14, characterized in that, The method further includes: Obtaining a training text, and performing word segmentation processing on multiple sample sentences in the training text to obtain multiple sample words; Determining the grammatical structure to which each of the multiple sample words belongs to generate a first label of the training text; Training a basic model using the multiple sample words and the first label to obtain the first detection model, where the model parameters of the first detection model include first state transition information and first observation probability information, where the first state transition information is used to represent the probability of transitioning from a first candidate grammatical structure to a second candidate grammatical structure, and the first observation probability information is used to represent the probabilities of multiple target sample words belonging to the second candidate grammatical structure, and the first candidate grammatical structure and the second candidate grammatical structure belong to at least one of the multiple grammatical structures to which the multiple sample words belong.
16. The method according to claim 15, wherein Performing a first detection on the multiple words based on the first detection model, and obtaining multiple first probabilities that each of the words belongs to multiple candidate grammatical structures includes: Determining multiple first probabilities that each of the words belongs to multiple candidate grammatical structures using the first state transition information and the first observation probability information.
17. The method according to claim 14, wherein The method further includes: Obtaining a training text, and performing sentence splitting processing on multiple sample paragraphs in the training text to obtain multiple sample sentences; Determining the logical structure to which each of the multiple sample sentences belongs to generate a second label of the training text; Train a basic model using the multiple sample statements and the second label to obtain the second detection model, where the model parameters of the second detection model include second state transition information and second observation probability information, where the second state transition information is used to represent the probability of transitioning from a first candidate logical structure to a second candidate logical structure, and the second observation probability information is used to represent the probabilities of the multiple target sample statements belonging to the second candidate logical structure respectively, and the first candidate logical structure and the second candidate logical structure belong to at least one of the multiple logical structures to which the multiple sample statements belong.
18. The method according to claim 14, wherein The method further includes: Obtain a training text, and perform segmentation processing on the training text to obtain a plurality of sample paragraphs; Determine the discourse structure to which each of the multiple sample paragraphs belongs to generate a third label for the training text; Train a basic model using the multiple sample paragraphs and the third label to obtain the third detection model, where the model parameters of the third detection model include third state transition information and third observation probability information, where the third state transition information is used to represent the probability of transitioning from a first candidate discourse structure to a second candidate discourse structure, and the third observation probability information is used to represent the probabilities of the multiple target sample paragraphs belonging to the second candidate discourse structure respectively, and the first candidate discourse structure and the second candidate discourse structure belong to at least one of the multiple discourse structures to which the multiple sample paragraphs belong.
19. A text processing device, characterized in that, The apparatus includes: A detection module, configured to perform text structure detection on a text to be processed to obtain text structure information of the text to be processed, where the text structure information includes at least one of the following: the syntactic structures to which the multiple words in the sentences of the text to be processed belong respectively, the logical structures to which the multiple sentences in the paragraphs of the text to be processed belong respectively, the discourse structures to which the multiple paragraphs in the chapters of the text to be processed belong respectively; A determination module, configured to determine at least one embedding position for adding watermark information in the text to be processed based on the relevance between the text structure information and the context semantics in the text to be processed; An adding module, configured to add target watermark information to be added at the at least one embedding position.
20. An electronic device, characterized in that, The electronic device includes: One or more processors; A memory, configured to store one or more computer programs, Characterized in that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 18.
21. A non-volatile computer-readable storage medium having a computer program or instructions stored thereon, characterized in that, The computer program or instruction, when executed by a processor, implements the steps of the method according to any one of claims 1 to 18.
22. A computer program product comprising a computer program, characterized in that, The computer program, when executed by a processor, implements the method according to any one of claims 1 to 18.
Citation Information
Patent Citations
Embedding method and extracting method for text watermark based on semantic role position mapping
CN105205355A
Watermark adding method and device, storage medium and electronic equipment
CN113538198A
Data generation method, text generation method and electronic equipment
CN119323015A
Text watermark embedding method in Logits generation period based on large language model
CN119337344A
Watermark generation method and system based on large language model
CN119577707A