Text watermark embedding method, text watermark extracting method and electronic equipment
By determining the word tuples and redundant embedding positions in the text and controlling the target model to select characters, the robustness and accuracy problems of text watermarks in the existing technology under color-aware substitution attacks are solved, and efficient embedding and extraction of text watermarks are achieved.
Patent Information
- Application Number
- CN202510688776.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-09-19
AI Technical Summary
Existing text watermark embedding methods lack effective resistance to color-aware substitution attacks, making it difficult to ensure the integrity and extraction accuracy of text watermarks.
By obtaining the watermark information to be embedded, determining the word tuple corresponding to each data unit, and selecting multiple redundant embedding positions in the target text, the target model is controlled to select characters at these positions, and the target text is generated to represent the watermark information. At the same time, verification is performed based on the word tuple and redundant embedding positions during extraction.
The robustness of text watermarks against color-aware substitution attacks is enhanced, ensuring the integrity of text watermarks and the accuracy of extraction.
Smart Images

Figure CN120672550A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and in particular to a text watermark embedding method, an extraction method, and an electronic device. Background Art
[0002] To prevent unauthorized use or tampering of text data, verify its origin, or track its dissemination path, it is necessary to embed text watermarks in the data. Commonly used text watermark embedding methods include the green list method, the sampling perturbation method, and the language modeling constraint method. These methods primarily embed text watermark information by controlling the probability of selecting tokens in the output of a large model. However, these methods lack effective resistance to color-perceived substitution attacks, making it difficult to ensure the integrity and extraction accuracy of the text watermark. Summary of the Invention
[0003] The present disclosure provides a text watermark embedding method, an extraction method and an electronic device to at least solve the above technical problems existing in the prior art.
[0004] According to a first aspect of the present disclosure, a text watermark embedding method is provided, comprising: obtaining watermark information to be embedded; determining a word tuple corresponding to each data unit in the watermark information to be embedded; the word tuple includes a plurality of word grammes; determining a plurality of redundant embedding positions corresponding to each data unit in the watermark information to be embedded in a target text; the number of the redundant embedding positions is the same as the number of word grammes corresponding to the data unit; controlling a target model to select characters at the plurality of redundant embedding positions corresponding to the data unit in the word tuple corresponding to the data unit; the target text contains the watermark information to be embedded.
[0005] In one possible implementation manner, determining the word tuple corresponding to each data unit in the watermark information to be embedded includes: obtaining input text, wherein the input text is used to instruct the target model to generate the target text; generating multiple word tuples based on multiple characters in the input text; and assigning a corresponding word tuple to each data unit in the watermark information to be embedded in the multiple word tuples.
[0006] In one possible implementation, the control target model selects characters at multiple redundant embedding positions corresponding to the data unit in the word tuple corresponding to the data unit, including: determining candidate characters at the redundant embedding position corresponding to the data unit based on the target model; the candidate characters are characters whose sampling probability satisfies a first threshold; in response to the existence of a target candidate character that overlaps with a word in the word tuple corresponding to the data unit, the control target model determines the character at the redundant embedding position corresponding to the data unit among the target candidate characters.
[0007] In one possible implementation, the determining of candidate characters at the redundant embedding position corresponding to the data unit based on the target model includes: obtaining all predicted characters at the redundant embedding position output by the target model and the sampling probabilities of the predicted characters; adjusting the sampling probabilities of target predicted characters that coincide with word tuples in the word tuple group corresponding to the data unit; and determining the predicted characters whose sampling probabilities satisfy a first threshold as the candidate characters.
[0008] In one possible implementation, a text watermark embedding method further includes: in response to the absence of a target candidate character that overlaps with a word in the word tuple corresponding to the data unit, determining the semantic similarity between the candidate character and the word tuple; and determining the character at the redundant embedding position corresponding to the data unit among all target candidate characters whose semantic similarity meets a second threshold.
[0009] In one possible implementation, a text watermark embedding method further includes: adding the characters at the redundant embedding position corresponding to the data unit or all the target candidate characters to the word tuple corresponding to the data unit.
[0010] In one possible implementation, the control target model selects characters at multiple redundant embedding positions corresponding to the data unit in the word tuple corresponding to the data unit, including: the control target model determines the sampling probability of the word in the word tuple corresponding to the data unit based on the context information at the redundant embedding position; and determines the word with the highest sampling probability in the word tuple corresponding to the data unit as the character at the redundant embedding position.
[0011] In one possible implementation manner, determining the multiple redundant embedding positions corresponding to each data unit in the to-be-embedded watermark information in the target text includes: determining the total number of word units corresponding to the to-be-embedded watermark information; randomly generating the total number of embedding positions; and allocating a corresponding redundant embedding position to each of the data units in turn based on the number of word units corresponding to the data units.
[0012] In one embodiment, the word tuple satisfies at least one of the following: the semantic similarity between the word tuples in the word tuple satisfies a third threshold; the sentiment tendencies of the word tuples in the word tuple are the same; and the probability of the word tuples in the word tuple appearing in the same context satisfies a fourth threshold.
[0013] According to a second aspect of the present disclosure, a text watermark extraction method is provided, comprising: obtaining a target text; embedding text watermark information in the target text; obtaining a word tuple corresponding to each data unit in the text watermark information; the word tuple includes multiple word grammars; obtaining multiple redundant embedding positions corresponding to each data unit in the text watermark information in the target text; the text watermark information is embedded in the target text by selecting characters at multiple redundant embedding positions corresponding to the data unit in the word tuple corresponding to the data unit by a target model; extracting character groups corresponding to the data unit at the multiple redundant embedding positions; determining target watermark information in the target text based on the degree of overlap between the character group corresponding to the data unit and the word tuple corresponding to the data unit; and comparing the target watermark information with the text watermark information to obtain a verification result of the target text.
[0014] According to a third aspect of the present disclosure, there is provided an electronic device, including:
[0015] at least one processor; and
[0016] a memory communicatively connected to the at least one processor; wherein,
[0017] The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the method described in the present disclosure.
[0018] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause the computer to execute the method described in the present disclosure.
[0019] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The above and other objects, features and advantages of the exemplary embodiments of the present disclosure will become readily understood by reading the detailed description below with reference to the accompanying drawings, in which several embodiments of the present disclosure are shown by way of example and not limitation, wherein:
[0021] In the drawings, the same or corresponding reference numerals denote the same or corresponding parts.
[0022] Figure 1 The following is a flow chart showing a text watermark embedding method according to an embodiment of the present disclosure. Figure 1 ;
[0023] Figure 2 The following is a flow chart showing a text watermark embedding method according to an embodiment of the present disclosure. Figure 2 ;
[0024] Figure 3 The following is a flow chart showing a text watermark embedding method according to an embodiment of the present disclosure. Figure 3 ;
[0025] Figure 4 A flow chart of a text watermark extraction method according to an embodiment of the present disclosure is shown;
[0026] Figure 5 A schematic diagram of the structure of an electronic device according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0027] To make the purposes, features, and advantages of the present disclosure more apparent and understandable, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of the present disclosure without creative work shall fall within the scope of protection of the present disclosure.
[0028] Figure 1 The following is a flow chart showing a text watermark embedding method according to an embodiment of the present disclosure. Figure 1 ,like Figure 1 As shown, a text watermark embedding method includes:
[0029] Step S101: Acquire watermark information to be embedded.
[0030] In this embodiment, the watermark information to be embedded must first be obtained. The watermark information to be embedded is text watermark information. Compared with traditional image watermark information, text watermark information can be embedded through subtle word replacement and sentence structure adjustment, which does not affect the readability of the embedded text and is more concealed. The watermark information to be embedded can be a pre-determined binary sequence, such as "11001". Of course, the watermark information to be embedded can also be data in other bases. During the embedding process, the watermark information in other bases can be converted into a binary sequence.
[0031] Step S102: determining the word tuple corresponding to each data unit in the watermark information to be embedded.
[0032] In this embodiment, each data unit in the watermark information to be embedded is the smallest independent unit in the watermark information to be embedded. For example, if the watermark information to be embedded is a binary sequence 11001, each data unit in the watermark information to be embedded is each bit in the binary sequence, for example, "1", "1", "0", "0", and "1" are each a data unit. The word tuple corresponding to each data unit includes multiple word tuples for redundantly representing the corresponding data unit. The word tuples corresponding to the same data unit can be different. For example, the word tuple corresponding to each data unit in the watermark information to be embedded 11001 can be:
[0033] The word tuple corresponding to the first "1" can be ["happy", "great", "interesting"];
[0034] The word tuple corresponding to the second “1” can be [“can”, “mild”, “general”];
[0035] The word tuple corresponding to the first “0” can be [“sad”, “depressed”, “painful”];
[0036] The word tuple corresponding to the second “0” can be [“forest”, “river”, “mountain”];
[0037] The word tuple corresponding to the third “1” can be [“rabbit”, “squirrel”, “hedgehog”].
[0038] It should be emphasized that the number of word units in the word tuples corresponding to different data units can be the same or different, and the number of characters in different word units in the same word tuple can be the same or different.
[0039] Step S103: determining a plurality of redundant embedding positions corresponding to each data unit in the to-be-embedded watermark information in the target text.
[0040] In this embodiment, for each data unit to be embedded in the watermark information, it is necessary to determine its multiple redundant embedding positions in the target text. The number of redundant embedding positions is the same as the number of word-grams corresponding to the data unit. For example, for the word-gram group example in step S102, it can be determined that there are three redundant embedding positions corresponding to each data unit in the target text. In one example, the multiple redundant embedding positions corresponding to each data unit can be determined by pseudo-random encoding, or the multiple redundant embedding positions corresponding to each data unit can be determined based on contextual information. For example, the data unit is embedded only after a verb or an adjective. That is, if the previous word generated by the target model is a verb or an adjective, the current position is considered to be a redundant embedding position. The multiple redundant embedding positions for each data unit can be determined sequentially based on the number of word-grams corresponding to each data unit.
[0041] Step S104 , controlling the target model to select characters at multiple redundant embedding positions corresponding to the data unit in the word tuple corresponding to the data unit, and generating a target model.
[0042] In this embodiment, the target model can be a large language model (LLM). In the process of the target model generating the target text, when the target model generates a redundant embedding position corresponding to a certain data unit, the target model can be controlled to select the character at the redundant embedding position in the word tuple corresponding to the data unit. The characters at all redundant embedding positions in the target text finally generated by the target model can represent the watermark information to be embedded, that is, the target text contains the watermark information to be embedded.
[0043] In one example, for the word-tuple example at step S102, assuming that the redundant embedding positions do not count punctuation marks, and the redundant embedding positions corresponding to the first "1" are [9, 23, 54], and the redundant embedding positions corresponding to the second "1" are [65, 87, 92]..., the target text finally generated by the target model can be:
[0044] One classmate said: "Today is a (happy) day. The sun is shining and everything feels (great). I spent the day with my friends. It is really a perfect day and makes me feel (happy)."
[0045] Another classmate said, "Today was okay. Nothing particularly exciting happened, just average overall. The weather was mild, not too hot or too cold, and I spent most of the day indoors reading and relaxing. It was a quiet and satisfying day."
[0046] Among them, the starting positions of happy, great, and happy in brackets in the first paragraph of the target text (used to mark characters at redundant embedding positions, and there are no brackets in the actual generated target text) correspond to positions 9, 23, and 54 respectively, that is, when the target model generates the characters at positions 9, 23, and 54, it will select the characters at these three positions from the word tuple ["happy", "great", "interesting"] corresponding to the first "1"; the starting positions of okay, general, and mild in brackets in the second paragraph of the target text correspond to positions 65, 87, and 92 respectively, that is, when the target model generates the characters at positions 65, 87, and 92 When the characters at the position are selected, the characters at these three positions will be selected from the word tuple corresponding to the second "1" ["can", "mild", "general"]. ..., and so on. The characters at the redundant embedding position corresponding to the first "0" will be selected from the word tuple corresponding to the first "0", the characters at the redundant embedding position corresponding to the second "0" will be selected from the word tuple corresponding to the second "0", and the characters at the redundant embedding position corresponding to the third "1" will be selected from the word tuple corresponding to the third "1". Finally, there will be 15 characters or words in the target text to represent the watermark information to be embedded. It should be emphasized that the redundant embedding position can represent the position of a character or the starting position of a word containing multiple characters. For example, when the target model generates the character at position 9, "happy" is selected from the word tuple corresponding to the first "1". At this time, the characters at positions 9 and 10 are simultaneously determined.
[0047] In the present disclosure, the watermark information to be embedded is obtained, the word tuple corresponding to each data unit in the watermark information to be embedded is determined, the word tuple includes multiple word grammes, and the multiple redundant embedding positions corresponding to each data unit in the watermark information to be embedded in the target text are determined, the number of redundant embedding positions is the same as the number of word grammes corresponding to the data unit, and then in the process of the target model generating the target text, the target model is controlled to select the characters at the multiple redundant embedding positions corresponding to the data unit in the word tuple corresponding to the data unit, and the characters at all redundant embedding positions in the target text can be used to represent the text of the watermark information to be embedded. Thus, in the generated target text, a data unit in the watermark information to be embedded is represented by multiple characters. As long as most of the characters corresponding to the data unit are not destroyed by the attacker, the correct data unit can be restored, thereby enhancing the robustness of the text watermark in the face of color perception substitution attacks and ensuring the integrity of the text watermark and the accuracy of extraction.
[0048] In another embodiment, step S101 of “determining the word tuple corresponding to each data unit in the watermark information to be embedded” includes:
[0049] An input text is obtained, where the input text is used to instruct a target model to generate a target text; a plurality of word tuples are generated based on a plurality of characters in the input text; and a corresponding word tuple is assigned to each data unit to be embedded in the watermark information among the plurality of word tuples.
[0050] In this embodiment, it is first necessary to obtain input text. The input text is used to instruct the target model to generate target text. The input text can be a prompt or example for generating target text, etc., which is used to guide the target model to generate target text that meets the requirements. Then, multiple word tuples can be generated based on the characters or semantic information in the input text. For example, the input text is segmented to obtain key characters such as nouns, adjectives or verbs in the data text, and multiple word tuples are generated based on these key characters. Then, among the multiple word tuples generated, a corresponding word tuple is assigned to each data unit to be embedded in the watermark information.
[0051] In one example, if the input information is "Generate a short article about summer that must include sunshine, beach, and ice cream," multiple word tuples can be generated based on keywords such as summer, sunshine, beach, and ice cream in the input information. The generated multiple word tuples can be:
[0052] The word tuple generated based on summer can be ["hot", "watermelon", "cicada chirping", "short sleeves"];
[0053] The word tuple generated based on sunlight can be ["bright", "golden", "warm", "ultraviolet", "sunburn"];
[0054] The word tuple generated based on the beach can be ["beach", "waves", "shells", "parasols", "coconut trees", "swimming"];
[0055] The word tuple generated based on ice cream can be ["cone", "chocolate", "melted", "cold", "cream"].
[0056] Then, among these generated word tuples, corresponding word tuples can be randomly assigned to each data unit to be embedded in the watermark information. In this way, word tuples that are closer to the input information can be obtained, and the target text with text watermark generated based on the input information can be more natural.
[0057] Figure 2 The following is a flow chart showing a text watermark embedding method according to an embodiment of the present disclosure. Figure 2 ,like Figure 2 As shown, a text watermark embedding method includes:
[0058] Step S201: Acquire watermark information to be embedded.
[0059] Step S202: determining the word tuple corresponding to each data unit in the watermark information to be embedded.
[0060] Step S203: determining a plurality of redundant embedding positions corresponding to each data unit in the to-be-embedded watermark information in the target text.
[0061] The specific implementation details of steps S201 to S203 are similar to those of steps S101 to S103 and are not repeated here.
[0062] Step S204, determining the candidate characters at the redundant embedding position corresponding to the data unit based on the target model, and in response to the existence of a target candidate character that overlaps with a word in the word tuple corresponding to the data unit, controlling the target model to determine the character at the redundant embedding position corresponding to the data unit among the target candidate characters.
[0063] In this embodiment, in the process of the target model generating the target text, when generating each character, the target model will predict multiple predicted characters based on the context information and determine the sampling probability of each predicted character, and determine the characters whose sampling probability meets the first threshold as candidate characters. Therefore, when generating the characters at the redundant embedding position corresponding to the data unit, it is necessary to first determine the candidate characters at the redundant embedding position corresponding to the data unit. If there is a target candidate character in the candidate characters that overlaps with the word in the word group corresponding to the data unit, the target model is controlled to determine the character at the redundant embedding position corresponding to the data unit in the target candidate characters. The target model can randomly determine the character at the redundant embedding position corresponding to the data unit among the target candidate characters, or it can determine the target candidate character with the largest sampling probability as the character at the redundant embedding position corresponding to the data unit.
[0064] In one example, for the target text example at step S104, assuming that the target model is generated to the 9th position, the target model determines that the candidate characters corresponding to position 9 include ["superb", "perfect", "happy", "great"], and position 9 is one of the redundant embedding positions of the first "1" in the watermark information 11001 to be embedded. It is necessary to determine whether the candidate characters ["superb", "perfect", "happy", "great"] overlap with the word tuple ["happy", "great", "interesting"] corresponding to the first "1". Obviously, there are overlapping target candidate characters ["happy", "great"], so it is necessary to determine the character at position 9 among the target candidate characters ["happy", "great"]. The target model finally selects "happy" as the character at position 9.
[0065] In the present disclosure, during the process of generating a target text from a target model, candidate characters at the redundant embedding position corresponding to a data unit are determined based on the target model, and in response to the presence of a target candidate character that coincides with a word in a word tuple corresponding to the data unit, the target model is controlled to determine the character at the redundant embedding position corresponding to the data unit among the target candidate characters. The candidate character is more consistent with the contextual semantics at the redundant embedding position, and the target candidate character that coincides with the word tuple corresponding to the data unit among the candidate characters is selected as the character at the redundant embedding position. This not only embeds the watermark information to be embedded in the target text, but also ensures the naturalness of the semantics of the target text.
[0066] In another embodiment, the step S204 of “determining candidate characters at redundant embedding positions corresponding to the data unit based on the target model” includes:
[0067] Obtain all predicted characters and sampling probabilities of the predicted characters at the redundant embedding positions output by the target model; adjust the sampling probabilities of the target predicted characters that overlap with the word tuples in the word tuple group corresponding to the data unit; and determine the predicted characters whose sampling probabilities meet the first threshold as candidate characters.
[0068] In this embodiment, during the process of generating the target text by the target model, when generating characters at the redundant embedding position, multiple predicted characters are predicted based on contextual information and the sampling probability of each predicted character is determined. Then, the sampling probability of the target predicted character that coincides with the word in the word tuple corresponding to the data unit is adjusted, such as by increasing the sampling probability of the target predicted character by a certain target threshold, and then determining the character whose sampling probability meets the first threshold as a candidate character, thereby increasing the probability of the target predicted character being determined as a candidate character, and further increasing the probability that the character at the redundant embedding position is a word in the corresponding word tuple. It should be emphasized that the value of the target threshold should be small to avoid affecting the naturalness of the target text due to an excessively high sampling probability of the target predicted character.
[0069] In one example, if the predicted characters and corresponding sampling probabilities at the redundant embedding positions generated by the target model include [“awesome” -0.3, “perfect” -0.25, “happy” -0.09, “great” -0.09, “beautiful” -0.1, “interesting” -0.1, “different” -0.05, “unexpected” -0.02], and the word tuple corresponding to the redundant embedding position is [“happy”, “great”, “interesting”], then the target predicted character is [“happy”, “great”], and then the sampling probability of the target predicted character is increased by 0.02, that is, the sampling probability of “happy” is adjusted to 0.11, and the sampling probability of “great” is adjusted to 0.11, and then the sampling probabilities of all predicted characters are normalized, and the characters that meet the first threshold in the normalized sampling probabilities are determined as candidate characters. For example, the candidate characters can be [“awesome”, “perfect”, “happy”, “great”].
[0070] In another embodiment, after step S204, a text watermark embedding method further includes:
[0071] In response to the absence of a target candidate character that overlaps with a word in a word tuple corresponding to the data unit, the semantic similarity between the candidate character and the word tuple is determined; and the character at the redundant embedding position corresponding to the data unit is determined among all target candidate characters whose semantic similarity meets a second threshold.
[0072] In this embodiment, if there is no character in the candidate characters that overlaps with a word in the word tuple corresponding to the data unit, the semantic similarity between the candidate character and the word tuple is calculated. For example, the average semantic vector of all words in the word tuple can be determined, and then the similarity between the semantic vector of the candidate character and the average semantic vector of the word tuple is calculated to obtain the semantic similarity between the candidate character and the word tuple, and the character at the redundant embedding position corresponding to the data unit is selected from all candidate characters whose semantic similarity meets the second threshold. Finally, the character at the redundant embedding position corresponding to the data unit or all target candidate characters can be added to the word tuple corresponding to the data unit. Thus, it is equivalent to making the watermark embedding process more flexible and ensuring the accuracy of watermark embedding by calculating the semantic similarity and dynamically adjusting the composition of the word tuple.
[0073] In one example, if the candidate characters at the redundant embedding position generated by the target model include ["awesome", "perfect", "happy", "different"], the word tuple corresponding to the redundant embedding position is ["happy", "great", "interesting"], and the candidate characters and the word tuple do not overlap, then the average semantic vector of all word elements in the word tuple can be determined, and then the similarity between the semantic vector of the candidate character and the average semantic vector of the word tuple can be calculated in turn. The character at the redundant embedding position corresponding to the data unit can be selected from all candidate characters whose similarity is greater than the second threshold. If only "happy" among the candidate characters has a semantic similarity with the word tuple greater than the second threshold, then "happy" can be determined as the character at the redundant embedding position, and "happy" can be added to the word tuple corresponding to the redundant embedding position to obtain ["happy", "great", "interesting", "happy"].
[0074] In one possible implementation, if there is no candidate character whose semantic similarity with the word tuple is greater than a second threshold value among the candidate characters, then a character at a redundant embedding position corresponding to the data unit is selected from the candidate characters.
[0075] Figure 3 The following is a flow chart showing a text watermark embedding method according to an embodiment of the present disclosure. Figure 3 ,like Figure 3 As shown, a text watermark embedding method includes:
[0076] Step S301: Acquire watermark information to be embedded.
[0077] Step S302: determining the word tuple corresponding to each data unit in the watermark information to be embedded.
[0078] Step S303: determining a plurality of redundant embedding positions corresponding to each data unit in the to-be-embedded watermark information in the target text.
[0079] The specific implementation details of steps S301 to S303 are similar to those of steps S101 to S103 and are not repeated here.
[0080] Step S304, controlling the target model to determine the sampling probability of the word in the word tuple corresponding to the data unit based on the context information at the redundant embedding position; determining the word with the highest sampling probability in the word tuple corresponding to the data unit as the character at the redundant embedding position.
[0081] In this embodiment, in the process of generating the target text by the target model, when generating the characters at the redundant embedding position, the target model does not predict multiple predicted characters based on the context information, but only selects the characters at the redundant embedding position in the word tuple corresponding to the redundant embedding position. The target model can determine the sampling probability of the word in the word tuple corresponding to the data unit based on the context information at the redundant embedding position. For example, assuming that the word tuple is ["happy", "great", "interesting"], and the context tends to describe a relaxed atmosphere, then "interesting" may be given a higher sampling probability, and then the word with the highest sampling probability in the word tuple corresponding to the data unit is determined as the character at the redundant embedding position, thereby simplifying the watermark embedding operation and improving the watermark embedding efficiency. Moreover, since the characters at all redundant embedding positions come from the corresponding word tuple, the anti-attack ability of the watermark information can be enhanced.
[0082] In another embodiment, step S103 of “determining multiple redundant embedding positions corresponding to each data unit in the to-be-embedded watermark information in the target text” includes:
[0083] Determine the total number of word units corresponding to the watermark information to be embedded; randomly generate a total number of embedding positions; and assign corresponding redundant embedding positions to each data unit based on the number of word units corresponding to the data unit.
[0084] In this embodiment, the total number of word elements corresponding to the watermark information to be embedded is also the total number of word elements in the word element group corresponding to all data units of the watermark information to be embedded. For example, if the watermark information to be embedded is a binary sequence 11001, there are 3 word elements in the word element group corresponding to each data unit, and the total number of word elements corresponding to the watermark information to be embedded is 15. Then, a total number of embedding positions, such as 15 embedding positions, can be randomly generated, and based on the number of word elements corresponding to the data units, corresponding redundant embedding positions can be assigned to each data unit in turn, such as assigning 3 embedding positions to each data unit in the watermark information to be embedded 11001 as redundant embedding positions of the data unit.
[0085] In another embodiment, the word tuple satisfies at least one of the following:
[0086] The semantic similarity between the word-tuples in the word-tuple group meets the third threshold;
[0087] The sentiment tendency of the lexical elements in the lexical element group is the same;
[0088] The probability that the word-tuples in the word-tuple group appear in the same context satisfies a fourth threshold.
[0089] In this embodiment, the word-tuples in a word-tuple should have a certain degree of semantic similarity to ensure that the embedded watermark information is semantically natural. For example, the word-tuples in the word-tuple ["tranquility", "calm", "tranquility"] are all semantically related to tranquility, meeting the third threshold of semantic similarity.
[0090] In this embodiment, the word-grams in the word-gram group should have the same sentiment tendency, for example, all positive or all negative. For example, the word-grams in the word-gram group ["excited", "passionate", "eager"] all have a positive sentiment tendency.
[0091] In this embodiment, the word-grams in a word-tuple should have a high probability of appearing in the same context. For example, the word-grams in the word-tuple ["sunny", "bright", "clear"] have a high co-occurrence probability in the context of describing sunny weather and can naturally fit into the context of text describing weather or outdoor scenes.
[0092] Figure 4 A flow chart of a text watermark extraction method according to an embodiment of the present disclosure is shown. Figure 4 As shown, a text watermark extraction method includes:
[0093] Step S401: Obtain target text.
[0094] In this embodiment, the target text refers to a text in which text watermark information has been embedded. The text watermark information may be embedded into the target text based on the text watermark embedding method disclosed in the present invention.
[0095] Step S402: Obtain the word tuple corresponding to each data unit in the text watermark information.
[0096] In this embodiment, when text watermark information is embedded, the text watermark information and the word tuple corresponding to each data unit in the text watermark information are saved. A word tuple includes multiple word tuples. Therefore, when extracting the text watermark information, the text watermark information saved during embedding and the word tuple corresponding to each data unit in the text watermark information can be obtained. For example, the text watermark information can be 11001, and the word tuple corresponding to each data unit in the text watermark information 11001 can be:
[0097] The word tuple corresponding to the first "1" can be ["happy", "great", "interesting"];
[0098] The word tuple corresponding to the second “1” can be [“can”, “mild”, “general”];
[0099] The word tuple corresponding to the first “0” can be [“sad”, “depressed”, “painful”];
[0100] The word tuple corresponding to the second “0” can be [“forest”, “river”, “mountain”];
[0101] The word tuple corresponding to the third “1” can be [“rabbit”, “squirrel”, “hedgehog”].
[0102] It should be emphasized that the number of word units in the word tuples corresponding to different data units can be the same or different, and the number of characters in different word units in the same word tuple can be the same or different.
[0103] Step S403: obtaining a plurality of redundant embedding positions corresponding to each data unit in the text watermark information in the target text.
[0104] In this embodiment, when embedding text watermark information, multiple redundant embedding positions corresponding to each data unit in the text watermark information in the target text will also be saved. The number of redundant embedding positions is the same as the number of word units corresponding to the data unit. Moreover, when embedding text watermark information, the target model will select characters at multiple redundant embedding positions corresponding to the data unit in the word unit corresponding to the data unit. That is, the text watermark information is embedded in the target text by the target model selecting characters at multiple redundant embedding positions corresponding to the data unit in the word unit corresponding to the data unit. Therefore, when extracting text watermark information, it is also necessary to obtain multiple redundant embedding positions corresponding to each data unit in the text watermark information in the target text.
[0105] Step S404: extracting character groups corresponding to the data units at the multiple redundant embedding positions.
[0106] In this embodiment, characters at multiple redundant embedding positions corresponding to each data unit in the target text may be extracted to obtain a character group corresponding to each data unit.
[0107] In one example, if the target text is:
[0108] One classmate said: "Today is a (happy) day. The sun is shining and everything feels (great). I spent the day with my friends. It is really a perfect day and makes me feel (happy)."
[0109] Another classmate said, "Today was okay. Nothing particularly exciting happened, just average overall. The weather was mild, not too hot or too cold, and I spent most of the day indoors reading and relaxing. It was a quiet and satisfying day."
[0110] The redundant embedding positions corresponding to the first "1" in the text watermark information 11001 are [9, 23, 54], the redundant embedding positions corresponding to the second "1" are [65, 87, 92]... Then, the characters at positions 9, 23, and 54 in the target text can be used to form the character group corresponding to the first "1". The characters at positions 9, 23, and 54 are "happy", "great", and "happy" with brackets (used to mark the characters at the redundant embedding positions, and there are no brackets in the actually generated target text) in the first paragraph of the target text. So, the character group corresponding to the first "1" is ["happy", "great", "happy"]; the characters at positions 65, 87, and 92 in the target text are used to form the character group corresponding to the second "1". The characters at positions 65, 87, and 92 are "ok", "average", and "mild" with brackets in the second paragraph of the target text. So, the character group corresponding to the second "1" is ["ok", "average", "mild"],... and so on. 15 characters or words can be extracted, forming a total of 5 character groups. It should be emphasized that the redundant embedding position can represent the position of a character or the starting position of a word containing multiple characters. For example, when extracting the character at position 9, the number of characters to be extracted starting from position 9 can be selected according to the context information. If the character "开" at position 9 can form a word with the character "心" at position 10, then "开心" is taken as the character at position 9.
[0111] Step S405: Determine the target watermark information in the target text based on the coincidence degree between the character group corresponding to the data unit and the token group corresponding to the data unit.
[0112] Step S406: Compare the target watermark information with the text watermark information to obtain the verification result of the target text.
[0113] In this embodiment, after extracting the character group corresponding to each data unit, it is necessary to determine the degree of overlap between the character group corresponding to each data unit and the word tuple corresponding to the data unit. If the degree of overlap is greater than a preset threshold, it is considered that the watermark bit corresponding to the extracted character group is the data unit. For example, for the first "1" in the text watermark information 11001, its corresponding word tuple is ["happy", "great", "interesting"], and the extracted character group is ["happy", "great", "happy"]. Each character in the extracted character group belongs to the corresponding word tuple. Therefore, it can be determined that the watermark bit corresponding to the first extracted character group is 1; the second "1" in the text watermark information 11001, its corresponding word tuple is ["ok", "gentle", "general"], and the extracted character group is ["ok", "general", "gentle"]. ], each character in the extracted character group belongs to the corresponding word tuple, therefore, it can be determined that the watermark bit corresponding to the second extracted character group is 1, ..., and so on, the watermark bits corresponding to all the extracted character groups can be determined, thereby obtaining the target watermark information in the target text, and then the target watermark information can be compared with the text watermark information to obtain the verification result of the target text, that is, if the target watermark information is consistent with the text watermark information, it proves that the target text has not been tampered with, if the target watermark information is inconsistent with the text watermark information, it proves that the target text may have been tampered with.
[0114] In the present disclosure, a data unit to be embedded in the watermark information is represented by multiple characters in the generated target text. As long as most of the characters corresponding to the data unit are not destroyed by the attacker, the correct data unit can be restored, thereby enhancing the robustness of the text watermark in the face of color perception replacement attacks and ensuring the integrity of the text watermark and the accuracy of extraction.
[0115] In another embodiment, if the degree of overlap between the character group corresponding to each data unit and the word tuple corresponding to the data unit is not greater than a preset threshold, the watermark position corresponding to the extracted character group cannot be determined. At this time, the semantic similarity between the characters in the character group and the word tuple can be determined. If the semantic similarity is greater than the similarity threshold, it can be determined that the character overlaps with the word in the word tuple. Therefore, the degree of overlap between the character group corresponding to each data unit and the word tuple corresponding to the data unit can be further determined, thereby further improving the accuracy of text watermark extraction.
[0116] In one embodiment, the semantic similarity between the characters in the character group and the word tuple can be determined by at least one of the sampling probability ranking, semantic vector and context entropy, wherein the sampling probability ranking is the softmax ranking, the semantic vector is the embedding vector, and the context entropy is calculated by counting the frequency distribution of the characters in different contexts. If a character appears with similar probability in a variety of different contexts, the context entropy of the character is high, and its context complexity and diversity are high. The average value of the sampling probability ranking, the average value of the semantic vector or the average value of the context entropy of the word tuple in the word tuple can be determined, and then the difference between the sampling probability ranking of the characters in the character group and the average value of the sampling probability ranking of the word tuple is determined in sequence, or the similarity between the semantic vector of the characters in the character group and the average value of the semantic vector of the word tuple is determined in sequence, or the difference between the context entropy of the characters in the character group and the average value of the context entropy of the word tuple is determined in sequence to determine the semantic similarity between the characters in the character group and the word tuple.
[0117] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device and a readable storage medium.
[0118] Figure 5 A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0119] like Figure 5 As shown, the device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. Various programs and data required for the operation of the device 800 can also be stored in the RAM 803. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0120] Various components in device 800 are connected to I / O interface 805, including an input unit 806, such as a keyboard, mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, optical disk, etc.; and a communication unit 809, such as a network card, modem, wireless communication transceiver, etc. The communication unit 809 allows device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0121] The computing unit 801 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above, such as a text watermark embedding method or a text watermark extraction method. For example, in some embodiments, a text watermark embedding method or a text watermark extraction method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as a storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the text watermark embedding method or the text watermark extraction method described above can be performed. Alternatively, in other embodiments, the computing unit 801 may be configured to execute a text watermark embedding method or a text watermark extraction method in any other appropriate manner (for example, by means of firmware).
[0122] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0123] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0124] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0125] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0126] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0127] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.
[0128] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.
[0129] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one such feature. Throughout the present disclosure, "plurality" means two or more, unless otherwise specifically defined.
[0130] The above description is merely a specific embodiment of the present disclosure, but the scope of protection of the present disclosure is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this disclosure should be included in the scope of protection of the present disclosure. Therefore, the scope of protection of the present disclosure should be based on the scope of protection of the claims.
Claims
1. A text watermark embedding method, comprising: Obtain the watermark information to be embedded; Determining a word tuple corresponding to each data unit in the to-be-embedded watermark information; The word tuple includes a plurality of word tuples; Determining a plurality of redundant embedding positions corresponding to each data unit in the to-be-embedded watermark information in the target text; The number of the redundant embedding positions is the same as the number of word units corresponding to the data units; Control the target model to select characters at multiple redundant embedding positions corresponding to the data unit in the word tuple corresponding to the data unit to generate a target text; The target text includes the watermark information to be embedded.
2. The method according to claim 1, wherein determining the word tuple corresponding to each data unit in the to-be-embedded watermark information comprises: Acquire input text, where the input text is used to instruct the target model to generate the target text; Generating a plurality of word tuples based on a plurality of characters in the input text; Among the plurality of word-tuples, a corresponding word-tuple is allocated to each data unit to be embedded in the watermark information.
3. The method according to claim 1, wherein the control target model selects characters at multiple redundant embedding positions corresponding to the data unit in the word tuple corresponding to the data unit, comprising: Determine, based on the target model, candidate characters at the redundant embedding position corresponding to the data unit; The candidate characters are characters whose sampling probability satisfies a first threshold; In response to the presence of a target candidate character that coincides with a word in the word tuple corresponding to the data unit, the target model is controlled to determine a character at a redundant embedding position corresponding to the data unit among the target candidate characters.
4. The method according to claim 3, wherein determining the candidate characters at the redundant embedding position corresponding to the data unit based on the target model comprises: Obtain all predicted characters at the redundant embedding positions output by the target model and the sampling probabilities of the predicted characters; Adjusting the sampling probability of the target predicted character that coincides with the word in the word group corresponding to the data unit; The predicted characters whose sampling probability meets the first threshold are determined as the candidate characters.
5. The method according to claim 3, further comprising: In response to the absence of a target candidate character that overlaps with a word in the word tuple corresponding to the data unit, determining a semantic similarity between the candidate character and the word tuple; The character at the redundant embedding position corresponding to the data unit is determined among all target candidate characters whose semantic similarity meets a second threshold.
6. The method according to claim 1, wherein the control target model selects characters at multiple redundant embedding positions corresponding to the data unit in the word tuple corresponding to the data unit, comprising: Controlling the target model to determine a sampling probability of a word in the word tuple corresponding to the data unit based on context information at the redundant embedding position; The word-gram with the highest sampling probability in the word-gram group corresponding to the data unit is determined as the character at the redundant embedding position.
7. The method according to claim 1, wherein determining a plurality of redundant embedding positions corresponding to each data unit in the to-be-embedded watermark information in the target text comprises: Determining the total number of word units corresponding to the watermark information to be embedded; Randomly generating the total number of embedding positions; Based on the number of word units corresponding to the data units, corresponding redundant embedding positions are sequentially allocated to each of the data units.
8. The method according to claim 1, wherein the word tuple satisfies at least one of the following: The semantic similarity between the word-tuples in the word-tuple group meets a third threshold; The sentiment tendencies of the lexical elements in the lexical element group are the same; The probability that the word-grams in the word-gram group appear in the same context satisfies a fourth threshold.
9. A text watermark extraction method comprising: Acquire a target text; wherein the target text has text watermark information embedded therein; Obtaining a word tuple corresponding to each data unit in the text watermark information; the word tuple includes multiple word tuples; Obtaining a plurality of redundant embedding positions corresponding to each data unit in the text watermark information in the target text; The text watermark information is embedded into the target text by selecting characters at multiple redundant embedding positions corresponding to the data unit in the word tuple corresponding to the data unit by the target model; Extracting character groups corresponding to the data units at the plurality of redundant embedding positions; determining target watermark information in the target text based on a degree of overlap between a character group corresponding to the data unit and a word tuple corresponding to the data unit; The target watermark information is compared with the text watermark information to obtain a verification result of the target text.
10. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, wherein the instructions are executed by the at least one processor to enable the at least one processor to perform: Obtain the watermark information to be embedded; Determining a word tuple corresponding to each data unit in the to-be-embedded watermark information; the word tuple includes a plurality of word tuples; Determining a plurality of redundant embedding positions corresponding to each data unit in the to-be-embedded watermark information in the target text; the number of the redundant embedding positions is the same as the number of word units corresponding to the data units; The target model is controlled to select characters at multiple redundant embedding positions corresponding to the data unit in the word tuple corresponding to the data unit to generate a target text; the target text includes the watermark information to be embedded.