Watermark embedding method, device, storage medium and computer readable storage medium

By dynamically adjusting the watermark embedding position through the generation of stroke count sequences and employing binary sequence and zero-width character techniques, the problem of insufficient anti-attack capability of existing watermark embedding methods is solved, thereby improving the concealment and robustness of watermark information.

CN120805113BActive Publication Date: 2026-01-23VIPSHOP (GUANGZHOU) SOFTWARE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511311894.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-15
Publication Date
2026-01-23
Estimated Expiration
2045-09-15

AI Technical Summary

Technical Problem

In existing watermark embedding methods, watermark information is easily discovered and removed by attackers through statistical analysis, resulting in poor resistance to attacks.

Method used

By obtaining the number of strokes in the text to be processed, a stroke count sequence is generated, the watermark embedding position is dynamically adjusted, and the watermark information is converted into a binary sequence for embedding. Zero-width characters and multi-layer embedding technology are used to improve the concealment and robustness of the embedding position.

Benefits of technology

It effectively improves the anti-attack capability of watermark data, making it difficult to discover the watermark location through statistical analysis, thus enhancing the concealment and robustness of watermark information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120805113B_ABST
    Figure CN120805113B_ABST
Patent Text Reader

Abstract

The application discloses a watermark embedding method and device, a storage medium and a computer readable storage medium, and relates to the technical field of digital watermarking. The watermark embedding method comprises the following steps: obtaining a to-be-processed text, and generating a stroke number sequence based on the stroke numbers of characters in the to-be-processed text; determining a watermark embedding position according to the stroke number sequence; and performing watermark embedding processing on the to-be-processed text based on the watermark embedding position to obtain a target watermark text. The application solves the technical problem that the existing watermark embedding method has poor attack resistance for embedded watermark data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of digital watermarking, and more particularly to a watermark embedding method, device, storage medium, and computer-readable storage medium. Background Technology

[0002] Digital watermarking technology, as a core supporting means of digital copyright protection, has formed an application system covering the entire multimedia field.

[0003] Traditional watermark embedding methods typically embed watermark information in fixed locations. Therefore, watermarks embedded in fixed locations are easily detected and removed by attackers through statistical analysis. In other words, existing watermark embedding methods have poor resistance to attacks on embedded watermark data.

[0004] The above content is only used to help understand the technical solution of this application and does not represent an admission that the above content is prior art. Summary of the Invention

[0005] The main objective of this application is to provide a watermark embedding method, device, storage medium, and computer-readable storage medium, aiming to solve the technical problem that existing watermark embedding methods have poor resistance to attacks on embedded watermark data.

[0006] To achieve the above objectives, this application proposes a watermark embedding method, which includes:

[0007] Obtain the text to be processed, and generate a stroke count sequence based on the number of strokes of each character in the text to be processed;

[0008] The watermark embedding position is determined based on the stroke count sequence;

[0009] Based on the watermark embedding position, the text to be processed is subjected to watermark embedding processing to obtain the target watermark text.

[0010] In one embodiment, the step of determining the watermark embedding position based on the stroke count sequence includes:

[0011] Obtain the text features of the text to be processed, and determine the derived values ​​of each stroke number in the stroke number sequence;

[0012] Based on the derived values ​​and the text features, a sequence of derived values ​​is constructed, and the text position corresponding to each derived value in the sequence is used as the watermark embedding position.

[0013] In one embodiment, the text features include text complexity, and the step of constructing a sequence of derived values ​​based on the derived values ​​and the text features includes:

[0014] Based on the text complexity, determine the corresponding number of samples, wherein the number of samples is positively correlated with the text complexity;

[0015] From the derived values ​​of the number of strokes, the derived values ​​of the sample number are selected respectively, and the sequence of the selected derived values ​​is taken as the derived value sequence.

[0016] In one embodiment, the text features further include text usage scenarios, and the step of extracting derived values ​​of the sampled number from the derived values ​​of the stroke count, and using the sequence of extracted derived values ​​as the derived value sequence, further includes:

[0017] After the text usage scenario is a word segmentation scenario, the derived values ​​of the sample number are extracted from the derived values ​​of the stroke count, and the sequence of the extracted derived values ​​is used as the initial sequence.

[0018] The position of punctuation marks in the text to be processed is identified, and the value corresponding to the position preceding the punctuation mark is inserted into the initial sequence to obtain a derived value sequence.

[0019] In one embodiment, the step of performing watermark embedding processing on the text to be processed based on the watermark embedding position to obtain the target watermark text includes:

[0020] Obtain the watermark information to be added, and convert the watermark information into a binary sequence;

[0021] The binary sequence is divided into multiple binary subsequences, wherein the number of binary subsequences is consistent with the number of target watermark layers;

[0022] Based on each of the binary subsequences, watermark embedding processing is performed on the watermark embedding positions of the text to be processed to obtain the text to be processed with multiple watermark layers as the target watermark text.

[0023] In one embodiment, the step of performing watermark embedding processing on the watermark embedding positions of the text to be processed according to each of the binary sub-sequences to obtain the text to be processed with multiple watermark layers as the target watermark text includes:

[0024] Based on predetermined mapping rules, each of the binary subsequences is mapped to a zero-width string;

[0025] Each zero-width character in the zero-width string is sequentially embedded into the watermark embedding position in the text to be processed, resulting in a text with multiple watermark layers, which is then used as the target watermark text.

[0026] In one embodiment, the predetermined mapping rules between different binary subsequences and the zero-width string are different.

[0027] In addition, to achieve the above objectives, this application also proposes a watermark embedding device, the device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the watermark embedding method as described above.

[0028] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the watermark embedding method described above.

[0029] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the watermark embedding method described above.

[0030] One or more technical solutions proposed in this application have at least the following technical effects:

[0031] This application obtains the text to be processed and generates a stroke count sequence based on the number of strokes of each character in the text. Based on this stroke count sequence, the application determines the watermark embedding position, which changes as the number of strokes of each character in the text changes. Furthermore, based on the watermark embedding position, the application performs watermark embedding processing on the text to obtain the target watermarked text. By changing the embedding position according to the number of strokes of each character in the text, this application achieves real-time dynamic adjustment of the watermark embedding position during the watermark embedding process. Compared to existing watermark embedding methods, the embedding method of this application cannot be detected by attackers through statistical analysis, effectively improving the anti-attack capability of the embedded watermark data. Attached Figure Description

[0032] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0033] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0034] Figure 1 This is a flowchart illustrating an embodiment of the watermark embedding method of this application.

[0035] Figure 2This is a schematic diagram illustrating a scenario involving the watermark embedding location in an embodiment of this application.

[0036] Figure 3 This is a flowchart illustrating Embodiment 2 of the watermark embedding method of this application;

[0037] Figure 4 This is a flowchart illustrating Embodiment 3 of the watermark embedding method of this application;

[0038] Figure 5 This is a schematic diagram of the watermark embedding device in the embodiments of this application.

[0039] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0040] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.

[0041] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.

[0042] The main solution of this application embodiment is: to obtain the text to be processed, and to generate a stroke count sequence based on the number of strokes of each character in the text to be processed; to determine the watermark embedding position according to the stroke count sequence; and to perform watermark embedding processing on the text to be processed based on the watermark embedding position to obtain the target watermark text.

[0043] Traditional watermark embedding methods typically embed watermark information in fixed locations. Therefore, watermarks embedded in fixed locations are easily detected and removed by attackers through statistical analysis. In other words, existing watermark embedding methods have poor resistance to attacks on embedded watermark data.

[0044] This application provides a solution that can change the embedding position according to the number of strokes of each character in the text to be processed, thereby realizing real-time dynamic adjustment of the watermark embedding position during the watermark embedding process. Compared with existing watermark embedding methods, the embedding method of this application cannot be discovered by attackers through statistical analysis or other means, effectively improving the anti-attack capability of the embedded watermark data.

[0045] Based on this, embodiments of this application provide a watermark embedding method, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the watermark embedding method of this application.

[0046] In this embodiment, the watermark embedding method includes steps S10 to S30:

[0047] Step S10: Obtain the text to be processed, and generate a stroke count sequence based on the number of strokes of each character in the text to be processed;

[0048] It should be noted that the text to be processed is the text data for which watermark embedding is desired.

[0049] Additionally, it should be noted that the term "characters" refers to the general term for various scripts and symbols, including national languages, punctuation marks, graphic symbols, and numbers.

[0050] This embodiment obtains the text to be processed and analyzes the stroke count of each character in the text to obtain the stroke count of each character. Then, the stroke counts are sorted sequentially to form a stroke count sequence. For example, this embodiment can query a predetermined database to obtain the stroke count of each character in the text to be processed. The predetermined database stores the correspondence between each character and its stroke count. For example, assuming the text to be processed is "This is a test text.", the stroke count of each character can be obtained by querying a Chinese character stroke count database, and then the stroke counts are sorted to obtain a stroke count sequence [7, 9, 1, 3, 9, 8, 4, 5, 1]. It is understood that the stroke counts constituting the stroke count sequence can be the stroke counts of all characters in the text to be processed, or they can be the stroke counts of some characters.

[0051] Step S20: Determine the watermark embedding position based on the stroke count sequence;

[0052] It should be noted that the watermark embedding position can be described in numerical form, such as... Figure 2 As shown, the position between adjacent characters in the text to be processed is taken as an embedding position, and then each embedding position is encoded sequentially to obtain the value corresponding to each embedding position (i.e., Figure 2 (The numbers 1, 2, 3, 4, 5, 6, 7, 8...).

[0053] As an example, in this embodiment, the stroke count sequence can be directly used as the sequence representing the watermark embedding positions, that is, the embeddable positions corresponding to each stroke count in the stroke count sequence are used as the watermark embedding positions. As another example, in this embodiment, the derivative values of each stroke count in the stroke count sequence can be determined. Exemplarily, the derivative value can be the result value after operations such as multiplication / summation between the stroke count and a predetermined number sequence (such as a prime number sequence, a Fibonacci sequence, etc.). Then, in this embodiment, a specified number of derivative values can be selected from the derivative values to form a derivative value sequence. The embeddable positions corresponding to each derivative value in the derivative value sequence are used as the watermark embedding positions. Taking the prime number sequence as the predetermined number sequence as an example, for the first character "这" (with 7 strokes), the derivative values are the products of the stroke count and each prime number in the prime number sequence, that is, 7*2 = 14, 7*3 = 21, 7*5 = 35... For the second character "是" (with 9 strokes), the derivative values are 9*2 = 18, 9*3 = 27, 9*5 = 45... and so on. In this embodiment, the first two derivative values of each stroke count can be selected to form a derivative value sequence [14, 21, 18, 27, 14, 21, 18, 27, 14]. As another example, in order to further improve the concealment and robustness of the watermark information, in this embodiment, the text features of the text to be processed can also be obtained, and the derivative values of each stroke count in the stroke count sequence can be determined. Then, based on the derivative values and the text features, a derivative value sequence is formed, and the text positions corresponding to each derivative value in the derivative value sequence are used as the watermark embedding positions. Thus, in this embodiment, the selection of the derivative values (i.e., the watermark embedding positions) is associated with the text features of the text to be processed, further adding a new perturbation factor to the selection of the watermark embedding positions, and can further improve the concealment and robustness of the watermark information.

[0054] Step S30, based on the watermark embedding positions, perform watermark embedding processing on the text to be processed to obtain a target watermark text.

[0055] As an example, this embodiment can obtain the watermark information to be added, convert the watermark information into a binary sequence, and perform watermark embedding processing on the watermark embedding position of the text to be processed according to the binary sequence to obtain the target watermark text. Exemplarily, this embodiment can convert the binary sequence into an invisible marker string, and then embed the invisible marker characters in the invisible marker string sequentially into the watermark embedding position to obtain the target watermark text. The invisible marker characters are special characters that do not occupy visible space, such as one or a combination of zero-width characters, word connectors, invisible multiplication signs, and invisible separators. Furthermore, to further improve the security of the watermark information, after converting the watermark information into a binary sequence, the binary sequence can be encrypted, and the following step can be performed based on the encrypted binary sequence: performing watermark embedding processing on the watermark embedding position of the text to be processed according to the binary sequence to obtain the target watermark text. As another example, this embodiment can also acquire the watermark information to be added, convert the watermark information into a binary sequence, and then divide the binary sequence into multiple binary sub-sequences, wherein the number of binary sub-sequences is consistent with the number of target watermark layers; according to each binary sub-sequence, watermark embedding processing is performed on the watermark embedding position of the text to be processed, resulting in text to be processed with multiple watermark layers as the target watermark text. This enables the superposition of multiple watermarks to improve the complexity and anti-attack capability of the watermark. Furthermore, to further improve the security of the watermark information, after converting the watermark information into a binary sequence, the binary sequence can be encrypted, and the step of dividing the binary sequence into multiple binary sub-sequences can be performed based on the encrypted binary sequence.

[0056] The first embodiment of this application provides a watermark embedding method. It obtains the text to be processed and generates a stroke count sequence based on the number of strokes of each character in the text. Therefore, this embodiment can determine the watermark embedding position based on the stroke count sequence; that is, the watermark embedding position changes with the number of strokes of each character in the text. Furthermore, this embodiment can perform watermark embedding processing on the text to be processed based on the watermark embedding position to obtain the target watermarked text. This embodiment achieves real-time dynamic adjustment of the watermark embedding position by correspondingly changing the embedding position according to the number of strokes of each character in the text. Compared with existing watermark embedding methods, the embedding method of this embodiment cannot be detected by attackers through statistical analysis or other methods, effectively improving the anti-attack capability of the embedded watermark data.

[0057] Based on the first embodiment of this application, in the second embodiment of this application, the content that is the same as or similar to that in the first embodiment described above can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 3 Step S20 includes steps S21 to S22:

[0058] Step S21: Obtain the text features of the text to be processed and determine the derived value of each stroke number in the stroke number sequence;

[0059] Step S22: Based on the derived values ​​and the text features, a sequence of derived values ​​is constructed, and the text position corresponding to each derived value in the sequence of derived values ​​is used as the watermark embedding position.

[0060] It should be noted that the text features of the text to be processed may include the text's own features (such as font, font size, language type, etc.) as well as related features (such as usage scenarios).

[0061] It should also be noted that the derived value is a value calculated based on the number of strokes through a predetermined operation method. For example, the derived value can be the result of operations such as multiplication / summation between the number of strokes and a predetermined sequence (such as a prime number sequence, Fibonacci sequence, etc.).

[0062] In this embodiment, the text features of the text to be processed can be obtained by extracting feature information from the text to be processed. It can be understood that the way of extracting the feature information can be determined according to the text features to be extracted. Exemplarily, taking the language type as the text feature, since different languages use specific character sets, the language type of the text to be processed can be determined through the character set features of the text to be processed. In addition, features such as font size and font of the text to be processed can be obtained by reading the format information of the text to be processed, and related features such as the usage scenario of the text to be processed can be obtained by reading the associated information of the text to be processed. Furthermore, in this embodiment, the derivative value of each stroke count in the stroke count sequence can be determined. Exemplarily, the derivative value can be the result value after operations such as multiplication / summation between the stroke count and a predetermined number sequence (such as a prime number sequence, a Fibonacci sequence, etc.). Taking the prime number sequence as the predetermined number sequence, for the first character "这": 7 strokes, the derivative value is the product of the stroke count and each prime number in the prime number sequence, that is, 7*2 = 14, 7*3 = 21, 7*5 = 35... For the second character "是": 9 strokes, the derivative value is 9*2 = 18, 9*3 = 27, 9*5 = 45... And so on. Furthermore, in this embodiment, the derivative value can be adjusted according to the text features to form a derivative value sequence. Thus, the embeddable positions corresponding to each derivative value in the derivative value sequence can be used as watermark embedding positions. As an example, the text features include text complexity. In this embodiment, the corresponding sampling number can be determined according to the text complexity, where the sampling number is positively correlated with the text complexity; the derivative values of the sampling number are respectively selected from the derivative values of the stroke count, and the sequence of the selected derivative values is used as the derivative value sequence. Since the larger the sampling number, the more watermark embedding positions there are, this embodiment can provide a higher density of embedding positions for more complex texts to be processed. As another example, the text features further include the text usage scenario. In this embodiment, after the text usage scenario is a word segmentation scenario (such as scenarios that require word segmentation in natural language processing, search engine indexing, etc.), the derivative values of the sampling number are respectively selected from the derivative values of the stroke count, and the sequence of the selected derivative values is used as the initial sequence. Furthermore, the symbol positions of punctuation marks in the text to be processed can be identified, and the value corresponding to the previous position of the symbol position (that is, the text position between the punctuation mark and the previous character) is inserted into the initial sequence to obtain the derivative value sequence. This can increase the redundant watermark embedding positions and is beneficial to protecting the integrity of the watermark information.

[0063] In some embodiments, the text features include text complexity, and step S22 includes steps A10~A20:

[0064] Step A10: Determine the corresponding sampling number according to the text complexity, where the sampling number is positively correlated with the text complexity;

[0065] Step A20: Select the derived values of the sampling number from the derived values of the stroke numbers respectively, and use the sequence of the selected derived values as the derived value sequence.

[0066] It should be noted that the text complexity is a characterization value representing the content complexity of the text to be processed. Exemplarily, different complexity values can be assigned to different language types, and corresponding complexity values can be assigned to text contents with different data volumes. Thus, the text complexity of the text to be processed is obtained through weighted calculation.

[0067] In this embodiment, the corresponding sampling number can be determined according to the text complexity. The sampling number is the number of derived values selected from the derived values of a stroke number. Among them, the sampling number is positively correlated with the text complexity, that is, the higher the text complexity, the larger the sampling number, which means more watermark embedding positions are added. For example, in Chinese text, the number of strokes of characters is relatively large, and Chinese is more complex than English. Therefore, the text complexity of the text to be processed with the language type being Chinese and / or a larger text data volume is relatively high, which is suitable for increasing the watermark embedding density, and can effectively enhance the concealment and robustness of the watermark information. Furthermore, in this embodiment, the derived values of the sampling number can be selected from the derived values of the stroke numbers respectively, and the sequence of the selected derived values is used as the derived value sequence. Taking the stroke number sequence [7, 9, 1, 3, 9, 8, 4, 5, 1] as an example, for the 1st character "这": 7 strokes, the derived value is the product of the stroke number and each prime number in the prime number sequence, that is, 7*2 = 14, 7*3 = 21, 7*5 = 35... For the 2nd character "是": 9 strokes, the derived value is 9*2 = 18, 9*3 = 27, 9*5 = 45... and so on. According to the text complexity, if the corresponding sampling number is determined to be 2, then this embodiment can select the first two derived values of each stroke number to form the derived value sequence [14, 21, 18, 27, 14, 21, 18, 27, 14].

[0068] In some embodiments, the text feature further includes the text usage scenario. In step A20, it includes steps B10~B20:

[0069] Step B10: After the text usage scenario is the word segmentation scenario, select the derived values of the sampling number from the derived values of the stroke numbers respectively, and use the sequence of the selected derived values as the initial sequence;

[0070] Step B20: Identify the symbol positions of the punctuation marks in the text to be processed, and insert the value corresponding to the previous position of the symbol position into the initial sequence to obtain the derived value sequence.

[0071] It should be noted that the word segmentation scenarios mentioned are those that require word segmentation, such as natural language processing and search engine indexing.

[0072] This embodiment can determine whether the text usage scenario is a word segmentation scenario. If the text usage scenario is not a word segmentation scenario, step A10 can be executed. If the text usage scenario is a word segmentation scenario, this embodiment can determine the corresponding sampling number based on the text complexity, and then extract the derived values ​​of the sampling number from the derived values ​​of the stroke count, using the sequence of extracted derived values ​​as the initial sequence. Furthermore, this embodiment can identify the symbol positions of punctuation marks in the text to be processed, inserting the value corresponding to the position preceding the symbol position (i.e., the text position between the punctuation mark and the preceding character) into the initial sequence to obtain the derived value sequence. This allows for the addition of redundant watermark embedding positions before punctuation marks such as question marks and periods, thereby protecting the integrity of the watermark information by leveraging the strong separating properties of punctuation marks.

[0073] In the second embodiment of this application, the text features of the text to be processed are obtained, and the derived values ​​of each stroke number in the stroke number sequence are determined. Then, based on the derived values ​​and the text features, a derived value sequence is constructed, and the text position corresponding to each derived value in the derived value sequence is used as the watermark embedding position. Therefore, this embodiment can adaptively adjust the derived value sequence according to different text features of the text to be processed, further enhancing the concealment and robustness of the embedded watermark information.

[0074] Based on the first embodiment of this application, in the second embodiment of this application, the content that is the same as or similar to that in the first embodiment described above can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 4 Step S30 includes steps S31 to S33:

[0075] Step S31: Obtain the watermark information to be added and convert the watermark information into a binary sequence;

[0076] Step S32: Divide the binary sequence into multiple binary subsequences, wherein the number of binary subsequences is consistent with the number of target watermark layers;

[0077] Step S33: According to each of the binary sub-sequences, watermark embedding processing is performed on the watermark embedding position of the text to be processed to obtain the text to be processed with multiple watermark layers as the target watermark text.

[0078] It should be noted that the watermark information may include preset static identification information, such as the warning text "Confidential document, do not distribute", as well as dynamic identification information acquired in real time, such as the identity information of the web browser (e.g., account, employee number, etc.).

[0079] It should also be noted that the target number of watermark layers refers to the desired number of watermark layers to be superimposed on the text to be processed. The target number of watermark layers can be a pre-set fixed number of layers, or it can be determined based on at least one of the following: security requirements, text features of the text to be processed, and the amount of watermark information.

[0080] In this embodiment, the watermark information to be added can be obtained and converted into a binary sequence. Exemplarily, in this embodiment, the watermark information can be converted into a binary sequence through encoding methods such as binary plaintext encoding, error correction encoding, encryption / disturbance encoding, check encoding, spread spectrum encoding, synchronization code, graphical encoding, compression encoding, etc. In addition, this embodiment can also perform encryption processing on the binary sequence, and based on the encrypted binary sequence, step S32 is executed, where the encryption processing can be implemented by encryption algorithms such as AES (Advanced Encryption Standard) and RSA (Rivest–Shamir–Adleman). Furthermore, this embodiment can divide the binary sequence into multiple binary subsequences, where the number of the binary subsequences is the same as the number of target watermark layers. Exemplarily, the number of target watermark layers can be set corresponding layers according to security requirements, and the number of target watermark layers is positively correlated with the level of security requirements. For example, in a high-security requirement scenario (such as copyright protection), more layers can be selected as the number of target watermark layers to enhance the anti-attack ability; in a low-security requirement scenario (such as advertising tracking), the number of layers can be reduced, and fewer layers can be selected as the number of target watermark layers. The number of target watermark layers can also be determined based on the text features of the text to be processed, and the number of target watermark layers is positively correlated with the text complexity in the text features. That is, a long text has more information, so more layers can be selected as the number of target watermark layers. To avoid information overload in a short text, fewer layers can be selected as the number of target watermark layers. For a language type with rich character strokes such as Chinese, more layers can be selected as the number of target watermark layers, while for a language type with simple characters such as English, fewer layers can be selected as the number of target watermark layers. The number of target watermark layers can also be determined based on the amount of watermark information of the watermark information, and the number of target watermark layers is positively correlated with the amount of watermark information. That is, the larger the amount of watermark information, the more information it has, so more layers can be selected as the number of target watermark layers. The smaller the amount of watermark information, the fewer layers can be selected as the number of target watermark layers. For example, the watermark information is "genuine". First, "genuine" is converted into a binary sequence. Using UTF-8 encoding: the UTF-8 encoding of "正" is E6ADA3, and the corresponding binary numerical string is 111001101010110110100011. The UTF-8 encoding of "版" is E78988, and the binary numerical string is 111001111000100110001000. The binary sequence obtained after merging is: 111001101010110110100011111001111000100110001000.The binary sequence is encrypted using the AES encryption algorithm, generating the encrypted binary sequence: 10101010 11001100 11110000101010 11001100 11110000. Since the target watermark has two layers, the encrypted binary sequence can be divided into two sub-sequences corresponding to the watermark layers: the first layer is 10101010 11001100 11110000; the second layer is 10101010 11001100 11110000. Furthermore, this embodiment can perform watermark embedding processing on the watermark embedding positions of the text to be processed according to each of the binary sub-sequences, obtaining the text to be processed with multiple watermark layers as the target watermark text.

[0081] In some embodiments, step S33 includes steps C10 to C20:

[0082] Step C10: Based on a predetermined mapping rule, map each of the binary subsequences to a zero-width string;

[0083] Step C20: Each zero-width character in the zero-width string is sequentially embedded into the watermark embedding position in the text to be processed, resulting in a text to be processed with multiple watermark layers as the target watermark text.

[0084] It should be noted that this embodiment uses zero-width characters as hidden marker characters. These zero-width characters include U+200B (zero-width space), U+200C (zero-width non-connector), U+200D (zero-width connector), and U+FEFF (zero-width non-breaking space). The predetermined mapping rules include the mapping relationship between binary values ​​("1" and "0") and zero-width characters. It can be understood that a binary value ("1" or "0") can correspond to a single zero-width character or a combination of zero-width characters. That is, for expansion of watermark layers, a combination of multiple zero-width characters can be used to represent a binary value in the predetermined mapping rules. For example, U+200B+U+200C can represent the binary value "1", and U+200D+U+FEFF can represent the binary value "0", thereby expanding the mapping space. Combinations of multiple zero-width characters can also be used in conjunction with text position markers to achieve the mapping of binary subsequences. For example, U+200B is used to represent the first watermark layer in even-numbered positions, and the second watermark layer in odd-numbered positions.

[0085] In this embodiment, each binary subsequence can be mapped to a zero-width string based on a predetermined mapping rule. For example, the predetermined mapping rule is that the binary value "1" is represented by U+200B (zero-width space), and the binary value "0" is represented by U+200C (zero-width non-connector). Taking the binary subsequence 10101010 as an example, the zero-width string is U+200B U+200C U+200B U+200C U+200B U+200C. Then, each zero-width character in the zero-width string is sequentially embedded into the watermark embedding position in the text to be processed. Taking the watermark embedding position [14,21,18,27,14,21,18,27] as an example, U+200B (zero-width space) corresponding to the binary value "1" is embedded at text position 14, U+200C (zero-width non-connector) corresponding to the binary value "0" is embedded at text position 21, U+200B (zero-width space) corresponding to the binary value "1" is embedded at text position 18, and so on. Thus, after embedding the zero-width characters obtained by mapping each binary subsequence into the text to be processed, the text to be processed with the target watermark layer is obtained as the target watermark text. Therefore, in this embodiment, the target watermark text after watermark embedding is completely identical to the original text to the naked eye, with no visible changes.

[0086] In one feasible embodiment, the predetermined mapping rules between different binary subsequences and the zero-width string are different.

[0087] This embodiment can pre-construct multiple different mapping rules. Before step C10, this embodiment can select a different mapping rule for each binary subsequence as a predetermined mapping rule, and then map each binary subsequence to a zero-width string based on the predetermined mapping rule. For example, the mapping rule selected for the binary subsequence of the first layer watermark is as follows: the binary value "1" is represented by U+200B (zero-width space); the binary value "0" is represented by U+200C (zero-width non-connector). The mapping rule selected for the binary subsequence of the second layer watermark is as follows: the binary value "1" is represented by U+200D (zero-width connector); the binary value "0" is represented by U+FEFF (zero-width non-breaking space). In this embodiment, the predetermined mapping rules between different binary subsequences and the zero-width strings are different. Therefore, each watermark layer on the target watermark text after watermark embedding uses a different zero-width character type, improving the complexity and anti-attack capability of the watermark.

[0088] Following this, in the watermark extraction and verification steps, this embodiment can extract the zero-width string from the target watermark text and, according to the predetermined mapping rules, reverse-map out a binary subsequence, which is then concatenated to form a binary sequence. The binary sequence is then decrypted and decoded to obtain the watermark information, thereby identifying the integrity and authenticity of the watermark information.

[0089] In the third embodiment of this application, watermark information to be added is obtained and converted into a binary sequence; the binary sequence is divided into multiple binary subsequences, wherein the number of binary subsequences is consistent with the number of target watermark layers; according to each binary subsequence, watermark embedding processing is performed on the watermark embedding position of the text to be processed, resulting in a text to be processed with multiple watermark layers as the target watermark text. Thus, this embodiment achieves the superposition of multiple watermarks on the target watermark text, with each watermark layer employing different mapping rules, improving the complexity and anti-attack capability of the watermark.

[0090] This application provides a watermark embedding device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, which are executed by the at least one processor to enable the at least one processor to perform the watermark embedding method in the first embodiment described above.

[0091] The following is for reference. Figure 5 The diagram illustrates a structural schematic of a watermark embedding device suitable for implementing embodiments of this application. The watermark embedding device in the embodiments of this application may include, but is not limited to, terminals such as mobile phones, laptops, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), and desktop computers. Figure 5 The watermark embedding device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.

[0092] like Figure 5As shown, the watermark embedding device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 1002 or a program loaded from storage device 1003 into random access memory (RAM) 1004. The random access memory 1004 also stores various programs and data required for the operation of the watermark embedding device. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An I / O (input / output) interface 1006 is also connected to the bus. Typically, the following systems can be connected to I / O interface 1006: input devices 1007 including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 1008 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 1003 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1009. Communication device 1009 allows the watermark embedding device to communicate wirelessly or wiredly with other devices to exchange data. Although watermark embedding devices with various systems are shown in the figures, it should be understood that it is not required to implement or possess all the systems shown. More or fewer systems can be implemented alternatively.

[0093] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from read-only memory 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.

[0094] The watermark embedding device provided in this application, employing the watermark embedding method described in the above embodiments, can solve the technical problem that existing watermark embedding methods have poor resistance to attacks on embedded watermark data. Compared with the prior art, the beneficial effects of the watermark embedding device provided in this application are the same as those of the watermark embedding method described in the above embodiments, and other technical features of this watermark embedding device are the same as those disclosed in the previous embodiment method, and will not be repeated here.

[0095] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.

[0096] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0097] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the watermark embedding method in the above embodiments.

[0098] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0099] The aforementioned computer-readable storage medium may be included in the watermark embedding device; or it may exist independently and not assembled into the watermark embedding device.

[0100] The aforementioned computer-readable storage medium carries one or more programs, which, when executed by a watermark embedding device, cause the watermark embedding device to: acquire text to be processed and generate a stroke count sequence based on the stroke count of each character in the text to be processed; determine the watermark embedding position according to the stroke count sequence; and perform watermark embedding processing on the text to be processed based on the watermark embedding position to obtain the target watermark text.

[0101] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0102] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0103] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.

[0104] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described watermark embedding method, which can solve the technical problem that existing watermark embedding methods have poor resistance to attacks on embedded watermark data. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the watermark embedding method provided in the above embodiments, and will not be repeated here.

[0105] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the watermark embedding method described above.

[0106] The computer program product provided in this application can solve the technical problem that existing watermark embedding methods have poor resistance to attacks on embedded watermark data. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the watermark embedding method provided in the above embodiments, and will not be repeated here.

[0107] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.

Claims

1. A watermark embedding method, characterized in that, The watermark embedding method includes: Obtain the text to be processed, and generate a stroke count sequence based on the number of strokes of each character in the text to be processed; Obtain the text features of the text to be processed, and determine the derived values ​​of each stroke number in the stroke number sequence; Based on the derived values ​​and the text features, a sequence of derived values ​​is constructed, and the text position corresponding to each derived value in the sequence is used as the watermark embedding position; Based on the watermark embedding position, the text to be processed is subjected to watermark embedding processing to obtain the target watermark text.

2. The watermark embedding method as described in claim 1, characterized in that, The text features include text complexity, and the step of constructing a sequence of derived values ​​based on the derived values ​​and the text features includes: Based on the text complexity, determine the corresponding number of samples, wherein the number of samples is positively correlated with the text complexity; From the derived values ​​of the number of strokes, the derived values ​​of the sample number are selected respectively, and the sequence of the selected derived values ​​is taken as the derived value sequence.

3. The watermark embedding method as described in claim 2, characterized in that, The text features also include text usage scenarios. The step of extracting derived values ​​of the sampled number from the derived values ​​of the stroke count and using the sequence of extracted derived values ​​as the derived value sequence further includes: After the text usage scenario is a word segmentation scenario, the derived values ​​of the sample number are extracted from the derived values ​​of the stroke count, and the sequence of the extracted derived values ​​is used as the initial sequence. The position of punctuation marks in the text to be processed is identified, and the value corresponding to the position preceding the punctuation mark is inserted into the initial sequence to obtain a derived value sequence.

4. The watermark embedding method as described in claim 1, characterized in that, The step of performing watermark embedding processing on the text to be processed based on the watermark embedding position to obtain the target watermark text includes: Obtain the watermark information to be added, and convert the watermark information into a binary sequence; The binary sequence is divided into multiple binary subsequences, wherein the number of binary subsequences is consistent with the number of target watermark layers; Based on each of the binary subsequences, watermark embedding processing is performed on the watermark embedding positions of the text to be processed to obtain the text to be processed with multiple watermark layers as the target watermark text.

5. The watermark embedding method as described in claim 4, characterized in that, The step of performing watermark embedding processing on the watermark embedding positions of the text to be processed according to each of the binary sub-sequences to obtain the text to be processed with multiple watermark layers as the target watermark text includes: Based on predetermined mapping rules, each of the binary subsequences is mapped to a zero-width string; Each zero-width character in the zero-width string is sequentially embedded into the watermark embedding position in the text to be processed, resulting in a text with multiple watermark layers, which is then used as the target watermark text.

6. The watermark embedding method as described in claim 5, characterized in that, The predetermined mapping rules between different binary subsequences and the zero-width string are different.

7. A watermark embedding device, characterized in that, The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the watermark embedding method as described in any one of claims 1 to 6.

8. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the watermark embedding method as described in any one of claims 1 to 6.

9. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the steps of the watermark embedding method as described in any one of claims 1 to 6.