Watermark embedding method and device, storage medium and computer readable storage medium

By generating a stroke number sequence and a derivative value sequence to dynamically adjust the watermark embedding position and performing multi-layer binary embedding, the problem of insufficient anti-attack ability of existing watermark embedding methods is solved, and the high concealment and robustness of the watermark information are achieved.

CN120805113AActive Publication Date: 2025-10-17VIPSHOP (GUANGZHOU) SOFTWARE CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511311894.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-15
Publication Date
2025-10-17
Estimated Expiration
2045-09-15

AI Technical Summary

Technical Problem

In existing watermark embedding methods, watermark information can be easily discovered and removed by attackers through statistical analysis, and the anti-attack capability is poor.

Method used

By obtaining the number of strokes of each character in the text to be processed, a stroke number sequence is generated, and the watermark embedding position is determined based on the sequence. The derived value of the stroke number and the text features are used to form a derived value sequence. The watermark embedding position is dynamically adjusted as the text changes, and the watermark information is converted into a binary sequence for multi-layer embedding. Zero-width characters and encryption processing are used to improve the concealment.

Benefits of technology

The real-time dynamic adjustment of the watermark embedding position is realized, the anti-attack capability of the watermark data is improved, and the concealment and robustness of the watermark information are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120805113A_ABST
    Figure CN120805113A_ABST
Patent Text Reader

Abstract

The invention discloses a watermark embedding method and device, a storage medium and a computer readable storage medium, and relates to the technical field of digital watermarking. The watermark embedding method comprises the following steps: acquiring a to-be-processed text, and generating a stroke number sequence based on the stroke number of each character in the to-be-processed text; determining a watermark embedding position according to the stroke number sequence; and based on the watermark embedding position, carrying out watermark embedding processing on the to-be-processed text to obtain a target watermark text. The technical problem that an existing watermark embedding method is poor in anti-attack capacity for embedded watermark data is solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of digital watermarking, and particularly relates to a watermark embedding method and device, a storage medium and a computer readable storage medium. BACKGROUND

[0002] As a core supporting means of digital copyright protection, the digital watermarking technology has formed an application system covering all fields of multimedia.

[0003] The traditional watermark embedding method usually embeds watermark information in a fixed position, so the watermark information embedded in the fixed position is easy to be found and removed by attackers through statistical analysis. That is, the existing watermark embedding method has poor attack resistance of embedded watermark data.

[0004] The above content is only used to assist in understanding the technical solutions of the present application and does not represent the acknowledgement of the above content as prior art. SUMMARY

[0005] The main purpose of the present application is to provide a watermark embedding method, device, storage medium and computer readable storage medium, which aims to solve the technical problem of poor attack resistance of existing watermark embedding method for embedded watermark data.

[0006] To achieve the above purpose, the present application provides a watermark embedding method, which comprises: Obtaining a to-be-processed text, and generating a stroke number sequence based on the stroke numbers of characters in the to-be-processed text; Determining a watermark embedding position according to the stroke number sequence; Performing watermark embedding processing on the to-be-processed text based on the watermark embedding position to obtain a target watermark text.

[0007] In an embodiment, the step of determining the watermark embedding position according to the stroke number sequence comprises: Obtaining a text feature of the to-be-processed text, and determining a derivative value of each stroke number in the stroke number sequence; Based on the derivative value and the text feature, a derivative value sequence is constituted, and a text position corresponding to each derivative value in the derivative value sequence is taken as the watermark embedding position.

[0008] In an embodiment, the text feature comprises a text complexity, and the step of constituting a derivative value sequence based on the derivative value and the text feature comprises: According to the text complexity, a corresponding sampling number is determined, wherein the sampling number is positively correlated with the text complexity; The sampled number of derivative values are selected from the derivative values of the number of strokes respectively, and a sequence of the selected derivative values is taken as a derivative value sequence.

[0009] In an embodiment, the text feature further comprises a text usage scenario, and the step of selecting the sampled number of derivative values from the derivative values of the number of strokes respectively, and taking a sequence of the selected derivative values as a derivative value sequence further comprises: After the text usage scenario is a word segmentation scenario, the sampled number of derivative values are selected from the derivative values of the number of strokes respectively, and a sequence of the selected derivative values is taken as an initial sequence. The symbol position of a punctuation mark in the text to be processed is identified, a value corresponding to a position before the symbol position is inserted into the initial sequence, and a derivative value sequence is obtained.

[0010] In an embodiment, the step of performing watermark embedding processing on the text to be processed based on the watermark embedding position to obtain a target watermark text comprises: Watermark information to be added is obtained, and the watermark information is converted into a binary sequence. The binary sequence is divided into a plurality of binary subsequences, and the number of the binary subsequences is consistent with the number of target watermark layers. According to each binary subsequence, watermark embedding processing is performed on the watermark embedding position of the text to be processed respectively, and the text to be processed with a plurality of watermark layers is obtained as a target watermark text.

[0011] In an embodiment, the step of performing watermark embedding processing on the text to be processed based on the watermark embedding position to obtain a target watermark text comprises: Each binary subsequence is mapped into a zero-width string based on a predetermined mapping rule. Each zero-width character in the zero-width string is sequentially embedded into the watermark embedding position in the text to be processed, and the text to be processed with a plurality of watermark layers is obtained as a target watermark text.

[0012] In an embodiment, the predetermined mapping rule between different binary subsequences and the zero-width string is different.

[0013] In addition, to achieve the above-mentioned purpose, the present application further provides a watermark embedding device, which comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, and the computer program is configured to implement the steps of the watermark embedding method as described above.

[0014] In addition, to achieve the above object, the present application also provides a storage medium, which is a computer readable storage medium, and a computer program is stored on the storage medium, and the computer program is executed by a processor to implement the steps of the watermark embedding method.

[0015] In addition, to achieve the above object, the present application also provides a computer program product, which comprises a computer program, and the computer program is executed by a processor to implement the steps of the watermark embedding method.

[0016] The one or more technical solutions provided by the present application have at least the following technical effects: The present application can obtain a to-be-processed text, and generate a stroke number sequence based on the stroke numbers of characters in the to-be-processed text. Therefore, the present application can determine a watermark embedding position according to the stroke number sequence, that is, the watermark embedding position changes with the stroke numbers of the characters in the to-be-processed text. Further, the present application can perform watermark embedding processing on the to-be-processed text based on the watermark embedding position to obtain a target watermark text. The present application can change the embedding position according to the stroke numbers of the characters in the to-be-processed text, thereby realizing real-time dynamic adjustment of the watermark embedding position in the watermark embedding process. Compared with the existing watermark embedding method, the embedding method of the present application cannot discover the watermark position by statistical analysis or other methods, thereby effectively improving the anti-attack ability of the embedded watermark data. BRIEF DESCRIPTION OF DRAWINGS

[0017] The accompanying drawings, which are incorporated into and form part of the specification, illustrate embodiments consistent with the present application and, together with the specification, serve to explain the principles of the present application.

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the accompanying drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, for those skilled in the art, other drawings can also be obtained without creative labor based on these drawings.

[0019] Figure 1 Flowchart provided for the first embodiment of the watermark embedding method of the present application; Figure 2 Scenario diagram of the watermark embedding position related to the embodiments of the present application; Figure 3 Flowchart provided for the second embodiment of the watermark embedding method of the present application; Figure 4 Flowchart provided for the third embodiment of the watermark embedding method of the present application; Figure 5This is a schematic diagram of the structure of the watermark embedding device in an embodiment of the present application.

[0020] The purpose, features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0021] It should be understood that the specific embodiments described herein are merely used to explain the technical solutions of the present application and are not intended to limit the present application.

[0022] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.

[0023] The main solution of the embodiment of the present application is: obtain the text to be processed, and generate a stroke number sequence based on the number of strokes of each character in the text to be processed; determine the watermark embedding position according to the stroke number sequence; based on the watermark embedding position, perform watermark embedding processing on the text to be processed to obtain the target watermark text.

[0024] Traditional watermark embedding methods usually embed watermark information in a fixed position. Therefore, these fixed-position embedded watermark information can be easily discovered and removed by attackers through statistical analysis. In other words, the existing watermark embedding methods have poor anti-attack capabilities against the embedded watermark data.

[0025] The present application provides a solution that can change the embedding position accordingly with the number of strokes of each character in the text to be processed, thereby realizing real-time dynamic adjustment of the watermark embedding position during the watermark embedding process. Compared with the existing watermark embedding method, the embedding method of the present application cannot be used by attackers to discover the watermark position through statistical analysis and other methods, effectively improving the anti-attack capability of the embedded watermark data.

[0026] Based on this, the embodiment of the present application provides a watermark embedding method, referring to Figure 1 , Figure 1 This is a flow chart of the first embodiment of the watermark embedding method of the present application.

[0027] In this embodiment, the watermark embedding method includes steps S10 to S30: Step S10, obtaining a text to be processed, and generating a stroke number sequence based on the number of strokes of each character in the text to be processed; It should be noted that the text to be processed is text data that is expected to be watermark embedded.

[0028] In addition, it should be noted that the characters refer to a general term for various characters and symbols such as national characters, punctuation marks, graphic symbols, and numbers.

[0029] The embodiment can analyze strokes of each character in the to-be-processed text after obtaining the to-be-processed text, and obtain stroke numbers of each character in the to-be-processed text. Then, the stroke numbers are sequentially sorted to form a stroke number sequence. For example, the embodiment can query the stroke numbers corresponding to each character through a predetermined database to obtain the stroke numbers of each character in the to-be-processed text. The predetermined database stores a corresponding relationship between each character and the stroke number. For example, assuming that the to-be-processed text is "This is a test text.", the stroke numbers of each character can be obtained by querying a Chinese character stroke number database, and then the stroke numbers are sorted to obtain a stroke number sequence [7, 9, 1, 3, 9, 8, 4, 5, 1]. It can be understood that the stroke numbers constituting the stroke number sequence can be stroke numbers of all characters in the to-be-processed text, or stroke numbers of part of the characters.

[0030] In step S20, a watermark embedding position is determined according to the stroke number sequence. It should be noted that the watermark embedding position can be described in the form of a numerical value, such as Figure 2 As shown in the figure, positions between adjacent characters in the to-be-processed text are taken as one embeddable position, and then each embeddable position is sequentially coded to obtain a numerical value corresponding to each embeddable position (i.e. 1, 2, 3, 4, 5, 6, 7, 8, … in Figure 2 ).

[0031] As an example, the embodiment can directly take the stroke number sequence as a sequence representing the watermark embedding positions, i.e., take the embeddable positions corresponding to each stroke number in the stroke number sequence as the watermark embedding positions. As another example, the embodiment can determine a derived value of each stroke number in the stroke number sequence. For example, the derived value can be a result value after performing an operation such as multiplication / addition between the stroke number and a predetermined number sequence (such as a prime number sequence, a Fibonacci sequence, etc.). Then the embodiment can select a specified number of derived values from the derived values to form a derived value sequence. Take the prime number sequence as an example, for the first character "this": 7 strokes, the derived value is the product of the stroke number and each prime number in the prime number sequence, i.e., 7*2=14, 7*3=21, 7*5=35, etc. For the second character "is": 9 strokes, the derived value is 9*2=18, 9*3=27, 9*5=45, etc. In this way, the embodiment can select the first two derived values of each stroke number to form a derived value sequence [14, 21, 18, 27, 14, 21, 18, 27, 14]. As another example, in order to further improve the concealment and robustness of the watermark information, the embodiment can further obtain a text feature of the to-be-processed text, determine a derived value of each stroke number in the stroke number sequence, and then form a derived value sequence based on the derived value and the text feature, and take the text positions corresponding to each derived value in the derived value sequence as the watermark embedding positions. Thus, the selection of the derived value (i.e., the watermark embedding position) is associated with the text feature of the to-be-processed text, further adding a new disturbance factor to the selection of the watermark embedding position, and the concealment and robustness of the watermark information can be further improved.

[0032] In step S30, based on the watermark embedding positions, the to-be-processed text is subjected to watermark embedding processing to obtain a target watermark text.

[0033] As an example, the embodiment can obtain watermark information to be added, and convert the watermark information into a binary sequence, and perform watermark embedding processing on watermark embedding positions of the to-be-processed text according to the binary sequence to obtain a target watermark text. As an example, the embodiment can convert the binary sequence into a hidden mark string, and then embed hidden mark characters in the hidden mark string into the watermark embedding positions in sequence to obtain the target watermark text. The hidden mark character is a special character that does not occupy visible space, for example, one or a combination of multiple of zero-width characters, word joiners, invisible multiplication marks, and invisible separators. In addition, in order to further improve the security of the watermark information, the binary sequence can be encrypted after the watermark information is converted into the binary sequence, and the step of performing watermark embedding processing on watermark embedding positions of the to-be-processed text according to the binary sequence to obtain a target watermark text is performed based on the encrypted binary sequence. As another example, the embodiment can also obtain watermark information to be added, and convert the watermark information into a binary sequence, and then divide the binary sequence into a plurality of binary subsequences, where the number of the binary subsequences is consistent with the number of target watermark layers; watermark embedding processing is performed on watermark embedding positions of the to-be-processed text according to each of the binary subsequences, and the to-be-processed text with multiple watermark layers is obtained as a target watermark text. In this way, the superposition of multiple layers of watermark can be realized to improve the complexity and attack resistance of the watermark. In addition, in order to further improve the security of the watermark information, the binary sequence can be encrypted after the watermark information is converted into the binary sequence, and the step of dividing the binary sequence into a plurality of binary subsequences is performed based on the encrypted binary sequence.

[0034] The first embodiment of the present application provides a watermark embedding method, which obtains a to-be-processed text, and generates a stroke number sequence based on stroke numbers of characters in the to-be-processed text. In this way, the embodiment can determine watermark embedding positions according to the stroke number sequence, that is, the watermark embedding positions change with the stroke numbers of the characters in the to-be-processed text. Then, the embodiment can perform watermark embedding processing on the to-be-processed text based on the watermark embedding positions to obtain a target watermark text. The embodiment can change the embedding positions according to the stroke numbers of the characters in the to-be-processed text, thereby realizing real-time dynamic adjustment of the watermark embedding positions in the watermark embedding process. Compared with existing watermark embedding methods, the embedding method of the embodiment cannot be used by attackers to find watermark positions through statistical analysis, thereby effectively improving the attack resistance of the embedded watermark data.

[0035] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar contents as the above embodiment one can refer to the above description, and the subsequent will not be described. On this basis, please refer to Figure 3 , step S20 includes steps S21-S22: Step S21, obtaining the text features of the text to be processed, and determining the derivative value of each stroke number in the stroke number sequence; Step S22, based on the derivative value and the text features, a derivative value sequence is formed, and the text position corresponding to each derivative value in the derivative value sequence is taken as the watermark embedding position.

[0036] It should be noted that the text features of the text to be processed can include the features of the text to be processed itself (such as font, font size, language type, etc.), and the associated features (such as the use scene).

[0037] It should also be noted that the derivative value is a value calculated based on the stroke number through a predetermined operation mode. For example, the derivative value can be the result value after multiplication / addition operation between the stroke number and a predetermined number sequence (such as prime number sequence, Fibonacci number sequence, etc.).

[0038] The embodiment can extract feature information of the to-be-processed text to obtain text features of the to-be-processed text. It can be understood that the manner of extracting the feature information can be determined according to the text features to be extracted. For example, taking the language type as an example, since different languages use specific character sets, the language type of the to-be-processed text can be determined by the character set features of the to-be-processed text. In addition, the font size, font, and other features of the to-be-processed text can be obtained by reading the format information of the to-be-processed text, and the use scenario and other related features of the to-be-processed text can be obtained by reading the associated information of the to-be-processed text. Furthermore, the embodiment can determine the derived values of each stroke number in the stroke number sequence. For example, the derived value can be the result value after multiplication / addition operation between the stroke number and a predetermined number sequence (such as a prime number sequence, a Fibonacci sequence, etc.). Taking the prime number sequence as an example, the first character "this" has 7 strokes, and the derived value is the product of the stroke number and each prime number in the prime number sequence, that is, 7*2=14, 7*3=21, 7*5=35, and so on. The second character "is" has 9 strokes, and the derived value is 9*2=18, 9*3=27, 9*5=45, and so on. Furthermore, the embodiment can adjust the derived values according to the text features to form a derived value sequence. Thus, the embeddable positions corresponding to each derived value in the derived value sequence can be used as watermark embedding positions. As an example, the text features include text complexity, and the embodiment can determine the corresponding sampling number according to the text complexity, wherein the sampling number is positively correlated with the text complexity; and the derived values of the sampling number are selected from the derived values of the stroke number, and the sequence of the selected derived values is used as the derived value sequence. Since the greater the sampling number is, the more watermark embedding positions are, the embodiment can provide higher-density embedding positions for more complex to-be-processed texts. As another example, the text features also include a text use scenario, and the embodiment can select the derived values of the sampling number from the derived values of the stroke number after the text use scenario is a word segmentation scenario (such as a natural language processing, search engine indexing, or other scenario requiring word segmentation processing), and the sequence of the selected derived values is used as an initial sequence. Furthermore, the symbol positions of punctuation marks in the to-be-processed text can be identified, a value corresponding to a position before the symbol position (i.e., a text position between a punctuation mark and a previous character) is inserted into the initial sequence to obtain a derived value sequence. Thus, redundant watermark embedding positions can be increased, which is beneficial to protecting the integrity of the watermark information.

[0039] In some embodiments, the text features include text complexity, and step S22 includes steps A10-A20. Step A10, determining a corresponding sampling number according to the text complexity, wherein the sampling number is positively correlated with the text complexity. Step A20, respectively sampling the derivative values of the sampling number from the derivative values of the stroke number, and taking the sequence of the sampled derivative values as the derivative value sequence.

[0040] It should be noted that the text complexity is a representation value representing the content complexity of the to-be-processed text. For example, different language types can be assigned different complexity values, and text contents with different data amounts can be assigned corresponding complexity values, so as to obtain the text complexity of the to-be-processed text through weighted calculation.

[0041] According to the text complexity, the embodiment can determine a corresponding sampling number, which is the number of sampled derivative values in a derivative value of a stroke number. The sampling number is positively correlated with the text complexity, that is, the higher the text complexity, the greater the sampling number, that is, more watermark embedding positions are increased. For example, in Chinese text, the number of character strokes is large, and Chinese is more complex than English. Therefore, the text complexity of the to-be-processed text with a language type of Chinese and / or a larger amount of text data is higher, and the watermark embedding density can be increased to effectively enhance the concealment and robustness of the watermark information. Furthermore, the embodiment can respectively sample the derivative values of the sampling number from the derivative values of the stroke number, and take the sequence of the sampled derivative values as the derivative value sequence. Taking the stroke number sequence [7, 9, 1, 3, 9, 8, 4, 5, 1] as an example, 1 character “this”: 7 strokes, the derivative value is the product of the stroke number and each prime number in the prime number sequence, that is, 7*2=14, 7*3=21, 7*5=35…… The second character “is”: 9 strokes, the derivative value is 9*2=18, 9*3=27, 9*5=45…… and so on. According to the text complexity, the corresponding sampling number is 2, and the embodiment can select the first two derivative values of each stroke number to form the derivative value sequence [14, 21, 18, 27, 14, 21, 18, 27, 14].

[0042] In some embodiments, the text features further include a text use scenario, and step A20 includes steps B10-B20: Step B10, after determining that the text use scenario is a word segmentation scenario, respectively sampling the derivative values of the sampling number from the derivative values of the stroke number, and taking the sequence of the sampled derivative values as an initial sequence. Step B20, identifying a symbol position of a punctuation symbol in the to-be-processed text, inserting a value corresponding to a position before the symbol position into the initial sequence to obtain a derivative value sequence.

[0043] It should be noted that the word segmentation scenario is a scenario that needs to be segmented, such as natural language processing and search engine indexing.

[0044] The embodiment can determine whether the text use scenario is a word segmentation scenario. If the text use scenario is not a word segmentation scenario, step A10 can be performed. After determining that the text use scenario is a word segmentation scenario, the embodiment can determine a corresponding number of samples according to the text complexity, and then select a number of derived values from the derived values of the number of strokes, and take the sequence of the selected derived values as an initial sequence. Furthermore, the embodiment can identify the symbol positions of punctuation marks in the text to be processed, insert the value corresponding to the position before the symbol position (i.e., the position of the text between the punctuation mark and the previous character) into the initial sequence, and obtain a derived value sequence. In this way, redundant watermark embedding positions can be added before question marks, periods, and other punctuation marks, so that the integrity of the watermark information can be protected by the strong separation of punctuation marks.

[0045] In the second embodiment of the present application, the text features of the text to be processed are obtained, and the derived values of the strokes in the stroke sequence are determined. Furthermore, based on the derived values and the text features, a derived value sequence is formed, and the text positions corresponding to the derived values in the derived value sequence are taken as the watermark embedding positions. In this way, the derived value sequence can be adjusted adaptively according to different text features of the text to be processed, further enhancing the concealment and robustness of the embedded watermark information.

[0046] Based on the first embodiment of the present application, the same or similar contents as the above embodiment one can be referred to the above description, and will not be repeated hereinafter. On this basis, please refer to Figure 4 , step S30 includes steps S31-S33: Step S31, obtaining the watermark information to be added, and converting the watermark information into a binary sequence; Step S32, dividing the binary sequence into a plurality of binary subsequences, wherein the number of binary subsequences is consistent with the target number of watermark layers; Step S33, performing watermark embedding processing on the watermark embedding positions of the text to be processed according to each binary subsequence, and obtaining a text to be processed with multiple watermark layers as a target watermark text.

[0047] It should be noted that the watermark information can include pre-set static identification information, such as warning text "confidential file, prohibited external transmission", and real-time acquired dynamic identification information, such as the identity information of a web browser (e.g., account number, employee number, etc.).

[0048] It should be further noted that the target number of watermark layers is the number of watermark layers expected to be superimposed on the text to be processed. The target number of watermark layers can be a fixed number of layers preset in advance, or can be determined based on at least one of a security requirement, a text feature of the text to be processed, and a watermark information amount of the watermark information.

[0049] In this embodiment, the watermark information to be added can be obtained and converted into a binary sequence. For example, the watermark information can be converted into a binary sequence by using binary plaintext encoding, error correction encoding, encryption / disturbance encoding, check encoding, spread spectrum encoding, synchronization code, graphical encoding, compression encoding, etc. In addition, the binary sequence can be encrypted, and step S32 can be performed based on the encrypted binary sequence. The encryption can be implemented by using an encryption algorithm such as AES (Advanced Encryption Standard), RSA (Rivest-Shamir-Adleman), etc. Furthermore, the binary sequence can be divided into a plurality of binary subsequences, and the number of the binary subsequences is consistent with the target number of watermark layers. For example, the target number of watermark layers can be set according to security requirements. The target number of watermark layers can be positively correlated with the level of security requirements. For example, in a high-security requirement scenario (such as copyright protection), more layers can be selected as the target number of watermark layers to enhance the attack resistance. In a low-security requirement scenario (such as advertisement tracking), fewer layers can be selected as the target number of watermark layers. The target number of watermark layers can also be determined based on the text features of the text to be processed. The target number of watermark layers is positively correlated with the text complexity in the text features. For example, more layers can be selected as the target number of watermark layers for long texts with more information, and fewer layers can be selected as the target number of watermark layers for short texts to avoid information overload. For example, more layers can be selected as the target number of watermark layers for Chinese characters with rich strokes, and fewer layers can be selected as the target number of watermark layers for English characters with simple strokes. The target number of watermark layers can also be determined based on the amount of watermark information. The target number of watermark layers is positively correlated with the amount of watermark information. For example, more layers can be selected as the target number of watermark layers for more information, and fewer layers can be selected as the target number of watermark layers for less information. For example, the watermark information is “original”, which is first converted into a binary sequence. For example, using UTF-8 encoding, the UTF-8 encoding of “original” is E6ADA3, and the corresponding binary value string is 1110011010101 101 10100011. The UTF-8 encoding of “original” is E78988, and the binary value string is 11100111 10001001 10001000. The combined binary sequence is: 11100110 10101101 10100011 11100111 10001001 10001000.The binary sequence is encrypted using the AES encryption algorithm to generate an encrypted binary sequence: 10101010 11001100 11110000 10101010 11001100 11110000. The target number of watermark layers is two, so the encrypted binary sequence can be divided into two binary subsequences corresponding to the two watermark layers. The first layer is: 10101010 11001100 11110000, and the second layer is: 10101010 11001100 11110000. Then, the embodiment can perform watermark embedding processing on the watermark embedding positions of the text to be processed according to the binary subsequences, respectively, to obtain the text to be processed with multiple watermark layers as the target watermark text.

[0050] In some embodiments, step S33 includes steps C10-C20: Step C10: mapping each binary subsequence into a zero-width string based on a predetermined mapping rule; Step C20: embedding each zero-width character in the zero-width string into a watermark embedding position in the text to be processed in sequence to obtain the text to be processed with multiple watermark layers as the target watermark text.

[0051] It should be noted that the embodiment uses zero-width characters as invisible mark characters, and the zero-width characters include U+200B (zero-width space), U+200C (zero-width non-joiner), U+200D (zero-width joiner), U+FEFF (zero-width non-breaking space), etc. The predetermined mapping rule includes the mapping relationship between binary values (“1” and “0”) and zero-width characters. It can be understood that a binary value (“1” or “0”) can correspond to a single zero-width character or a combination of zero-width characters. That is, for more watermark layers, a combination of multiple zero-width characters can be used to represent a binary value in the predetermined mapping rule. For example, U+200B+U+200C represents binary value “1”, and U+200D+U+FEFF represents binary value “0”, thereby expanding the mapping space. A combination of multiple zero-width characters can also be used in combination with a text position marker to implement the mapping of binary subsequences. For example, U+200B is used to represent the first layer of watermark layers at even positions, and U+200D is used to represent the second layer of watermark layers at odd positions.

[0052] In this embodiment, each binary subsequence can be mapped to a zero-width string based on a predetermined mapping rule. For example, the predetermined mapping rule can be that binary value "1" is represented by U+200B (zero-width space) and binary value "0" is represented by U+200C (zero-width non-joiner). Taking the binary subsequence 10101010 as an example, the zero-width string is U+200B U+200C U+200B U+200C U+200B U+200C U+200B U+200C. Then, each zero-width character in the zero-width string is sequentially embedded into the watermark embedding position in the text to be processed. For example, if the watermark embedding position is [14, 21, 18, 27, 14, 21, 18, 27], then binary value "1" corresponding to U+200B (zero-width space) is embedded at text position 14, binary value "0" corresponding to U+200C (zero-width non-joiner) is embedded at text position 21, binary value "1" corresponding to U+200B (zero-width space) is embedded at text position 18, and so on. Thus, after each binary subsequence is mapped to a zero-width character, the text to be processed with the watermark layer with the target number of layers is obtained as the target watermark text. Thus, the target watermark text after embedding the watermark in this embodiment is completely consistent with the original text and has no visible changes.

[0053] In a feasible embodiment, the predetermined mapping rule between different binary subsequences and the zero-width string is different.

[0054] In this embodiment, a plurality of different mapping rules can be constructed in advance. Before step C10, a different mapping rule can be selected for each binary subsequence as a predetermined mapping rule, and then each binary subsequence can be mapped to a zero-width string based on the predetermined mapping rule. For example, the mapping rule selected for the binary subsequence of the first layer of watermark is as follows: binary value "1" is represented by U+200B (zero-width space); and binary value "0" is represented by U+200C (zero-width non-joiner). The mapping rule selected for the binary subsequence of the second layer of watermark is as follows: binary value "1" is represented by U+200D (zero-width joiner); and binary value "0" is represented by U+FEFF (zero-width non-breaking space). In this embodiment, the predetermined mapping rule between different binary subsequences and the zero-width string is different, so that different types of zero-width characters are used for each watermark layer on the target watermark text after embedding the watermark, thereby improving the complexity and attack resistance of the watermark.

[0055] After that, in the steps of watermark extraction and verification, the embodiment can extract the zero-width string in the target watermark text, and according to the predetermined mapping rule, reversely map the binary sub-sequences to form a binary sequence. Then the binary sequence is decrypted and decoded to obtain the watermark information, and then the integrity and authenticity of the watermark information are identified.

[0056] In the third embodiment of the present application, the watermark information to be added is obtained, and the watermark information is converted into a binary sequence. The binary sequence is divided into a plurality of binary sub-sequences, wherein the number of the binary sub-sequences is consistent with the number of target watermark layers. According to each binary sub-sequence, watermark embedding processing is performed on the watermark embedding position of the to-be-processed text, and the to-be-processed text with multiple watermark layers is obtained as the target watermark text. Thus, the embodiment realizes the superposition of multiple layers of watermarks on the target watermark text, each layer of watermark uses different mapping rules, and the complexity and attack resistance of the watermark are improved.

[0057] The present application provides a watermark embedding device, which comprises at least one processor and a memory connected with the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the watermark embedding method in the above-mentioned embodiment one.

[0058] Reference will be made to the following description Figure 5 which shows a structural schematic diagram of a watermark embedding device suitable for being used to implement the embodiments of the present application. The watermark embedding device in the embodiments of the present application can include but is not limited to terminals such as mobile phones, notebook computers, PDAs (Personal Digital Assistant: personal digital assistants), PADs (Portable Application Description: tablet computers), desktop computers, etc. Figure 5 The watermark embedding device shown is only an example, and should not bring any limitation to the functions and use range of the embodiments of the present application.

[0059] As Figure 5As shown, the watermark embedding device can include a processing device 1001 (e.g., a central processing unit, a graphics processing unit, etc.) that can perform various appropriate actions and processes according to a program stored in a read only memory (ROM) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM) 1004. In the random access memory 1004, various programs and data required for the operation of the watermark embedding device are also stored. The processing device 1001, the read only memory 1002, and the random access memory 1004 are connected to each other through a bus 1005. An I / O (input / output) interface 1006 is also connected to the bus. Generally, the following systems can be connected to the I / O interface 1006: input devices 1007 including, for example, a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; the storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the watermark embedding device to communicate wirelessly or wired with other devices to exchange data. Although the watermark embedding device with various systems is shown in the figure, it should be understood that all the shown systems are not required to be implemented or possessed. More or less systems can be alternatively implemented or possessed.

[0060] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, the embodiments disclosed in the present application include a computer program product comprising a computer program carried on a computer readable medium, the computer program containing program codes for executing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network through the communication device, or installed from the storage device 1003, or installed from the read only memory 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the methods of the embodiments disclosed in the present application are performed.

[0061] The watermark embedding device provided in the present application adopts the watermark embedding method in the above-mentioned embodiments, and can solve the technical problem that the existing watermark embedding method has poor attack resistance for embedded watermark data. Compared with the prior art, the watermark embedding device provided in the present application has the same beneficial effects as the watermark embedding method provided in the above-mentioned embodiments, and other technical features in the watermark embedding device are the same as the features disclosed in the previous embodiment method, which will not be repeated here.

[0062] It should be understood that various aspects of the disclosure can be implemented in hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials or characteristics can be combined in any appropriate manner in any one or more embodiments or examples.

[0063] The above description is only specific embodiments of the present application, but the protection scope of the present application is not limited thereto, and any person skilled in the art can easily think of changes or replacements within the technical range disclosed by the present application, which should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

[0064] The present application provides a computer readable storage medium having stored thereon computer readable program instructions (i.e. computer programs) for performing the watermark embedding method in the above embodiments.

[0065] The computer readable storage medium provided by the present application may, for example, be a U disk, but is not limited to an electric, magnetic, optical, electromagnetic, infrared, or semiconductor system or device, or any combination thereof. More specific examples of the computer readable storage medium can include, but are not limited to, an electric connection with one or more conductive wires, a portable computer disk, a hard disk, a random access memory (RAM), a read only memory (ROM), an erasable programmable read only memory (EPROM or flash memory), an optical fiber, a portable compact disk read only memory (CD-ROM), an optical storage device, a magnetic storage device, or any appropriate combination thereof. In the present embodiment, the computer readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system or device. The program code contained on the computer readable storage medium can be transmitted by any appropriate medium, including but not limited to an electric wire, an optical cable, an RF (Radio Frequency), etc., or any appropriate combination thereof.

[0066] The above computer readable storage medium can be contained in a watermark embedding device; or can exist separately without being assembled into the watermark embedding device.

[0067] The computer readable storage medium carries one or more programs, when the one or more programs are executed by the watermark embedding device, the watermark embedding device is caused to: acquire a to-be-processed text, and generate a stroke number sequence based on stroke numbers of characters in the to-be-processed text; determine a watermark embedding position according to the stroke number sequence; and perform watermark embedding processing on the to-be-processed text based on the watermark embedding position, to obtain a target watermark text.

[0068] Computer program code for carrying out operations of the present application can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).

[0069] The flow diagrams and the block diagrams in the drawings are illustrations of architectures, functionalities, and operations of possible implementations of systems, methods, and computer program products according to various embodiments of present application. In this regard, each block in the flow diagrams or block diagrams can represent a module, a segment, or a portion of code, which comprises one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flow diagrams, and combinations thereof, can be implemented by special purpose hardware-based systems that perform the specified functions or operations, or combinations of special purpose hardware and computer instructions.

[0070] The modules involved in the embodiments of the present application can be implemented in the manner of software or in the manner of hardware. In some cases, the name of the module does not constitute a limitation on the module itself.

[0071] The readable storage medium provided by the application is a computer readable storage medium, which stores computer readable program instructions (i.e. computer program) for executing the watermark embedding method described above, and can solve the technical problem that the existing watermark embedding method has poor attack resistance to embedded watermark data. Compared with the prior art, the computer readable storage medium provided by the application has the same beneficial effects as the watermark embedding method provided by the above-mentioned embodiments, and will not be repeated here.

[0072] The application also provides a computer program product comprising a computer program, which, when executed by a processor, implements the steps of the watermark embedding method as described above.

[0073] The computer program product provided by the application can solve the technical problem that the existing watermark embedding method has poor attack resistance to embedded watermark data. Compared with the prior art, the computer program product provided by the application has the same beneficial effects as the watermark embedding method provided by the above-mentioned embodiments, and will not be repeated here.

[0074] The above-mentioned is only part of the embodiments of the application, and does not limit the patent scope of the application, and any equivalent structural transformation, direct / indirect application in other related technical fields within the technical concept of the application, and the contents of the specification and drawings are included in the patent protection scope of the application.

Claims

1. A watermark embedding method, characterized in that: The watermark embedding method comprises: Acquire a text to be processed, and generate a stroke number sequence based on the number of strokes of each character in the text to be processed; Determining a watermark embedding position according to the stroke number sequence; Based on the watermark embedding position, watermark embedding processing is performed on the text to be processed to obtain a target watermark text.

2. The watermark embedding method according to claim 1, wherein: The step of determining the watermark embedding position according to the stroke number sequence includes: Acquiring text features of the text to be processed, and determining a derivative value of each stroke number in the stroke number sequence; Based on the derived values ​​and the text features, a derived value sequence is constructed, and the text position corresponding to each derived value in the derived value sequence is used as the watermark embedding position.

3. The watermark embedding method according to claim 2, wherein: The text feature includes text complexity, and the step of forming a derived value sequence based on the derived value and the text feature includes: Determining a corresponding sampling number according to the text complexity, wherein the sampling number is positively correlated with the text complexity; The derivative values ​​of the sampling number are respectively selected from the derivative values ​​of the stroke number, and the sequence of the selected derivative values ​​is used as the derivative value sequence.

4. The watermark embedding method according to claim 3, wherein: The text feature further includes a text usage scenario, and the step of respectively selecting the derived values ​​of the sampled number from the derived values ​​of the number of strokes and using the sequence of the selected derived values ​​as the derived value sequence further includes: After the text usage scenario is a word segmentation scenario, respectively selecting derived values ​​of the sampled number from the derived values ​​of the number of strokes, and using a sequence of the selected derived values ​​as an initial sequence; The symbol position of the punctuation mark in the text to be processed is identified, and the numerical value corresponding to the position before the symbol position is inserted into the initial sequence to obtain a derived value sequence.

5. The watermark embedding method according to claim 1, wherein: The step of performing watermark embedding processing on the text to be processed based on the watermark embedding position to obtain a target watermark text includes: Obtaining watermark information to be added, and converting the watermark information into a binary sequence; Splitting the binary sequence into a plurality of binary subsequences, wherein the number of the binary subsequences is consistent with the number of target watermark layers; According to each of the binary subsequences, watermark embedding processing is performed on the watermark embedding position of the text to be processed respectively, and the text to be processed with multiple watermark layers is obtained as the target watermark text.

6. The watermark embedding method according to claim 5, wherein: The step of performing watermark embedding processing on the watermark embedding position of the to-be-processed text according to each of the binary subsequences to obtain the to-be-processed text with multiple watermark layers as the target watermark text includes: Mapping each of the binary subsequences into a zero-width character string based on a predetermined mapping rule; Each zero-width character in the zero-width character string is sequentially embedded in the watermark embedding position of the to-be-processed text to obtain the to-be-processed text with multiple watermark layers as the target watermark text.

7. The watermark embedding method according to claim 6, wherein: The predetermined mapping rules between different binary subsequences and the zero-width character string are different.

8. A watermark embedding device, characterized in that: The device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the watermark embedding method according to any one of claims 1 to 7.

9. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the watermark embedding method according to any one of claims 1 to 7 are implemented.

10. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the steps of the watermark embedding method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Form text anti-counterfeiting watermark generation method and system and computer storage medium

    CN115082281A

  • Method for embedding and extracting printing and scanning resistant digital watermark for text image

    CN116977149A

  • Method for embedding and extracting watermark in English texts

    CN1700205A