Text watermark embedding method and device, text watermark extraction method and device, electronic equipment, storage medium and computer program product

By selectively embedding watermark encoding information based on the total number of strokes and tone values ​​in the text data, the problem of insufficient redundancy space in text watermark embedding is solved, achieving highly concealed and robust watermark embedding, thus improving text quality and security.

CN121935898APending Publication Date: 2026-04-28CHINA MOBILE COMM LTD RES INST +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA MOBILE COMM LTD RES INST
Filing Date
2025-12-19
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing technologies for embedding watermarks in text data suffer from problems such as insufficient redundancy space, difficulty in balancing watermark information capacity and robustness with text quality, resulting in obvious embedding traces and affecting text quality.

Method used

By adjusting the total number of strokes and tone values ​​of each sub-text data, watermark encoding information is selectively embedded. The encoding range and density of the watermark information are controlled by dilution parameters and embedding amount parameters. The watermark information is encrypted using a preset key to generate a watermark encoding information sequence, which is then adjusted in the text to conform to the rules.

Benefits of technology

It achieves watermark embedding with good concealment and high robustness in text data, reduces the impact on the original text, and improves the quality and security of the embedded text.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121935898A_ABST
    Figure CN121935898A_ABST
Patent Text Reader

Abstract

The invention provides a text watermark embedding method and device, a text watermark extraction method and device, electronic equipment, a storage medium and a computer program product, and relates to the technical field of information security. Wherein the text data comprises a plurality of pieces of first sub-text data; based on the total number of strokes corresponding to each piece of first sub-text data, determining an embedding detection result; wherein the embedding detection result is used for representing whether corresponding first watermark coding information is embedded in the first sub-text data or not; if the embedding detection result represents embedding, determining first watermark coding information corresponding to the first sub-text data in the watermark coding information sequence, and adjusting text information in the first sub-text data until the total number of strokes and the total number of tone values of the first sub-text data conform to corresponding rules to obtain embedded text data; wherein the watermark coding information sequence is formed based on watermark information embedded in the text data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of information security technology, and in particular to a text watermark embedding method, a text watermark extraction method, an apparatus, an electronic device, a storage medium, and a computer program product. Background Technology

[0002] Digital watermarking technology is an important tool in the field of information security, used to embed invisible information into digital content to achieve functions such as copyright protection and content tracking. Watermarking embedding methods for text data can be divided into two main categories: format-oriented and content-oriented.

[0003] In related technologies, text-format watermarking techniques can be implemented through various means such as modifying font size, font type, color, background shading, and full / half-width characters. However, text-content-oriented watermarking techniques can only embed watermarks into the content itself. The challenges are twofold: first, limited redundancy space—due to the smaller data volume of text and the scarcity of redundant information, and the fact that text content consists mainly of meaningful characters, making it difficult to embed watermark information; and second, the balance between watermark capacity, robustness, and text quality. Because of the limited redundancy space, improving watermark embedding capacity and robustness requires further modification of the original text content, impacting text quality. Current technical solutions still suffer from low embedded text quality and noticeable embedding artifacts in practical applications. Summary of the Invention

[0004] This application provides a text watermark embedding method, a text watermark extraction method, an apparatus, an electronic device, a storage medium, and a computer program product.

[0005] The technical solution of this application is implemented as follows: This application provides a text watermark embedding method, including: Acquire text data; wherein the text data includes: multiple first sub-text data; Based on the total number of strokes corresponding to each of the first sub-text data, an embedding detection result is determined; wherein, the embedding detection result is used to characterize whether the corresponding first watermark encoding information is embedded in the first sub-text data; If the embedding detection result indicates embedding, then the first watermark encoding information corresponding to the first sub-text data is determined in the watermark encoding information sequence, and the text information in the first sub-text data is adjusted until the total number of strokes and the total number of tone values ​​of the first sub-text data meet the corresponding rules to obtain embedded text data; wherein, the watermark encoding information sequence is formed based on the watermark information embedded in the text data.

[0006] The method in the above scheme further includes: Obtain watermark information, embedding amount parameters, and dilution parameters; The embedding amount parameter is used to determine the proportion of the first sub-text data in which the first watermark encoding information is embedded among multiple first sub-text data; the dilution parameter is used to determine the encoding range of the watermark information.

[0007] In the above scheme, the embedding parameters include: a first parameter and a second parameter; wherein, the first parameter is used to take the remainder of the total number of strokes of the first sub-text data; and the second parameter is a preset maximum remainder value. The dilution parameter includes a third parameter and a fourth parameter; wherein the third parameter is used to take the remainder of the total number of tone values ​​of the first sub-text data; and the fourth parameter is used to determine the number of segments for the tone values.

[0008] The method in the above scheme further includes: The watermark information is converted into first encoded information in the encoding base characterized by the dilution parameter; wherein the encoding base is determined based on the number of segments determined by the fourth parameter; Convert the preset key into second encoded information in the encoded base; The first encoded information is encrypted based on the second encoded information to determine the watermark encoded information sequence.

[0009] In the above scheme, determining the embedding detection result based on the total number of strokes corresponding to each of the first sub-text data includes: The first remainder is taken from the total number of strokes of the first sub-text data based on the first parameter; The embedding detection result is determined based on the comparison between the first remainder and the second parameter.

[0010] In the above scheme, determining the first watermark encoding information corresponding to the first sub-text data in the watermark encoding information sequence includes: Determine the third encoded information resulting from the combination of the preset key and the preset position text information in the first sub-text data; The second remainder is taken based on the sequence length of the watermark encoded information sequence according to the third encoded information; The first watermark encoding information is determined in the watermark encoding information sequence based on the first offset information represented by the second remainder.

[0011] In the above scheme, before adjusting the text information in the first sub-text data until the total number of strokes and the total number of tone values ​​of the first sub-text data meet the corresponding rules, the method further includes: The third remainder is taken based on the total number of tone values ​​of the first sub-text data according to the third parameter. The tone segment corresponding to the third remainder and the encoded character corresponding to the tone segment are determined; wherein each tone segment is determined based on the number of segments determined by the fourth parameter; Based on the consistency between the first watermark encoding information and the encoded character, an adjustment detection result is determined; wherein, the adjustment detection result is used to characterize whether to adjust the text information in the first sub-text data until the total number of strokes and the total number of tone values ​​of the first sub-text data meet the corresponding rules.

[0012] In the above scheme, adjusting the text information in the first sub-text data until the total number of strokes and the total number of tone values ​​of the first sub-text data conform to the corresponding rules includes: In the first sub-text data, text information is replaced or inserted to obtain embedded sub-text data; Specifically, the remainder of the total number of strokes in the embedded sub-text data based on the first parameter is less than the second parameter, and the encoded character corresponding to the remainder of the total number of tone values ​​in the embedded sub-text data based on the third parameter is consistent with the first watermark encoding information.

[0013] This application also provides a text watermark extraction method, including: Obtain embedded text data; wherein, the embedded text data includes: multiple second sub-text data; The embedded text data is obtained by determining the first watermark encoding information corresponding to the first sub-text data in the watermark encoding information sequence during the embedding detection result representation, and adjusting the text information in the first sub-text data until the total number of strokes and the total number of tone values ​​of the first sub-text data meet the corresponding rules; the watermark encoding information sequence is formed based on the watermark information embedded in the text data; the embedding detection result is determined based on the total number of strokes corresponding to each first sub-text data. Based on the total number of strokes corresponding to each second sub-text data, an extraction detection result is determined; wherein, the extraction detection result is used to characterize whether there is a corresponding second watermark encoding information in the second sub-text data; If the extracted detection result representation exists, then the corresponding second watermark encoding information and second offset information are extracted based on the total number of tone values ​​of the second sub-text data. The watermark information is determined based on multiple sets of the second watermark encoding information and the second offset information extracted from the embedded text data.

[0014] In the above scheme, the step of extracting the corresponding second watermark encoding information and second offset information based on the total number of tone values ​​of the second sub-text data includes: The fourth remainder is taken based on the total number of tone values ​​of the second sub-text data according to the third parameter in the dilution parameter. The tone segment corresponding to the fourth remainder is determined, and the second watermark encoding information corresponding to the tone segment is determined; wherein, each tone segment is determined based on the number of segments determined by the fourth parameter in the dilution parameter; Determine the fourth encoded information, which is a combination of the preset key and the preset position text information in the second sub-text data; The sequence length is calculated by taking a fifth remainder based on the fourth encoding information, and the second offset information is determined based on the fifth remainder; wherein the sequence length is determined based on the number of the second watermark encoding information extracted from the embedded text data.

[0015] In the above scheme, determining the watermark information based on multiple sets of the second watermark encoding information and the second offset information extracted from the embedded text data includes: Based on multiple second offset information, multiple second watermark encoding information are statistically analyzed to determine the first intermediate watermark encoding information sequence; Using the fifth encoding information corresponding to the preset key, the first intermediate watermark encoding information sequence is decrypted to obtain the second intermediate watermark encoding information sequence; The watermark information is determined by restoring the second intermediate watermark encoding information sequence.

[0016] This application also provides a text watermark embedding device, including: The first data acquisition unit is used to acquire text data; wherein the text data includes: a plurality of first sub-text data; The first determining unit is used to determine the embedding detection result based on the total number of strokes corresponding to each of the first sub-text data; wherein the embedding detection result is used to characterize whether the corresponding first watermark encoding information is embedded in the first sub-text data; An embedding adjustment unit is configured to, if the embedding detection result indicates embedding, determine the first watermark encoding information corresponding to the first sub-text data in the watermark encoding information sequence, and adjust the text information in the first sub-text data until the total number of strokes and the total number of tone values ​​of the first sub-text data meet the corresponding rules, thereby obtaining embedded text data; wherein, the watermark encoding information sequence is formed based on the watermark information embedded in the text data.

[0017] This application also provides a text watermark extraction device, including: The second data acquisition unit is used to acquire embedded text data; wherein, the embedded text data includes: a plurality of second sub-text data; The embedded text data is obtained by determining the first watermark encoding information corresponding to the first sub-text data in the watermark encoding information sequence during the embedding detection result representation, and adjusting the text information in the first sub-text data until the total number of strokes and the total number of tone values ​​of the first sub-text data meet the corresponding rules; the watermark encoding information sequence is formed based on the watermark information embedded in the text data; the embedding detection result is determined based on the total number of strokes corresponding to each first sub-text data. The second determining unit is used to determine the extraction detection result based on the total number of strokes corresponding to each of the second sub-text data; wherein, the extraction detection result is used to characterize whether the second sub-text data has corresponding second watermark encoding information; The extraction unit is configured to, if the extraction detection result indicates extraction, extract the corresponding second watermark encoding information and second offset information based on the total number of tone values ​​of the second sub-text data. The restoration unit is used to determine the watermark information based on multiple sets of the second watermark encoding information and the second offset information extracted from the embedded text data.

[0018] This application also provides a first electronic device, including a first memory and a first processor. The first memory stores a computer program that can run on the first processor. When the first processor executes the computer program, it implements the steps in the text watermark embedding method.

[0019] This application also provides a second electronic device, including a second memory and a second processor. The second memory stores a computer program that can run on the second processor. When the second processor executes the computer program, it implements the steps in the text watermark extraction method.

[0020] This application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a first processor, implements the steps in the text watermark embedding method.

[0021] This application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a second processor, implements the steps in the text watermark extraction method.

[0022] This application also provides a computer program product, including a computer program that, when executed by a first processor, implements the steps in the text watermark embedding method.

[0023] This application also provides a computer program product, including a computer program that, when executed by a second processor, implements the steps in the text watermark extraction method.

[0024] In this embodiment, text data is acquired, including multiple first sub-text data. An embedding detection result is determined based on the total number of strokes corresponding to each first sub-text data. This embedding detection result indicates whether the corresponding first watermark encoding information is embedded in the first sub-text data. If the embedding detection result indicates embedding, the first watermark encoding information corresponding to the first sub-text data is determined in the watermark encoding information sequence, and the text information in the first sub-text data is adjusted until the total number of strokes and the total number of tone values ​​of the first sub-text data conform to the corresponding rules, thus obtaining embedded text data. The watermark encoding information sequence is formed based on the watermark information embedded in the text data. This method, by counting the total number of strokes of each first sub-text data to determine whether the first watermark encoding information is embedded, avoids indiscriminate modification of all content, reduces the impact on the original text, and minimizes embedding traces. Furthermore, selecting specific watermark encoding information from the watermark encoding information sequence for embedding improves the flexibility and security of watermark embedding. Simultaneously, during the adjustment process, the matching of the total number of strokes and the total number of tone values ​​is considered, making the embedded text more grammatically and semantically natural, thus improving the quality of the watermarked text. Attached Figure Description

[0025] Figure 1 A flowchart illustrating the text watermark embedding method provided in this application embodiment. Figure 1 ; Figure 2 A flowchart illustrating the text watermark embedding method provided in this application embodiment. Figure 2 ; Figure 3 A flowchart illustrating the text watermark embedding method provided in this application embodiment. Figure 3 ; Figure 4 A flowchart illustrating the text watermark embedding method provided in this application embodiment. Figure 4 ; Figure 5 A flowchart illustrating the text watermark extraction method provided in this application embodiment. Figure 5 ; Figure 6 A flowchart illustrating the text watermark extraction method provided in this application embodiment. Figure 6 ; Figure 7 An interactive schematic diagram of the text watermark embedding method provided in the embodiments of this application; Figure 8 This is a schematic diagram of the structure of the text watermark embedding device provided in the embodiments of this application; Figure 9 A schematic diagram of a hardware entity of the first electronic device provided in an embodiment of this application; Figure 10 This is a schematic diagram of the structure of the text watermark extraction device provided in the embodiments of this application; Figure 11 This is a schematic diagram of a hardware entity of a second electronic device provided in an embodiment of this application.

[0026] It should be noted that the terms "first" and "second" mentioned above are only used to distinguish between different options and do not represent the degree of superiority or inferiority of the options or their priority in the implementation process. Detailed Implementation

[0027] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application are further described in detail below with reference to the accompanying drawings and embodiments. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0028] In the following description, terms such as first sub-text data, total number of strokes, total number of tone values, embedding detection result, watermark encoding information sequence, embedding amount parameter, and dilution parameter are all core concepts in this application for achieving text watermark embedding and extraction, and need to be defined to ensure consistent understanding. The specific explanations of these terms are as follows: 1) First sub-text data: refers to each sentence or paragraph in the text data that has been divided into independent processing units. In the embodiments of this application, each first sub-text data is used as a basic unit to determine whether a watermark is embedded, and serves as the basis for subsequent adjustments.

[0029] 2) Total number of strokes: This refers to the sum of the number of strokes of all Chinese characters in a first sub-text data. The number of strokes is determined according to the standard Chinese character writing specifications and is often used to measure the complexity of characters. In this invention, it is used to determine whether watermark information needs to be embedded.

[0030] 3) Total Tone Values: This refers to the sum of the tone values ​​(i.e., the four tones of Mandarin) of all Chinese characters in a first sub-text data. The first tone is 1, the second tone is 2, the third tone is 3, and the fourth tone is 4. In this invention, tone values ​​are used to assist in determining the mapping relationship of the watermark encoding and to control the semantic impact during the embedding process.

[0031] 4) Embedding Detection Result: This result is calculated based on the total number of strokes in the first sub-text data and is used to determine whether the first sub-text data is suitable for embedding watermark information. If the total number of strokes meets specific conditions, it is determined that it can be embedded.

[0032] 5) Watermark Encoded Information Sequence: This is a set of encoded data generated from user-input watermark information after encoding and encryption. This sequence is used to represent the watermark content and is mapped to the target text according to specific rules.

[0033] 6) Embedding Density Parameters: These include a first parameter and a second parameter, which control the embedding density. The first parameter is used to take the remainder after dividing by the total number of strokes, and the second parameter is the maximum allowed remainder value. The combination of these two parameters determines which first sub-text data will be selected for watermark embedding.

[0034] 7) Dilution parameter: This includes the third and fourth parameters, which are used for modulo operation on the total number of tone values ​​and segmented encoding, respectively. This parameter determines the encoding base of the watermark information, thus affecting the watermark capacity and concealment.

[0035] 8) Key: A cryptographic parameter used to encrypt and decrypt watermark information. In this invention, the key participates in the generation of watermark encoded information and is used to locate the offset position of the watermark encoded information in the sequence, thereby improving security.

[0036] 9) Adjust the detection result: This is the result of judging the sum of tone values ​​of the first sub-text data before embedding the watermark. It is used to confirm whether the text needs to be replaced or inserted to meet the watermark encoding requirements.

[0037] 10) Embedded text data: This is the text data after the watermark has been embedded. The total number of strokes and the total number of tone values ​​have been adjusted according to the watermark encoding information, so that the watermark information can be implicitly embedded without significantly changing the original text content.

[0038] The text watermark embedding methods provided in the embodiments of this application can be executed by a computing device, which can be a server, a terminal, or a cloud platform. That is, the text watermark embedding methods in the embodiments of this application can be executed by a server, by a terminal device, or by interaction between a cloud platform and a terminal device to complete the execution.

[0039] This application provides a text watermark embedding method. Please refer to [link to relevant documentation]. Figure 1 The following is a flowchart illustrating the text watermark embedding method provided in this application embodiment. Figure 1 , will combine Figure 1 The steps shown are explained below: S101. Obtain text data; wherein, the text data includes: multiple first sub-text data.

[0040] In this embodiment, text data typically refers to the original input content, such as articles, news articles, technical text data, etc. First sub-text data refers to each sentence or paragraph segmented from the text data as an independent processing unit. Each first sub-text data unit serves as a basic unit, used to determine whether a watermark should be embedded, and as the basis for subsequent adjustments.

[0041] In this embodiment of the application, a first sub-text data can be a sentence in the text data. For example, the first sub-text data may include "Digital watermarking technology is a technology that embeds specific identification information into digital media through an algorithm, which can ensure that the quality of the carrier is not significantly affected, and can also realize traceability verification."

[0042] S102. Based on the total number of strokes corresponding to each of the first sub-text data, determine the embedding detection result; wherein, the embedding detection result is used to characterize whether the corresponding first watermark encoding information is embedded in the first sub-text data.

[0043] In this embodiment, the total number of strokes refers to the sum of the stroke counts of all Chinese characters in a first sub-text data. The embedding detection result is calculated based on the total number of strokes of the first sub-text data and is used to determine whether the first sub-text data is suitable for embedding watermark information. If the total number of strokes meets a specific condition (such as the remainder being less than a preset threshold), it is determined to be embeddable. Specifically, the embedding detection result depends on the embedding amount parameter, which determines which first sub-text data will be selected for embedding the corresponding first watermark encoding information.

[0044] In this embodiment, the embedding parameters include a first parameter and a second parameter, which are used to control the embedding density. The first parameter is used to take the remainder after dividing by the total number of strokes, and the second parameter is the maximum allowed remainder value. The combination of the two determines which first sub-text data will be selected for watermark embedding.

[0045] In this embodiment, the embedding detection result is used as the criterion for determining whether watermark encoding information is embedded in a certain first sub-text data. For example, if the remainder of the total number of strokes is less than a set threshold, it is determined that embedding is possible; otherwise, it is not possible. The embedding detection result ensures that the embedding of the watermark is controllable and random, thereby improving its concealment.

[0046] S103. If the embedding detection result indicates embedding, then the first watermark encoding information corresponding to the first sub-text data is determined in the watermark encoding information sequence, and the text information in the first sub-text data is adjusted until the total number of strokes and the total number of tone values ​​of the first sub-text data meet the corresponding rules, thereby obtaining embedded text data; wherein, the watermark encoding information sequence is formed based on the watermark information embedded in the text data.

[0047] In this embodiment, if the embedding detection result indicates that corresponding first watermark encoding information needs to be embedded in the first sub-text data, then the first watermark encoding information corresponding to the first sub-text data needs to be determined in the watermark information encoding sequence. Specifically, the offset in the watermark encoding information sequence can be determined based on the text information at a predetermined position of the first sub-text data, and then the corresponding first watermark encoding information can be determined in the watermark encoding information sequence based on the offset. The watermark encoding information sequence is a set of encoded data generated by encrypting and encoding the watermark information input by the user. The watermark encoding information sequence is used to represent the watermark information and is mapped to the text data according to specific rules. The watermark information refers to the identification information that is to be embedded in the text data for purposes such as marking copyright ownership and verifying source.

[0048] The first watermark encoding information refers to the specific encoding selected from the watermark encoding information sequence and used to embed it into the first sub-text data. For example, the encoding 1 with an offset value of 2 in the watermark encoding sequence can be used as the first watermark encoding information of the first sub-text data.

[0049] In this embodiment, after determining the first watermark encoding information corresponding to the first sub-text data, it is further possible to determine whether the text number information in the first sub-text data needs to be adjusted based on whether the total number of tone values ​​of the first sub-text data meets the corresponding rules. If the corresponding rules are not met, the text information of the first sub-text data is adjusted. Adjusting the text information of the first sub-text data involves segmenting the sentence and performing synonym retrieval, replacing or inserting words that conform to the stroke and tone rules. Specifically, words that conform to the stroke and tone rules refer to sentences where the total number of strokes after replacement still meets the rule requirements for the total number of strokes, and the total number of tone values ​​also meets the corresponding rule requirements. Thus, after all the first sub-text data of the text data is embedded, the embedded text data is obtained.

[0050] The total tone value refers to the sum of the tone values ​​(i.e., the four tones of Mandarin) of all Chinese characters in a first sub-text data. The first tone is 1, the second tone is 2, the third tone is 3, and the fourth tone is 4. In this invention, tone values ​​are used to assist in determining the mapping relationship of the watermark encoding and to control the semantic impact during the embedding process.

[0051] In this embodiment of the application, the rule requirement for meeting the total number of strokes may include: the remainder after taking the total number of strokes is less than a predetermined threshold. The rule requirement for meeting the total number of tone values ​​may also include: the remainder after taking the total number of tone values ​​is less than a predetermined threshold.

[0052] In this embodiment, text data is acquired, including multiple first sub-text data. An embedding detection result is determined based on the total number of strokes corresponding to each first sub-text data. This embedding detection result indicates whether the corresponding first watermark encoding information is embedded in the first sub-text data. If the embedding detection result indicates embedding, the first watermark encoding information corresponding to the first sub-text data is determined in the watermark encoding information sequence, and the text information in the first sub-text data is adjusted until the total number of strokes and the total number of tone values ​​of the first sub-text data conform to the corresponding rules, thus obtaining embedded text data. The watermark encoding information sequence is formed based on the watermark information embedded in the text data. This method, by counting the total number of strokes of each first sub-text data to determine whether the first watermark encoding information is embedded, avoids indiscriminate modification of all content, reduces the impact on the original text, and minimizes embedding traces. Furthermore, selecting specific watermark encoding information from the watermark encoding information sequence for embedding improves the flexibility and security of watermark embedding. Simultaneously, during the adjustment process, the matching of the total number of strokes and the total number of tone values ​​is considered, making the embedded text more grammatically and semantically natural, thus improving the quality of the watermarked text.

[0053] In this embodiment of the application, step S201 may also be included, which will be described in conjunction with the steps: S201. Obtain watermark information, embedding amount parameter, and dilution parameter; wherein, the embedding amount parameter is used to determine the proportion of the first sub-text data in which the first watermark encoding information is embedded in multiple first sub-text data; the dilution parameter is used to determine the encoding range of the watermark information.

[0054] In this embodiment, watermark information refers to identifying content that is intended to be embedded in the target text, such as copyright information, source markers, or specific strings. Watermark information can be in plaintext or an encrypted or transformed binary sequence, depending on the embedding strategy and security requirements. The purpose of embedding watermark information is to achieve functions such as tracking the source of the text, copyright protection, or anti-counterfeiting verification.

[0055] In this embodiment of the application, the embedding parameters include: a first parameter and a second parameter; wherein, the first parameter is used to take the remainder of the total number of strokes of the first sub-text data; and the second parameter is a preset maximum remainder value.

[0056] The embedding parameters (M, N) refer to the remainder value M (first parameter) of the sum of stroke counts for each sentence and the maximum remainder N (second parameter). The embedding parameters determine the proportion of sentences with a 2N / M ratio that will have their first watermark encoded information embedded; that is, what percentage of the original text will be modified. A larger M and a smaller N result in a smaller impact on the text data. The number of strokes in Chinese characters is generally between 1 and 30, with an average of about 11 strokes. In a specific embodiment, the embedding parameters (M, N) can be (200, 40).

[0057] In this embodiment of the application, the dilution parameter includes a third parameter and a fourth parameter; wherein, the third parameter is used to take the remainder of the total number of tone values ​​of the first sub-text data; and the fourth parameter is used to determine the number of segments of the tone values.

[0058] The dilution parameters (P, Q) refer to the remainder of the sum of tone counts for each sentence (P, the third parameter) and the number of interval segments (Q, the fourth parameter). The dilution parameters determine the encoding range of the watermark information and the base number to which it needs to be encoded, thus affecting the degree of dilution. P must be a multiple of Q, and Q determines the base that can be encoded. In a specific embodiment, the dilution parameters (P, Q) can be (4, 2), and the watermark information will be encoded as a binary value.

[0059] In this embodiment, the embedding amount parameter and dilution parameter can be obtained manually. There is a certain correlation between the embedding amount parameter and the dilution parameter. The embedding amount parameter mainly controls the breadth of watermark information embedding, i.e., the frequency of watermark information appearing in the text data; the dilution parameter controls the depth of watermark embedding, i.e., the complexity and precision of each embedded watermark information. Therefore, when designing a watermark embedding system, designers need to comprehensively consider the settings of the embedding amount parameter and the dilution parameter to achieve the best embedding effect and concealment.

[0060] In this embodiment, by introducing embedding amount and dilution parameters, the application can flexibly control embedding density and encoding precision while ensuring effective embedding of watermark information. The parameters can be adjusted in different application scenarios to achieve precise control and optimized embedding effect of watermark information, thereby improving the concealment, invisibility, and robustness of the watermark, while reducing interference with the original text content.

[0061] In this embodiment of the application, the scheme for determining the watermark encoding information sequence may include S202 to S204, which will be described in conjunction with the steps: S202, The watermark information is converted into first encoded information in the encoding base characterized by the dilution parameter; wherein the encoding base is determined based on the number of segments determined by the fourth parameter.

[0062] In this embodiment, the encoding base is determined by the value of the fourth parameter in the dilution parameter. For example, if the fourth parameter is 2, the encoding base is binary; if the fourth parameter is 4, the encoding base is quaternary. The encoding base determines how the watermark information is converted into digital form for processing during the embedding process. By using different encoding bases, the capacity and embedding density of the watermark information can be flexibly adjusted.

[0063] In this embodiment, the first encoded information refers to the digital sequence after converting the original watermark information according to a specific encoding base. For example, if the watermark information is nine days, the UTF8 encoding of the watermark information is a string of binary bits, and the UTF8 encoding of the watermark information is organized into a form suitable for subsequent encryption processing according to the encoding base (such as binary).

[0064] S203. Convert the preset key into second encoded information in the encoded base.

[0065] In this embodiment of the application, the preset key can be converted into second encoded information in the same base as the first encoded information.

[0066] The preset key refers to the encryption key used in the watermark embedding process, which is usually preset by the user or the system. This key is used to enhance the security of the watermark information and prevent unauthorized access or tampering. The key can be a string, a number, or other form of data, depending on the requirements of the encryption algorithm.

[0067] In this embodiment, the second encoded information refers to the digital sequence obtained by converting the preset key according to the same encoding base. For example, if the key is cmri, the UTF8 encoding of the key is also a stream of binary bits. The binary bit stream is organized into the same data format as the first encoded information according to the encoding base (such as binary). This organization method ensures the compatibility between the key and the watermark information and provides a basis for subsequent encryption operations.

[0068] S204. Encrypt the first encoded information based on the second encoded information to determine the watermark encoded information sequence. In this embodiment of the application, the first encoding information is encrypted based on the second encoding information, and the watermark encoding information sequence is determined according to the encryption result.

[0069] In the embodiments of the present application, generating a watermark coding information sequence based on the watermark information and the key means performing radix coding processing on the watermark information according to the value of the interval segmentation number Q of the dilution influence factor. Then, a specific transformation is performed on the watermark information and the key. The specific transformation is not limited to exclusive OR operation, left or right shift of specified bits, cryptographic algorithm encryption, etc. The transformation method is not the content concerned in this solution. For example, if the watermark information is "Jiutian", the dilution parameter is (4, 2), and the watermark coding method is UTF8-based binary coding, it is converted into the binary bit stream "111001001011100110011101111001011010010010101001". Assuming the key is "cmri", it is converted into the binary coding "1100011110110111100101101001", and using the exclusive OR operation as the transformation method, it performs an ordered exclusive OR operation on the watermark coding information in a cyclic manner, and the generated watermark coding information sequence is "001000110000111000001011011110011101111111010000".

[0070] Among them, encryption refers to the process of transforming the first coding information using the second coding information. Common encryption methods include exclusive OR operation, displacement operation, cryptographic algorithms (such as AES), etc. The purpose of encryption is to increase the unpredictability and anti-attack ability of the watermark information, so that the watermark information is difficult to be detected or extracted. For example, in this embodiment, the exclusive OR operation is used to perform an ordered exclusive OR of the second coding information bit by bit with the first coding information to generate the final watermark coding information sequence.

[0071] Among them, the watermark coding information sequence refers to the digital representation of the final watermark information formed after encryption processing. The watermark coding information sequence contains all necessary watermark information and has been encrypted by the key, with high security and concealment. In the subsequent text embedding process, the watermark coding information sequence will be used to replace or insert words that conform to the stroke and tone rules to complete the watermark embedding operation.

[0072] In the embodiments of the present application, by combining the watermark information with a preset key and performing encryption processing on the watermark information using the coding radix based on the dilution parameter, a watermark coding information sequence is generated. The above encryption method can improve the concealment and security of the watermark information, thus effectively preventing unauthorized watermark extraction or tampering behavior, and further enhancing the overall robustness and practicality of the Chinese text watermark.

[0073] Please refer to Figure 2 , which is the flow schematic of the text watermark embedding method provided by the embodiments of the present application Figure 2 , Figure 1 shown in S102 to S103 can also be implemented by S301 to S305, which will be combined with Figure 2 The steps shown are described as follows: S301. Take a first remainder of the total number of strokes of the first sub - text data based on the first parameter.

[0074] In the embodiments of the present application, by taking the remainder of the total number of strokes Tb with respect to the first parameter M, a first remainder can be obtained. According to the value of the first remainder, it is determined whether the first sub - text data is selected for embedding the watermark. If the remainder falls within a certain specific interval or is less than a certain threshold, it indicates that the first sub - text data is suitable for embedding watermark information.

[0075] S302. Determine the embedding detection result based on the comparison between the first remainder and the second parameter.

[0076] In the embodiments of the present application, if the first remainder is less than the second parameter N, it indicates that the corresponding first sub - text data is suitable for embedding watermark information; otherwise, no embedding is performed. The judgment mechanism based on the comparison between the first remainder and the second parameter N ensures that only part of the first sub - text data will be modified, thus reducing the impact on the quality of the original text.

[0077] Exemplarily, for example, the content of this sentence is "Digital watermarking technology is a technology that embeds specific identification information into digital media through algorithms, which can not only ensure that the quality of the carrier is not significantly affected but also achieve traceability verification.", and the number of strokes of each character is as follows: number

[13] , character [6], water [4], mark [5], technology [7], art [5], is [9], one [1], kind [9], will [9], special

[10] , definite [8], standard [9], identification [7], information

[10] , through

[10] , by [6], algorithm

[14] , embed

[12] , into [2], to [8], number

[13] , character [6], medium

[12] , body [7], in [4], of [8], technology [7], art [5], both [9], can

[10] , guarantee [9], carrier

[10] , body [7], quality [8], quantity

[12] , not [4], affected [8], obvious [8], impact

[15] , influence [9], also [2], can [5], realize [8], trace

[13] , source

[13] , verify

[10] , prove [7]. By calculation, the total number of strokes Tb is 436. The embedding amount parameter is (200, 40). The first remainder obtained by taking the remainder of 436 with respect to 200 is 36, which is less than 40. Therefore, watermark information needs to be embedded.

[0078] In this embodiment, by taking the remainder of the total number of strokes based on the first parameter, the system can flexibly control the watermark embedding ratio, thereby reducing the impact of the watermark embedding ratio set by this operation on the original text and improving the concealment of the watermark embedding according to this ratio. By comparing the first remainder with the second parameter, the first sub-text data suitable for watermark embedding can be accurately selected. This operation can avoid unnecessary modifications to the text. This operation improves the naturalness and readability of the text while ensuring the watermark capacity.

[0079] S303. Determine the third encoded information after combining the preset key and the preset position text information in the first sub-text data.

[0080] In this embodiment, the preset key refers to a set of encrypted strings pre-defined by the user for generating watermark encoding information. It is typically associated with an embedding algorithm to ensure the security of the watermark embedding process. The preset position text information refers to characters selected at specific positions in the first sub-text data according to fixed rules, such as the first character of a sentence or Chinese characters at other predefined positions. After combining the preset key and the preset position text information, a unique identifier (third encoding information) is generated using a predetermined algorithm.

[0081] S304. Take the second remainder of the sequence length of the watermark encoding information sequence based on the third encoding information.

[0082] In this embodiment, the second remainder refers to the remainder obtained by dividing the third encoded information into a numerical value and then dividing that numerical value by the total length of the watermark encoded information sequence. This remainder is used to determine which element in the watermark encoded information sequence to select as the offset, so as to locate the watermark encoded information in subsequent steps. For example, in the example mentioned in the technical disclosure, the last two digits of the MD5 value generated by combining the preset key with the first character of the first sub-text data are "8C". After converting "8C" into a decimal value, the remainder is calculated by dividing the value by the length of the watermark encoded information sequence (e.g., 48), and finally the offset value is calculated to be 2.

[0083] S305. The first watermark encoding information is determined in the watermark encoding information sequence based on the first offset information represented by the second remainder.

[0084] In this embodiment, the first offset information is an index value represented by the second remainder, used to locate specific watermark encoding information in the watermark encoding information sequence. The first watermark encoding information is the watermark data actually embedded in the first sub-text data. The first watermark encoding information may be a binary bit, a set of specific characters, or other forms of information. In specific implementation, the corresponding data is extracted from the watermark encoding information sequence according to the first offset information and embedded into the first sub-text data. For example, taking the code '1' of the offset value 2 in the watermark encoding information sequence means that the data at the second position in the watermark encoding information sequence is 1, and the code '1' of the offset value 2 will be used as the first watermark encoding information actually embedded in the first sub-text data.

[0085] In this embodiment of the application, by using the method of selecting the first watermark encoding information based on the dynamically generated first offset information, the system can accurately select the first watermark encoding information according to the dynamically generated first offset information, thereby avoiding repeated embedding or omission of the first watermark encoding information, thereby improving the accuracy and stability of the watermark information, and further enhancing the robustness and anti-attack capability of the watermark.

[0086] In this embodiment, a third encoding information is generated by combining a preset key with text information at a preset position in the first sub-text data. Then, based on this encoding information, a second remainder is taken from the length of the watermark encoding information sequence to determine the first offset information, from which the first watermark encoding information is extracted. This allows for dynamic selection and positioning of the watermark encoding information, thereby improving the security and flexibility of watermark embedding and effectively preventing unauthorized modification and forgery.

[0087] In this embodiment of the application, steps S401 to S403 may be included before adjusting the text information in the first sub-text data, which will be described in conjunction with the steps: S401. Take the third remainder of the total number of tone values ​​of the first sub-text data based on the third parameter.

[0088] In this embodiment, the total tone value refers to the sum of the tone values ​​of all Chinese characters in the first sub-text data. The tone values ​​are defined according to the four tones of Mandarin Chinese: the first tone is 1, the second tone is 2, the third tone is 3, and the fourth tone is 4. By counting the total tone values, the overall speech characteristics of the first sub-text data can be reflected, and it can serve as one of the basic data for watermark embedding.

[0089] In this embodiment, by using a third parameter modulo the total number of tone values, continuous tone features can be discretized, facilitating subsequent encoding and comparison. This modulo operation helps achieve a uniform distribution of watermark encoding while reducing the impact on the semantics of the original text.

[0090] S402. Determine the tone segment corresponding to the third remainder, and the encoded character corresponding to the tone segment.

[0091] In this embodiment, tone segmentation is based on a fourth parameter to divide the tone into several segments, each segment corresponding to an encoded character. For example, when the fourth parameter is 2, the tone segmentation can be set to two segments: [0, 1] and [2, 3], corresponding to 0 and 1 in binary encoding, respectively. Then, the encoded character corresponding to that segment can be determined based on which segment the third remainder belongs to.

[0092] There is a one-to-one correspondence between tone segments and encoded characters. Each tone segment represents a specific encoded character, thus forming a mapping mechanism. In practice, the mapping relationship between tone segments and encoded characters can be flexibly configured according to the specific watermarking encoding strategy. For example, the system can dynamically adjust the segment boundaries based on the output of the encryption algorithm, thereby enhancing the system's security and adaptability.

[0093] S403. Based on the consistency between the first watermark encoding information and the encoded character, determine to adjust the detection result.

[0094] In this embodiment, consistency refers to whether the first watermark encoding information is the same as the encoded character corresponding to the current first sub-text data. If the first watermark encoding information is the same as the encoded character corresponding to the current first sub-text data, the adjustment detection result indicates that no further adjustment is needed; if the first watermark encoding information is inconsistent with the encoded character corresponding to the current first sub-text data, the adjustment detection result indicates that words in the first sub-text data need to be replaced or inserted so that the total number of tone values ​​meets the requirements of the watermark encoding. The adjustment detection result is used to characterize whether to adjust the text information in the first sub-text data until the total number of strokes and the total number of tone values ​​of the first sub-text data meet the corresponding rules.

[0095] There is a crucial data logic relationship between the first watermark encoding information and the encoded character. The first watermark encoding information determines the expected encoded character, which is determined by the total number of tone values ​​and tone segments. The comparison result between the first watermark encoding information and the encoded character directly affects the formulation of the watermark embedding strategy; this comparison result is the core judgment basis in the watermark embedding process.

[0096] In this embodiment, a third parameter is introduced and modulo the total number of tone values ​​to divide the text into tone segments and map them to encoded characters. Then, a consistency judgment is performed based on the first watermark encoding information to determine whether the text information in the first sub-text data needs adjustment. This method effectively controls the proportion and precision of the watermark embedding, thereby reducing the impact on the semantics of the original text and improving the concealment and robustness of the watermark.

[0097] Please see Figure 3 The following is a flowchart illustrating the text watermark embedding method provided in this application embodiment. Figure 3 , Figure 1 The shown S103 can also be implemented via S501, combining Figure 3 The steps shown are explained below: S501. Replace or insert text information in the first sub-text data to obtain embedded sub-text data.

[0098] In this embodiment, if adjusting the detection result representation requires adjusting the text information in the first sub-text data, then word segmentation and synonym retrieval are performed on the first sub-text data. Words that conform to the stroke and tone rules are replaced or inserted to obtain the embedded sub-text data after embedding the first sub-text data. Specifically, the remainder of the total number of strokes in the embedded sub-text data based on the first parameter is less than the second parameter, and the encoded character corresponding to the remainder of the total number of tone values ​​in the embedded sub-text data based on the third parameter is consistent with the first watermark encoding information.

[0099] In this embodiment, if the total number of strokes in the embedded sub-text data modulo M is less than N, the embedding is considered appropriate. Furthermore, the sum of tone values ​​in the embedded sub-text data is calculated and divided into several intervals based on the dilution factor (P, Q), with each interval corresponding to an encoded character. If the sum of tone values ​​modulo P falls within a specific interval, the encoded character corresponding to that interval is the same as the current first watermark encoding information, indicating that the embedding is appropriate.

[0100] Exemplarily, the total number Ts of the tone values of the first sub-text data is 193, and the dilution parameter is (4, 2). 4 is divided into two intervals of [0-1] and [2-3], corresponding to binary 0 and 1 respectively. The remainder of Ts divided by 4 is 1, which falls in the interval of [0-1], and the corresponding code is 0, which is different from the first watermark coding information "1" to be embedded. Therefore, synonym replacement is required. The replacement compliance stroke rule is that the number of strokes is not reduced by more than 36 and not increased by more than 4 (the remainder of the total number of strokes 436 divided by 200 is 36, and the maximum value of the remainder is 40). The replacement compliance tone rule is that tones 1, 2, and 5 need to be increased, and tones 2, 3, and 6 need to be reduced, that is, the result of taking the remainder of 4 is converted from the interval [0-1] to the interval [2-3]. The words "bao[9], zheng[7]" in the sentence can be replaced with "shi[8], de

[11] ", and the number of increased strokes is 3, which complies with the stroke rule. At the same time, the tones are replaced from the third tone and the fourth tone to the third tone and the second tone, reducing the tone value by 2, which complies with the tone rule.

[0101] In the embodiments of the present application, by replacing or inserting text information in the first sub-text data, an embedded sub-text data is obtained, and based on the first parameter, the remainder of the total number of strokes is less than the second parameter, and based on the third parameter, the encoded character corresponding to the remainder of the total number of tone values is consistent with the first watermark coding information. In this way, the embedding position and quantity can be flexibly controlled, and this embedding method that does not change the content of the text information can improve the concealment and robustness of watermark embedding, and achieve high-quality Chinese text watermark embedding and extraction.

[0102] Please refer to Figure 4 , which is a schematic flow chart of the text watermark embedding method provided by the embodiments of the present application Figure 4 , and will be described in combination with Figure 4 the steps shown below: S11. Configure the influencing factors of the embedding algorithm, and generate a watermark coding information sequence together with the key.

[0103] In the embodiments of the present application, the configured influencing factors of the embedding algorithm include an embedding amount influencing factor and a dilution influencing factor. Generating a watermark coding information sequence together with the key by the watermark information means performing a base encoding process on the watermark information according to the value of the number of interval segments Q of the dilution influencing factor. Then, specific transformations are performed on the watermark information and the key to generate a watermark coding information sequence. The specific transformations are not limited to exclusive OR operations, left and right shifts of specified bits, cryptographic algorithm encryption, etc. The transformation method is not the focus of this solution.

[0104] S12. For each sentence in the text content, count the total number of strokes, determine whether to embed and the embedded watermark encoding information; if information needs to be embedded, determine the watermark encoding information to be embedded in the sentence, perform word segmentation and synonym retrieval on the sentence, and replace or insert words that meet the stroke and tone rules.

[0105] In this embodiment, the total number of strokes is calculated by adding the strokes of all Chinese characters in the sentence to obtain the total number of strokes, Tb. The process of determining whether to embed watermark encoding information and the embedded watermark encoding information is performed. Specifically, based on the embedding influence factor (M, N), if Tb modulo M is less than N, the subsequent steps continue; otherwise, the process for that sentence ends without further operation. The watermark encoding information to be embedded in the sentence is determined by performing a specific transformation on the key and the first character of the sentence, then taking the remainder of the watermark encoding information sequence length to obtain the sequence offset value, from which the watermark encoding information to be embedded is obtained. The sum of the tone values ​​Ts of all Chinese characters in the sentence is obtained, and the dilution influence factor (P, Q) is used to determine whether synonym replacement is needed. The tone value is defined as 1 for the first tone, 2 for the second tone, 3 for the third tone, and 4 for the fourth tone. The determination method is to check whether the remainder of Ts modulo P falls within the corresponding interval. If it does not fall within the corresponding interval, the sentence is segmented and synonyms are searched, and words that conform to the stroke and tone rules are replaced or inserted. Words that conform to the rules of stroke count and tone count specifically refer to sentences whose total number of strokes and total tone count still conform to the rules after replacement.

[0106] The text watermark extraction methods provided in the embodiments of this application can be executed by a computing device, which can be a server, a terminal, or a cloud platform. That is, the text watermark extraction methods in the embodiments of this application can be executed by a server, by a terminal device, or by interaction between a cloud platform and a terminal device.

[0107] Please see Figure 5 The above is a flowchart illustrating the text watermark extraction method provided in this application embodiment. Figure 1 , will combine Figure 5 The steps shown are explained below: S601. Obtain embedded text data; wherein, the embedded text data includes: multiple second sub-text data.

[0108] In this embodiment of the application, embedded text data refers to text content that has been embedded with watermark information, which is usually composed of several sentences, each of which is called a second sub-text data.

[0109] The embedded text data is obtained by determining the first watermark encoding information corresponding to the first sub-text data in the watermark encoding information sequence during the embedding detection result representation, and adjusting the text information in the first sub-text data until the total number of strokes and the total number of tone values ​​of the first sub-text data meet the corresponding rules; the watermark encoding information sequence is formed based on the watermark information embedded in the text data; and the embedding detection result is determined based on the total number of strokes corresponding to each first sub-text data.

[0110] S602. Based on the total number of strokes corresponding to each of the second sub-text data, determine the extraction detection result; wherein, the extraction detection result is used to characterize whether there is a corresponding second watermark encoding information in the second sub-text data.

[0111] In this embodiment, the total number of strokes refers to the sum of the total number of strokes of all Chinese characters in a certain second sub-text data. For the second sub-text data, the total number of strokes Tb is obtained by adding the strokes of all Chinese characters in the second sub-text data. The remainder of the total number of strokes Tb is then taken, and the extraction detection result is determined based on the relationship between the remainder and the second parameter. If the remainder is less than the second parameter, the detection result indicates that the corresponding second watermark encoding information exists in the second sub-text data. If the remainder is not less than the second parameter, the detection result indicates that the corresponding second watermark encoding information does not exist in the second sub-text data.

[0112] For example, for each sentence in the text content, the total number of strokes of all Chinese characters in the sentence is added together to obtain the total number of strokes Tb. b) Determine whether there is watermark encoding information. Specifically, based on the embedding influence factor (M, N), if Tb modulo M is less than N, then continue the steps; otherwise, the process for that sentence ends and no further operations are performed.

[0113] S603. If the extracted detection result representation exists, then extract the corresponding second watermark encoding information and second offset information based on the total number of tone values ​​of the second sub-text data.

[0114] In this embodiment, the sum of tone values ​​is divided into intervals according to a dilution parameter to determine the corresponding binary second watermark encoding information (such as 0 or 1). Simultaneously, an offset is generated by combining the key with text information at a predetermined position in the second sub-text data, used to locate the second offset information in the watermark encoding sequence.

[0115] S604. Based on the multiple sets of second watermark encoding information and second offset information extracted from the embedded text data, determine the watermark information.

[0116] In this embodiment, the watermark information is the final extracted identification information, which may be a string, a number sequence, or other identifiable information. The watermark information is obtained by arranging multiple sets of second watermark encoded information in the order of the second offset information, and then decoding and restoring this information in combination with a key.

[0117] In this embodiment, the determination of the watermark information relies on a statistical filtering and restoration algorithm for the second watermark encoding information. By statistically counting the second watermark encoding information extracted from multiple second sub-text data, using methods such as election to determine the most likely encoding sequence, and combining this with a key to perform an inverse transformation on the second watermark encoding information, the original watermark information can be restored.

[0118] In this embodiment, embedded text data is acquired and divided into multiple second sub-text data. Then, the presence of a watermark is determined based on the total number of strokes. Next, the second watermark encoding information and second offset information are extracted by combining the total number of tone values. Finally, the watermark information is restored through statistical filtering and using a key. This process enables efficient extraction and restoration of watermark information, enhancing the robustness and concealment of the watermark. This enhancement meets the technical requirements for tracing and verifying the source of generated content.

[0119] In this embodiment of the application, S603 shown can also be implemented by S6031 to S6035, which will be described in conjunction with the steps: S6031. Take the fourth remainder of the total number of tone values ​​of the second sub-text data based on the third parameter in the dilution parameter.

[0120] In this embodiment of the application, the fourth remainder can be taken from the total number of tone values ​​of the second sub-text data according to the third parameter in the dilution parameter.

[0121] S6032. Determine the tone segment corresponding to the fourth remainder, and determine the second watermark encoding information corresponding to the tone segment; wherein, each tone segment is determined based on the number of segments determined by the fourth parameter in the dilution parameter.

[0122] In this embodiment, the sum of tone values ​​Ts of all Chinese characters in the second sub-text data is obtained. The tone values ​​are 1 for the first tone, 2 for the second tone, 3 for the third tone, and 4 for the fourth tone. The extraction method is to take the remainder of Ts with respect to P to obtain the fourth remainder, determine the tone segment corresponding to the fourth remainder, and the second watermark encoding information corresponding to the segment.

[0123] The tone segmentation is a process of dividing the tone into several segments based on the fourth parameter, with each segment corresponding to a coded character. For example, when the fourth parameter is 2, the tone segmentation can be set to two segments: [0, 1] and [2, 3], which correspond to 0 and 1 in binary encoding, respectively.

[0124] For example, the preset key is "cmri", the embedding parameter is (200, 40), the dilution parameter is (4, 2), and the partitions [0-1] and [2-3] correspond to binary 0 and 1 respectively. The sentence "Digital watermarking technology is a technology that embeds specific identification information into digital media through an algorithm, which can ensure that the quality of the carrier is not significantly affected, and can also achieve traceability verification." The total number of strokes Tb for all Chinese characters is 439. The remainder of 439 divided by 200 is 39, indicating the presence of watermark encoding information. Therefore, the sum of the tone values ​​Ts of all Chinese characters in the sentence is 191, and the remainder when divided by 4 is 3, falling within the interval [2-3]. The extracted second watermark encoding information is 1.

[0125] S6033. Determine the fourth encoded information after combining the preset key and the preset position text information in the second sub-text data; In this embodiment, the preset key is a set of fixed characters or strings used to encrypt or obfuscate watermark information. It is typically set by the user or system administrator and remains unchanged throughout the watermark embedding and extraction process. In this embodiment, the preset key is combined with preset position text information in the second sub-text data to generate fourth encoded information. This combination method can employ various encryption algorithms, such as XOR, bitwise shift, and hashing, to ensure that the generated fourth encoded information has high randomness and unpredictability.

[0126] The preset position text information refers to the characters or strings at a specified position in the second sub-text data. This preset position text information can be the first or last character of a sentence, or a word at a fixed position. By combining the preset key with the preset position text information, a unique fourth encoding information can be generated. This fourth encoding information is used to further control the generation and extraction process of the watermark encoding information. This method not only increases the security of the watermark but also prevents unauthorized third parties from easily modifying or forging the watermark content.

[0127] S6034. Take the fifth remainder on the sequence length based on the fourth encoding information, and determine the second offset information based on the fifth remainder; wherein, the sequence length is determined based on the number of the second watermark encoding information extracted from the embedded text data.

[0128] In this embodiment, the sequence length refers to the total amount of second watermark encoding information extracted from the embedded text data. By using the fourth encoding information as a divisor and performing a modulo operation on the sequence length, a fifth remainder within a finite range can be obtained. This fifth remainder will be used as the second offset information to guide further processing and extraction of the watermark information.

[0129] Exemplarily, the preset key is "cmri", which is combined with the first character "数" to form "cmri数" to obtain the MD5 value "2D847FDAF24CE916FD95681257C1828C". Take the last two digits "8C" and take the remainder of the sequence length 48, getting 2, that is, the offset value of the watermark information coding sequence is 2, forming the coding information combination (2, 1).

[0130] In the embodiment of the present application, by taking the fourth remainder of the total number of tone values of the second sub-text data based on the third parameter in the dilution parameter system, and then determining the corresponding tone segmentation and watermark coding information according to the operation result. Further combine the preset key with the preset position text information, and generate the fourth coding information based on this. Subsequently, use the fourth coding information to take the fifth remainder of the sequence length to determine the second offset information. Through the above method, the efficient coding and extraction of watermark information are realized. This operation method improves the flexibility and security of watermark embedding, so that the system can meet diverse watermark application scenarios.

[0131] In the embodiment of the present application, S604 shown can also be implemented through S6041 to S6043, and the combination steps will be described: S6041. Statistically analyze the multiple second watermark coding information based on the multiple second offset information to determine the first intermediate watermark coding information sequence.

[0132] In the embodiment of the present application, the first intermediate watermark coding information sequence can be determined by sorting and statistically analyzing each second watermark coding information according to the offset position specified by the corresponding second offset information.

[0133] In the embodiment of the present application, all the extracted second watermark coding information will be classified and statistically analyzed according to their corresponding second offset information to form the first intermediate watermark coding information sequence. The first intermediate watermark coding information sequence can be regarded as a preliminary restoration structure of the watermark information, preparing for the next decryption.

[0134] S6041. Use the fifth coding information corresponding to the preset key to decrypt the first intermediate watermark coding information sequence to obtain the second intermediate watermark coding information sequence.

[0135] In the embodiment of the present application, the fifth coding information is a set of coding data generated by performing a specific transformation on the preset key, and is used to perform an exclusive OR or other decryption operations with the first intermediate watermark coding information sequence to obtain the second intermediate watermark coding information sequence.

[0136] S6043. Restore the second intermediate watermark coding information sequence to determine the watermark information.

[0137] In this embodiment, restoration refers to the process of converting the second intermediate watermark encoded information sequence back into watermark information after decryption. The restoration process typically includes operations such as sorting the encoded information, statistical analysis using an election method, and final character mapping or binary restoration.

[0138] In this embodiment, the first intermediate watermark encoding information sequence is determined by statistically analyzing the second watermark encoding information based on the second offset information; the first intermediate watermark encoding information sequence is decrypted using the fifth encoding information corresponding to the preset key to obtain the second intermediate watermark encoding information sequence; and the watermark information is further determined by restoring the second intermediate watermark encoding information sequence. This method effectively improves the accuracy of watermark information extraction, thereby enhancing the robustness and confidentiality of the watermark information, and ultimately enabling efficient traceability and anti-counterfeiting of embedded text content.

[0139] Please see Figure 6 The above is a flowchart illustrating the text watermark extraction method provided in this application embodiment. Figure 2 , will combine Figure 6 The steps shown are explained below: S21. For each sentence in the text content, count the total number of strokes, and determine whether there is watermark encoding information based on the influence factor of the embedding algorithm. If so, extract the encoding based on the sum of tone values.

[0140] In this embodiment, for each sentence of the text content, the total number of strokes is counted. Based on the influence factor of the embedding algorithm, it is determined whether watermark encoding information exists. If it exists, encoding is extracted based on the sum of tone values. The rules for the influence factor and the encoding corresponding to the partition are consistent with those during embedding. For each sentence of the text content, the total number of strokes Tb is obtained by adding the strokes of all Chinese characters in the sentence. To determine whether second watermark encoding information exists, specifically, based on the embedding influence factor (M, N), if Tb modulo M is less than N, the process continues; otherwise, the process for that sentence ends, and no further operations are performed.

[0141] S22. Perform statistical filtering on all extracted watermark encoding information to restore the original watermark information.

[0142] In this embodiment, all extracted watermark encoding information is statistically filtered and restored to the original watermark information. Specifically, the statistical method for watermark encoding information involves counting the encoding information corresponding to each sequence offset value and then obtaining the final watermark encoding information sequence through an election method. Using the same transformation method as during embedding, combined with the key, the watermark encoding information sequence is restored to the original watermark information.

[0143] Please see Figure 7 This is an interactive schematic diagram of the text watermark embedding method provided in the embodiments of this application.

[0144] S701, The text watermark embedding device acquires text data; wherein, the text data includes: multiple first sub-text data.

[0145] In this embodiment, the implementation steps of S701 can be referred to S101, and will not be described in detail here.

[0146] S702, the text watermark embedding device determines the embedding detection result based on the total number of strokes corresponding to each of the first sub-text data; wherein, the embedding detection result is used to characterize whether the corresponding first watermark encoding information is embedded in the first sub-text data.

[0147] In this embodiment, the implementation steps of S702 can be referred to S102, and will not be described in detail here.

[0148] S703. Text watermark embedding device: If the embedding detection result indicates embedding, the first watermark encoding information corresponding to the first sub-text data is determined in the watermark encoding information sequence, and the text information in the first sub-text data is adjusted until the total number of strokes and the total number of tone values ​​of the first sub-text data meet the corresponding rules to obtain embedded text data; wherein, the watermark encoding information sequence is formed based on the watermark information embedded in the text data.

[0149] In this embodiment, the implementation steps of S703 can be referred to S103, and will not be described in detail here.

[0150] S704. The text watermark extraction device acquires embedded text data; wherein, the embedded text data includes: multiple second sub-text data.

[0151] In this embodiment, the implementation steps of S704 can be referred to S601, and will not be described in detail here.

[0152] S705. The text watermark extraction device determines the extraction detection result based on the total number of strokes corresponding to each second sub-text data; wherein, the extraction detection result is used to characterize whether there is corresponding second watermark encoding information in the second sub-text data.

[0153] In this embodiment, the implementation steps of S705 can be referred to S602, and will not be described in detail here.

[0154] S706. If the extraction detection result indicates that the watermark extraction device exists, it extracts the corresponding second watermark encoding information and second offset information based on the total number of tone values ​​of the second sub-text data.

[0155] In this embodiment, the implementation steps of S706 can be referred to S603, and will not be described in detail here.

[0156] S707. The text watermark extraction device determines the watermark information based on multiple sets of the second watermark encoding information and the second offset information extracted from the embedded text data.

[0157] In this embodiment, the implementation steps of S707 can be referred to S604, and will not be described in detail here.

[0158] Please see Figure 8 This is a schematic diagram of the text watermark embedding device provided in the embodiments of this application.

[0159] In this application embodiment, a text watermark embedding device 600 is provided, including: a first data acquisition unit 601, a first determination unit 602, and an embedding adjustment unit 603.

[0160] The first data acquisition unit 601 is used to acquire text data; wherein, the text data includes: a plurality of first sub-text data; The first determining unit 602 is used to determine the embedding detection result based on the total number of strokes corresponding to each of the first sub-text data; wherein, the embedding detection result is used to characterize whether the corresponding first watermark encoding information is embedded in the first sub-text data; The embedding adjustment unit 603 is used to determine the first watermark encoding information corresponding to the first sub-text data in the watermark encoding information sequence if the embedding detection result indicates embedding, and adjust the text information in the first sub-text data until the total number of strokes and the total number of tone values ​​of the first sub-text data meet the corresponding rules to obtain embedded text data; wherein, the watermark encoding information sequence is formed based on the watermark information embedded in the text data.

[0161] In this embodiment of the application, the first data acquisition unit 601 in the text watermark embedding device 600 is used to acquire watermark information, embedding amount parameters and dilution parameters; The embedding amount parameter is used to determine the proportion of the first sub-text data in which the first watermark encoding information is embedded among multiple first sub-text data; the dilution parameter is used to determine the encoding range of the watermark information.

[0162] In this embodiment of the application, the embedding parameters include: a first parameter and a second parameter; wherein, the first parameter is used to take the remainder of the total number of strokes of the first sub-text data; and the second parameter is a preset maximum remainder value; The dilution parameter includes a third parameter and a fourth parameter; wherein the third parameter is used to take the remainder of the total number of tone values ​​of the first sub-text data; and the fourth parameter is used to determine the number of segments for the tone values.

[0163] In this embodiment of the application, the first determining unit 602 in the text watermark embedding device 600 is used to convert the watermark information into first encoded information in the encoding base characterized by the dilution parameter; wherein, the encoding base is determined based on the number of segments determined by the fourth parameter; Convert the preset key into second encoded information in the encoded base; The first encoded information is encrypted based on the second encoded information to determine the watermark encoded information sequence.

[0164] In this embodiment of the application, the first determining unit 602 in the text watermark embedding device 600 is used to take the first remainder of the total number of strokes of the first sub-text data based on the first parameter. The embedding detection result is determined based on the comparison between the first remainder and the second parameter.

[0165] In this embodiment of the application, the embedding adjustment unit 603 in the text watermark embedding device 600 is used to determine the third encoding information after combining the preset key and the preset position text information in the first sub-text data; The second remainder is taken based on the sequence length of the watermark encoded information sequence according to the third encoded information; The first watermark encoding information is determined in the watermark encoding information sequence based on the first offset information represented by the second remainder.

[0166] In this embodiment of the application, the first determining unit 602 in the text watermark embedding device 600 is used to take the third remainder of the total number of tone values ​​of the first sub-text data based on the third parameter. The tone segment corresponding to the third remainder and the encoded character corresponding to the tone segment are determined; wherein each tone segment is determined based on the number of segments determined by the fourth parameter; Based on the consistency between the first watermark encoding information and the encoded character, an adjustment detection result is determined; wherein, the adjustment detection result is used to characterize whether to adjust the text information in the first sub-text data until the total number of strokes and the total number of tone values ​​of the first sub-text data meet the corresponding rules.

[0167] In this embodiment of the application, the embedding adjustment unit 603 in the text watermark embedding device 600 is used to replace or insert text information in the first sub-text data to obtain embedded sub-text data. Specifically, the remainder of the total number of strokes in the embedded sub-text data based on the first parameter is less than the second parameter, and the encoded character corresponding to the remainder of the total number of tone values ​​in the embedded sub-text data based on the third parameter is consistent with the first watermark encoding information.

[0168] It should be noted that, in the embodiments of this application, if the above-described text watermark embedding method is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiments of this application, or the part that contributes to the related technology, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a text watermark embedding device (which may be a personal computer, etc.) to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), magnetic disks, or optical disks. Thus, the embodiments of this application are not limited to any specific hardware and software combination.

[0169] Correspondingly, embodiments of this application provide a computer-readable storage medium storing a computer program thereon, which, when executed by a first processor, implements the steps in the text watermark embedding method.

[0170] It should be noted that the descriptions of the storage medium and device embodiments above are similar to those of the method embodiments above, and have similar beneficial effects. For technical details not disclosed in the storage medium and device embodiments of this application, please refer to the descriptions of the method embodiments of this application for understanding.

[0171] It should be noted that, Figure 9 A schematic diagram of a hardware entity of the first electronic device provided in the embodiments of this application, such as... Figure 9 As shown, this application embodiment provides a first electronic device 700, including a first memory 702 and a first processor 701. The first memory 702 stores a computer program that can run on the first processor 701. When the first processor 701 executes the program, it implements the steps in the above-described method, wherein; The first processor 701 typically controls the overall operation of the electronic device 700.

[0172] The first memory 702 is configured to store instructions and applications executable by the first processor 701, and can also cache data to be processed or already processed by the first processor 701 and the various modules in the first electronic device 700 (e.g., image data, audio data, voice communication data and video communication data), which can be implemented by flash memory or random access memory (RAM).

[0173] Correspondingly, this application embodiment also provides a computer program product, including a computer program that can be executed by a first processor 701 of a first electronic device 700 to complete the steps in the method of the text watermark embedding device 600.

[0174] Please see Figure 10 This is a schematic diagram of the text watermark extraction device provided in the embodiments of this application.

[0175] This application embodiment also provides a text watermark extraction device 800, including: a second data acquisition unit 801, a second determination unit 802, an extraction unit 803, and a restoration unit 804.

[0176] The second data acquisition unit 801 is used to acquire embedded text data; wherein, the embedded text data includes: a plurality of second sub-text data; The embedded text data is obtained by determining the first watermark encoding information corresponding to the first sub-text data in the watermark encoding information sequence during the embedding detection result representation, and adjusting the text information in the first sub-text data until the total number of strokes and the total number of tone values ​​of the first sub-text data meet the corresponding rules; the watermark encoding information sequence is formed based on the watermark information embedded in the text data; the embedding detection result is determined based on the total number of strokes corresponding to each first sub-text data. The second determining unit 802 is used to determine the extraction detection result based on the total number of strokes corresponding to each of the second sub-text data; wherein, the extraction detection result is used to characterize whether the second sub-text data has corresponding second watermark encoding information; Extraction unit 803 is used to extract the corresponding second watermark encoding information and second offset information based on the total number of tone values ​​of the second sub-text data if the extraction detection result indicates extraction. The restoration unit 804 is used to determine watermark information based on multiple sets of the second watermark encoding information and the second offset information extracted from the embedded text data.

[0177] In this embodiment of the application, the extraction unit 803 in the text watermark extraction device 800 is used to take the fourth remainder of the total number of tone values ​​of the second sub-text data based on the third parameter in the dilution parameter. The tone segment corresponding to the fourth remainder is determined, and the second watermark encoding information corresponding to the tone segment is determined; wherein, each tone segment is determined based on the number of segments determined by the fourth parameter in the dilution parameter; Determine the fourth encoded information, which is a combination of the preset key and the preset position text information in the second sub-text data; The sequence length is calculated by taking a fifth remainder based on the fourth encoding information, and the second offset information is determined based on the fifth remainder; wherein the sequence length is determined based on the number of the second watermark encoding information extracted from the embedded text data.

[0178] In this embodiment of the application, the restoration unit 804 in the text watermark extraction device 800 is used to statistically analyze multiple second watermark encoding information based on multiple second offset information to determine the first intermediate watermark encoding information sequence. Using the fifth encoding information corresponding to the preset key, the first intermediate watermark encoding information sequence is decrypted to obtain the second intermediate watermark encoding information sequence; The watermark information is determined by restoring the second intermediate watermark encoding information sequence.

[0179] Correspondingly, embodiments of this application provide a computer-readable storage medium storing a computer program thereon, which, when executed by a second processor, implements the steps in the text watermark embedding method.

[0180] It should be noted that the descriptions of the storage medium and device embodiments above are similar to those of the method embodiments above, and have similar beneficial effects. For technical details not disclosed in the storage medium and device embodiments of this application, please refer to the descriptions of the method embodiments of this application for understanding.

[0181] It should be noted that, Figure 11 This is a schematic diagram of a hardware entity of the second electronic device provided in an embodiment of this application, such as... Figure 11 As shown, this application embodiment provides a second electronic device 900, including a second memory 902 and a second processor 901. The second memory 902 stores a computer program that can run on the second processor 901. When the second processor 901 executes the program, it implements the steps in the above-described method, wherein; The second processor 901 typically controls the overall operation of the electronic device 900.

[0182] The second memory 902 is configured to store instructions and applications executable by the second processor 901, and can also cache data to be processed or already processed by the second processor 901 and the various modules in the second electronic device 900 (e.g., image data, audio data, voice communication data and video communication data), which can be implemented by flash memory or random access memory (RAM).

[0183] Correspondingly, this application embodiment also provides a computer program product, including a computer program that can be executed by a second processor 901 of a second electronic device 900 to complete the steps in the method of the text watermark embedding device 800.

[0184] The above description is merely an embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.

Claims

1. A method for embedding text watermarks, characterized in that, include: Acquire text data; wherein the text data includes: multiple first sub-text data; Based on the total number of strokes corresponding to each of the first sub-text data, an embedding detection result is determined; wherein, the embedding detection result is used to characterize whether the corresponding first watermark encoding information is embedded in the first sub-text data; If the embedding detection result indicates embedding, then the first watermark encoding information corresponding to the first sub-text data is determined in the watermark encoding information sequence, and the text information in the first sub-text data is adjusted until the total number of strokes and the total number of tone values ​​of the first sub-text data meet the corresponding rules to obtain embedded text data; wherein, the watermark encoding information sequence is formed based on the watermark information embedded in the text data.

2. The text watermark embedding method according to claim 1, characterized in that, The method further includes: Obtain watermark information, embedding amount parameters, and dilution parameters; The embedding amount parameter is used to determine the proportion of the first sub-text data in which the first watermark encoding information is embedded among multiple first sub-text data; the dilution parameter is used to determine the encoding range of the watermark information.

3. The text watermark embedding method according to claim 2, characterized in that, The embedding parameters include: a first parameter and a second parameter; wherein, the first parameter is used to take the remainder of the total number of strokes of the first sub-text data; and the second parameter is a preset maximum remainder value. The dilution parameter includes a third parameter and a fourth parameter; wherein the third parameter is used to take the remainder of the total number of tone values ​​of the first sub-text data; and the fourth parameter is used to determine the number of segments for the tone values.

4. The text watermark embedding method according to claim 3, characterized in that, The method further includes: The watermark information is converted into first encoded information in the encoding base characterized by the dilution parameter; wherein the encoding base is determined based on the number of segments determined by the fourth parameter; Convert the preset key into second encoded information in the encoded base; The first encoded information is encrypted based on the second encoded information to determine the watermark encoded information sequence.

5. The text watermark embedding method according to any one of claims 1 to 4, characterized in that, The step of determining the embedding detection result based on the total number of strokes corresponding to each of the first sub-text data includes: The first remainder is taken from the total number of strokes of the first sub-text data based on the first parameter; The embedding detection result is determined based on the comparison between the first remainder and the second parameter.

6. The text watermark embedding method according to any one of claims 1 to 4, characterized in that, Determining the first watermark encoding information corresponding to the first sub-text data in the watermark encoding information sequence includes: Determine the third encoded information resulting from the combination of the preset key and the preset position text information in the first sub-text data; The second remainder is taken based on the sequence length of the watermark encoded information sequence according to the third encoding information; The first watermark encoding information is determined in the watermark encoding information sequence based on the first offset information represented by the second remainder.

7. The text watermark embedding method according to any one of claims 1 to 4, characterized in that, Before adjusting the text information in the first sub-text data until the total number of strokes and the total number of tone values ​​in the first sub-text data meet the corresponding rules, the method further includes: The third remainder is taken based on the total number of tone values ​​of the first sub-text data according to the third parameter. The tone segment corresponding to the third remainder and the encoded character corresponding to the tone segment are determined; wherein each tone segment is determined based on the number of segments determined by the fourth parameter; Based on the consistency between the first watermark encoding information and the encoded character, an adjustment detection result is determined; wherein, the adjustment detection result is used to characterize whether to adjust the text information in the first sub-text data until the total number of strokes and the total number of tone values ​​of the first sub-text data meet the corresponding rules.

8. The text watermark embedding method according to claim 7, characterized in that, The step of adjusting the text information in the first sub-text data until the total number of strokes and the total number of tone values ​​in the first sub-text data conform to the corresponding rules includes: In the first sub-text data, text information is replaced or inserted to obtain embedded sub-text data; Specifically, the remainder of the total number of strokes in the embedded sub-text data based on the first parameter is less than the second parameter, and the encoded character corresponding to the remainder of the total number of tone values ​​in the embedded sub-text data based on the third parameter is consistent with the first watermark encoding information.

9. A method for extracting text watermarks, characterized in that, include: Obtain embedded text data; wherein, the embedded text data includes: multiple second sub-text data; The embedded text data is obtained by determining the first watermark encoding information corresponding to the first sub-text data in the watermark encoding information sequence during the embedding detection result representation, and adjusting the text information in the first sub-text data until the total number of strokes and the total number of tone values ​​of the first sub-text data meet the corresponding rules; the watermark encoding information sequence is formed based on the watermark information embedded in the text data; the embedding detection result is determined based on the total number of strokes corresponding to each first sub-text data. Based on the total number of strokes corresponding to each second sub-text data, an extraction detection result is determined; wherein, the extraction detection result is used to characterize whether there is a corresponding second watermark encoding information in the second sub-text data; If the extracted detection result representation exists, then the corresponding second watermark encoding information and second offset information are extracted based on the total number of tone values ​​of the second sub-text data. The watermark information is determined based on multiple sets of the second watermark encoding information and the second offset information extracted from the embedded text data.

10. The text watermark extraction method according to claim 9, characterized in that, The extraction of the corresponding second watermark encoding information and second offset information based on the total number of tone values ​​of the second sub-text data includes: The fourth remainder is taken based on the total number of tone values ​​of the second sub-text data according to the third parameter in the dilution parameter. The tone segment corresponding to the fourth remainder is determined, and the second watermark encoding information corresponding to the tone segment is determined; wherein, each tone segment is determined based on the number of segments determined by the fourth parameter in the dilution parameter; Determine the fourth encoded information, which is a combination of the preset key and the preset position text information in the second sub-text data; The sequence length is calculated by taking a fifth remainder based on the fourth encoding information, and the second offset information is determined based on the fifth remainder; wherein the sequence length is determined based on the number of the second watermark encoding information extracted from the embedded text data.

11. The text watermark extraction method according to claim 9, characterized in that, The determination of watermark information based on multiple sets of the second watermark encoding information and the second offset information extracted from the embedded text data includes: Based on multiple second offset information, multiple second watermark encoding information are statistically analyzed to determine the first intermediate watermark encoding information sequence; Using the fifth encoding information corresponding to the preset key, the first intermediate watermark encoding information sequence is decrypted to obtain the second intermediate watermark encoding information sequence; The watermark information is determined by restoring the second intermediate watermark encoding information sequence.

12. A text watermark embedding device, characterized in that, include: The first data acquisition unit is used to acquire text data; wherein the text data includes: a plurality of first sub-text data; The first determining unit is used to determine the embedding detection result based on the total number of strokes corresponding to each of the first sub-text data; wherein the embedding detection result is used to characterize whether the corresponding first watermark encoding information is embedded in the first sub-text data; An embedding adjustment unit is configured to, if the embedding detection result indicates embedding, determine the first watermark encoding information corresponding to the first sub-text data in the watermark encoding information sequence, and adjust the text information in the first sub-text data until the total number of strokes and the total number of tone values ​​of the first sub-text data meet the corresponding rules, thereby obtaining embedded text data; wherein, the watermark encoding information sequence is formed based on the watermark information embedded in the text data.

13. A text watermark extraction device, characterized in that, include: The second data acquisition unit is used to acquire embedded text data; wherein, the embedded text data includes: a plurality of second sub-text data; The embedded text data is obtained by determining the first watermark encoding information corresponding to the first sub-text data in the watermark encoding information sequence during the embedding detection result representation, and adjusting the text information in the first sub-text data until the total number of strokes and the total number of tone values ​​of the first sub-text data meet the corresponding rules; the watermark encoding information sequence is formed based on the watermark information embedded in the text data; the embedding detection result is determined based on the total number of strokes corresponding to each first sub-text data. The second determining unit is used to determine the extraction detection result based on the total number of strokes corresponding to each of the second sub-text data; wherein, the extraction detection result is used to characterize whether the second sub-text data has corresponding second watermark encoding information; The extraction unit is configured to, if the extraction detection result indicates extraction, extract the corresponding second watermark encoding information and second offset information based on the total number of tone values ​​of the second sub-text data. The restoration unit is used to determine the watermark information based on multiple sets of the second watermark encoding information and the second offset information extracted from the embedded text data.

14. An electronic device, characterized in that, The method includes a memory and a processor, the memory storing a computer program executable on the processor, the processor executing the computer program to implement the steps of the method according to any one of claims 1 to 8, or to implement the steps of the method according to any one of claims 9 to 11.

15. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 8, or the steps of the method according to any one of claims 9 to 11.

16. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 8, or the steps of the method according to any one of claims 9 to 11.