Method for resisting printing and scanning digital watermarking for text image

By adjusting the average skeleton quality (ASM) of text images for watermark embedding and extraction, the problems of low watermark capacity and poor robustness in text images are solved, achieving efficient watermark embedding and extraction under printing and scanning attacks, applicable to text images with multiple languages ​​and fonts.

CN120997026AActive Publication Date: 2025-11-21浣江实验室

Patent Information

Application Number
CN202511511411.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-22
Publication Date
2025-11-21
Estimated Expiration
2045-10-22

AI Technical Summary

Technical Problem

Existing digital watermarking technologies have low embedding capacity in text images and are not suitable for print scanning attacks, especially in formal documents where text content cannot be altered, resulting in poor robustness.

Method used

The blind watermarking algorithm is used to embed and extract watermarks by adjusting the average skeleton quality (ASM) of each character in the text image. The specific steps include character segmentation, ASM adjustment and image correction, which is applicable to text images of different languages ​​and fonts.

Benefits of technology

It enables efficient embedding and extraction of watermarks in text images, resists printing and scanning attacks, maintains the integrity of document content, is applicable to multiple languages ​​and font types, and improves the robustness and capacity of watermarks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120997026A_ABST
    Figure CN120997026A_ABST
Patent Text Reader

Abstract

The invention discloses an anti-printing and anti-scanning digital watermark method for a text image, which comprises the following steps of: embedding and extracting a watermark on the basis of a stroke-like thought, segmenting characters and determining the average skeleton quality ASM of each character during embedding, and forming a group by every two characters, changing the ASM of each group of characters according to the to-be-embedded information to embed the watermark and replace the corresponding characters of the source document; during extraction, two characters are also taken as a group, ASM is extracted, and watermark information is extracted by comparing ASM among the characters, so that 1-bit information can be embedded into the two characters, the language type and font of the characters are not limited, the watermark information amount of the document is ensured, and meanwhile, the watermark information is extracted. The problems that algorithms among different languages are limited, and the robustness is poor and the capacity is small when digital watermarks in a space domain and a transform domain resist printing and scanning attacks are effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of digital watermarking, in particular to a method for anti-printing and scanning digital watermarking of text images. BACKGROUND

[0002] Text data, such as medical records, contracts and identity documents, play a key role in people's daily life. Digital watermarking and data hiding technology emerged as the times require, which can embed copyright information imperceptibly into multimedia data. However, most previous image watermarking schemes are for color and grayscale images, which are not suitable for text images because text images lack redundant information, especially after printing and scanning. Therefore, it is crucial to study watermarking technology that can survive printing and scanning. Current anti-printing and scanning watermarking embedding can be roughly divided into three categories: image-based methods, format-based methods and language-based methods.

[0003] The existing watermarking technology has the following shortcomings: 1) most image-based methods are based on pixel flipping, which has low capacity and requires high requirements for printing and scanning equipment; 2) most format-based methods are not blind watermarking, and also have the disadvantage of low embedding capacity; 3) existing language-based methods embed watermarking by changing the syntax and semantic properties of text, but are not suitable for some formal documents because some formal documents do not allow changes to the text content, so the scope of application is not very large. SUMMARY

[0004] In order to achieve the above purpose, the present application is realized by the following technical scheme: The present application is a method for anti-printing and scanning digital watermarking of text images, which comprises a watermark embedding process and a watermark extraction process. The watermark embedding process comprises the following steps: Sa-1: converting the target document into a picture and performing binaryzation processing to generate a binary text image I; Sa-2: dividing each character in the text image I by projection method, denoted as C1, C2,..., Cn, where n is the total number of characters; n where n is the total number of characters; Sa-3: selecting the first 2L characters and extracting their ASM, where ASM is the average skeleton mass, and L is the number of watermark bits; Sa-4: taking every two characters as a group, denoted as , , until starting from the first character group, performing watermark bits w1, w2,..., wn, where n is the total number of characters. LL is the number of watermark bits; ASM of the first character in each character group is denoted as ASM1, and ASM of the last character is denoted as ASM2; if the current watermark bit is 0, the group is changed to the group in which ASM1 is increased and ASM2 is decreased until ASM1>ASM2+th; if the current watermark bit is 1, the group is changed to the group in which ASM1 is decreased and ASM2 is increased until ASM1<ASM2-th; wherein th is a preset value in the range of 0.1-0.5; the character group after embedding the watermark is denoted as , and the character is replaced by , and the operation is repeated until L watermark bits are embedded completely; Sa-5: generating the text image after embedding the watermark; The watermark extraction process comprises the following steps: Sb-1: correcting the direction of the image after printing and scanning, and removing isolated noise points; Sb-2: converting the corrected image into a binary text image I w ; Sb-3: separating each character in the binary text image I w by projection method, denoted as , n is the total number of characters, and is the same as that in embedding; Sb-4: screening the first 2L embeddable characters, and extracting the ASM of each character, and each two characters are a group denoted as , ASM of the first character in each character group is denoted as , and ASM of the last character is denoted as ; Sb-5: starting from the first character group to extract the watermark: if , the extracted watermark information is 0; if , the extracted watermark information is 1; until the extraction of bit information is completed; wherein, the extraction algorithm of ASM is: assuming that the size of the current character is (M, N), M is the number of character height pixels, and N is the number of character width pixels, the number of black pixels in the character is counted ; the skeleton length of the character is extracted by skeleton extraction algorithm , that is, the total length of the skeleton pixel points, and the formula is:

[0005] . Preferably, the projection method in step Sa-2 is an optimization projection method based on the traditional projection method, and the specific process of the optimization projection method is as follows:First, the line information is extracted, and the character blocks are divided based on the character connected domain characteristics; for the binary text image I, the white background is set as 1, and the black text is set as 0; the pixel sum is calculated by row, and the region with a sum of 0 is divided into a text area, so as to obtain the number of lines and the uppermost and lowermost boundaries of each line; Then, the column pixel sum of each line is calculated, and the left and right boundaries of each character block are obtained based on the character connectivity; subsequently, the width of the character block and the height of the black pixels are judged, if the width is less than 0.4 times the line height and the height of the black pixels is less than 0.6 times the line height, it is determined to be a punctuation mark; the method for judging whether the character block is a character is as follows: the width of the character block is checked, if the width is between a times the line height and b times the line height, and the height of the black pixels is between 0.6-0.7 times the line height, b is 1-1.1, it is considered to be a Chinese character; after discarding the punctuation mark, it is sequentially judged whether three, two or one continuous character blocks are a Chinese character, so as to determine the character position. The width of the character block is checked, if the width is between a times the line height and b times the line height, and the height of the black pixels is between 0.6-0.7 times the line height, b is 1-1.1, it is considered to be a Chinese character; after discarding the punctuation mark, it is sequentially judged whether three, two or one continuous character blocks are a Chinese character, so as to determine the character position.

[0006] Preferably, the ASM in steps Sa-3, Sa-4, Sb-4 and Sb-5 is the average skeleton mass, and the stroke contains the pixel width.

[0007] Preferably, the embedding condition in step Sa-4 is divided into the following four kinds, and the current character group is , and the specific steps are as follows: When the watermark bit is 1, if ASM1≥ASM2-th, ASM1 is reduced and ASM2 is increased until ASM1<ASM2-th; When the watermark bit is 1, if ASM1<ASM2-th, it remains unchanged; When the watermark bit is 0, if ASM1≤ASM2+th, ASM1 is increased and ASM2 is reduced until ASM1>ASM2+th; When the watermark bit is 0, if ASM1>ASM2-th, it remains unchanged.

[0008] Preferably, the ASM adjustment of the current character group contains two operations of “reducing ASM1 and increasing ASM2” and “increasing ASM1 and reducing ASM2”: The operation steps of “reducing ASM1 and increasing ASM2” are as follows: Si-1: flip the black pixels in the first character and the white pixels in the last character, initially reduce ASM1 and increase ASM2, and compare the size of ASM1 and ASM2-th; if ASM1<ASM2-th, the adjustment is completed; ​Si-2: If ASM1≥ASM2-th, make the thickest stroke in the head character thinner and the thinnest stroke in the tail character thicker, and then compare the size of ASM1 and ASM2-th again; if ASM1<ASM2-th, the adjustment is completed; Si-3: If ASM1≥ASM2-th, re-execute step Si-1. The operation steps of "increasing ASM1 and decreasing ASM2" are as follows: Sii-1: Flip the white pixels in the head character and the black pixels in the tail character, and preliminarily increase ASM1 and decrease ASM2, and then compare the size of ASM1 and ASM2+th; if ASM1>ASM2+th, the adjustment is completed. Sii-2: If ASM1≤ASM2+th, make the thinnest stroke in the head character thicker and the thickest stroke in the tail character thinner, and then compare the size of ASM1 and ASM2+th again; if ASM1>ASM2+th, the adjustment is completed. Sii-3: If ASM1≥ASM2-th, re-execute step Sii-1.

[0009] Preferably, the process of adjusting the stroke thickness in steps Si-2 and Sii-2 comprises: S6-1: Extract the outer contour of the character as ext_edge, and the inner contour as int_edge, extract the skeleton of part of the strokes in the skeleton by a straight line searching algorithm, and mark as ; S6-2: Increase the number of times of the and perform "AND" operation with the inner and outer contours of the character to obtain the inner and outer contours of the strokes, wherein the inner contour is marked as , and the outer contour is marked as ; S6-3: Increase the number of times of the and perform "AND" operation with the character itself to obtain the strokes, calculate the "ASM" of each stroke, and determine the strokes corresponding to the maximum value and the minimum value; S6-4: Take the inner contour of the stroke corresponding to the maximum value, record the positions of all pixel points on the same side of the axial direction in the character according to the axial direction of the stroke, and exclude the points where the pixel points corresponding to the positions in the character are black on both sides perpendicular to the axial direction; change the pixel points corresponding to the positions in the character to white, and complete the thinning operation. S6-5: Take the outer contour of the stroke corresponding to the minimum value, record the positions of all pixel points on the same side of the axial direction in the character according to the axial direction of the stroke; change the pixel points corresponding to the positions in the character to black, and complete the thickening operation.

[0010] Preferably, the method for image direction correction in step Sb-1 is as follows: First, the image tilt angle is detected by the existing algorithm, then the image is rotated by the corresponding angle, and then the four edges of the image are cut until there is no black pixel point at the edge, and the black edge generated by rotation is removed; isolated noise is removed to ensure the accuracy of character segmentation.

[0011] Preferably, the skeleton extraction algorithm is a morphological-based skeleton extraction method, which extracts the center skeleton of the character through multiple morphological erosion and reconstruction operations, and obtains the skeleton length .

[0012] Preferably, the binarization processing adopts Otsu threshold method, automatically determines the binarization threshold, and converts the target document into a black and white binary text image I.

[0013] The application of the method for resisting printed scanning digital watermark of a text image includes application to text images of different language types and different font types, and the different language types include Chinese characters and English.

[0014] Beneficial effects: the application can embed 1bit information in two characters, and the language type and font of the characters are not limited. While ensuring the amount of hidden watermark information in a document, the application effectively solves the problem of algorithm limitation between different languages, and also solves the problems of poor robustness against printed scanning attacks and small capacity when adding digital watermark in the spatial domain and the transform domain, and has high robustness against printed scanning attacks. BRIEF DESCRIPTION OF DRAWINGS

[0015] Fig. 1 is the embedding process flowchart of the application.

[0016] Fig. 2 is the extraction process flowchart of the application.

[0017] Fig. 3 is the difference diagram of the embedding effect of different watermark information of the application. DETAILED DESCRIPTION

[0018] The technical solutions in the embodiments of the application will be clearly and completely described below with reference to the drawings in the embodiments of the application Figs. 1 to 3 It is obvious that the described embodiments are only part of the embodiments of the application, not all the embodiments. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor are within the protection scope of the application.

[0019] As Fig. 1The flowchart shown illustrates the steps of watermark embedding: First, the target document is converted into an image and binarized; then, the position of each character is found using the projection method; next, the first 2L characters are selected and their "Average Skeleton Quality (ASM)" is extracted; then, each pair of characters is grouped together, and the "ASM" is changed according to the watermark information to be embedded to perform watermark embedding; finally, the watermarked text image is converted into a PDF document to obtain the watermarked document.

[0020] like Fig. 2 The flowchart shown illustrates the watermark extraction process: First, the document containing the watermark is printed and scanned; then, the scanned image is oriented, noise is removed, and binarized; next, the position of each character is found using the projection method; then, the first 2L characters are selected and their "ASM" is extracted; after that, every two characters are grouped together, and the "ASM" is modified according to the watermark information to be embedded for watermark embedding; finally, the text image containing the watermark is converted into a PDF document to obtain the watermarked document.

[0021] like Fig. 3 The diagram showing the difference in the effects of embedding different watermark information specifically illustrates the two words "watermark". Fig. 3 The watermark on the left side of the middle is embedded with "1". Fig. 3 The watermark on the right side of the middle is embedded with "0", that is... Fig. 3 The left side shows the effect of embedding "1", and the right side shows the effect of embedding "0". This visually demonstrates the difference in text appearance when embedding different watermark information, reflecting the result of embedding different watermark bits by changing "ASM".

[0022] Technical solution / principle: The watermarking algorithm of this invention is a blind watermarking algorithm, which can embed and extract watermarks without the original document. The embedding and extraction of watermarks are achieved by adjusting the average skeleton quality (ASM). The specific principle is as follows: I. Creativity includes Average Skeleton Quality (ASM) "ASM" is defined as Average Skeleton Quality, which is the pixel width contained in a stroke. It is extracted based on the number of black pixels and the skeleton length of the character, using the following formula: .

[0023] Among them, SUM black : Counts the total number of black pixels in a character. Black pixels are the pixels that represent text after binarization in a text image (pixel value is 0).

[0024] The skeleton length of a character is extracted using a skeleton extraction algorithm (such as a morphology-based skeleton extraction method, multiple morphological erosion and reconstruction operations). The skeleton is the central structure of the character and reflects its shape features.

[0025] II. Watermark embedding principle Grouping every two characters, adjust the "ASM" of the two characters in each group according to the watermark bit (0 or 1) to be embedded: When the embedded bit is 1, make the "ASM" of the first character in the group smaller than that of the last character (by reducing the "ASM" of the first character and increasing the "ASM" of the last character.

[0026] The anti-printing and scanning digital watermarking method for text images provided by the present application is a blind watermarking algorithm, including two processes of watermark embedding and watermark extraction.

[0027] Watermark embedding process: Step Sa-1: Convert the target document into a picture and perform binaryzation processing to generate a binaryzation text image I.

[0028] Step Sa-2: Divide each character in the text image I by projection method, denoted as C1, C2,..., Cn. n , where n is the total number of characters. The projection method specifically includes the following steps: first, perform line information extraction, divide the character blocks based on the character connected domain characteristics; calculate the pixel sum of the binaryzation text image I (white background is 1 and black text is 0) by row, divide the region with a sum of 0 into a text area, and thereby obtain the number of rows and the uppermost and lowermost boundaries of each row; then, calculate the column pixel sum of each row, and obtain the left and right boundaries of each character block based on the character connectivity; subsequently, judge the width and black pixel height of the character block, if the width is less than 0.4 times the row height and the black pixel height is less than 0.6 times the row height, it is determined as a punctuation symbol; the method for judging whether the character block is a character is to check the width of the character block, if the width is between 0.6 times the row height and 1.1 times the row height (usually 0.6-0.7, usually 1-1.1), it is considered as a Chinese character; after discarding the punctuation symbols, judge whether 3, 2 or 1 continuous character blocks are a Chinese character, thereby determining the character position.

[0029] Step Sa-3: Screen the first 2L characters and extract the "average skeleton mass (ASM)" thereof (L is the number of watermark bits). "ASM" is defined as the average skeleton mass, i.e. the pixel width contained by the strokes, and the extraction algorithm is as follows: assuming that the size of the current character is (M, N) (M is the number of character height pixels and N is the number of character width pixels), count the number of black pixels SUM in the character black ; extract the skeleton length Ls (the total length of skeleton pixel points) of the character by a skeleton extraction method based on morphology (multiple morphological erosion and reconstruction operations), then .

[0030] ​​​​Step Sa-4: Group every two characters, denoted as , and start embedding watermark bits from the first character group (L is the number of watermark bits). Denote the ASM of the first character of each character group as ASM1 and the ASM of the last character as ASM2. If the current watermark bit is 0, increase ASM1 and decrease ASM2 in the group until ASM1 > ASM2 + th (th is a preset value, ranging from 0.1 to 0.5); if the current watermark bit is 1, decrease ASM1 and increase ASM2 in the group until ASM1 < ASM2 - th. The embedding situations are divided into the following four cases (assuming the current character group is ): When the watermark bit is 1, if ASM1 ≥ ASM2 - th, decrease ASM1 and increase ASM2 until ASM1 < ASM2 - th; When the watermark bit is 1, if ASM1 < ASM2 - th, keep it unchanged; When the watermark bit is 0, if ASM1 ≤ ASM2 + th, increase ASM1 and decrease ASM2 until ASM1 > ASM2 + th; When the watermark bit is 0, if ASM1 > ASM2 - th, keep it unchanged.

[0031] The adjustment of the ASM of the current character group includes two operations: "decrease ASM1 and increase ASM2" and "increase ASM1 and decrease ASM2"; Among them, the operation steps of "decrease ASM1 and increase ASM2" are as follows: Si-1: Flip the black pixels in the first character and the white pixels in the last character, initially decrease ASM1 and increase ASM2, and compare the sizes of ASM1 and ASM2 - th; if ASM1 < ASM2 - th, the adjustment is completed; Si-2: If ASM1 ≥ ASM2 - th, make the thickest stroke in the first character thinner and the thinnest stroke in the last character thicker, and compare the sizes of ASM1 and ASM2 - th again; if ASM1 < ASM2 - th, the adjustment is completed; Si-3 If ASM1 ≥ ASM2 - th, re-execute step Si-1; The operation steps of "increase ASM1 and decrease ASM2" are as follows: Sii-1: Flip the white pixels in the first character and the black pixels in the last character, initially increase ASM1 and decrease ASM2, and compare the sizes of ASM1 and ASM2 + th; if ASM1 > ASM2 + th, the adjustment is completed; Sii-2: If ASM1≤ASM2+th, thicken the thinnest stroke in the first character and thin the thickest stroke in the last character, then compare ASM1 and ASM2+th again; if ASM1>ASM2+th, the adjustment is complete. Sii-3: If ASM1≥ASM2-th, repeat step Sii-1.

[0032] The process of adjusting stroke thickness in steps Si-2 and Si-2 includes: S6-1: Extract the outer contour of the character, denoted as ext_edge, and the inner contour, denoted as int_edge. Extract the skeleton of some strokes from the skeleton using a line-finding algorithm, denoted as... ; S6-2: Will After widening the stroke several times, an AND operation is performed with the inner and outer contours of the character to obtain the inner and outer contours of the stroke. The inner contour is denoted as... The outer contour is recorded as ; S6-3: Will After widening the character by several times, perform an AND operation with the character itself to obtain the number of strokes, calculate the ASM of each stroke, and determine the strokes corresponding to the maximum and minimum values; S6-4: Take the inner contour of the stroke corresponding to the maximum value, and record the position of all pixels on the same side of the stroke in the character according to the axis of the stroke. Exclude the points where "the corresponding pixel in the character is black on both sides perpendicular to the axis direction"; turn the corresponding pixel in the character white to complete the thinning operation. S6-5: Take the outer contour of the stroke corresponding to the minimum value, and record the position of all pixels on the same side of the stroke axis in the character according to the axis axis; turn the corresponding pixels in the character into black to complete the thickening operation.

[0033] The character group after embedding the watermark is denoted as and characters Replace with Repeat the operation until All watermark bits are embedded.

[0034] Step Sa-5: Generate a text image with an embedded watermark.

[0035] In addition, the watermark extraction process of this invention is as follows: Step Sb-1: Correct the orientation of the printed and scanned image and remove isolated noise. The image orientation correction method is as follows: first, detect the image tilt angle using an existing algorithm, then rotate the image by the corresponding angle, and then crop the four edges of the image until there are no black pixels at the edges (remove the black edges caused by rotation); remove isolated noise to ensure the accuracy of character segmentation.

[0036] Step Sb-2: converting the corrected image into a binary text image I w (Otsu threshold method is used to automatically determine the binary threshold).

[0037] Step Sb-3: each character in the text image I is segmented by the same projection method as when the watermark is embedded, denoted as The total number of characters is the same as when embedding.

[0038] Step Sb-4: screening the first 2L embeddable characters, extracting their ASM, and each two characters as a group denoted as , the ASM of the first character of each character group is denoted as , and the ASM of the last character is denoted as .

[0039] Step Sb-5: starting from the first character group, extract the watermark: if , the extracted watermark information is 0; if , the extracted watermark information is 1; until bit information is extracted and the process ends.

[0040] The present application is applicable to text images of different language types including but not limited to Chinese characters, English, etc., and is not limited by font type.

[0041] Embodiment 1: The technical scheme of the present application is described in detail below with specific examples I. Watermark embedding process (take embedding "101" 3-bit watermark information as an example, and the target document is a Chinese text image containing "digital watermark technology is important") Step Sa-1: convert the target document containing "digital watermark technology is important" into a picture, and use Otsu threshold method for binary processing to generate a binary text image I. At this time, the image presents a clear black and white text and white background effect, and the text part is black (pixel value 0) and the background is white (pixel value 1).

[0042] Step Sa-2: segment each character in the text image I by projection method.

[0043] Line information extraction: calculate the pixel sum of the binary text image I by line, and divide the region with a sum of 0 into a text area. For example, the pixel sum of the "number" line is 0, so it is determined that this line is a text line, and the uppermost boundary (top pixel position) and the lowermost boundary (bottom pixel position) of the line are obtained.

[0044] ​Column pixel sum and calculation of character block boundary determination: column pixel sum is calculated for the row where the "number" word is located, and the left and right boundaries of the "number" word character block are obtained based on character connectivity (left pixel position and right pixel position).

[0045] Punctuation and character judgment: the width and black pixel height of the character block are judged, and if the width is less than 0.4 times the row height and the black pixel height is less than 0.6 times the row height, it is determined to be a punctuation mark. Here, "Digital watermark technology is important" has no punctuation, so it is not discarded. The width of the "number" word character block is between 0.6-0.7 times the row height, and it is determined to be a Chinese character. Then the same operation is performed on the "word", "water", "print", "technology", "very", "important", "important" and "important" words, and finally the 9 characters are segmented, recorded as C1 (number), C2 (word), C3 (water), C4 (print), C5 (technology), C6 (technology), C7 (very), C8 (important), and C9 (important), and n=9 is the total number of characters.

[0046] Step Sa-3: It is known that 3-bit watermark information is to be embedded, i.e. L=3, so the first 2L=6 characters, i.e. C1 (number), C2 (word), C3 (water), C4 (print), C5 (technology), and C6 (technology), are selected, and their average skeleton mass ASM is extracted.

[0047] Taking character C1 (number) as an example, assuming its size is (M=30, N=25) (M is the number of character height pixels, and N is the number of character width pixels), the number of black pixels in it is counted ; its skeleton length Ls=40 (total length of skeleton pixel points) is extracted by a skeleton extraction method based on morphology (multiple morphological erosion and reconstruction operations), and then according to the formula , the ASM of C1 can be obtained as . Similarly, the ASM2 of C2 (word) is calculated as 2.8, the ASM3 of C3 (water) is calculated as , the ASM4 of C4 (print) is calculated as 2.5, the ASM5 of C5 (technology) is calculated as 3.5, and the ASM6 of C6 (technology) is calculated as 2.7.

[0048] Step Sa-4: Each two characters form a group, a total of , , three groups, and the watermark bits , (L=3 is the number of watermark bits) are embedded from the first group.

[0049] Group (embedding ): Group The ASM of the first character c1 is recorded as ASM11 = 3, the ASM of the tail character C2 is denoted as ASM 12 = 2.8.

[0050] Since the current watermark bit is 1, check ASM 11 ≥ ASM 12 -th (set th = 2), ASM 12 -th = 2.8 - 0.2 = 2.6, 3 > 2.6, so ASM 11 needs to be reduced 12 , and ASM 11 needs to be increased 12 , until ASM 11 < ASM 12 -th = 2.6.

[0051] Perform the operation of "reducing ASM 11 , increasing ASM 12 ": Step Si-1: flip the black pixels in C1 and the white pixels in C2, after preliminary adjustment, the SUM black of C1 becomes 110, and Ls becomes 42, ; the SUM black of C2 becomes 130, and Ls becomes 38, . Compare and , 2.62 < 3.22, the adjustment is completed. At this time becomes the after embedding the watermark (including the adjusted and ).

[0052] Group (embedding ): Group The ASM of the head character C3 is denoted as ASM 21 = 3.2, and the ASM of the tail character C4 is denoted as ASM 22 = 2.5.

[0053] Since the current watermark bit is 0, check ASM 21 ≤ ASM 22 + th = 2.5 + 0.2 = 2.7, 3.2 > 2.7, so ASM 21 needs to be increased 22 , and ASM 21 needs to be reduced 22 , until ASM 21 > ASM 22 + th = 2.7.

[0054] Perform the operation of "increasing ASM 21 , reducing ASM 22 ": Step Sii-1: flip the white pixel in C3 and the black pixel in C4, after the preliminary adjustment, SUM of C3 black becomes 140, Ls becomes 39, ; SUM of C4 black becomes 115, Ls becomes 45, . Compare with , 3.59 > 2.76, adjustment is completed. At this time becomes the after embedding the watermark (including the adjusted and ).

[0055] Group (embedding ): Group The first character C5 is recorded as ASM 31 = 3.5, and the last character C6 is recorded as ASM 32 = 2.7.

[0056] Because the current watermark bit is 1, check ASM 31 ≥ ASM 32 -th = 2.7 - 0.2 = 2.5, 3.5 ≥ 2.5, so ASM 31 needs to be reduced and ASM 32 needs to be increased until ASM

[0057] Perform the "reduce ASM 31 , increase ASM 32 " operation: Step Si-1: flip the black pixel in C5 and the white pixel in C6, after the preliminary adjustment, SUM of C5 black becomes 125, Ls becomes 50, ; SUM of C6 black becomes 135, Ls becomes 36, . Compare with , 2.5 < 3.55, adjustment is completed. At this time becomes the after embedding the watermark (including the adjusted and ).

[0058] Replace the character group after embedding the watermark with the original character group , and the embedding of 3 watermark bits is completed.

[0059] Step Sa-5: Generate the watermarked text image, at this time, the "number", "word", "water", "print", "technology", and "art" characters in the image have been adjusted according to the "ASM" of the strokes of the watermarked information.

[0060] II. Watermark extraction process (extracting the watermark after printing and scanning the above-mentioned watermarked text image) Step Sb-1: Correct the direction of the printed and scanned image, and remove isolated noise points.

[0061] Direction correction: The existing algorithm detects that the image is tilted by 2 degrees, rotates the image by 2 degrees, and then cuts the four edges of the image until there are no black pixel points on the edges, and removes the black edges generated by rotation.

[0062] Remove isolated noise points: Use the connected domain analysis method to set the black area (isolated noise points) with an area less than 5 pixels to white, to ensure the accuracy of character segmentation.

[0063] Step Sb-2: Convert the corrected image into a binary text image I using the Otsu threshold method w .

[0064] Step Sb-3: Divide each character in the text image I by the same projection method as when embedding the watermark, denoted as 、 、 、 、 、 、 、 、 , n=9 is the total number of characters (the same as when embedding).

[0065] Step Sb-4: Select the first 2L=6 characters that can be embedded , extract their "ASM", and each two characters form a group denoted as , the "ASM" of the first character of each character group is denoted as , and the "ASM" of the last character is denoted as .

[0066] Calculate , ; , ; , .

[0067] Step Sb-5: Extract the watermark from the first group of characters: Group , and the extracted watermark information is 1.

[0068] Group The extracted watermark information is 0.

[0069] group The extracted watermark information is 1.

[0070] Finally, 3-bit information "101" is extracted, which is consistent with the embedded watermark information, and the extraction is completed.

[0071] The present application is applicable to text images of different language types including but not limited to Chinese characters, English, etc., and is not limited by font types.

[0072] Finally, it should be noted that the present application is not limited to the above embodiments, and there can be many variations. All variations that can be directly derived or inferred from the disclosed content by those of ordinary skill in the art should be considered within the scope of the present application.

Claims

1. A method of counter-print-scan digital watermarking for a text image, characterized in that The method comprises a watermark embedding process and a watermark extracting process; The watermark embedding process comprises the following steps: Sa-1: converting the target document into a picture and performing a binaryzation process to generate a binaryzation text image I; Sa-2: segmenting each character in the text image I by a projection method; Sa-3: screening the first 2L characters and extracting the ASM thereof, L being the number of watermark bits; Sa-4: every two characters as a group, from the first character group to start the watermark bits w1, w2 until w L embedded; the first character of each character group ASM is recorded as ASM1, and the last character of each character group ASM is recorded as ASM2; if the current watermark bit is 0, the ASM1 of the group is increased, and the ASM2 is decreased, until ASM1> ASM2+th; if the current watermark bit is 1, the ASM1 of the group is decreased, and the ASM2 is increased, until ASM1< ASM2-th; the character group after embedding the watermark is recorded as , and the character is replaced with , and the operation is repeated until all L watermark bits are embedded; Sa-5: generating a text image with embedded watermark.

2. The method of claim 1, wherein, The watermark extracting process comprises the following steps: Sb-1: performing a direction correction on the image after printing and scanning and removing isolated noise points; Sb-2: converting the corrected image into a binarized text image I w ; Sb-3: binarized text image I is segmented into individual characters by the projection method w Sb-3: binarized text image I is segmented into individual characters by the projection method Sb-4: screening the first 2L embeddable characters and extracting the ASM thereof, two characters being a group; Sb-5: Extract the watermark from the 1st group of characters: if , the extracted watermark information is 0; if , the extracted watermark information is 1; until the extraction of bit information is finished. Wherein, the extraction algorithm of ASM is: assuming that the size of the current character is (M, N), M is the number of character height pixels, and N is the number of character width pixels, the number of black pixels in the character is counted ; the skeleton length of the character is extracted through the skeleton extraction algorithm , and the formula is: .

3. The method of claim 2, wherein, The projection method in step Sa-2 is an optimized projection method based on a traditional projection method and removing punctuation symbols, and the specific process of the optimized projection method is as follows: First, perform line information extraction and divide character blocks based on the character connected domain characteristics; for the binaryzation text image I, set the white background as 1 and the black text as 0; calculate the pixel sum by line, and divide the region with a sum of 0 into a text area, thereby obtaining the number of lines and the uppermost and lowermost boundaries of each line; Next, calculate the column pixel sum for each row, and obtain the left and right boundaries of each character block based on character connectivity; then determine the width and black pixel height of the character block. If the width is less than 0.4 times the row height and the black pixel height is less than 0.6 times the row height, it is determined to be a punctuation mark. The method to determine if a character block is a single character is: check the width of the character block; if the width is within... Between double row height and b double row height; If the value of b is 0.6-0.7 and the value of b is 1-1.1, then it is considered to be a Chinese character. After discarding the punctuation marks, we sequentially determine whether 3, 2, and 1 consecutive character blocks are a Chinese character, thereby determining the character position.

4. The method of claim 2, wherein, The ASM in steps Sa-3, Sa-4, Sb-4 and Sb-5 is the average skeleton quality, and the stroke contains the pixel width.

5. The method of claim 2, wherein, The embedding case in step Sa-4 is divided into the following four cases, assuming that the current character group is , and the details are as follows: When the watermark bit is 1, if ASM1≥ASM2-th, decrease ASM1 and increase ASM2 until ASM1<ASM2-th; When the watermark bit is 1, if ASM1<ASM2-th, keep it unchanged; When the watermark bit is 0, if ASM1≤ASM2+th, increase ASM1 and decrease ASM2 until ASM1>ASM2+th; When the watermark bit is 0, if ASM1>ASM2-th, keep it unchanged.

6. The method of claim 5, wherein, ASM adjustment of the current character group contains two operations: "decrease ASM1, increase ASM2" and "increase ASM1, decrease ASM2": The operation steps of "decreasing ASM1 and increasing ASM2" are as follows: Si-1: flip the black pixels in the first character and the white pixels in the tail character, preliminarily decrease ASM1 and increase ASM2, and compare the size of ASM1 and ASM2-th; if ASM1<ASM2-th, the adjustment is completed; Si-2: if ASM1≥ASM2-th, make the thickest stroke in the first character thinner and the thinnest stroke in the tail character thicker, and compare the size of ASM1 and ASM2-th again; if ASM1<ASM2-th, the adjustment is completed; Si-3: if ASM1≥ASM2-th, execute step Si-1 again; The operation steps of "increasing ASM1 and decreasing ASM2" are as follows: Sii-1: flip the white pixels in the first character and the black pixels in the tail character, preliminarily increase ASM1 and decrease ASM2, and compare the size of ASM1 and ASM2+th; if ASM1>ASM2+th, the adjustment is completed; Sii-2: if ASM1≤ASM2+th, make the thinnest stroke in the first character thicker and the thickest stroke in the tail character thinner, and compare the size of ASM1 and ASM2+th again; if ASM1>ASM2+th, the adjustment is completed; Sii-3: If ASM1≥ASM2-th, re-perform step Sii-1.

7. The method of claim 6, wherein, The process of adjusting the stroke thickness in steps Si-2 and Sii-2 includes: S6-1: The outer contour extraction of the character is denoted as ext_edge, the inner contour extraction is denoted as int_edge, and the skeleton of part of strokes is extracted in the skeleton by a straight line searching algorithm, denoted as ; S6-2: "and" operation is performed between the increased number of times and the inner and outer contours of the character, to obtain the inner and outer contours of the stroke, wherein the inner contour is denoted as , and the outer contour is denoted as . ; S6-3 : Calculate the "ASM" of each stroke after the stroke is multiplied by the number of times of the character itself, and determine the stroke corresponding to the maximum value and the minimum value. S6-3 : Calculate the "ASM" of each stroke after the stroke is multiplied by the number of times of the character itself, and determine the stroke corresponding to the maximum value and the minimum value. S6-4: Taking the inner contour of the stroke corresponding to the maximum value, recording the positions of all pixel points on the same side of the axial direction in the character according to the axial direction of the stroke, excluding the points that the pixel points corresponding to the positions in the character are black pixels on both sides perpendicular to the axial direction; changing the pixel points corresponding to the positions in the character to white, and completing the thinning operation; S6-5: Taking the outer contour of the stroke corresponding to the minimum value, recording the positions of all pixel points on the same side of the axial direction in the character according to the axial direction of the stroke; changing the pixel points corresponding to the positions in the character to black, and completing the thickening operation.

8. The method of claim 2, wherein, The method of image direction correction in step Sb-1 is: First, the existing algorithm is used to detect the image tilt angle, then the image is rotated by the corresponding angle, and then the four edges of the image are cut until there is no black pixel point on the edge, and the black edge generated by rotation is removed; isolated noise points are removed to ensure the accuracy of character segmentation.

9. The method of claim 2, wherein, The skeleton extraction algorithm is a morphological-based skeleton extraction method, which extracts the center skeleton of the character through multiple morphological erosion and reconstruction operations to obtain the skeleton length .

10. The method of claim 2, wherein, The binarization processing adopts Otsu threshold method to automatically determine the binarization threshold, and converts the target document into a black and white binary text image I.

Citation Information

Patent Citations

  • Watermark method capable of resisting printing scanning attack and based on character refinement

    CN103761700A

  • Image and text mixing digital watermark embedding and extracting method of resisting to printing and scanning

    CN103985078A

  • Text digital watermark embedding / extracting method and device

    CN111738898A

  • Anti-print-scanning text image digital watermarking method

    CN112651879A

  • Text image printing and scanning resisting method based on pixel invariance

    CN115239605A

Cited By

  • Document file digital watermark generation method, extraction method and electronic device

    CN122636390A