Root-based English Electronic Document Watermark Embedding and Extraction Method and System
Through the embedded and extracting method of English electronic document watermarks based on roots, the problem that the watermarks of Chinese and English electronic document watermarks are easily destroyed after taking pictures and screenshots in the prior art are solved, and the integrity and anti-destructiveness of watermark information are achieved.
Patent Information
- Application Number
- CN202211020501.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-24
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2042-08-24
AI Technical Summary
The prior art cannot ensure the integrity of watermark information after English electronic documents are taken or screenshot.
The watermark embedding and extraction method of English electronic document based on roots is adopted, and the watermark information is ensured that the watermark information remains complete after operation by replacing font files, word segmentation processing, root extraction and encoding, and watermark information embedding.
After taking photos, screenshots and other operations, the integrity of the watermark information in English electronic documents can be ensured, and the anti-destruction ability of the watermark can be improved.
Smart Images

Figure CN115408669B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of electronic document security management, and particularly relates to a method and system for embedding and extracting watermarks for English electronic documents based on word roots. Background Art
[0002] As a means of copyright protection, watermarks can embed information representing a specific identity in a document. Currently, for the method of embedding watermarks in English documents, it cannot resist the damage to the watermark caused by means such as photographing, screenshotting, printing, etc.
[0003] Therefore, it is very important to design a method and system for embedding and extracting watermarks for English electronic documents based on word roots that can still ensure the integrity of watermark information after operations such as photographing and screenshotting.
[0004] For example, a method for embedding and extracting watermarks in an English text described in a Chinese patent document with the application number CN200510077471.3 converts the copyright information of the copyright owner into a binary bit string; reads the text, filters out spaces and special characters, performs a hash operation on the obtained string and the private key of the copyright owner to obtain an integer Z; if Z is divisible by the embedding ratio, the next sentence is the watermark information sentence; takes the remainder of Z divided by the length of the copyright information bit string to determine the watermark information bit to be embedded therein; takes the remainder of Z divided by the number of characters in the watermark information sentence to determine the position of the watermark information bit, so that the size relationship of the encodings of two adjacent letters at the position represents 0 and 1, which is the same as the watermark information bit to be embedded, until the end of the text. The extraction of watermark information is the reverse process of the embedding process. Although the watermark has good concealment and high security, especially has complete anti-attack ability against format conversion attacks, and the text will not degrade in quality due to the existence of watermark information, its disadvantage is that this method reads the first sentence of the English text, filters out spaces and special characters to obtain a string containing only English characters, and performs a separate hash operation on the string and the private key information to achieve the purpose of embedding the watermark, and has certain limitations in ensuring the integrity of the watermark. Summary of the Invention
[0005] The present invention is to overcome the problem in the prior art that after embedding watermarks in English electronic documents, it cannot resist the damage to watermark information caused by operations such as photographing and screenshotting of the document, and provides a method and system for embedding and extracting watermarks for English electronic documents based on word roots that can still ensure the integrity of watermark information after operations such as photographing and screenshotting.
[0006] To achieve the above invention purpose, the present invention adopts the following technical solutions:
[0007] A method for embedding and extracting watermarks for English electronic documents based on word roots, including watermark embedding and watermark extraction;
[0008] The watermark embedding includes the following steps:
[0009] S1, replace the font file:
[0010] The original computer English font file is replaced with a new English font file through deformation processing;
[0011] S2, word segmentation processing:
[0012] Perform word segmentation processing on the new English font file to generate a set of words containing the entire document;
[0013] S3, root word extraction:
[0014] For all words in the set of words generated in step S2, extract the root words contained in the words;
[0015] S4, root word encoding:
[0016] Encode the extracted root words to carry the bit information "0" or "1";
[0017] S5, watermark information embedding:
[0018] Use the root words that have been encoded in step S4 to replace the original root words of the English words in the English font file to complete the embedding of the watermark information;
[0019] The watermark extraction includes the following steps:
[0020] S6, image acquisition:
[0021] Obtain the text image of the English electronic document from which the watermark is to be extracted;
[0022] S7, text recognition:
[0023] Perform text recognition processing on the text image of the English electronic document to obtain all the English words contained in the image;
[0024] S8, image processing:
[0025] Perform image processing on the text image of the English electronic document to obtain the position information of each word in the image, and correspond the words obtained in step S7 with the position information one by one;
[0026] S9, root word extraction:
[0027] Extract the root words of the words in step S7 to obtain the root words contained in each word;
[0028] S10, root word matching:
[0029] Perform segmentation processing on the word pictures obtained in step S7 to obtain root pictures, match them with the encoded root pictures, and extract the carried bit information;
[0030] S11, extract the watermark:
[0031] Find the start code and end code, and extract the watermark information of the English electronic document.
[0032] Preferably, step S1 includes the following steps:
[0033] S11, obtain the computer English font file, modify the structure of the English letters, and generate deformed English letters;
[0034] Among them, each English letter has two glyphs;
[0035] S12, re-encode all the deformed English letters, and generate a new English font file to replace the English font file of the terminal.
[0036] Preferably, step S2 includes the following steps:
[0037] S21, use spaces as delimiters to extract an English electronic document into a word set, remove the punctuation marks in the document at the same time, and do not process the repeated English words in the document.
[0038] Preferably, step S3 includes the following steps:
[0039] S31, perform root extraction processing on the words obtained after word segmentation;
[0040] The roots are divided into the following three cases:
[0041] The first case, when the number of roots contained in a word is greater than 1, query the root frequency table and extract the root with the highest frequency as the root of the word;
[0042] The root frequency table is a table formed by counting the number of times the root appears in the corpus and arranging it in descending order of root usage frequency;
[0043] The second case, when a root contains another root, select the root with the longest length as the root of the word;
[0044] The third case, for words that do not contain roots, no root extraction processing is performed.
[0045] S32, form a set of the roots after root extraction.
[0046] Preferably, step S4 includes the following steps:
[0047] S41. For a root word with an odd number of English letters, according to the root word letter usage frequency table, select the letters with higher usage frequencies. The number of selected letters is half of the number of root word letters, and perform encoding;
[0048] Among them, each root word corresponds to 1-bit information in the binary string; the letter usage frequency table is a table formed by counting the number of occurrences of English letters in the corpus and arranging them in descending order of English letter usage frequency.
[0049] Preferably, in step S6, the picture acquisition method is to take a photo or scan or capture a screen or print the text picture of the English electronic document.
[0050] Preferably, step S7 includes the following steps:
[0051] S71. Use the PaddleOCR text recognition technology to perform text recognition processing on the text picture of the English electronic document to obtain all the English words contained in the picture.
[0052] Preferably, step S8 includes the following steps:
[0053] S81. Use the segmentation algorithm based on edge detection to perform character segmentation on the text picture of the English electronic document, obtain the position information of each word in the picture, cut out rectangular word image blocks according to the position information, and make the rectangular word image blocks correspond one by one to the words recognized by the text.
[0054] Preferably, step S10 includes the following steps:
[0055] S101. Cut out the root word image block according to the position information of the root word, perform image matching on the root word image block and the encoded standard root word image, and extract the carried bit information.
[0056] The present invention also provides a watermark embedding and extraction system for English electronic documents based on root words, including:
[0057] The letter deformation module is used to modify the structures of 52 English letters to generate deformed English letters capable of carrying information;
[0058] The font file replacement module is used to generate a new English font file with the deformed English letters and automatically replace the English font file in the terminal;
[0059] The word segmentation processing module is used to perform word segmentation on the new English font file to generate a word set containing the entire document;
[0060] The first root word extraction module is used to perform root word extraction processing on all the English words obtained by the word segmentation processing module;
[0061] The root word encoding module is used to encode root words, select the first variant letter to represent 0, and select the second variant letter to represent 1;
[0062] The root word replacement module is used to replace the root words in the English font file to complete the embedding of watermark information;
[0063] The text recognition module is used to perform text recognition processing on the text picture of the English electronic document to obtain all the English words contained in the picture;
[0064] The image processing module is used to perform image processing on the text picture of the English electronic document, obtain the position information of each word in the picture, and correspond the words obtained in the text recognition module with the position information one by one;
[0065] The second root word extraction module is used to perform root word extraction processing on all the English words obtained in the text recognition module;
[0066] The root word matching module is used to perform segmentation processing on the word pictures obtained in the text recognition module to obtain root word pictures, match them with the encoded root word pictures, and extract the carried bit information;
[0067] The watermark extraction module is used to extract the watermark information contained in the English electronic document.
[0068] Compared with the prior art, the beneficial effects of the present invention are: (1) After the document is segmented in the present invention, the root words of the words are extracted, and the root words are encoded according to the root word frequency table and the letter usage frequency table to complete the embedding of watermark information, and the process is simple and convenient; (2) Using the method provided by the present invention, the integrity of the watermark information of the English electronic document can still be ensured after operations such as taking pictures and screenshots. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] Figure 1 It is a flowchart of a watermark embedding process provided by an embodiment of the present invention;
[0070] Figure 2 It is a schematic diagram of a system module of a watermark embedding process provided by an embodiment of the present invention;
[0071] Figure 3 It is a flowchart of a watermark extraction process provided by an embodiment of the present invention;
[0072] Figure 4 It is a schematic diagram of a system module of a watermark extraction process provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0073] To more clearly illustrate the embodiments of the present invention, the specific implementation manners of the present invention will be described below with reference to the accompanying drawings. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings, and other implementation manners can also be obtained.
[0074] Embodiment:
[0075] As Figure 1 and Figure 3 shown, the present invention provides a method for embedding and extracting English electronic document watermarks based on word roots, including watermark embedding and watermark extraction:
[0076] The watermark embedding process of the present invention is carried out according to the following steps:
[0077] Step 101, replace the font file. Obtain the English font file of the computer, modify the structure of the English letters to generate deformed English letters, and each English letter has two glyphs. For example, for the English letter "f", move the position of the horizontal line in the "f" letter to generate two different English letters "f", denoted as f1 and f2. Re-encode all the deformed English letters. For example, f1 is encoded as CEF0 and f2 is encoded as CEF1, and at the same time, generate a new English font file to replace the English font file of the terminal.
[0078] Step 102, word segmentation processing. Using a space as the delimiter, extract an English electronic document into a set of words, and at the same time remove the punctuation marks in the document. For the English words that appear repeatedly in the document, no processing is done. For example, after the word segmentation processing of "Love is acarefully designed lie", the generated set of words is {"Love", "is", "a", "carefully", "designed", "lie"}.
[0079] Step 103, word root extraction. Perform word root extraction operations on the words obtained after the word segmentation processing. The word roots can be divided into three cases:
[0080] The first case is that the number of word roots contained in a word may be greater than 1. For example, in "commercial", it contains two word roots, "com-" and "-cial". By querying the word root frequency table (a table formed by counting the number of times a word root appears in a corpus and arranging them in descending order of the usage frequency of the word roots, which is called the word root frequency table), the frequency of "com-" is greater than the frequency of "-cial", so "com-" is selected as the word root of "commercial".
[0081] The second case: One root contains another root. For example, in the root "pre-", it contains the root "re-". In this case, select the root with the longest length. For example, the length of "pre-" is greater than that of "re", so select "pre-" as the root of the word.
[0082] The third case: For those without roots, such as "a", "is", no root extraction is performed in this case.
[0083] After root extraction, the roots form a map set, {"commercial": "com", "i s": "", "prefer": "pre",...}.
[0084] Step 104, root encoding. For roots with an odd number of English letters, such as "com", c1 represents the first deformed letter, and c2 represents the second deformed letter. According to the root letter usage frequency table (a table formed by counting the number of occurrences of English letters in the corpus and arranging them in descending order of usage frequency of English letters, called the letter usage frequency table), select the letters with higher usage frequency, and the number is half of the number of root letters for encoding. For example, for "prep", the usage frequencies of the letters "r" and "e" are higher than that of "p", and the number of root letters is 4. Then select "r" and "e" as the encoding combination. p2r1e1p2 represents the bit information 0, and p1r2e2p1 represents the bit information 1. Among them, each root corresponds to 1 bit of information in the binary string.
[0085] Step 105, replace the word roots in the original document with the encoded roots to complete the embedding of the watermark information.
[0086] The watermark extraction process in the present invention is carried out according to the following steps:
[0087] Step 301, picture acquisition. Obtain the text picture of the English electronic document from which the watermark is to be extracted. The word roots in the text picture of the English words carry the watermark information. The way to obtain the text picture can be to take a photo, scan, take a screenshot, print, etc. of the English electronic document file and input it into the system.
[0088] Step 302, text recognition. Use text recognition technologies such as PaddleOCR to perform text recognition processing on the text picture of the electronic document, and obtain all the English words contained in the picture.
[0089] Step 303, Image processing. Using a segmentation algorithm based on edge detection, perform character segmentation on the text image to obtain the position information of each word in the image. Cut out rectangular word image blocks according to the position information, and make the rectangular word image blocks correspond one by one to the words recognized by the text. The specific method is as follows: Select the upper left point of the image as the anchor point, and select the upper left point A(x1, y1) and the lower right point B(x2, y2) of the rectangle as the position information of the image block. x is the horizontal distance from the anchor point, and y is the vertical distance from the anchor point. If the position information of the anchor point is (0, 0), the horizontal distance and vertical distance of the upper left point A of the word "Love" from the anchor point are 300 and 400 respectively, and the horizontal distance and vertical distance of the lower right point B from the anchor point are 316 and 430 respectively, then the position information of the word "Love" is recorded as (300, 400) and (316, 430).
[0090] Step 304, Root extraction. The same as step 103, extract the root of each word.
[0091] Step 305, Root matching. Cut out the root image block according to the position information of the root, perform image matching on the root image block and the encoded standard root image, and extract the bit information carried by it. If it is similar to the first type, the extracted bit information is "0", and if it is similar to the second type, the extracted bit information is "1".
[0092] Step 306, Extract watermark. According to the bit information extracted in step 305, form a bit information string and extract the watermark information of the document.
[0093] As Figure 2 and Figure 4 shown, the present invention also provides a system for embedding and extracting watermarks for English electronic documents based on roots, including:
[0094] An alphabet deformation module, used to modify the structures of 52 English letters to generate deformed English letters capable of carrying information;
[0095] A font file replacement module, used to generate a new English font file from the deformed English letters and automatically replace the English font file in the terminal;
[0096] A word segmentation processing module, used to perform word segmentation processing on the new English font file to generate a word set containing the entire document;
[0097] A first root extraction module, used to perform root extraction processing on all English words obtained in the word segmentation processing module;
[0098] A root encoding module, used to encode the roots, select the first type of deformed letter to represent 0, and select the second type of deformed letter to represent 1;
[0099] A root word replacement module, which is used to replace the root words of words in an English font file to complete the embedding of watermark information;
[0100] A text recognition module, which is used to perform text recognition processing on the text picture of the English electronic document to obtain all the English words contained in the picture;
[0101] An image processing module, which is used to perform image processing on the text picture of the English electronic document to obtain the position information of each word in the picture, and to correspond the words obtained in the text recognition module with the position information one by one;
[0102] A second root word extraction module, which is used to perform root word extraction processing on all the English words obtained in the text recognition module;
[0103] A root word matching module, which is used to perform segmentation processing on the word pictures obtained in the text recognition module to obtain root word pictures, match them with the encoded root word pictures, and extract the carried bit information;
[0104] A watermark extraction module, which is used to extract the watermark information contained in the English electronic document.
[0105] After the document is segmented in the present invention, the root words of the words are extracted, and the root words are encoded according to the root word frequency table and the letter usage frequency table to complete the embedding of the watermark information. The process is simple and convenient; by using the method provided by the present invention, the integrity of the watermark information of the English electronic document can still be ensured after operations such as taking pictures and screenshots.
[0106] The above is only a detailed description of the preferred embodiments and principles of the present invention. For those of ordinary skill in the art, according to the idea provided by the present invention, there will be changes in the specific implementation manners, and these changes should also be regarded as the protection scope of the present invention.
Claims
1. A method for embedding and extracting English electronic document watermarks based on word roots, characterized in that, It includes watermark embedding and watermark extraction; The watermark embedding includes the following steps: S1, Replace the font file: The original computer English font file is processed by deformation and replaced with a new English font file; S2, Word segmentation processing: Perform word segmentation processing on the new English font file to generate a word set containing the entire document; S3, Root extraction: For all words in the word set generated in step S2, extract the roots contained in the words; S4, Root encoding: Encode the extracted roots to carry the bit information "0" or "1"; S5, Watermark information embedding: Use the roots that have been encoded in step S4 to replace the original roots of English words in the English font file to complete the embedding of watermark information; The watermark extraction includes the following steps: S6, Image acquisition: Obtain the text image of the English electronic document from which the watermark is to be extracted; S7, Text recognition: Perform text recognition processing on the text image of the English electronic document to obtain all English words contained in the image; S8, Image processing: Perform image processing on the text image of the English electronic document to obtain the position information of each word in the image, and correspond the words obtained in step S7 with the position information one by one; S9, Root extraction: Extract the roots of the words in step S7 to obtain the roots contained in each word; S10, Root matching: Perform segmentation processing on the word images obtained in step S7 to obtain root images, match them with the encoded root images, and extract the carried bit information; S11, Extract watermark: Find the start code and end code, and extract the watermark information of the English electronic document; Step S3 includes the following steps: S31, Perform root extraction processing on the words obtained after word segmentation processing; The roots are divided into the following three cases: The first case, when the number of roots contained in a word is greater than 1, extract the root with the highest frequency as the root of the word by querying the root frequency table; The root frequency table is a table formed by counting the number of times the roots appear in the corpus and arranging them in descending order of root usage frequency; The second case, when a root contains another root, select the root with the longest length as the root of the word; The third case, for words that do not contain roots, no root extraction processing is performed; S32, Form a set of the roots after root extraction; Step S4 includes the following steps: S41, For roots with an odd number of English letters, according to the letter usage frequency table, select the letters with higher usage frequency, and the number of selected letters is half of the number of root letters, and perform encoding; Among them, each root corresponds to 1 bit of information in the binary string; the letter usage frequency table is a table formed by counting the number of times the English letters appear in the corpus and arranging them in descending order of English letter usage frequency.
2. The method for embedding and extracting watermarks in English electronic documents based on word roots according to claim 1, characterized in that Step S1 includes the following steps: S11, Obtain the computer English font file, modify the structure of the English letters to generate deformed English letters; Among them, each English letter has two glyphs; S12, Re-encode all the deformed English letters and generate a new English font file to replace the English font file of the terminal.
3. The method for embedding and extracting watermarks in English electronic documents based on word roots according to claim 2, characterized in that Step S2 includes the following steps: S21: Split an English electronic document into a set of words using spaces as delimiters, removing punctuation marks from the document. For repeated English words in the document, no processing is done.
4. The method for embedding and extracting watermarks in English electronic documents based on word roots according to claim 1, characterized in that In step S6, the method for obtaining images is to take a photo of, scan, capture a screen shot of, or print the text image of the English electronic document.
5. The method for embedding and extracting watermarks in English electronic documents based on word roots according to claim 1, characterized in that Step S7 includes the following steps: S71: Use PaddleOCR text recognition technology to perform text recognition on the text image of the English electronic document to obtain all English words contained in the image.
6. The method for embedding and extracting watermarks in English electronic documents based on word roots according to claim 5, characterized in that Step S8 includes the following steps: S81: Use a segmentation algorithm based on edge detection to perform character segmentation on the text image of the English electronic document, obtain the position information of each word in the image, cut out rectangular word image blocks according to the position information, and make the rectangular word image blocks correspond one by one to the words recognized by text recognition.
7. The method for embedding and extracting watermarks in English electronic documents based on word roots according to claim 6, characterized in that Step S10 includes the following steps: S101: Cut out the root image block according to the position information of the root, perform image matching between the root image block and the encoded standard root image, and extract the carried bit information.
8. A system for embedding and extracting watermarks in English electronic documents based on word roots, used to implement the method for embedding and extracting watermarks in English electronic documents based on word roots described in any one of claims 1-7, characterized in that The system for watermark embedding and extraction of an English electronic document based on roots includes: An alphabet deformation module for modifying the structures of 52 English letters to generate deformed English letters capable of carrying information. A font file replacement module for generating a new English font file from the deformed English letters and automatically replacing the English font file in the terminal. A word segmentation processing module for performing word segmentation on the new English font file to generate a set of words containing the entire document. A first root extraction module for performing root extraction processing on all English words obtained from the word segmentation processing module. A root encoding module for encoding the roots, selecting the first type of deformed letter to represent 0 and the second type of deformed letter to represent 1. A root replacement module for replacing the word roots in the English font file to complete the embedding of watermark information. A text recognition module for performing text recognition on the text image of the English electronic document to obtain all English words contained in the image. An image processing module for performing image processing on the text image of the English electronic document to obtain the position information of each word in the image, and making the words obtained in the text recognition module correspond one by one to the position information. A second root extraction module for performing root extraction processing on all English words obtained from the text recognition module. A root matching module for performing segmentation processing on the word images obtained in the text recognition module to obtain root images, performing matching with the encoded root images, and extracting the carried bit information. A watermark extraction module for extracting the watermark information contained in the English electronic document.
Citation Information
Patent Citations
Method for embedding and extracting watermark in English texts
CN100367274C