Character grouping method based on binary character frequencies and security character library construction method
Through grouping methods and deformation design based on binary character frequency, the problem of unbalanced grouping of safe font library in the prior art is solved, the rationality of character grouping and the reliability of safe font library are improved, and the robustness and applicability of watermark algorithms are enhanced.
Patent Information
- Application Number
- CN202211416943.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-14
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-11-14
AI Technical Summary
The grouping method of secure font library in the prior art has problems such as uneven character count, insufficient word frequency, and complex and time-consuming calculations, resulting in low robust performance and poor applicability of the watermark algorithm.
A grouping method based on binary character frequency is adopted, by calculating the binary character frequency matrix and weight allocation, characters are reasonably allocated to different groups, and deformed design is carried out in combination with word frequency sorting to build a safe font library.
It realizes the rationality of character grouping and the reliability of the secure font library, reduces the number of characters required to extract security codes, and improves the robustness and applicability of the watermark algorithm.
Smart Images

Figure CN115618809B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of font invisible watermarking, and particularly relates to a character grouping method based on binary character frequencies and a secure font library construction method. Background Art
[0002] In existing text watermarking technologies, in order to improve the robust performance of watermark algorithms against malicious attacks such as printing and scanning, screen capture, and screen photography, text digital watermarking technologies based on character topology modification have become mainstream. That is, by deforming specific characters in different forms to correspond to different watermark information bit strings, the character deformation data is stored in a specific watermark font library, and during the printing output and screen display of electronic text documents, watermark information is embedded through font replacement. When different character deformation data is used for different users, for that user, the specific watermark font library constitutes their secure font library.
[0003] Existing secure font libraries have many defects. In order to solve the problems of poor generality of watermark loading, poor system stability, complex implementation process, and low robust performance of watermark algorithms in the prior art without changing any user usage habits, the patent "A General Text Watermarking Method and Device" (Publication No.: CN114708133A) applied by Beijing Guoyin Technology Co., Ltd. discloses the following solution: A general text watermarking method includes the following steps: grouping a certain number of characters in a selected font library according to a specific strategy; designing deformations for all characters in each group according to specific rules, and generating a temporary watermark character data file; generating user terminal watermark encoding data for identifying the identity authentication information of the user terminal; dynamically generating and real-time loading a watermark font library file based on the watermark encoding data, in combination with the temporary watermark character data file and the grouped characters; running an electronic format text file, and embedding watermark information in real time in the document content data during file printing output and screen display using the watermark font library file.
[0004] In this solution, it is necessary to group characters. When grouping characters, theoretically, characters with higher character frequencies should be in different groups; characters that often appear together should be in different groups. The security character library generated by meeting these two requirements requires less text content when extracting security codes. Therefore, the extraction effect and accuracy are better. There are many deficiencies in the character grouping method in this solution: First, the number of characters in each group is basically equal, which conflicts with the above requirements; Second, only the character frequency is considered during grouping, and the word frequency is not fully considered. Theoretically, the characters corresponding to frequently occurring words should be in different groups, so that more groups can appear in a shorter content and less content is required when extracting security codes; Third, the calculation process for optimizing the grouping in this solution is too complex and consumes a large amount of time and computing power. Summary of the Invention
[0005] The primary object of the present invention is to provide a character grouping method and a security character library construction method based on binary character frequencies, which can group characters more reasonably.
[0006] To achieve the above object, the technical solution adopted by the present invention is: A character grouping method based on binary character frequencies, including the following steps: Traverse the corpus, and count the occurrence times of any two characters among the N characters to be grouped to obtain a binary character frequency matrix , the elements of the binary character frequency matrix represent the frequency of character followed by character ; Traverse the characters one by one from high to low according to the character frequency, and calculate the weight of the character c to be assigned to the kth group according to the following formula:
[0007]
[0008] In the formula, A is the set composed of the grouped characters and the character c to be assigned, and are constants greater than 0 and ; Add the character c to be assigned to the group with the maximum weight, and so on until all characters are grouped.
[0009] Compared with the prior art, the present invention has the following technical effects: This grouping scheme mainly groups characters based on the association between binary characters. For two characters that often appear together, they are preferably assigned to different groups. The binary character frequency matrix reflects the frequency of two characters appearing together. Then, through the weight calculation formula, the weight increases when two characters that often appear together are in different groups. In this way, we can make the characters that appear together in different groups as much as possible by selecting the group with the largest weight, thus achieving a reasonable grouping of characters. This grouping method does not limit the number of characters in each group, making it more reasonable.
[0010] The second object of the present invention is to provide a method for constructing a secure font library based on the above character grouping method, so as to improve the applicability and reliability of the secure font library.
[0011] To achieve the above object, the first technical solution adopted by the present invention is: A method for constructing a secure font library includes the following steps: Select the top N characters according to the word frequency ranking, perform deformation design on each of the N characters to obtain deformed characters, and the standard character and the deformed character of each character represent 0 and 1 respectively; Divide the N characters into K groups according to the steps in claim 1, and each character belongs to only one group, where K is the number of bits of the binary string encoded by the security code represented by the secure font library; For any security code, select the standard character or the deformed character corresponding to the character according to the binary number corresponding to the group where each character is located. The standard characters or deformed characters of the selected N characters and the standard characters of the other unselected characters constitute the secure font library corresponding to the security code.
[0012] To achieve the above object, the second technical solution adopted by the present invention is: A method for constructing a secure font library includes the following steps: Select the top N characters according to the word frequency ranking, and perform deformation design on each of the N characters to obtain deformed characters; Perform binary encoding on the standard character and its deformed character of each character, and the number of bits x of the binary encoding and the number of deformed characters of the character Satisfy the following formula: ; Divide the N characters into K groups according to the steps in claim 1, and the number of groups where each character is located is equal to the number of bits x of the binary encoding corresponding to the character, where K is the number of bits of the binary string encoded by the security code represented by the secure font library; For any security code, select the standard character or the deformed character corresponding to the character according to the binary number corresponding to the group where each character is located as the binary encoding. The standard characters or deformed characters of the selected N characters and the standard characters of the other unselected characters constitute the secure font library corresponding to the security code.
[0013] Compared with the prior art, the above two security character library construction methods have the following technical effects: since the above character grouping is more reasonable, the security character library constructed based on the above character grouping method is bound to be more reliable. At the same time, the security character library composed of each character placed in a certain group has higher reliability; the security character library composed of the above single characters placed in multiple groups requires fewer characters on average when extracting security codes, and is more applicable. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 is a flow chart of the character grouping method of the present invention;
[0015] Figure 2 is a flow chart of a first embodiment of a method for constructing a security character library in the present invention;
[0016] Figure 3 It is a flow chart of Embodiment 2 of the method for constructing a secure character library in the present invention. DETAILED DESCRIPTION
[0017] Combine the following Figures 1 to 3 , the present invention is further described in detail.
[0018] See also Figure 1 The present invention discloses a character grouping method based on binary character frequency, comprising the following steps: traversing a corpus, counting the number of occurrences of any two characters in N characters to be grouped, and obtaining a binary character frequency matrix , the binary character frequency matrix Elements Representative characters After the character The frequency of binary characters reflects the frequency of characters. After the character The frequency of characters is not considered here. and characters Whether it belongs to a certain word is only considered from the position relationship. For example, for a sentence "Hefei Gaowei Data Technology Co., Ltd. is a new type of network security company, founded in May 2014 by famous professors and alumni of the School of Network Security of University of Science and Technology of China", the frequency of "合" followed by "肥" is recorded as 1, the frequency of "肥" followed by "高" is recorded as 1, and so on. Even if "肥高" is not a word, it still needs to be counted. When doing specific statistics, you can ignore the symbol. In this case, for "司,由", the frequency of "司" followed by "由" is increased by 1; you can also consider the symbol. In this case, for "司,由", the character "司" is followed by a comma, and the character "由" is preceded by a comma, so there is no need to increase the frequency of "司" followed by "由".
[0019] The selection of the corpus can also be based on the user's needs, that is, a general corpus can be selected, or an internal corpus of a certain enterprise or organization can be selected. For different corpora, the resulting character groupings are also different.
[0020] Traverse the characters one by one from highest to lowest according to the character frequency, and calculate the weight of the character c to be assigned to the k-th group according to the following formula:
[0021]
[0022] In the formula, A is the set composed of the grouped characters and the character c to be assigned, and are constants greater than 0 and . By analyzing the formula of the weight , we know that when the number of grouped characters is less than the number of groups, assign the character to be assigned to the empty group, is the largest; when there are characters assigned in all groups, when two characters that often appear together are assigned to different groups, is the largest; and both of these situations are the desired assignment results. Therefore, we only need to add the character c to be assigned to the group with the largest weight to put the character into the best grouping, and so on until all characters are grouped. The introduction of the weight converts the grouping problem into specific calculations and judgments, making the problem easier to solve, which is very convenient.
[0023] This grouping scheme mainly groups characters based on the association between binary characters. For two characters that often appear together, try to assign them to different groups. The binary character frequency matrix reflects the frequency of two characters appearing together. Then, through the weight calculation formula, the weight increases when two characters that often appear together are assigned to different groups. In this way, we can make the characters that appear together in different groups as much as possible by selecting the group with the largest weight, thus achieving reasonable grouping of characters. This grouping method does not limit the number of characters in each group, making it more reasonable.
[0024] Since the statistical frequency may cause the values of some elements to be very large, which is not convenient for subsequent calculations. Preferably, in the present invention, before the step of traversing the characters one by one from highest to lowest according to the character frequency, the following steps are further included: normalizing the binary character frequency matrix to obtain the binary character frequency matrix ; the weight of the character c to be assigned to each group is calculated according to the following formula:
[0025]
[0026] According to the binary character frequency matrix to calculate the weights, the amount of calculation will be much smaller and it is more convenient to calculate.
[0027] Furthermore, for the binary character frequency matrix to be normalized to obtain the binary character frequency matrix in, normalization is performed through any of the following formulas:
[0028] Or .
[0029] The above two formulas represent two different normalization methods respectively. In the first formula, the sum of the probabilities of each row of the normalized binary character frequency matrix is 1, which means that the sum of the probabilities that a certain character is followed by other characters is 1; in the second formula, the sum of the probabilities of each column of the normalized binary character frequency matrix is 1, which means that the sum of the probabilities that a certain character is preceded by other characters is 1, or it can be understood as processing all characters in reverse. The specific normalization method can be selected according to actual needs, because the normalization operation does not affect the grouping result.
[0030] During calculation, there is a very low probability that there are multiple groups with the largest weights. Therefore, the following judgment logic is added in the present invention to ensure that each character can be reliably grouped. The step of adding the character c to be assigned to the group with the largest weight includes the following steps: if there is only one group with the largest weight, add the character c to be assigned to this group; if there are multiple groups with the largest weight, select the group with the fewest characters among all the groups with the largest weight; if there is only one group with the fewest characters, add the character c to be assigned to this group; if there are multiple groups with the fewest characters, randomly add the character c to any one of them.
[0031] And There are many values for the selection of. Preferably, in the present invention, the , . Such a selection is more conducive to calculating the difference between , and the size of can be conveniently judged through the difference. The specific process of how to optimize the calculation of is not the focus of this case and will not be elaborated here.
[0032] The present invention also discloses two methods for constructing a secure font library, which are as follows.
[0033] Refer to Figure 2, Embodiment 1, The method for constructing a secure font library includes the following steps: Select the top N characters according to the character frequency ranking, perform deformation design on each of the N characters to obtain deformed characters, and the standard character and the deformed character of each character represent 0 and 1 respectively; Divide the N characters into K groups according to the steps in Claim 1, and each character belongs to only one group, where K is the number of bits of the binary string encoded by the security code represented by the secure font library; For any security code, select the standard character or the deformed character corresponding to each character according to the binary number corresponding to the group where each character is located, and the standard characters of the selected N characters or deformed characters and the standard characters of the other unselected characters constitute the secure font library corresponding to this security code.
[0034] Refer to Figure 3 , Embodiment 2, The method for constructing a secure font library includes the following steps: Select the top N characters according to the character frequency ranking, and perform deformation design on each of the N characters to obtain deformed characters; Perform binary encoding on the standard character and its deformed character of each character, and the number of bits x of this binary encoding and the number of deformed characters of this character Satisfy the following formula: ; Divide the N characters into K groups according to the steps in Claim 1, and the number of groups where each character is located is equal to the number of bits x of the binary encoding corresponding to this character, where K is the number of bits of the binary string encoded by the security code represented by the secure font library; For any security code, select the standard character or the deformed character corresponding to each character according to the binary number corresponding to the group where each character is located as the binary encoding, and the standard characters of the selected N characters or deformed characters and the standard characters of the other unselected characters constitute the secure font library corresponding to this security code.
[0035] The basic principles of these two solutions for the method of constructing a secure font library are the same. The only difference is that in Embodiment 1, each character is only divided into one group; in Embodiment 2, characters with a high character frequency can be divided into multiple groups. Since the above character grouping is more reasonable, the secure font library constructed based on the above character grouping method is necessarily more reliable. At the same time, the secure font library formed by placing each character in one group has higher reliability; for the secure font library formed by placing a single character in multiple groups, the average number of characters required when extracting the security code is less, and the applicability is stronger.
[0036] The present invention also discloses a computer-readable storage medium and an electronic device. Among them, a computer-readable storage medium has a computer program stored thereon, and when the computer program is executed by a processor, it implements the method for grouping characters based on binary character frequencies as described above or implements the method for constructing a secure font library as described above. An electronic device includes a memory, a processor, and a computer program stored on the memory. When the processor executes the computer program, it implements the method for grouping characters based on binary character frequencies as described above or implements the method for constructing a secure font library as described above.
Claims
1. A character grouping method based on the frequency of binary characters, characterized in that: It includes the following steps: Traverse the corpus and count the occurrence times of any two characters among the N characters to be grouped to obtain a binary character frequency matrix , the binary character frequency matrix of the elements represents the character followed by the character frequency; Traverse the characters one by one from the highest to the lowest character frequency, and calculate the weight of the character c to be assigned to the k-th group according to the following formula: Wherein, A is a set composed of the grouped characters and the character c to be assigned, and are constants greater than 0 and ; Add the character c to be assigned to the group with the largest weight, and so on until all characters are grouped.
2. The character grouping method based on binary character frequencies according to claim 1, characterized in that: Before the step of traversing characters one by one from high to low according to the character frequency, the following steps are further included: performing normalization on the binary character frequency matrix to obtain a binary character frequency matrix ; the weight of the character c to be assigned after being assigned to each group is calculated according to the following formula: 。 3. The character grouping method based on binary character frequencies according to claim 2, wherein: The aforementioned binary character frequency matrix is normalized to obtain a binary character frequency matrix by normalizing with any of the following formulas: or 。 4. The character grouping method based on the binary character frequency according to claim 1, wherein: The step of adding the character c to be assigned to the group with the largest weight includes the following steps: If there is only one group with the largest weight, add the character c to this group; If there are multiple groups with the largest weight, select the group with the fewest number of characters among all groups with the largest weight; If there is only one group with the fewest number of characters, add the character c to this group; If there are multiple groups with the fewest number of characters, randomly add the character c to any one of them.
5. The character grouping method based on binary character frequencies according to claim 1, wherein: The described , .
6. A method for constructing a secure font library, characterized in that: It includes the following steps: Select the first N characters according to the character frequency sorting, perform deformation design on each of the N characters to obtain deformed characters, and the standard character and the deformed character of each character represent 0 and 1 respectively; Divide the N characters into K groups according to the steps in claim 1, and each character belongs to only one group. K is the number of bits of the binary string encoded by the security code represented by the security character library; For any security code, select the standard character or the deformed character corresponding to the character according to the binary number corresponding to the group where each character is located. The standard characters or deformed characters of the selected N characters and the standard characters of the other unselected characters constitute the security character library corresponding to the security code.
7. A method for constructing a secure font library, characterized in that: It includes the following steps: Select the first N characters according to the character frequency sorting, and perform deformation design on each of the N characters to obtain deformed characters; Perform binary encoding on the standard character and its deformed characters for each character. The number of bits x of the binary encoding and the number of deformed characters of the character Satisfy the following formula: ; Divide the N characters into K groups according to the steps in claim 1. The number of groups where each character is located is equal to the number of bits x of the binary code corresponding to the character. K is the number of bits of the binary string encoded by the security code represented by the security character library; For any security code, select the standard character or the deformed character corresponding to the character according to the binary number corresponding to the group where each character is located as the binary code. The standard characters or deformed characters of the selected N characters and the standard characters of the other unselected characters constitute the security character library corresponding to the security code.
8. A computer-readable storage medium, characterized in that: It stores a computer program, and when the computer program is executed by a processor, it implements the character grouping method based on binary character frequencies described in any one of claims 1-5 or implements the security character library construction method described in claim 6 or 7.
9. An electronic device, characterized in that: It includes a memory, a processor, and a computer program stored on the memory. When the processor executes the computer program, it implements the character grouping method based on binary character frequencies described in any one of claims 1-5 or implements the security character library construction method described in claim 6 or 7.
Citation Information
Patent Citations
Universal text watermarking method and device
CN114708133A
Security font library construction method and security code extraction method thereof
CN115455966A