A text carrier-free information hiding method based on the combination of Chinese character components
Through the method based on the combination of Chinese characters, the index generation algorithm and label form are improved, the success rate and capacity of text without carrier information is improved, the hidden problem of very use of characters is solved, and the security and concealment of communication are enhanced.
Patent Information
- Application Number
- CN202211467880.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2022-01-25
- Filing Date
- 2022-11-22
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-11-22
AI Technical Summary
The existing text-free information hiding method has low success rate when hiding very words, and it is impossible to achieve complete secret information transmission.
Using a method based on the combination of Chinese characters, the keywords are split into radicals and independent Chinese characters, and new Chinese characters are generated through combination, and a Chinese character component combination mechanism is introduced in the index generation algorithm to improve the label form to distinguish keywords from generated Chinese characters. The features of the carrier text library are extracted using a multi-layer RNN model, and information hiding and extraction are combined with Huffman encoding.
It improves the hidden success rate of Chinese characters and the hidden capacity of single-story carrier text, enhances the security and concealment of communication, making it difficult for attackers to crack and encode, and ensures high hidden success rate and high hidden capacity under a small text library.
Smart Images

Figure CN115758415B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a text carrierless information hiding method based on Chinese character component combination, belonging to the technical field of information security. Background Art
[0002] In order to ensure the security of the communication process, information hiding, as one of the most commonly used security technologies, is widely applied. It embeds information into a public carrier in a specific way, and the carrier form can be common digital media such as text, image, video and audio in the network. The information hiding technology based on modification has reached a relatively mature stage. At the same time, the development of deep learning has promoted the maturity of steganalysis algorithms, and it is difficult for traditional modification-based hiding methods to resist such detections. Secondly, as one of the mainstream carriers of information hiding, text has a relatively low embedding efficiency of secret information due to its small file size. Therefore, improving the information embedding rate of text carriers on the premise of ensuring the anti-detection of secret information has become the focus of current researchers. In this context, the carrierless information hiding technology has been proposed and quickly attracted wide attention.
[0003] The carrierless information hiding technology does not mean that there is no need for a carrier, but rather uses a method of not modifying the carrier or directly generating a carrier containing secrets to transmit secret information. This information hiding technology is different from traditional information hiding methods in terms of the hiding principle. Since the transmitted information is all natural text, it can resist various steganalysis algorithms and has strong concealment. Currently, the carrierless information hiding methods mainly include two types: search-based and generation-based. Some researchers have combined these two methods, effectively improving the hiding capacity of a single carrier text. However, when the secret information contains some rare characters, this method still cannot achieve the complete transmission of secret information. Summary of the Invention
[0004] Aiming at the problem that some rare characters cannot be hidden in the text carrierless information hiding method, the present invention proposes a text carrierless information hiding method based on Chinese character component combination, which improves the traditional search-based carrierless information hiding mode of "positioning label + keyword". Each Chinese character in the keyword is split into "radical + independent Chinese character", and these Chinese character components are saved in a set. The components in the set are combined pairwise to generate new Chinese characters. The method of the present invention effectively improves the hiding success rate of rare Chinese characters, and can still ensure a high hiding success rate and a high hiding capacity on the premise of using a small text library.
[0005] In order to achieve the above object, the following technical solutions are adopted by the present invention:
[0006] A text carrierless information hiding method based on Chinese character component combination, comprising the following steps:
[0007] Step 1. Determine the search-based carrierless information hiding method. According to the selected method, construct the corresponding carrier text library, determine the form of the positioning tag, and the information hiding and extraction algorithm. Improve the index generation algorithm of the search-based carrierless information hiding method, introduce the Chinese character component combination mechanism, and at the same time improve the tag form to distinguish keywords from generated Chinese characters.
[0008] Step 2. The sender segments the secret information to obtain a keyword set. Use the information hiding method selected in Step 1 and combine it with the improved tag to embed the keywords into multiple carrier texts and send them to the receiver to complete the secret communication.
[0009] Step 3. The receiver receives all the texts in sequence, uses the extraction algorithm selected in Step 1 combined with the improved tag to extract keywords from multiple carriers, and finally forms the original secret information by arranging the keywords in sequence.
[0010] The method of the present invention further improves the communication security on the basis of the inherent strong concealment technology of the carrierless information hiding technology. First, after introducing the Chinese character component combination mechanism, the original positioning tag may point to a keyword or a recombined Chinese character after the keyword is split and recombined. Therefore, additional flag bits and coding bits are required in the positioning tag to indicate this situation, which makes it more difficult for attackers to brute-force guess the coding. Second, the communication parties can customize the Chinese character decomposition method. Even if the attacker analyzes the coding format, they still need to decompose and recombine the original keywords to obtain the recombined Chinese characters embedded. And the coding result after decomposing and recombining a single Chinese character is determined by the constructed Chinese character component library, decomposition algorithm, recombination order, and the coding method of the recombined Chinese character. Therefore, after the attacker cracks the coding file, they still need to construct the same component library as the sender and use exactly the same decomposition algorithm and coding to restore the embedded information.
[0011] The method of the present invention has a certain probability of generating Chinese characters that do not exist in the original text, so as to increase the success rate of embedding the secret information, and at the same time improve the probability of embedding multiple keywords in a single text. At the same time, the present invention has relatively low requirements for the carrier text library, and can still ensure a high hiding success rate and high hiding capacity under the premise of using a small text library. Description of the Drawings
[0012] Figure 1 It is the information hiding framework diagram introducing the Chinese character component combination mechanism provided by the embodiment of the present invention;
[0013] Figure 2 It is the example diagram of keyword splitting and recombination provided by the embodiment of the present invention;
[0014] Figure 3 It is the example diagram of the coding after Chinese character splitting and recombination provided by the embodiment of the present invention;
[0015] Figure 4 This is a binary parameter format diagram provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0016] In order to better understand the technical solution of the present invention, the embodiments of the present invention are described in detail below in conjunction with the accompanying drawings. It should be clear that the described embodiments and all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0017] The embodiment of the present invention introduces a Chinese character component mechanism based on an existing search-type carrier-free information hiding method. Figure 1 This is a diagram of an information hiding framework for introducing a Chinese character component combination mechanism provided by an embodiment of the present invention. Figure 1 As shown, the method comprises the following steps:
[0018] Step 1, determine the existing search-based carrier-free information hiding method, build the corresponding carrier text library according to the selected method, and determine the positioning label form and information hiding extraction algorithm (different search-based carrier-free information hiding methods have different label forms, so the index generation method, information hiding process and extraction process are all different), improve the index generation algorithm of the search-based carrier-free information hiding method, introduce the Chinese character component combination mechanism, and improve the label form to distinguish between keywords and generated Chinese characters. The detailed process is as follows:
[0019] Step 1.1, determine the length n of the positioning tag. Take a text T from the carrier text library, remove non-Chinese characters in T, count the total number of Chinese characters W, and set the starting position IP of T to 0.
[0020] Step 1.2, select n Chinese characters starting from IP in the text T, and convert the n Chinese characters into a binary sequence as the label L according to the parity of the GBK encoding. Segment the four Chinese characters after the label, take the first word after segmentation as the keyword K, create a hash table and name it L, and store the keyword and text path in the hash table named L. If the file named L already exists, store it directly.
[0021] Step 1.3, execute the Chinese character component combination algorithm for the keyword K to generate a recombined Chinese character set H.
[0022] Step 1.3.1, split each Chinese character in the keyword K to obtain a radical set P = {p1, p2, ..., p i} and independent Chinese character set C = {c1, c2, ..., c j}.
[0023] Step 1.3.2: If \(i + j\leq8\), continue to split the Chinese characters in the independent Chinese character set \(C\) to obtain a radical set \(P'=\{p'_1, p'_2,\cdots\}\) and an independent Chinese character set \(C'=\{c'_1, c'_2,\cdots\}\). Perform a union operation on the \(P\) set and the \(P'\) set, and assign the final result to \(P\). Perform a union operation on the \(C\) set and the \(C'\) set, and assign the final result to \(C\). Otherwise, directly execute Step 1.3.3.
[0024] Step 1.3.3: Combine the radicals in the radical set \(P\) in pairs in sequence. If a Chinese character is successfully formed and the keyword \(K\) does not contain this Chinese character, add this Chinese character to the generated Chinese character set \(H\). Combine the radicals in the radical set \(P\) with the independent Chinese characters in the set \(C\) in pairs in sequence. If a Chinese character is successfully formed and the keyword \(K\) does not contain this Chinese character, add this Chinese character to the recombined Chinese character set \(H\). Specific examples are as Figure 2 shown.
[0025] Step 1.3.4: If the length of the recombined Chinese character set \(H\) is greater than 8, only randomly retain 8 of them to obtain the final recombined Chinese character set \(H\). Store each Chinese character in the set \(H\) in a hash table named \(L\).
[0026] Step 1.4: \(IP = IP + 1\), repeat Step 1.2 until \(IP + n + 4>W\).
[0027] Step 1.5: Take out another piece of text from the carrier text library, repeat Steps 1.2 to 1.4 until all texts in the text library are traversed. Return the hash tables named after each label as the index file.
[0028] Step 1.6: Use a multi-layer RNN model to extract the text features of the carrier text library to obtain a language model that meets the sample features of the carrier text library.
[0029] Step 2: The sender segments the secret information to obtain a keyword set, and uses the information hiding method selected in Step 1 and combines the improved labels to embed the keywords into multiple carrier texts and send them to the receiver to complete the secret communication. The detailed process is as follows.
[0030] Step 2.1: Determine the secret information \(M\).
[0031] Step 2.2: Segment and remove stop words from the secret information \(M\) to obtain a keyword set \(KeywordSet\). Expand each keyword in the keyword set \(KeywordSet\) into a synonym set using a synonym forest, and then calculate the similarity using the following calculation formula:
[0032]
[0033] where \(\beta\) v(1 ≤ v ≤ 4 and v ∈ N) is an adjustment parameter, and the four adjustment parameters are as follows: β1 = 0.5, β2 = 0.2, β3 = 0.17, β4 = 0.13. Sim o (1 ≤ o ≤ v and o ∈ N) represents the similarity between specific descriptions in the semantic description formula, and the formula is as follows:
[0034]
[0035] Where p1 and p2 are two sememes, d is the shortest path length between p1 and p2 in the sememe hierarchy, and a is an adjustable parameter. Select words in the synonym set with a similarity above 0.5 to the original keyword, and obtain the final synonym expansion set S′ = {s1, s2, …, s n}(s k = {w1, w2, …}, s k is the finally expanded synonym set).
[0036] Step 2.3, for each synonym set s k in S′, traverse each synonym w in s k , query the texts that meet the conditions in all the hash tables obtained in Step 1.2 according to the synonym w, and store all the retrieved texts into the carrier text set t k corresponding to the synonym set s k . After the traversal, remove duplicates from the texts in the set t k . If t k is an empty set, then split the keyword corresponding to s k (s k is the synonym set after expanding the original keyword, and the corresponding keyword is the keyword before the synonym expansion) into single Chinese characters, query the texts that meet the conditions in all the hash tables obtained in Step 1.2 for each Chinese character as the keyword w, and store the final result into t k , and store the carrier text set t k into the text set collection T.
[0037] Step 2.4, construct a bag-of-words model for T, take out the text txt with the highest occurrence frequency, record all the hidden keywords in this text to form a keyword set K′ = {k′1, k′2, …}, the corresponding label set L′ = {l′1, l′2, …} and the position set U′ = {u′1, u′2, …} of the keywords in the secret information, and determine whether the keyword is the original keyword or a recombined Chinese character. If the keyword is the original keyword, in the text txt, retrieve the label position d′ x according to the keyword k′ x and the label l′ x , and the label l′ x, the set of positions m′ of keywords in the secret information x and the tag position d′ x are converted into binary bits e according to a fixed format and stored. If the keyword is a recombined Chinese character, in addition to retrieving the tag position d′ x in the text according to the keyword k′ x and the tag l′ x , it is also necessary to use the Chinese character component combination algorithm defined in steps 1.3.1 to 1.3.4 to split, recombine and encode each Chinese character in the keyword. Specific examples are Figure 3 as shown. The tag l′ x , the set of positions u′ of keywords in the secret information x and the tag position d′ x , as well as the encoding of the recombined Chinese characters, are converted into binary parameters e according to a fixed format and stored. The above fixed format includes the following parameters:
[0038] Number of word segments: According to the number n of keywords into which the secret information is segmented kws , calculate the value of the number of word segments a to satisfy 2 a-1 ≤n kws ≤2 a , and record the value of the number of word segments a with a fixed 6 bits;
[0039] Maximum number of hidden words: Select the text with the most hidden words and record the number max of hidden keywords kws , calculate the value of the maximum number of hidden words c to satisfy 2 c-1 ≤max kws ≤2 c , and record the value of the maximum number of hidden words c with a fixed 5 bits;
[0040] Number of keywords: Represents the number of keywords in a certain text, represented by cbit;
[0041] Tag: Positioning tag, represented by 5 bits;
[0042] Tag position: Used in combination with the tag to indicate which keyword under a certain tag in this text. Represented by 6 bits;
[0043] Flag: Indicates whether the positioning tag corresponds to a keyword or a recombined Chinese character, represented by 1 bit. Among them, "0" represents a keyword and "1" represents a recombined Chinese character;
[0044] Encoding: Used in combination with the flag. When the flag bit is "0", the encoding bit is 0 bit. When the flag bit is "1", the encoding bit is 3 bits, used to record the encoding of the recombined Chinese character;
[0045] Position of the secret information: Corresponding to the position of the keyword in the original secret information, represented by abit.
[0046] The specific format is as follows Figure 4 as shown
[0047] Step 2.5: Send the text txt to the recipient, remove the set of carrier texts that have been hidden in the above text from T, and repeat Step 2.4 until T is an empty set
[0048] Step 2.6: Randomly select several words to form a candidate pool, calculate the transition probabilities of the words in the candidate pool using the language model obtained in Step 1.6, encode these words according to the conditional probabilities using Huffman coding, select appropriate words as the next-round input according to the binary parameter e until the binary parameter e is completely embedded, and finally generate the text txt' and send it to the recipient
[0049] Step 3: The recipient receives all the texts in sequence, extracts keywords from multiple carriers using the extraction algorithm selected in Step 1 combined with the improved tags, and finally forms the original secret information by arranging the keywords in sequence. The detailed steps are as follows
[0050] Step 3.1: Use the same carrier text library as the sender and the multi-layer RNN model to extract the text library features, and obtain a language model that meets the sample features of the carrier text library
[0051] Step 3.2: Use the language model obtained in Step 3.1 to calculate the probability distribution of each word in txt' at each moment, encode the words in the text using Huffman coding method according to the calculated conditional probabilities, and solve the binary parameter e
[0052] Step 3.3: Parse the binary parameter e in a fixed format, and extract the keyword k' from the carrier according to each set of tags l′ x and the tag position d′ x If the flag bit is 0, directly record the keyword k' x and its position u′ in the secret information x x ; if the flag bit is 1, use the Chinese character component combination algorithm defined in Steps 1.3.1 to 1.3.4 to split the Chinese characters in the keyword, obtain a set of recombined Chinese characters and encode each Chinese character in it, and obtain the recombined Chinese characters as the keyword k' according to the encoding bits in e x x and record its position u′ in the secret information x
[0053] Step 3.4: According to u′ x Arrange the corresponding keyword k' x Put it in the corresponding position of the secret information M'. After all the parameters e are parsed, the complete secret information M' is obtained. The secret information M' may be different from the original secret information M. The main reason is that during the information hiding process, the original secret information M is split into individual keywords, and there is a certain probability that a synonym of a keyword will be used to replace the keyword and be embedded in the carrier text during the information hiding process. Therefore, finally, when extracting the secret information, a phenomenon will occur that a certain word in the original secret information M is replaced by a synonym of the word. Experiments have proved that this operation will not affect the semantics of the sentence.
[0054] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the scope of protection of the present invention.
Claims
1. A text carrierless information hiding method based on Chinese character component combination, characterized in that, The steps include: Step 1. Determine the search-based carrier-free information hiding method, build the corresponding carrier text library according to the selected method, determine the positioning label form and information hiding extraction algorithm, improve the index generation algorithm of the search-based carrier-free information hiding method, introduce the Chinese character component combination mechanism, and improve the label form to distinguish keywords from generated Chinese characters; Step 2. The sender segments the secret information to obtain a set of keywords, and uses the information hiding method selected in step 1 combined with the improved tags to embed the keywords into multiple carrier texts and send them to the receiver to complete the secret communication; Step 3. The receiver receives all the texts in order, uses the extraction algorithm selected in step 1 combined with the improved tags to extract keywords from multiple carriers, and finally composes the keywords into the original secret information in order; Step 1: Determine the existing search-based carrier-free information hiding method, build the corresponding carrier text library according to the selected method, determine the positioning label form and information hiding extraction algorithm, improve the index generation algorithm of the search-based carrier-free information hiding method, introduce the Chinese character component combination mechanism, and improve the label form to distinguish between keywords and generated Chinese characters. The detailed process is as follows: Step 1.1, determine the length n of the positioning tag; take a text T from the carrier text library, remove non-Chinese characters in T, count the total number of Chinese characters W, and set the starting position IP of T to 0; Step 1.2, select n Chinese characters starting from IP in the text T, convert the n Chinese characters into a binary sequence as the label L according to the parity of the GBK encoding; segment the four Chinese characters after the label, take the first word after segmentation as the keyword K, create a hash table and name it L, store the keyword and text path in the hash table named L; if the file named L already exists, store it directly; Step 1.3, executing the Chinese character component combination algorithm for the keyword K to generate a recombined Chinese character set H; Step 1.4, IP = IP + 1, repeat step 1.2 until IP + n + 4 > W; Step 1.5, take out another text from the carrier text library, repeat steps 1.2 to 1.4 until all texts in the text library have been traversed; return the hash table named by each tag as the index file; Step 1.6, using a multi-layer RNN model to extract text features of the carrier text library, and obtaining a language model that meets the sample features of the carrier text library; Step 2: The sender segments the secret information to obtain a set of keywords. It uses the information hiding method selected in step 1 and combines it with the improved tags to embed the keywords into multiple carrier texts and send them to the receiver to complete the secret communication. The detailed process is as follows; Step 2.1, determine the secret information M; Step 2.2, segment the secret information M, remove stop words, and obtain the keyword set KeywordSet. For each keyword in the keyword set KeywordSet, use the synonym forest to expand each keyword into a synonym set, and then use the following calculation formula to calculate the similarity: where β v , 1 ≤ v ≤ 4 and v ∈ N, is an adjustment parameter, and the four adjustment parameters are as follows: β1 = 0.5, β2 = 0.2, β3 = 0.17, β4 = 0.13; Sim o , 1 ≤ o ≤ v and o ∈ N, represents the similarity between specific descriptions in the semantic description formula, and the formula is as follows: Among them, p1 and p2 are two semantic primitives, d is the shortest path length between p1 and p2 in the semantic primitive hierarchy, and a is an adjustable parameter; words with a similarity above 0.5 to the original keyword in the synonym set are selected to obtain the final synonym expansion set S′ = {s1, s2, …, s n}, s k = {w1, w2, …}, s k is the finally expanded synonym set; Step 2.3, for each synonym set s in S′ k , traverse s k For each synonym w in s, query the texts that meet the conditions in all the hash tables obtained in Step 1.2 according to the synonym w, and store all the retrieved texts into the carrier text set t k corresponding to the synonym set s k ; after the traversal is completed, deduplicate the texts in the set t k ; if t k is an empty set, then split the keyword corresponding to s k into single Chinese characters, query the texts that meet the conditions in all the hash tables obtained in Step 1.2 for each Chinese character as the keyword w, and store the final results into t k , and store the carrier text set t k into the text set T; Step 2.4, construct a bag-of-words model for T, take out the text txt with the highest occurrence frequency, record all the hidden keywords in this text to form a keyword set K' = {k'1, k'2, …}, the corresponding label set L' = {l'1, l'2, …} and the position set U' = {u'1, u'2, …} of the keywords in the secret information, and determine whether the keyword is an original keyword or a recombined Chinese character; if the keyword is an original keyword, in the text txt, according to the keyword k' x and label l' x retrieve the label position d' x , and convert the label l' x , the position set m' of the keyword in the secret information x and the label position d' x into binary bits e in a fixed format and store them; if the keyword is a recombined Chinese character, in addition to retrieving the label position d' according to the keyword k' x and label l' x in the text x , it is also necessary to use a Chinese character component combination algorithm to split, recombine and encode each Chinese character in the keyword; convert the label l' x , the position set u' of the keyword in the secret information x and the label position d' x and the encoding of the recombined Chinese character into binary parameters e in a fixed format and store them; Step 2.5: Send the text txt to the receiver, and remove the set of carrier texts that have been hidden in the above text from T. Repeat Step 2.4 until T is an empty set; Step 2.6: Randomly select several words to form a candidate pool. Use the language model obtained in Step 1.6 to calculate the transition probabilities of the words in the candidate pool. Encode these words according to the conditional probabilities using Huffman coding. Select appropriate words as the next-round input according to the binary parameter e until the binary parameter e is completely embedded. Finally, generate the text txt' and send it to the receiver.
2. The text carrierless information hiding method based on Chinese character component combination according to claim 1, characterized in that, The steps of the Chinese character component combination algorithm described in Step 1.3 are as follows: Step 1.3.1: Split each Chinese character in keyword K to obtain a radical set P = {p1, p2, …, p i} and an independent Chinese character set C = {c1, c2, …, c j}; Step 1.3.2: If i + j ≤ 8, continue to split the Chinese characters in the independent Chinese character set C to obtain the radical set P' = {p’1, p'2, …} and the independent Chinese character set C' = {c’1, c'2, …}. Take the union of the P set and the P' set, and assign the final result to P. Take the union of the C set and the C' set, and assign the final result to C; otherwise, directly execute Step 1.3.3; Step 1.3.3: Combine the radicals in the radical set P in pairs in sequence. If a Chinese character is successfully combined and the keyword K does not contain this Chinese character, add this Chinese character to the generated Chinese character set H. Combine the radicals in the radical set P with the independent Chinese characters in the set C in pairs in sequence. If a Chinese character is successfully combined and the keyword K does not contain this Chinese character, add this Chinese character to the recombined Chinese character set H; Step 1.3.4: If the length of the recombined Chinese character set H is greater than 8, randomly retain only 8 of them to obtain the final recombined Chinese character set H; Store each Chinese character in the set H into a hash table named L.
3. A text carrierless information hiding method based on Chinese character component combination according to claim 2, characterized in that The fixed format described above contains the following parameters: Number of segmented words: The number n of keywords obtained by segmenting the secret information kws , calculate the value of the number of segmented words a such that 2 a-1 ≤n kws ≤2 a , and use a fixed 6 bits to record the value of the number of segmented words a; Maximum hidden number: Select the text with the most hidden content and record the number of hidden keywords as max kws , calculate the value of the maximum hidden number c such that 2 c-1 ≤ max kws ≤ 2 c , and use a fixed 5-bit to record the value of the maximum hidden number c; Number of keywords: Represents the number of keywords in a certain text, denoted by cbit; Label: Location label, denoted by 5bit; Label position: Used in combination with the label, indicating which keyword to take under a certain label in this text; denoted by 6bit; Flag: Indicates whether the positioning label corresponds to a keyword or a recombined Chinese character, denoted by 1bit; among them, "0” represents a keyword, and "1” represents a recombined Chinese character; Encoding: Used in combination with the flag; when the flag bit is "0”, the encoding bit is 0bit, and when the flag bit is "1”, the encoding bit is 3bit, used to record the encoding of the recombined Chinese character; Secret information position: Corresponds to the position of the keyword in the original secret information, denoted by abit.
4. A text carrierless information hiding method based on Chinese character component combination according to claim 2, characterized in that The specific method of Step 3 is as follows: The receiver receives all texts in sequence, uses the extraction algorithm selected in Step 1 and the improved label to extract keywords from multiple carriers, and finally forms the original secret information by arranging the keywords in sequence. The detailed steps are as follows: Step 3.1: Use the same carrier text library and multi-layer RNN model as the sender to extract the text library features to obtain a language model that meets the sample features of the carrier text library; Step 3.2: Use the language model obtained in Step 3.1 to calculate the probability distribution of each word in txt' at each moment. According to the calculated conditional probability, use the Huffman coding method to encode the words in the text and solve the binary parameter e; Step 3.3, parse the binary parameter e according to the fixed format, and according to each group of labels l' x and label position d' x Extract keyword k' from the vector x ; If the flag is 0, directly record the keyword k' x and its position in the secret message u' x ; If the flag is 1, use the Chinese character component combination algorithm defined in steps 1.3.1 to 1.3.4 to split the Chinese characters in the keyword, obtain a recombined Chinese character set and encode each Chinese character in it, and obtain the recombined Chinese character as the keyword k' according to the encoding bit in e x And record its position in the secret information u' x ; Step 3.4, according to u' x Put the corresponding keyword k' x into the corresponding position of the secret information M'. After all the parameters e are parsed, the complete secret information M' is obtained; The secret information M' may be different from the original secret information M. The reason is that during the information hiding process, the original secret information M is segmented into individual keywords. There is a certain probability that a synonym of a keyword will be used to replace the keyword and be embedded in the carrier text during the information hiding process. Therefore, when extracting the secret information finally, a phenomenon will occur where a certain word in the original secret information M is replaced by its synonym. This operation does not affect the semantics of the sentence.
Citation Information
Patent Citations
Text carrier-free information hiding method based on Chinese character component combination
CN114491597A