Data steganography method and device based on longest matching position coding of sliding window
By using sliding windows to find the longest matching string and embed data in the data compressed by the Deflate algorithm, the problem of low data steganography embedding capacity in the prior art is solved, and more efficient information hiding and stronger detection resistance are achieved.
Patent Information
- Application Number
- CN202510136466.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-05-13
AI Technical Summary
In the prior art, the data steganography method based on the LZS-77 algorithm has insufficient compression efficiency and information embedding capacity, especially when processing data with frequent changes or no obvious repeat sequences, the steganography embedding capacity is low and information hiding cannot be effectively realized.
The data steganography method based on the longest matching position encoding of the sliding window is adopted. By using the sliding window to find the longest matching string in the data compressed by the Deflate algorithm, the data is embedded according to the position and length information of these strings. The method includes obtaining the data to be embedded and the file data to be compressed, setting a sliding window, finding the longest matching string, storing its position and length information, and embeding the data to be embedded according to this information.
It improves the embedding capacity of data steganography, can hide information more effectively in compressed files, and has the characteristics of resisting longest matching detection.
Smart Images

Figure CN119989303A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer information hiding, and in particular to a data steganography method and device based on sliding window longest matching position coding. Background Art
[0002] Data compression is a technology that optimizes data storage and transmission by reducing the storage space or transmission bandwidth represented by the data. Data compression can be divided into two types: lossy compression and lossless compression. Lossless compression does not cause any information loss in the data during the compression and decompression process. Common lossless compression algorithms include LZW (Lempel-Ziv-Welch), Deflate, Gzip and Brotli algorithms. Among them, the Deflate algorithm is based on the combination of the LZ77 algorithm and Huffman coding, and is mainly used for compression of formats such as ZIP files and PNG images. Lossy compression will cause the loss of some details or redundant information in the data during the compression process. The decompressed data is not exactly the same as the original data, and there is a certain loss. Common lossy compression algorithms include JPEG, MPEG, MP3 and H.264 algorithms.
[0003] Data steganography is a technique that hides additional information in other data. It is often used in the fields of steganography, digital watermarking, and data hiding. Data steganography can make the additional information imperceptible in appearance, and the hidden information can only be extracted when there is a corresponding decoding method. It is based on the relevant knowledge base of data compression and information hiding that the method of data steganography in compressed files was born.
[0004] The LZS-77 algorithm is a steganographic method based on the LZ77 algorithm. Since the Deflate algorithm consists of the LZ77 algorithm and Huffman coding, the LZS-77 algorithm is also a steganographic method based on the Deflate algorithm. It uses a sliding window to find and replace repeated data sequences, which can achieve efficient data compression. However, its compression efficiency is affected by the characteristics of the data, and there is a shortcoming of low steganographic embedding capacity. For example, for data that changes frequently or has no obvious repeated sequences, the compression effect may not be as expected, and information embedding cannot be achieved, resulting in the problem of low data steganographic embedding capacity.
[0005] In view of this, the present invention is proposed. Summary of the invention
[0006] The technical problem solved by the present invention is to overcome the deficiencies of the prior art and provide a data steganography method based on the longest matching position encoding of a sliding window, which can be used to perform data steganography on compressed data based on the Deflate algorithm to achieve information hiding.
[0007] In a first aspect, the present invention proposes a data steganography method based on sliding window longest matching position encoding, comprising: S11. Obtain the data M to be embedded, and convert the data M to be embedded into a binary sequence M1; S12. Obtain the file data T to be compressed, compress the file data T to be compressed according to the Defalte algorithm and read the i-th character, and set a sliding window W, wherein the sliding window W is a character buffer before the i-th character whose length does not exceed |W|; S13. linearly search from left to right in the sliding window W for the longest matching string that matches the string starting with the i-th character, wherein the length of the longest matching string meets the minimum compression requirement, record the [position, length] of the longest matching string, preset q longest matching strings to be found, and record the set of q longest matching strings as S, that is, S={S1, S2, ..., Sj, ..., Sq}, where j=1, 2, ..., q; S14. Storing the [position, length] of the longest matching string and obtaining the compression coding information of the longest matching string, embedding the to-be-embedded data M into the to-be-compressed file data T according to the number q of the longest matching strings, and outputting the compression coding of the to-be-compressed file data T, specifically including: S141. If the number of the longest matching strings q = 0, output a single character at the current position i without embedding information; S142. If the number of the longest matching strings q=1, output the matching string information [position, length], compress the longest matching string starting at the current position i, and do not embed the information; S143. If the number of the longest matching character strings q>1, information embedding is performed according to the following method, including: a. Record the [position, length] information of the longest matching string Sj as [pos_j, len_j], and calculate the hash value H (Sj | len_j | pos_j) of the longest matching string Sj according to the hash function H; b. According to the hash value H (Sj | len_j | pos_j) of the longest matching string Sj, rearrange the order of the strings in the longest matching string set S to obtain the longest matching string set S1, where S1 = {S_1, S_2, ..., S_j, ..., S_q}, and the hash value corresponding to the string S_j is H (S_j | len_j | pos_j); c. Encoding the string S_j in the longest matching string set S1={S_1, S_2, ..., S_j, ..., S_q}, including: remember , establish a complete binary tree with a height of K layers, prune the leaf nodes of the complete binary tree so that the complete binary tree retains exactly q leaf nodes, record the left branch of the complete binary tree as '0', the right branch as '1', and each leaf node has a unique K-bit or K-1-bit binary code sequence from the root to the node, corresponding to the longest matching string of sequence numbers S_1, S_2, ..., S_j, ..., S_q in sequence; d. Select the longest coded leaf node with the same sequence as the current data in the binary sequence M1, output the LZ77 code of the longest matching string S_j of the corresponding sequence number, and complete the information embedding, where the string length of the embedded information is the code length of the longest coded leaf node; S5. compress the compressed file data T according to the Defalte algorithm and read the i+1th character, repeating the steps S12-S14 until the information embedding of the data to be embedded M or the compression encoding of the compressed file data T is completed; S6. Packing the compression code of the compressed file data T to generate a compressed file.
[0008] In a second aspect, the present invention proposes a method for extracting stego data, using the data steganography method based on sliding window longest matching position encoding as described in the first aspect, comprising: S21. Obtain compressed file data T, read the compressed file data T according to the Deflate decompression algorithm and decompress it to obtain the original data D of the compressed file data T; S22. Read the compressed file data T again, de-Huffman-code the compressed file data T according to the Deflate decompression algorithm, and obtain the LZ77 encoding information Li [position pos_i, length len_i] at the current decompressed data position i of the compressed file data T; S23. Set a sliding window W at the same position i as the original data D, wherein the sliding window W is a character buffer before the i-th character whose length does not exceed |W|, linearly search from left to right in the sliding window W for the longest matching string that matches the string starting with the i-th character, wherein the length of the longest matching string meets the minimum compression requirement, record the [position, length] of the longest matching string, preset to find q longest matching strings, record the set of q longest matching strings as S, that is, S={S1, S2, ..., Sj, ..., Sq}, wherein j=1, 2, ..., q, q≥1; S24. Decompress and encode the LZ77 encoded information Li[position pos_i, length len_i] and extract the embedded steganographic data information, specifically including: S241. If the number of the longest matching string q=1, the LZ77 encoded information Li[position pos_i, length len_i] is not embedded with the steganographic data information, and proceed to step S22; S242. If the number of the longest matching strings q>1, then store the string information such as [position, length] of the longest matching string, record the [position, length] information of the longest matching string Sj as [pos_j, len_j], calculate the hash value H(Sj | len_j | pos_j) of the longest matching string Sj according to the hash function H, rearrange the order of the strings in the longest matching string set S, and obtain the longest matching string set S1, where S1={S_1, S_2, ..., S_j, ..., S_q}, and the hash value corresponding to the string S_j is H(S_j | len_j | pos_j), perform complete binary tree encoding on the string S_j in the longest matching string set S1, perform leaf node pruning on the complete binary tree, so that the complete binary tree retains exactly q leaf nodes, record the left branch of the complete binary tree as '0', the right branch as '1', each leaf node has a unique K-bit or K-1-bit binary encoding sequence from the root to the node, search for the corresponding string S_j in the reordered S1 set according to the LZ77 encoding information Li[position pos_i, length len_i], the binary tree node code corresponding to the string S_j is the data information of the embedded data M to be extracted; S25. De-Huffman-code the compressed file data T according to the Deflate decompression algorithm, obtain the LZ77 encoding information Li+1 [position pos_i+1, length len_i+1] at the currently decompressed data position i+1 of the compressed file data T, and repeat steps S23-S24 until all the compressed file data T are decompressed or all the data information of the embedded data M is extracted.
[0009] In a third aspect, the present invention proposes a data system based on sliding window longest matching position coding, which adopts the data steganography method based on sliding window longest matching position coding as described in the first aspect, comprising: Embedded data acquisition module D31: acquires the data M to be embedded, and converts the data M to be embedded into a binary sequence M1; The module D32 for acquiring the file data to be compressed acquires the file data T to be compressed, compresses the file data T to be compressed according to the Defalte algorithm and reads the i-th character, and sets a sliding window W, wherein the sliding window W is a character buffer before the i-th character whose length does not exceed |W|; The longest match calculation module D33: linearly searches from left to right in the sliding window W for the longest matching string that matches the string starting with the i-th character, wherein the length of the longest matching string meets the minimum compression requirement, records the [position, length] of the longest matching string, presets that q longest matching strings are found, and records the set of q longest matching strings as S, that is, S={S1, S2, ..., Sj, ..., Sq}, wherein j=1, 2, ..., q; The longest matching position length information storage module D34 stores the [position, length] of the longest matching character string; Data embedding module D35: obtains the compression coding information of the longest matching character string, embeds the data to be embedded M into the file data to be compressed T according to the number q of the longest matching character strings, and outputs the compression coding of the file data to be compressed T; compresses the compressed file data T according to the Defalte algorithm and reads the i+1th character, repeating the above steps until the information embedding of the data to be embedded M or the compression coding of the file data to be compressed T is completed; Matching position length output module D36: Packs the compression code of the compressed file data T to generate a compressed file.
[0010] In a fourth aspect, the present invention provides a system for extracting stego data, using the method for extracting stego data as described in the second aspect, comprising: Compressed file data acquisition module D41: acquires compressed file data T, reads the compressed file data T according to the Deflate decompression algorithm, and decompresses the compressed file data T to obtain the original data D of the compressed file data T; LZ77 encoding information acquisition module D42: reads the compressed file data T again, de-Huffman-encodes the compressed file data T according to the Deflate decompression algorithm, and obtains the LZ77 encoding information Li [position pos_i, length len_i] at the current decompressed data position i of the compressed file data T; The longest matching string acquisition module D43: at the same position i as the original data D, a sliding window W is set, wherein the sliding window W is a character buffer before the i-th character whose length does not exceed |W|, and the longest matching string matching the string starting with the i-th character is linearly searched from left to right in the sliding window W, wherein the length of the longest matching string meets the minimum compression requirement, and the [position, length] of the longest matching string is recorded, and q longest matching strings are preset to be found, and the set of q longest matching strings is recorded as S, that is, S={S1, S2, ..., Sj, ..., Sq}, wherein j=1, 2, ..., q, q≥1; The steganographic data information extraction module D44: decompresses and encodes the LZ77 encoded information Li [position pos_i, length len_i] and extracts the embedded steganographic data information; de-Huffman encodes the compressed file data T according to the Deflate decompression algorithm, obtains the LZ77 encoded information Li+1 [position pos_i+1, length len_i+1] at the currently decompressed data position i+1 of the compressed file data T, and repeats the above steps until all the compressed file data T are decompressed or all the data information of the embedded data M is extracted.
[0011] In a fifth aspect, the present invention proposes an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the data steganography method based on the longest matching position encoding of the sliding window as described above is implemented.
[0012] In a sixth aspect, the present invention proposes a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the data steganography method based on sliding window longest matching position encoding as described above.
[0013] Compared with the prior art, the beneficial effects of the present invention are: it can solve the shortcoming of low steganographic embedding capacity of the LZS-77 steganographic method, can greatly improve the steganographic embedding capacity, and has the characteristic of resisting longest match detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] The accompanying drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present invention, and together with the specification are used to explain the principles of the present invention. Obviously, the accompanying drawings described below are only some embodiments of the present invention, and for those of ordinary skill in the art, other accompanying drawings can be obtained based on these accompanying drawings without creative work.
[0015] Figure 1It is a flow chart of the data steganography method based on sliding window longest matching position encoding proposed by the present invention.
[0016] Figure 2 It is a structural diagram of a data steganography system based on sliding window longest matching position coding proposed by the present invention.
[0017] Explanation of reference numerals: D31 embedded data acquisition module; D32 to-be-compressed file data acquisition module; D33 longest match calculation module; D34 longest match position length information storage module; D35 data embedding module; D36 matching position length output module. DETAILED DESCRIPTION
[0018] In order to make the purpose, technical scheme and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0019] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The singular forms "a", "said" and "the" used in the embodiments of the present invention and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings, and "multiple" generally includes at least two.
[0020] It should be understood that the term "and / or" used in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship.
[0021] As used herein, the words "if" and "if" may be interpreted as "at the time of" or "when" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrases "if it is determined" or "if (stated condition or event) is detected" may be interpreted as "when it is determined" or "in response to determining" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)", depending on the context.
[0022] It should also be noted that the term "includes", "comprising" or any other variation thereof is intended to cover non-exclusive inclusion, so that a commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprising a ..." do not exclude the existence of other identical elements in the commodity or device including the elements.
[0023] The optional embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0024] In a first aspect, the present invention proposes a data steganography method based on sliding window longest matching position encoding, comprising: S11. Obtain the data M to be embedded, and convert the data M to be embedded into a binary sequence M1; S12. Obtain the file data T to be compressed, compress the file data T to be compressed according to the Defalte algorithm and read the i-th character, and set a sliding window W, wherein the sliding window W is a character buffer before the i-th character whose length does not exceed |W|; S13. linearly search from left to right in the sliding window W for the longest matching string that matches the string starting with the i-th character, wherein the length of the longest matching string meets the minimum compression requirement, record the [position, length] of the longest matching string, preset q longest matching strings to be found, and record the set of q longest matching strings as S, that is, S={S1, S2, ..., Sj, ..., Sq}, where j=1, 2, ..., q; S14. Storing the [position, length] of the longest matching string and obtaining the compression coding information of the longest matching string, embedding the to-be-embedded data M into the to-be-compressed file data T according to the number q of the longest matching strings, and outputting the compression coding of the to-be-compressed file data T, specifically including: S141. If the number of the longest matching strings q = 0, output a single character at the current position i without embedding information; S142. If the number of the longest matching strings q=1, output the matching string information [position, length], compress the longest matching string starting at the current position i, and do not embed the information; S143. If the number of the longest matching character strings q>1, information embedding is performed according to the following method, including: a. Record the [position, length] information of the longest matching string Sj as [pos_j, len_j], and calculate the hash value H (Sj | len_j | pos_j) of the longest matching string Sj according to the hash function H; b. According to the hash value H (Sj | len_j | pos_j) of the longest matching string Sj, rearrange the order of the strings in the longest matching string set S to obtain the longest matching string set S1, where S1 = {S_1, S_2, ..., S_j, ..., S_q}, and the hash value corresponding to the string S_j is H (S_j | len_j | pos_j); c. Encoding the string S_j in the longest matching string set S1={S_1, S_2, ..., S_j, ..., S_q}, including: remember , establish a complete binary tree with a height of K layers, prune the leaf nodes of the complete binary tree so that the complete binary tree retains exactly q leaf nodes, record the left branch of the complete binary tree as '0', the right branch as '1', and each leaf node has a unique K-bit or K-1-bit binary code sequence from the root to the node, corresponding to the longest matching string of sequence numbers S_1, S_2, ..., S_j, ..., S_q in sequence; d. Select the longest coded leaf node with the same sequence as the current data in the binary sequence M1, output the LZ77 code of the longest matching string S_j of the corresponding sequence number, and complete the information embedding, where the string length of the embedded information is the code length of the longest coded leaf node; S5. compress the compressed file data T according to the Defalte algorithm and read the i+1th character, repeating the steps S12-S14 until the information embedding of the data to be embedded M or the compression encoding of the compressed file data T is completed; S6. Packing the compression code of the compressed file data T to generate a compressed file.
[0025] Furthermore, in S13, any one of the KMP algorithm, the Boyer-Moore algorithm, the Rabin-Karp algorithm, and the LCSS algorithm may be used to search for the longest matching string.
[0026] Furthermore, in S13, the character string length threshold THD that meets the minimum compression requirement is 3.
[0027] Furthermore, the string length threshold that meets the minimum compression requirement in S13 is , where the parameter α is related to the steganographic embedding capacity, 0<α<1, the smaller α is, the larger the steganographic embedding capacity is.L max It is the longest string in the longest matching string set S.
[0028] Furthermore, the hash function H in step a of S143 may be any one of MD, SHA-1, and SHA-256.
[0029] In a second aspect, the present invention proposes a method for extracting stego data, using the data steganography method based on sliding window longest matching position encoding as described in the first aspect, comprising: S21. Obtain compressed file data T, read the compressed file data T according to the Deflate decompression algorithm and decompress it to obtain the original data D of the compressed file data T; S22. Read the compressed file data T again, de-Huffman-code the compressed file data T according to the Deflate decompression algorithm, and obtain the LZ77 encoding information Li [position pos_i, length len_i] at the current decompressed data position i of the compressed file data T; S23. Set a sliding window W at the same position i as the original data D, wherein the sliding window W is a character buffer before the i-th character whose length does not exceed |W|, linearly search from left to right in the sliding window W for the longest matching string that matches the string starting with the i-th character, wherein the length of the longest matching string meets the minimum compression requirement, record the [position, length] of the longest matching string, preset to find q longest matching strings, record the set of q longest matching strings as S, that is, S={S1, S2, ..., Sj, ..., Sq}, wherein j=1, 2, ..., q, q≥1; S24. Decompress and encode the LZ77 encoded information Li[position pos_i, length len_i] and extract the embedded steganographic data information, specifically including: S241. If the number of the longest matching string q=1, the LZ77 encoded information Li[position pos_i, length len_i] is not embedded with the steganographic data information, and proceed to step S22; S242. If the number of the longest matching strings q>1, then store the string information such as [position, length] of the longest matching string, record the [position, length] information of the longest matching string Sj as [pos_j, len_j], calculate the hash value H(Sj | len_j | pos_j) of the longest matching string Sj according to the hash function H, rearrange the order of the strings in the longest matching string set S, and obtain the longest matching string set S1, where S1={S_1, S_2, ..., S_j, ..., S_q}, and the hash value corresponding to the string S_j is H(S_j | len_j | pos_j), perform complete binary tree encoding on the string S_j in the longest matching string set S1, perform leaf node pruning on the complete binary tree, so that the complete binary tree retains exactly q leaf nodes, record the left branch of the complete binary tree as '0', the right branch as '1', each leaf node has a unique K-bit or K-1-bit binary encoding sequence from the root to the node, search for the corresponding string S_j in the reordered S1 set according to the LZ77 encoding information Li[position pos_i, length len_i], the binary tree node code corresponding to the string S_j is the data information of the embedded data M to be extracted; S25. De-Huffman-code the compressed file data T according to the Deflate decompression algorithm, obtain the LZ77 encoding information Li+1 [position pos_i+1, length len_i+1] at the currently decompressed data position i+1 of the compressed file data T, and repeat steps S23-S24 until all the compressed file data T are decompressed or all the data information of the embedded data M is extracted.
[0030] Among them, in S23, since the LZ77 encoding information Li is [position pos_i, length len_i], the number of the longest matching character strings q≥1.
[0031] In a third aspect, the present invention proposes a data system based on sliding window longest matching position coding, which adopts the data steganography method based on sliding window longest matching position coding as described in the first aspect, comprising: Embedded data acquisition module D31: acquires the data M to be embedded, and converts the data M to be embedded into a binary sequence M1; The module D32 for acquiring the file data to be compressed acquires the file data T to be compressed, compresses the file data T to be compressed according to the Defalte algorithm and reads the i-th character, and sets a sliding window W, wherein the sliding window W is a character buffer before the i-th character whose length does not exceed |W|; The longest match calculation module D33: linearly searches from left to right in the sliding window W for the longest matching string that matches the string starting with the i-th character, wherein the length of the longest matching string meets the minimum compression requirement, records the [position, length] of the longest matching string, presets that q longest matching strings are found, and records the set of q longest matching strings as S, that is, S={S1, S2, ..., Sj, ..., Sq}, wherein j=1, 2, ..., q; The longest matching position length information storage module D34 stores the [position, length] of the longest matching character string; Data embedding module D35: obtains the compression coding information of the longest matching character string, embeds the data to be embedded M into the file data to be compressed T according to the number q of the longest matching character strings, and outputs the compression coding of the file data to be compressed T; compresses the compressed file data T according to the Defalte algorithm and reads the i+1th character, repeating the above steps until the information embedding of the data to be embedded M or the compression coding of the file data to be compressed T is completed; Matching position length output module D36: Packs the compression code of the compressed file data T to generate a compressed file.
[0032] Furthermore, the positions of each module in the data system based on the longest matching position encoding of the sliding window are: The to-be-compressed file data acquisition module D32 is connected to the longest match calculation module D33, and the longest match calculation module D33 is connected to the longest match position length information storage module D34; The embedded data acquisition module D31 and the longest matching position length information storage module D34 are both connected to the data embedding module D35; The data embedding module D35 is connected to the matching position length output module D36.
[0033] In a fourth aspect, the present invention provides a system for extracting stego data, using the method for extracting stego data as described in the second aspect, comprising: Compressed file data acquisition module D41: acquires compressed file data T, reads the compressed file data T according to the Deflate decompression algorithm, and decompresses the compressed file data T to obtain the original data D of the compressed file data T; LZ77 encoding information acquisition module D42: reads the compressed file data T again, de-Huffman-encodes the compressed file data T according to the Deflate decompression algorithm, and obtains the LZ77 encoding information Li [position pos_i, length len_i] at the current decompressed data position i of the compressed file data T; The longest matching string acquisition module D43: at the same position i as the original data D, a sliding window W is set, wherein the sliding window W is a character buffer before the i-th character whose length does not exceed |W|, and the longest matching string matching the string starting with the i-th character is linearly searched from left to right in the sliding window W, wherein the length of the longest matching string meets the minimum compression requirement, and the [position, length] of the longest matching string is recorded, and q longest matching strings are preset to be found, and the set of q longest matching strings is recorded as S, that is, S={S1, S2, ..., Sj, ..., Sq}, wherein j=1, 2, ..., q, q≥1; The steganographic data information extraction module D44: decompresses and encodes the LZ77 encoded information Li [position pos_i, length len_i] and extracts the embedded steganographic data information; de-Huffman encodes the compressed file data T according to the Deflate decompression algorithm, obtains the LZ77 encoded information Li+1 [position pos_i+1, length len_i+1] at the currently decompressed data position i+1 of the compressed file data T, and repeats the above steps until all the compressed file data T are decompressed or all the data information of the embedded data M is extracted.
[0034] In a fifth aspect, the present invention proposes an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the data steganography method based on the longest matching position encoding of the sliding window as described above is implemented.
[0035] In a sixth aspect, the present invention proposes a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the data steganography method based on sliding window longest matching position encoding as described above.
[0036] Specifically, a system or device equipped with a storage medium can be provided, on which software program code that implements the functions of any of the above-mentioned embodiments is stored, and a computer (or CPU or MPU) of the system or device can read and execute the program code stored in the storage medium.
[0037] In this case, the program code itself read from the storage medium can realize the function of any one of the above-mentioned embodiments, and thus the program code and the storage medium storing the program code constitute a part of the present invention.
[0038] The storage medium embodiments for providing the program code include a floppy disk, a hard disk, a magneto-optical disk, an optical disk (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RYM, DVD-RW, DVD+RW), a magnetic tape, a non-volatile memory card, and a ROM. Alternatively, the program code can be downloaded from a server computer via a communication network.
[0039] In addition, it should be clear that the functions of any of the above embodiments can be implemented not only by executing the program code read by the computer, but also by enabling an operating system operating on the computer to complete part or all of the actual operations based on instructions from the program code.
[0040] In addition, it can be understood that the program code read from the storage medium is written to a memory provided in an expansion board inserted into the computer or written to a memory provided in an expansion unit connected to the computer, and then based on the instructions of the program code, a CPU installed on the expansion board or the expansion unit is enabled to perform part or all of the actual operations, thereby realizing the functions of any of the above-mentioned embodiments.
[0041] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A data steganography method based on sliding window longest matching position encoding, characterized in that: include: S11. Obtain the data M to be embedded, and convert the data M to be embedded into a binary sequence M1; S12. Obtain the file data T to be compressed, compress the file data T to be compressed according to the Defalte algorithm and read the i-th character, and set a sliding window W, wherein the sliding window W is a character buffer before the i-th character whose length does not exceed |W|; S13. linearly search from left to right in the sliding window W for the longest matching string that matches the string starting with the i-th character, wherein the length of the longest matching string meets the minimum compression requirement, record the [position, length] of the longest matching string, preset q longest matching strings to be found, and record the set of q longest matching strings as S, that is, S={S1, S2, ..., Sj, ..., Sq}, where j=1, 2, ..., q; S14. Storing the [position, length] of the longest matching string and obtaining the compression coding information of the longest matching string, embedding the to-be-embedded data M into the to-be-compressed file data T according to the number q of the longest matching strings, and outputting the compression coding of the to-be-compressed file data T, specifically including: S141. If the number of the longest matching strings q = 0, output a single character at the current position i without embedding information; S142. If the number of the longest matching strings q=1, output the matching string information [position, length], compress the longest matching string starting at the current position i, and do not embed the information; S143. If the number of the longest matching character strings q>1, information embedding is performed according to the following method, including: a. Record the [position, length] information of the longest matching string Sj as [pos_j, len_j], and calculate the hash value H (Sj | len_j | pos_j) of the longest matching string Sj according to the hash function H; b. According to the hash value H (Sj | len_j | pos_j) of the longest matching string Sj, rearrange the order of the strings in the longest matching string set S to obtain the longest matching string set S1, where S1 = {S_1, S_2, ..., S_j, ..., S_q}, and the hash value corresponding to the string S_j is H (S_j | len_j | pos_j); c. Encoding the string S_j in the longest matching string set S1={S_1, S_2, ..., S_j, ..., S_q}, including: remember , establish a complete binary tree with a height of K layers, prune the leaf nodes of the complete binary tree so that the complete binary tree retains exactly q leaf nodes, record the left branch of the complete binary tree as '0', the right branch as '1', and each leaf node has a unique K-bit or K-1-bit binary code sequence from the root to the node, corresponding to the longest matching string of sequence numbers S_1, S_2, ..., S_j, ..., S_q in sequence; d. Select the longest coded leaf node with the same sequence as the current data in the binary sequence M1, output the LZ77 code of the longest matching string S_j of the corresponding sequence number, and complete the information embedding, where the string length of the embedded information is the code length of the longest coded leaf node; S5. compress the compressed file data T according to the Defalte algorithm and read the i+1th character, repeating the steps S12-S14 until the information embedding of the data to be embedded M or the compression encoding of the compressed file data T is completed; S6. Packing the compression code of the compressed file data T to generate a compressed file.
2. According to claim 1, a data steganography method based on sliding window longest matching position encoding is characterized in that: In S13, any one of the KMP algorithm, Boyer-Moore algorithm, Rabin-Karp algorithm, and LCSS algorithm may be used to search for the longest matching string.
3. According to claim 1, a data steganography method based on sliding window longest matching position encoding is characterized in that: The string length threshold THD that meets the minimum compression requirement in S13 is 3.
4. According to claim 1, a data steganography method based on sliding window longest matching position encoding is characterized in that: The string length threshold that meets the minimum compression requirement in S13 , where the parameter α is related to the steganographic embedding capacity, 0<α<1, the smaller α is, the larger the steganographic embedding capacity is. L max It is the longest string in the longest matching string set S.
5. According to claim 1, a data steganography method based on sliding window longest matching position encoding is characterized in that: The hash function H in step S143 a can be any one of MD, SHA-1, and SHA-256.
6. A method for extracting stego data, characterized in that: The data steganography method based on sliding window longest matching position encoding as claimed in claim 1 comprises: S21. Obtain compressed file data T, read the compressed file data T according to the Deflate decompression algorithm and decompress it to obtain the original data D of the compressed file data T; S22. Read the compressed file data T again, de-Huffman-code the compressed file data T according to the Deflate decompression algorithm, and obtain the LZ77 encoding information Li [position pos_i, length len_i] at the current decompressed data position i of the compressed file data T; S23. Set a sliding window W at the same position i as the original data D, wherein the sliding window W is a character buffer before the i-th character whose length does not exceed |W|, linearly search from left to right in the sliding window W for the longest matching string that matches the string starting with the i-th character, wherein the length of the longest matching string meets the minimum compression requirement, record the [position, length] of the longest matching string, preset to find q longest matching strings, record the set of q longest matching strings as S, that is, S={S1, S2, ..., Sj, ..., Sq}, wherein j=1, 2, ..., q, q≥1; S24. Decompress and encode the LZ77 encoded information Li[position pos_i, length len_i] and extract the embedded steganographic data information, specifically including: S241. If the number of the longest matching string q=1, the LZ77 encoded information Li[position pos_i, length len_i] is not embedded with the steganographic data information, and the process continues to step S22; S242. If the number of the longest matching strings q>1, then store the string information such as [position, length] of the longest matching string, record the [position, length] information of the longest matching string Sj as [pos_j, len_j], calculate the hash value H(Sj | len_j | pos_j) of the longest matching string Sj according to the hash function H, rearrange the order of the strings in the longest matching string set S, and obtain the longest matching string set S1, where S1={S_1, S_2, ..., S_j, ..., S_q}, and the hash value corresponding to the string S_j is H(S_j | len_j | pos_j), perform complete binary tree encoding on the string S_j in the longest matching string set S1, perform leaf node pruning on the complete binary tree, so that the complete binary tree retains exactly q leaf nodes, record the left branch of the complete binary tree as '0', the right branch as '1', each leaf node has a unique K-bit or K-1-bit binary encoding sequence from the root to the node, search for the corresponding string S_j in the reordered S1 set according to the LZ77 encoding information Li[position pos_i, length len_i], the binary tree node code corresponding to the string S_j is the data information of the embedded data M to be extracted; S25. De-Huffman-code the compressed file data T according to the Deflate decompression algorithm, obtain the LZ77 encoding information Li+1 [position pos_i+1, length len_i+1] at the currently decompressed data position i+1 of the compressed file data T, and repeat steps S23-S24 until all the compressed file data T are decompressed or all the data information of the embedded data M is extracted.
7. A data system based on sliding window longest matching position coding, using the data steganography method based on sliding window longest matching position coding according to any one of claims 1 to 5, comprising: Embedded data acquisition module D31: acquires the data M to be embedded, and converts the data M to be embedded into a binary sequence M1; The module D32 for acquiring the file data to be compressed acquires the file data T to be compressed, compresses the file data T to be compressed according to the Defalte algorithm and reads the i-th character, and sets a sliding window W, wherein the sliding window W is a character buffer before the i-th character whose length does not exceed |W|; The longest match calculation module D33: linearly searches from left to right in the sliding window W for the longest matching string that matches the string starting with the i-th character, wherein the length of the longest matching string meets the minimum compression requirement, records the [position, length] of the longest matching string, presets that q longest matching strings are found, and records the set of q longest matching strings as S, that is, S={S1, S2, ..., Sj, ..., Sq}, wherein j=1, 2, ..., q; The longest matching position length information storage module D34 stores the [position, length] of the longest matching character string; Data embedding module D35: obtains the compression coding information of the longest matching character string, embeds the to-be-embedded data M into the to-be-compressed file data T according to the number q of the longest matching character strings, and outputs the compression coding of the to-be-compressed file data T; Compress the compressed file data T according to the Defalte algorithm and read the i+1th character, repeating the above steps until the information embedding of the to-be-embedded data M or the compression encoding of the to-be-compressed file data T is completed; Matching position length output module D36: Packs the compression code of the compressed file data T to generate a compressed file.
8. A data system based on sliding window longest matching position coding according to claim 7, characterized in that: The locations of the modules are: The to-be-compressed file data acquisition module D32 is connected to the longest match calculation module D33, and the longest match calculation module D33 is connected to the longest match position length information storage module D34; The embedded data acquisition module D31 and the longest matching position length information storage module D34 are both connected to the data embedding module D35; The data embedding module D35 is connected to the matching position length output module D36.
9. A system for extracting stego data, using the method for extracting stego data as claimed in claim 6, comprising: Compressed file data acquisition module D41: acquires compressed file data T, reads the compressed file data T according to the Deflate decompression algorithm, and decompresses the compressed file data T to obtain the original data D of the compressed file data T; LZ77 encoding information acquisition module D42: reads the compressed file data T again, de-Huffman-encodes the compressed file data T according to the Deflate decompression algorithm, and obtains the LZ77 encoding information Li [position pos_i, length len_i] at the current decompressed data position i of the compressed file data T; The longest matching string acquisition module D43: at the same position i as the original data D, a sliding window W is set, wherein the sliding window W is a character buffer before the i-th character whose length does not exceed |W|, and the longest matching string matching the string starting with the i-th character is linearly searched from left to right in the sliding window W, wherein the length of the longest matching string meets the minimum compression requirement, and the [position, length] of the longest matching string is recorded, and q longest matching strings are preset to be found, and the set of q longest matching strings is recorded as S, that is, S={S1, S2, ..., Sj, ..., Sq}, wherein j=1, 2, ..., q, q≥1; The steganographic data information extraction module D44: decompresses and encodes the LZ77 encoded information Li [position pos_i, length len_i] and extracts the embedded steganographic data information; de-Huffman encodes the compressed file data T according to the Deflate decompression algorithm, obtains the LZ77 encoded information Li+1 [position pos_i+1, length len_i+1] at the currently decompressed data position i+1 of the compressed file data T, and repeats the above steps until all the compressed file data T are decompressed or all the data information of the embedded data M is extracted.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the data steganography method based on sliding window longest matching position encoding as described in any one of claims 1 to 5 is implemented.