Language steganography load enhancement method and device based on large language model
By constructing a shared semantic restorer and probabilistic index compression coding, the secret message processing of large language models is optimized, solving the payload bottleneck and redundancy problems in language steganography, and realizing efficient and secure secret message transmission.
Patent Information
- Application Number
- CN202511515786.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-22
- Publication Date
- 2026-01-23
AI Technical Summary
Existing language steganography methods suffer from load bottlenecks and redundant processing blind spots, resulting in a low ratio of secret message length to steganographic text length, increasing the risk of being suspected by third parties, and failing to effectively compress low-information semantic components in the text.
By constructing a shared semantic restorer, optimizing secret message processing using a large language model, performing dynamic semantic compression and key information protection, generating a pseudo-random bitstream by combining probabilistic index compression coding, embedding it into the carrier text, and transmitting it through a public channel, the receiving end can recover the original message without loss.
It achieves efficient compression and lossless recovery of secret messages, improves the load capacity and security of language steganography systems, and is suitable for covert communication over public channels.
Smart Images

Figure CN121397100A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of large language models and information hiding technology, and in particular to a language steganographic payload enhancement method and device based on a large language model. BACKGROUND
[0002] Steganography, as a core technology of information hiding, realizes covert communication through non-perceptual modification of carrier data, which is different from the logic of encrypted content in cryptography. Its security relies on "eliminating physical / statistical traces of hidden information" to avoid suspicion by third parties (such as enterprises transmitting commercial information through public channels, which need to embed secrets in normal text to evade competitor monitoring). Language steganography uses text as a carrier, and the traditional process is divided into two stages: message processing and message embedding. The message processing stage compresses, encrypts, and formats the secret message, but existing methods only use basic compression methods such as Huffman coding, without utilizing the semantic redundancy characteristics of text. The message embedding stage balances imperceptibility and payload through channel coding, but such modified methods are prone to statistical deviations between steganographic text and normal text, which can be detected by steganalysis tools.
[0003] With the development of large language models (LLM), language steganography has entered the era of generative models: LLMs can generate content that is highly consistent with human text distribution and output token-level sampling probability distribution, based on which a provably secure steganographic system can be built. However, existing generative steganography methods still have key limitations: 1. Payload bottleneck: Due to the low entropy characteristics of text generation probability, the ratio of secret message length to steganographic text length (payload) is low, which requires long steganographic text or multiple transmissions to ensure secret integrity, significantly increasing the risk of suspicion by third parties; 2. Processing blind area: Existing methods only optimize the message embedding stage, ignoring the redundancy elimination of the secret message itself. A large amount of low-information semantic components in the text (such as redundant modifiers and repetitive expressions) occupy transmission resources and are not effectively compressed.
[0004] Therefore, there is an urgent need for a technical solution that can utilize the capabilities of large language models (LLM) to optimize the processing of secret messages, breaking through the payload limitations of existing language steganography while ensuring security and processing efficiency. SUMMARY
[0005] The purpose of the present application is to provide a language steganographic payload enhancement method and device based on a large language model, which optimizes the processing of secret messages through a large language model, achieving efficient compression and lossless recovery of secret messages.
[0006] The purpose of the present application is achieved through the following technical solutions: A language steganographic payload enhancement method based on a large language model, the method comprising: Step 1, first construct a shared semantic restorer obtained by fine-tuning of a large language model LLM, which is shared by the sender and the receiver, for lossless restoration of the secret message after semantic pruning; Step 2, the sender performs dynamic semantic compression and key information protection on the secret message to be processed through the shared semantic restorer, obtaining a compressed message with low semantic redundancy and lossless restoration; Step 3, the obtained compressed message is converted into a compact binary information stream through probability-based index compression encoding, and then a corresponding pseudo-random bit stream is obtained by XOR with a key to adapt to the embedding requirements of the underlying steganography algorithm; Step 4, the pseudo-random bit stream is embedded into the carrier text generated by the large language model LLM through the underlying steganography algorithm to obtain a steganography text, which is transmitted through a public channel; Step 5, after receiving the steganography text, the receiver losslessly restores the original secret message through reverse operation based on the shared steganography algorithm, encoding method, and shared semantic restorer.
[0007] A language steganography payload enhancement device based on a large language model, the device comprising: a shared semantic restorer construction module for constructing a shared semantic restorer obtained by fine-tuning of a large language model LLM, which is shared by the sender and the receiver, for lossless restoration of the secret message after semantic pruning; a dynamic semantic redundancy pruning compression module for performing semantic compression and key information protection on the secret message to be processed through the shared semantic restorer, obtaining a compressed message with low semantic redundancy and lossless restoration; a pseudo-random bit stream generation module for converting the obtained compressed message into a compact binary information stream through probability-based index compression encoding, and then obtaining a corresponding pseudo-random bit stream by XOR with a key to adapt to the embedding requirements of the underlying steganography algorithm; a secret message embedding module for embedding the pseudo-random bit stream into the carrier text generated by the large language model LLM through the underlying steganography algorithm to obtain a steganography text, which is transmitted through a public channel; a message extraction and restoration module for losslessly restoring the original secret message through reverse operation based on the shared steganography algorithm, encoding method, and shared semantic restorer after receiving the steganography text.
[0008] The above technical solution provided by the present application can be seen that the above method and device optimize the processing flow of the secret message through the large language model, realize efficient compression and lossless restoration of the secret message, thereby improving the payload capacity and practical security of the language steganography system, and are suitable for covert communication scenarios under a public channel. BRIEF DESCRIPTION OF DRAWINGS
[0009] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0010] Figure 1 This is a schematic diagram of the language steganography payload enhancement method based on a large language model provided in an embodiment of the present invention; Figure 2 This is a comparison chart showing the time consumed by steganography at different stages in the examples given in this invention. Figure 3 Different pruning quantiles are key parameters in the examples given in this invention. Impact result diagram; Figure 4 The key parameter adaptive coefficients in the examples given in this invention The impact results are shown in the figure. Detailed Implementation
[0011] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments, and do not constitute a limitation of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the protection scope of the present invention.
[0012] like Figure 1 The diagram shows a flowchart of a language steganography payload enhancement method based on a large language model provided in an embodiment of the present invention. The method includes: Step 1: First, construct a shared semantic restorer obtained by fine-tuning the Large Language Model (LLM). This shared semantic restorer is shared by the sender and receiver and is used for lossless recovery of secret messages after semantic pruning. In this step, the shared semantic restorer It is obtained by fine-tuning the Large Language Model (LLM) through three stages, specifically including: Self-information computation: Constructing a paired dataset of "low semantic input - high semantic output" based on a public dataset, for each text sample , This represents a public dataset, broken down into lexical units. Each lexical unit is calculated using the Large Language Model (LLM). Self-information The informational importance of quantifiable vocabulary units is represented as follows: ; wherein , denotes the ith token, is a vocabulary unit contains the number of tokens, token represents the smallest processing unit into which the large language model Chinese text is divided; is the conditional probability of the tth token predicted by the large language model LLM, and the lower the probability, the higher the self-information (the more critical the information); Semantic pruning: pruning the quantile number by self-information, that is, The quantile pruning removes low self-information vocabulary units to construct a compressed text set , denoted as: ; wherein is an element multiplication; is an indicator function; is the proportion of self-information, and is the numerical value of the quantile, which satisfies the following relationship: ; wherein, inf is the lower bound, indicating the smallest value of all τ that satisfy the subsequent conditions; τ is a real number; denotes the number of elements in the set ; The formula indicates that in the set , the proportion of elements with self-information less than or equal to τ is the smallest real number τ; Instruction fine-tuning: pair the compressed text set with the public data set according to the template to construct a fine-tuning data set , fine-tune the large language model LLM so that it only restores the original semantics by filling in the specified special symbols, such as "[]", and does not modify the existing text structure, and finally obtain a shared semantic restorer for the sender and the receiver to share. Step 2, the sender performs dynamic semantic compression and key information protection on the secret message to be processed through the shared semantic restorer, and obtains a compressed message with low semantic redundancy and lossless recovery;
[0013] In this step, first, the large language model LLM detects the key entities (such as names, numbers, addresses) in the secret message to be processed , and sets the self-information of the key entities to , ensuring that the key entities are not pruned; Based on the average self-information of the secret message to be processed , Compared with public datasets Average self-information Dynamically adjust the pruning threshold To avoid over-pruning of short or high-entropy text, a pruning threshold is set. Represented as: ; in These are adaptive coefficients used to balance compression efficiency and semantic preservation. The proportion of self-information is The value of the quantile; The sender uses a shared semantic restorer News after pruning Perform a self-check; if any misaligned parts are found, replace those parts with the original content and enclose them in square brackets "[]" to mark them as special characters. For example, if the word "interest" cannot be recovered, place it in square brackets "[]". Represented as "[interest]", this ultimately yields a compressed message with low semantic redundancy and lossless recovery capability. .
[0014] Step 3: Convert the obtained compressed message into a compact binary information stream through probability-based index compression coding, and then obtain the corresponding pseudo-random bit stream by XORing it with the key to adapt to the embedding requirements of the underlying steganography algorithm. In this step, for compressed messages token sequence The index s represents the length of the token sequence, which is determined by the shared semantic restorer. Predict each token conditional probability ; Lexicon of large language models sorted by probability in descending order Sort, calculate sort index value Rank value, also known as: ; Vocabulary Remove word elements All lexical units after the part; by accumulating all ratios The number of words with high sampling probability, plus 1, gives the result. The Rank value; at this point, the token with a higher sampling probability corresponds to a lower Rank value, which can improve the subsequent compression efficiency; The Rank value sequence is Huffman encoded or arithmetic encoded to generate a binary stream B, achieving lossless compression. The pseudo-random bit stream S is obtained by performing an exclusive OR operation on the binary stream B using a pseudo-random key K generated by using the ChaCha20 algorithm, so that the bit stream before embedding is indistinguishable from a random sequence, and the security of the underlying steganography algorithm is maintained; the ChaCha20 algorithm is a high-efficiency stream cipher algorithm and is used for data encryption.
[0015] Step 4: The pseudo-random bit stream is embedded into the carrier text generated by the large language model LLM through the underlying steganography algorithm to obtain steganography text, and the steganography text is transmitted through a public channel. In a specific implementation, the pseudo-random bit stream can be embedded into the carrier text generated by the large language model LLM through an existing underlying steganography algorithm such as Discop, SparSamp, etc., to obtain steganography text.
[0016] Step 5: After receiving the steganography text, the receiver recovers the original secret message through reverse operations based on the steganography algorithm, the encoding mode and the shared semantic restorer shared with the sender.
[0017] In this step, the receiver first uses the steganography extraction algorithm negotiated in the embedding stage to extract the pseudo-random bit stream from the received steganography text; Then, the same key as that of the sender is used to perform reverse pseudo-randomization on the extracted pseudo-random bit stream to recover the binary stream before Huffman coding or arithmetic coding, and then the Rank value sequence is obtained by decoding through the shared codebook; The Rank value sequence is converted into a compressed message through reverse Rank-token mapping (based on the probability modeling capability of the shared semantic restorer ; Finally, the shared semantic restorer is used to fill in the special symbols in the compressed message , correct the part marked by the sender during self-checking, and recover the original secret message losslessly.
[0018] Based on the above method, an embodiment of the present application also provides a language steganography payload enhancement device based on a large language model, which comprises: A shared semantic restorer construction module is configured to construct a shared semantic restorer obtained by fine-tuning of a large language model LLM, the shared semantic restorer being shared by a sender and a receiver and being used for lossless recovery of a secret message after semantic pruning. A dynamic semantic redundancy pruning compression module is configured to perform semantic compression and key information protection on a secret message to be processed through the shared semantic restorer to obtain a compressed message with low semantic redundancy and capable of being recovered losslessly. A pseudo-random bit stream generation module converts the obtained compressed message into a compact binary information stream through probability-based index compression encoding, and obtains a corresponding pseudo-random bit stream through XOR operation with a key to adapt to the embedding requirements of the underlying steganography algorithm; A secret message embedding module is configured to embed the pseudo-random bit stream into a carrier text generated by a large language model (LLM) through the underlying steganography algorithm to obtain a steganography text, and transmit the steganography text through a public channel. A message extraction and recovery module is configured to, after receiving the steganography text, recover the original secret message through reverse operation based on the shared steganography algorithm, encoding mode, and shared semantic restorer.
[0019] The specific implementation process of each module in the above device is described in the method embodiment.
[0020] It is worth noting that the contents not described in detail in the embodiments of the present application belong to the prior art known to those skilled in the art.
[0021] In order to illustrate the effect of the method described in the embodiments of the present application, the following experimental examples are given: 1. Experimental environment and parameter configuration 1) Hardware platform: Intel(R) Xeon(R) Gold 6130 CPU (2.10 GHz), 256 GB RAM, NVIDIA A6000 GPU; 2) Model selection: base model for fine-tuning to obtain shared restorer: Qwen2.5-7B, DeepSeek-R1-Distill-Llama-8B, using LoRA fine-tuning (2 epochs, learning rate 3x10⁻ 5 ); steganography text generation model: LLaMA2-7B, random sampling (temperature = 0.9, no top-k / top-p limit); 3) Dataset: fine-tuning dataset includes AGNews (news dataset, 30k training / 1.9k testing, average length 241 characters) and IMDb (movie review dataset, 25k training / 25k testing, average length 1300 characters); steganography text generation dataset is WikiText-2-v1 (extract the first 2 sentences as generation prompt); 4) Baseline method: Discop, SparSamp, baseline message processing is "tokenize + Huffman encoding" (based on BPE tokenization); 5) Evaluation criteria include: Payload: the ratio of secret message length to stego-text length, which is a key indicator for evaluating compression and embedding efficiency; Processing time: the encoding time, which contains the total time spent from processing the secret message to generating the stego-text, and the decoding time, which contains the total time spent from extracting the bitstream from the stego-text to recovering the original message; Word-level similarity: measures the proportion of single-word overlap between the recovered / compressed message and the original message, the higher the better indicates the word-level preservation is more complete; Semantic-level similarity: calculates the semantic alignment between the recovered / compressed message and the original message based on BERT contextual embeddings, the higher the better indicates the semantic preservation is better.
[0022] 2. Experimental results verification 1) The main performance evaluation of the method described in this application is shown in Table 1 below: Table 1 Main performance evaluation results
[0023] The experimental results shown in Table 1 show that the method described in the embodiments of the application achieves significant performance improvement on the original steganography system, and the effective payload reaches 2.5 times that of the benchmark system. The core driving force of this breakthrough improvement comes from the deep integration of dynamic semantic redundancy pruning (DSRP) and probability-based index compression coding (ICC) two technologies. Through synergistic effect, they efficiently complete the compression processing of lexical units in the secret message. It is worth noting that when processing longer secret messages, the application exhibits more excellent compression performance due to its advantage of richer contextual information, further verifying its applicability in complex scenarios.
[0024] At the same time, Table 2 below evaluates the message fidelity, by comparing the recovered and compressed secret message with the original message, and quantitatively analyzing from two key dimensions of word-level similarity and semantic-level alignment: Table 2 The application can losslessly reconstruct secret information (achieve an indicator score of 1.0000)
[0025] In Table 2, (recovered / compressed) represents the degree of semantic preservation of the recovered / compressed secret information relative to the original information.
[0026] The evaluation results confirm that this application can achieve lossless reconstruction of secret messages at the receiving end, and the entire compression process completely preserves the core semantic information of the original content, ensuring the accuracy and integrity of message transmission. In summary, this application makes full use of the powerful semantic understanding capabilities of LLM to accurately prune low information density units in secret messages and encode them into dense and secure binary bit streams. This design not only achieves efficient data representation but also ensures that the lossless decoding process is not affected. Ultimately, in covert communication scenarios on public channels, it significantly reduces the data transmission load while simultaneously improving the security and efficiency of communication.
[0027] like Figure 2 The figure shown is a comparison of the time consumed by steganography processing at different stages in the examples of this invention. Although this application introduces additional preprocessing steps, including self-checking, ICC (probability-based indexed compression coding), ICC removal (probability-based indexed compression decoding), and recovery, these steps seem time-consuming, but Table 1 and... Figure 2 The results demonstrate that the optimized payload efficiency ensures the same steganography embedding and extraction time. This is attributed to the reduction in the number of binary bits required per steganographic text, thus significantly shortening the time required for the embedding and extraction processes. Furthermore, the higher payload efficiency enhances security by reducing the amount of steganographic text required for transmission. At the behavioral level, compared to traditional methods, this reduction minimizes the likelihood of arousing adversary suspicion under equivalent communication requirements. These advancements make this application a practical and effective solution for balancing capacity, efficiency, and security in language steganography systems.
[0028] 2) The influence of the main parameters in this application on the experimental results Quantiles: This embodiment studies the self-information pruning quantiles. The impact on steganographic payload and the degree of semantic preservation of compressed ciphertext. For example... Figure 3 The figure shows the key parameters of different pruning quantiles in the examples of this invention. The impact result diagram, in Within the quantile range, the key semantic information in the compressed text varies. Increase the pruning rate gradually and then decrease it. A moderate pruning rate can enhance the payload because the pruning effect becomes more pronounced; however, excessive pruning can disrupt the shared semantic restorer. The ability to effectively reconstruct the original semantics leads to an increased self-checking error rate and more extensive correction operations, ultimately reducing the payload. Therefore, precise balancing is necessary. The value is crucial, as it directly determines the trade-off between payload and semantic fidelity.
[0029] Adaptive coefficients Adaptive coefficient plays an important role in performance in the DSRP module. As Figure 4 Figure 6 shows the impact of the adaptive coefficient of the key parameters of the example of the present application on the results, as the increases, the model trades off between compression efficiency and semantic robustness. A lower value tends to aggressive compression, which maximizes efficiency but risks over-compression, which can lead to loss of key semantic information and hinder accurate restoration; in contrast, a higher value is more conservative, prioritizing semantic integrity at the expense of some compression efficiency.
[0030] In addition, different samples have different complexity and information density, and the value needs to be dynamically adjusted to meet the diversified compression needs. This adaptability ensures that the present application balances between compression efficiency and semantic restoration, thereby enhancing its versatility in various scenarios.
[0031] 3) Evaluation of the transferability of the present application Initially, fine-tuning of small language models is only intended to enable them to understand the restoration task, without considering the domain of secret information. However, given the particularity of covert communication, the interacting parties can usually predict the domain of the covert information. This embodiment analyzes the impact of domain differences between the data sets used to train the restorer and the test set on the payload, as shown in Table 3 below: Table 3 Experimental results of the transferability of the present application
[0032] In the above table: TrainSet is the training set; business is the business category in AGNews; World is the world category in AGNews; AGNews-Mixed indicates that the data sets of various categories in AGNews are mixed; Mixed indicates that the AGNews and IMDb data sets are mixed; This study compares the "business" and "world" categories (which have similarities) in the AGNews data set with the IMDb data set, which has significant differences in content and style. In addition, the present application tests a configuration scheme in which the two data sets are mixed in equal proportions, which slightly reduces the payload without affecting the lossless reconstruction of the message, while mixed training can significantly alleviate this impact. These results highlight the robustness of the present application across domains, even with domain transfer, mixed training can still maintain a high payload.
[0033] 4) Evaluation of the module ablation experiment of the present application This example conducts an ablation experiment on the AGNews data set to evaluate the importance of each component in the present application method, and the relevant results are shown in Table 4 below: Table 4. Experimental results of module compression of the present application
[0034] In the above table: the left side is the model name; w / o means not using; DSRP and ICC are short for dynamic semantic redundancy pruning and index compression coding based on probability of the method of the present application; Discop is the integrated underlying steganography algorithm; is the shared semantic restorer after fine-tuning proposed by the present application; instruction is the system prompt word instruction template in the fine-tuning process, which assists the index compression coding in ICC.
[0035] The experimental results prove that DSRP and ICC are crucial to the performance of the framework. It is worth noting that the introduction of the instruction template in ICC can make the shared semantic restorer more accurately predict the subsequent compressed content, and this improvement can be attributed to the provision of prior knowledge about the format and content of the compressed secret message for the restorer.
[0036] In summary, the method described in the embodiments of the present application introduces the semantic understanding and probability modeling capability of the large language model LLM into the secret message processing stage, and realizes efficient compression and lossless recovery of the secret message through the cooperation of the dynamic semantic redundancy pruning and index-based compression coding double modules, solving the steganography load bottleneck caused by low-entropy generated text.
[0037] The method described in the embodiments of the present application can be widely applied to scenarios requiring covert communication, such as secure transmission of enterprise business secrets (such as competitive analysis data, cooperation schemes) through public channels (email, social platforms), encryption interaction of sensitive government information, and covert sharing of personal privacy data (such as medical record summaries), etc., providing an efficient and secure practical solution for the information hiding field.
[0038] In addition, those skilled in the art can understand that all or part of the steps of the above-mentioned embodiment methods can be completed by programs instructing related hardware, and the corresponding programs can be stored in a computer readable storage medium. The storage medium mentioned above can be a read-only memory, a disk or an optical disk, etc.
[0039] The above is only a preferred specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Any changes or replacements within the technical scope disclosed in the present application can be easily thought of by those skilled in the art, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims. The information disclosed in the background section of this document is only intended to deepen the understanding of the overall background of the present application, and should not be regarded as acknowledging or implying in any form that the information constitutes prior art known to those skilled in the art.
Claims
1. A method for enhancing language steganography payload based on a large language model, characterized in that, The method includes: Step 1: First, construct a shared semantic restorer obtained by fine-tuning the Large Language Model (LLM). This shared semantic restorer is shared by the sender and receiver and is used for lossless recovery of secret messages after semantic pruning. Step 2: The sender performs dynamic semantic compression and key information protection on the secret message to be processed through a shared semantic restorer, resulting in a compressed message with low semantic redundancy and lossless recovery. Step 3: Convert the obtained compressed message into a compact binary information stream through probability-based index compression coding, and then obtain the corresponding pseudo-random bit stream by XORing it with the key to adapt to the embedding requirements of the underlying steganography algorithm. Step 4: Embed the pseudo-random bitstream into the carrier text generated by the Large Language Model (LLM) using a low-level steganography algorithm to obtain the steganographic text, and transmit it through a public channel; Step 5: After receiving the steganographic text, the receiver uses the steganography algorithm, encoding method, and shared semantic restorer shared with the sender to recover the original secret message losslessly through reverse operation.
2. The language steganography load enhancement method based on a large language model according to claim 1, characterized in that, In step 1, the shared semantic restorer It is obtained by fine-tuning the Large Language Model (LLM) through three stages, specifically including: Self-information computation: Constructing a paired dataset of "low semantic input - high semantic output" based on a public dataset, for each text sample , This represents a public dataset, broken down into lexical units. Each lexical unit is calculated using the Large Language Model (LLM). Self-information The informational importance of quantifiable vocabulary units is represented as follows: ; in , This represents the i-th token. For vocabulary units The number of tokens included, where a token represents the smallest processing unit into which text is segmented in a large language model; It is the conditional probability of the t-th token predicted by the Large Language Model (LLM). The lower the probability, the higher the self-information. Semantic pruning: pruning quantiles through self-information, i.e. Quantile pruning removes low-self-information vocabulary units to construct a compressed text set. , represented as: ; in Element-wise multiplication; For indicator functions; The proportion of self-information is The values of quantiles satisfy the following relationship: ; Where inf is the infimum, representing the smallest value among all τ that satisfy the subsequent conditions; τ is a real number; Represents a set The number of elements in the set; this formula represents the set. In the middle, self-information The proportion of elements less than or equal to τ The smallest real number τ; Command fine-tuning: Compress the text set Compared with public datasets Build a fine-tuning dataset by matching templates By fine-tuning the large language model LLM to recover the original semantics by filling in specified special symbols without modifying the existing text structure, a shared semantic restorer is obtained. It is shared by the sender and the receiver.
3. The language steganography load enhancement method based on a large language model according to claim 2, characterized in that, The process of step 2 is as follows: First, detect the secret message to be processed using a large language model (LLM). The key entity in the data is set to its self-information. Ensure that critical entities are not pruned; Based on the secret message to be processed Average self-information Compared with public datasets Average self-information Dynamically adjust the pruning threshold To avoid over-pruning of short or high-entropy texts, a pruning threshold is set. Represented as: ; in These are adaptive coefficients used to balance compression efficiency and semantic preservation. The proportion of self-information is The value of the quantile; The sender uses a shared semantic restorer News after pruning A self-check is performed. If any misaligned parts are found, those parts are replaced with the original content and enclosed in square brackets "[]" to mark them as special symbols. This results in a compressed message with low semantic redundancy and lossless recovery capability. .
4. The language steganography load enhancement method based on a large language model according to claim 3, characterized in that, The process of step 3 is as follows: For compressed messages token sequence The index s represents the length of the token sequence, which is determined by the shared semantic restorer. Predict each token conditional probability ; Vocabulary of large language models sorted in descending order of probability Sort, calculate sort index value Rank value, also known as: ; Vocabulary Remove word elements All lexical units after the part; by accumulating all ratios The number of words with high sampling probability, plus 1, gives the result. The Rank value; at this point, the token with a higher sampling probability corresponds to a lower Rank value; The Rank value sequence is Huffman encoded or arithmetic encoded to generate a binary stream B, achieving lossless compression. The ChaCha20 algorithm is then used to generate a pseudo-random key K. The binary stream B is XORed to obtain a pseudo-random bit stream S, ensuring that the bit stream before embedding is indistinguishable from the random sequence and maintaining the security of the underlying steganography algorithm. The ChaCha20 algorithm is an efficient stream cipher algorithm used for data encryption.
5. The language steganography load enhancement method based on a large language model according to claim 4, characterized in that, In step 5, the receiver first uses the steganalysis algorithm negotiated with the embedding stage to extract the pseudo-random bit stream from the received steganalysis text. Then, using the same key as the sender, the extracted pseudo-random bit stream is reversed and pseudo-randomized to recover the binary stream before Huffman coding or arithmetic coding. Finally, the Rank value sequence is obtained by decoding through a shared codebook. Then, through reverse Rank-token mapping, the Rank value sequence is converted into a compressed message. ; Finally, the shared semantic restorer is utilized. Fill in compressed messages By using special symbols in the code, the parts marked during the sender's self-check are corrected, and the original secret message is recovered without loss.
6. A language steganography payload enhancement device based on a large language model, characterized in that, The device includes: The shared semantic restorer building module is used to build a shared semantic restorer obtained by fine-tuning the large language model LLM. This shared semantic restorer is shared by the sender and receiver and is used for lossless recovery of secret messages after semantic pruning. The dynamic semantic redundancy pruning and compression module is used to perform semantic compression and key information protection on the secret messages to be processed through a shared semantic restorer, resulting in compressed messages with low semantic redundancy and lossless recovery. The pseudo-random bit stream generation module converts the obtained compressed message into a compact binary information stream through probability-based index compression encoding, and then obtains the corresponding pseudo-random bit stream by XORing it with the key, so as to adapt to the embedding requirements of the underlying steganography algorithm. The secret message embedding module is used to embed pseudo-random bit streams into the carrier text generated by the Large Language Model (LLM) through the underlying steganography algorithm to obtain steganographic text, which is then transmitted through a public channel. The message extraction and recovery module is used to recover the original secret message without loss by reverse operation after receiving steganographic text, based on the shared steganography algorithm, encoding method and shared semantic restorer.
Citation Information
Cited By
Cross-domain steganography text recognition and analysis method and system based on multi-adversarial domain adaptation
CN121960500A