Text steganography method for controlling text generation through ciphertext
By grouping fixed-length encoding of text classification labels and hierarchical candidate pools, combined with large language model fine-tuning and Bert-CNN model, the problems of semantic destruction and noise influence in generative text steganography methods are solved, and more concealed and robust steganographic text is generated.
Patent Information
- Application Number
- CN202510783504.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-09-23
AI Technical Summary
Existing generative text steganography methods easily destroy the semantic coherence of the text during the generation process, and the secret information is easily affected by noise and attacks, resulting in reduced concealment.
Text classification labels are used as the embedding space of secret information. A hierarchical candidate pool is constructed and grouped fixed-length encoding is performed. The steganographic text is generated and extracted by combining large language model fine-tuning and Bert-CNN text classification model.
The generated steganographic text is closer to natural text, which improves the concealment and noise resistance, and enhances the integrity and robustness of the secret information.
Smart Images

Figure CN120688511A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a text steganography method for generating secret text-controlled text, belonging to the technical field of information hiding. Background Art
[0002] Text steganography is an information hiding technique that conceals secret information within a text carrier, generating steganographic text that is indistinguishable from natural text. The steganographic text is transmitted over an open channel, and only authorized recipients can detect whether the text is steganographic and accurately extract the secret information. Current generative text steganography methods are primarily based on a paradigm: a language model is trained using a corpus to generate high-quality text. During the text generation process, specific encoding methods are used to alter word selection based on the secret information to generate the steganographic text. This steganographic scheme generates highly concealed steganographic text that is difficult to detect and detect.
[0003] The concealment of steganographic text determines the success of covert communication to a certain extent. According to the constraints, its concealment is mainly reflected in Stealthiness at the "perceptual," "statistical," and "semantic" levels. "Perceptual stealth" means the steganographic scheme can generate complete and natural-sounding stegotext. "Statistical stealth" means the generated stegotext must be as close as possible to the statistical distribution of natural text. "Semantic stealth" means the steganographic scheme can generate a complete and specific semantic representation of the text.
[0004] Current generative text steganography methods primarily embed secret information by editing the probability distribution of words output by language models. To improve the concealment of steganographic text, enhance its resistance to steganalysis, and minimize the difference between steganographic text and natural text, state-of-the-art language models, such as GPT (Generative Pre-trained Transformer) and LLaMA (Large Language Model Meta AI), have been introduced. Furthermore, various encoding methods, such as Huffman coding, arithmetic coding, and adaptive dynamic block coding, have been introduced to further enhance steganography. These methods attempt to achieve steganography without significantly altering the text's semantics by finely adjusting the probability distribution during text generation. However, the underlying principles of existing generative text steganography methods are flawed. The probability distribution of a language model inherently follows the statistical laws of natural language, while the stego-operation deviates from this original distribution. This results in local features of the generated text differing from those of normal text, disrupting the text's semantic coherence and reducing the concealment of the steganographic text. Furthermore, secret information extraction relies on the precise recovery of text sequences, which has poor robustness. Secret information in steganographic text can be easily destroyed and lost by both non-malicious noise and malicious attacks. Therefore, it is valuable to break away from the inherent generative steganography paradigm of word-level probabilistic editing and implement a generative text steganography algorithm based on sentence-level semantic tags. Summary of the Invention
[0005] The present invention aims to overcome the deficiencies in the prior art, solve the limitations of the current generative text steganography method in word-level probabilistic editing, and The problem that the integrity of secret information is damaged due to noise or attacks during the transmission of steganographic text.
[0006] To achieve the above objectives, the present invention proposes a text steganography method for controlling text generation using secret text, comprising: using the classification labels of the text as the embedding space of the secret information, encoding the classification labels, and constructing a candidate pool; mapping the secret information to be embedded to match the labels in the candidate pool as the secret text; fine-tuning and training a large language model, using the secret text to control the model text generation, so that the semantics of the generated stegotext contains the secret text information; and extracting the semantic labels from the stegotext using the Bert-CNN text classification model, and searching the candidate pool for the encoding corresponding to the labels to obtain the secret information.
[0007] The text steganography method proposed in this invention, which combines secret text to control text, is implemented using the following technical solutions:
[0008] Candidate pool construction: Utilize the hierarchical constraints between text labels to build a hierarchical candidate pool. Filter multiple field labels, sentiment labels, topic labels and other text classification labels to form a label set L = {l1,l2,l3,...,ln}, as a candidate pool element, and extract the hierarchical constraints of the labels using the tree-like inclusion relationship between the labels Build a hierarchical candidate pool. The hierarchical candidate pool effectively increases the embedding capacity of the steganographic algorithm. By leveraging the spatial topology of hierarchical labels, the embedding of a single underlying label can trigger the coordinated extraction of multiple layers of labels. The introduced hierarchical constraints also serve as a natural authentication mechanism. Illegally extracted label combinations violate the hierarchical relationships of the labels, preventing the correct secret information from being retrieved.
[0009] Grouped fixed-length encoding scheme: According to the hierarchical relationship of the labels, the labels in the candidate pool are grouped and fixed-length encoded, which can be formally expressed as: Enc(L,C)→D Where Enc represents the encoding scheme, L represents the set of labels in the candidate pool, C represents the hierarchical constraints between labels, and D represents the encoding of the labels in the candidate pool. The specific process is as follows: Group by label sibling relationship, select 2 in each group i tags, the length of each fixed-length code is determined by i, i satisfies: Where |L1| represents the number of labels in each group.
[0010] According to the encoding of the candidate pool, find the label that matches the secret information to be embedded, and use the bottom-level label as the secret text mystery.
[0011] Secret text control text generation: Screen and construct a multi-label dataset, and perform low-rank adaptation (LoRA) fine-tuning training on the LLaMA2 language model. To adapt to various label text generation tasks; Introduce a trainable soft prompt vector and design the task loss function L according to the probability output of the text classification model task : Where n represents the number of labels in the candidate pool, d i Indicates whether the text contains the current tag, p i Represents the predicted probability of the text classification model. Fine-tune ReSoftPrompt on large language models by regulating text generation through text classification feedback; Through the above two-step fine-tuning, the semantics of the generated steganographic text contains secret information, realizing the generation of secret control text.
[0012] Bert-CNN text classification model: The global context information of the stegotext is extracted based on the Bert model, and Conv1D convolution with three sizes of convolution kernels is used to extract local features of the stegotext at different spans, thereby achieving the extraction of local refined label features of the stegotext; Use the attention mechanism to measure global and local features and fuse the multi-scale features of steganographic text; The extracted features are further mapped to the label space through the fully connected layer to achieve accurate text label extraction.
[0013] The complete secret information steganography process is as follows: Select a set of text classification labels and construct a multi-level label as a candidate pool based on the hierarchical constraints C between labels; Encoding the candidate pool using the grouped fixed-length encoding scheme; Find the label that matches the secret information to be embedded from the candidate pool and put the bottom label l i As a secret text; The secret text is input into the text generation model for autoregressive generation to obtain the steganographic text.
[0014] The complete secret information extraction process is as follows: Implement the same candidate pool and encoding as the steganographic process; Input the stegotext into the text classification model to obtain the classification label l0 of the stegotext; Use the hierarchical constraints of the label to find the upper label of the label, recursively search until the top label, and get the hierarchical label path L=[l0,l1,...l k ]; Find the encoding enc that maps each label in the label path L from the candidate pool CP i , the secret information bit sequence is obtained by splicing and encoding.
[0015] The beneficial effects of the present invention are as follows: the text steganography method for secret text-controlled text generation proposed in the present invention uses the semantic label of the stegotext as the embedding space of the secret information, does not interfere with and deviate from the words in the text generation process, and the generated stegotext is closer to the natural text and has higher text quality. At the same time, embedding is achieved by using the classification label of the text, which can effectively enhance the ability of the stegotext to resist noise disturbance; by utilizing the hierarchical constraints between labels, the embedding of a single underlying label can trigger the linkage extraction of multi-level labels, effectively improving the embedding capacity of the steganography algorithm; the introduced hierarchical constraints can be used as a natural authentication mechanism, and illegal extraction cannot obtain the correct secret information. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1A structural diagram of the text steganography algorithm for generating a secret text control text according to the present invention;
[0017] Figure 2 A flowchart of secret information steganography using a text steganography algorithm combined with secret text control text generation provided in the first embodiment of the present invention;
[0018] Figure 3 A steganographic example diagram of a text steganographic algorithm combined with secret text control text generation provided in the first embodiment of the present invention;
[0019] Figure 4 This is a flowchart of a secret information extraction method using a text steganography algorithm combined with secret text control text generation provided in the second embodiment of the present invention. DETAILED DESCRIPTION
[0020] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention. This is not intended to limit the scope of protection of the present invention.
[0021] Example 1:
[0022] like Figure 2 As shown, an embodiment of the present invention provides a text steganography method for generating secret text control text, comprising the following steps: a. Obtain the text classification label set and secret information bit sequence; b. Extract the hierarchical constraint relationship of labels from the label set and construct a multi-level label as a candidate pool; c. Encode the candidate pool using a grouped fixed-length encoding scheme; d. Find a label from the candidate pool that matches the bit sequence of the secret information to be embedded, and use the bottom label as the secret text; e. Screen and construct the data set, and fine-tune the semantic model to achieve secret control text generation; f. Input the secret text into the fine-tuned language model and generate the steganographic text through autoregression.
[0023] Specific examples include Figure 3 As shown, the detailed process includes:
[0024] 1. The process of building the candidate pool includes: 1.1. Obtain preset text classification labels, such as football, anger, and other 56 labels in different classification systems; 1.2. Extract the hierarchical constraints and structured relationships of labels, and then construct multi-level labels as a candidate pool for embedding secret information; For example, football, its upper labels are sports and theme, forming a hierarchical constraint label path. {footbal<-->sports<-->theme}.
[0025] 2. The specific process of the group fixed-length coding scheme includes: 2.1. Group the tags based on the constructed multi-level tags, and divide the tags with sibling relationships into one group; 2.2. Select 2 from each set of labels i The tags are fixed-length encoded, and the encoding length is i, i satisfies That is, the maximum integer power of 2 that does not exceed the current group length. For example, if the sports sub-tag has 18 tags, such as football, basketball, swimming, and tennis, 16 of them are selected for encoding. The current level encoding length is 4, football is encoded as 0000, and basketball is encoded as 0001. 2.3. Encode all groups in the candidate pool.
[0026] 3. The mapping and embedding process of the secret information bit sequence includes: searching the candidate pool from top to bottom for a tag that matches the secret information bit sequence, and using the bottom-level tag as the secret text. For example, if the secret information bit sequence is 10010000, it matches the theme tag at the top level, then further matches the sports tag within the theme subtag, and finally matches the football tag within the sports subtag, using the football tag as the secret text.
[0027] 4. The process of building a secret control text generation model includes: 4.1. Construct a multi-label text dataset in the “question, answer” format; for example, we obtained 119,658 text data items from public datasets such as GoEmotions, MIND, and Wikipedia through data processing and screening. 4.2. Use this dataset to perform LoRA fine-tuning training on the text generation model; the text generation model uses the LLaMA2-7B large language model with 674 million parameters, the LoRA rank is set to 8, the Q and V matrices are fine-tuned in the module, the learning algorithm is AdamW, and the batch size is 16; 4.3. Fine-tune the fine-tuned text generation model using ReSoftPrompt: Introduce a trainable vector into Prompt as a constraint for stegotext generation. Use the classification results of the classifier as feedback to adjust the constrained text generation model and achieve label-controlled text generation. Set the trainable vector dimension to 256, the batch size to 16, and the dropout to 0.1.
[0028] 5. The process of generating steganographic text includes: 5.1. Construct the secret text into a prompt. For example, the prompt constructed by the football tag is: give the history of football. 5.2. Input into the text generation model, and generate the final steganographic text through autoregression; for example, "football, known as soccer..."
[0029] Example 2:
[0030] like Figure 4 As shown, based on embodiment 1, the present invention provides a secret information extraction method of a text steganography method for generating a secret text control text. The method comprises the following steps: (a) Obtain text classification label set and steganographic text; (b) Extract the hierarchical constraint relationship of labels from the label set and construct a multi-level label as a candidate pool; (c) encoding the candidate pool using a grouped fixed-length coding scheme; (d) According to the text classification label set and its hierarchical constraints, the label range is determined and the data set is selected to train the Bert-CNN text classification model; (e) Input the stegotext into the text classification model to obtain the label of the stegotext; (f) Using the hierarchical constraints of the label, find all the upper-level labels of the label and obtain the label path; (g) Find the encodings of all label mappings in the label path from the candidate pool and concatenate them in order to obtain the secret information bit sequence.
[0031] The candidate pool construction and encoding scheme are consistent with the first embodiment.
[0032] 1. The Bert-CNN text classification model construction process includes: 1.1. Construct a multi-label text dataset in the "text, label" format; for example, from public datasets such as GoEmotions, MIND, and Wikipedia, we obtained 119,658 text data through data processing and screening; 1.2. Use the Bert model to generate a high-dimensional vector representation for each word in the text. Leverage the bidirectional context of words to more comprehensively understand semantics. A multi-head self-attention mechanism is used to capture global dependencies between words and extract global contextual information from the text. The Bert-base-uncased model is used, with 12 Transformer encoder layers and a 768-dimensional hidden layer. The multi-head attention mechanism has 4 heads and a dimension of 128. 1.3. Set three different sizes of convolution kernels to output the global context information T of the Bert model. CLS Perform one-dimensional convolution to extract local refined features of different spans of text and improve the language comprehension ability of the text multi-classification model; use Conv1D convolution with kernel sizes of 3, 5, and 7, stride of 1, and same padding; 1.4. Map the global features extracted by the BERT model and the local features extracted by the convolution operation into the attention input query Q, key K, and value V. Dynamically assign weights to different features, adaptively fuse features at multiple scales, and accurately extract fine-grained classification labels. 1.5. Use a fully connected layer to map the fused features to the label space to achieve text multi-classification; set 2 hidden layers, 512 hidden layer units, and use sigmoid as the activation function; 1.6. Use the constructed dataset to train the Bert-CNN text classification model; the dropout is 0.1, the batch size is 16, and the number of training rounds is 3.
[0033] 2. The secret text extraction process includes: 2.1. Input the steganographic text into the Bert-CNN text classification model to obtain the label l of the steganographic text i , i.e. secret text; for example, the secret text of “football, known as soccer...” is football; 2.2. Using the constructed multi-level labels, according to the hierarchical constraint relationship between labels, find the label l i The upper label of the tag is recursively searched until the top label is found to obtain the tag path; for example, the parent label of football is sports, the top label is theme, and the tag path is {football-->sports-->theme}.
[0034] 3. The process of extracting the secret information bit sequence includes: 3.1. Find the encoding of each label mapping in the label path in the candidate pool; for example, theme is encoded as 10, sports is encoded as 01, and football is encoded as 0000; 3.2. Concatenate the codes of the labels at different levels in the order of top-middle-bottom to obtain the secret information bit sequence. For example, the secret information bit sequence extracted by the inverse mapping of the football label is 10,01,0000.
[0035] Except for the technical features described in the specification, the rest of the technologies involved are known technologies in the field of this professional and technical field. In order to avoid redundancy and unnecessary limitation of the present invention, the specification omits detailed descriptions of known components and known technologies. The implementation methods described in the above embodiments are only specific examples of the present invention and do not represent all implementation methods consistent with this application. On the basis of the technical solution of the present invention, those skilled in the art may make equivalent substitutions or various modifications to the present invention, which still fall within the scope of protection of the present invention.
Claims
1. A text steganography method for generating secret text controlled by text, characterized in that: The following steps are involved: (1) Using the classification labels of the text as the embedding space of the secret information, the classification labels are grouped and fixed-length encoded to construct a candidate pool; (2) According to the secret information to be embedded, a matching classification label is found from the candidate pool and used as the secret text; (3) The secret text is input into the fine-tuned large language model for autoregressive text generation to obtain a steganographic text containing the secret text in the semantics; (4) The multi-scale features extracted by Bert and CNN are integrated through the attention mechanism to realize text classification, and the secret text, that is, the corresponding classification label, is extracted from the steganographic text; (5) The corresponding coding bits, that is, the secret information, are obtained by inverse mapping based on the classification label.
2. The method according to claim 1, characterized in that Breaking through the limitations of existing generative text steganography methods based on word-level probabilistic editing, the method of controlling text generation through secret text is used to achieve the embedding of secret information with sentence-level semantic tags.
3. The method according to claim 2, characterized in that The candidate pool construction and encoding process includes: Select domain tags, sentiment tags, and topic tags to form a tag set; Utilize the inclusion relationship between labels to extract the hierarchical constraints of labels and form multi-level labels; Group by label sibling relationship, select 2 in each group i The labels are taken as the candidate pool CP, and the length of each fixed-length code is determined by i, and the labels in the candidate pool are grouped and coded with fixed length.
4. The method according to claim 2, characterized in that The secret text control text generation process includes: The secret information bit sequence to be embedded is matched to the corresponding label from the candidate pool CP as the secret text mystery; Based on the low-rank adaptation fine-tuning of the large language model, a trainable soft prompt vector is introduced, and the task loss function L is designed according to the probability output of the text classification model. task : Where n represents the number of labels in the candidate pool, d i Indicates whether the text contains the current tag, p i Represents the predicted probability of the text classification model. It uses text classification feedback to adjust text generation and achieve ReSoftPrompt fine-tuning of the large language model. The secret text mystery is input into the fine-tuned language model and the original autoregressive generation process is performed to obtain the steganographic text S.
5. The method according to claim 1, wherein Unlike existing generative steganography methods that rely on a text generation model with the same generation process, this method implements the secret information extraction process based on a text classification model, including: Design the Bert-CNN text classification model to classify the steganographic text S and obtain the classification label l0 of the steganographic text matching; Using the hierarchical constraints of labels, find the upper label of the classification label l0, recursively search until the top label, and obtain the label path L = [l0, l1, ... l k ]; Find the encoding enc that maps each label in the label path L from the candidate pool CP i , and concatenate them in order to obtain the secret information bit sequence.
6. The method according to claim 1 and the extraction process according to claim 5, characterized in that The multi-scale features of the steganographic text are integrated to realize the text classification model. Based on the multi-scale features of the steganographic text extracted by Bert and Conv1D convolution, the multi-scale features of the steganographic text are integrated using the attention mechanism, and the text classification is further realized through the fully connected layer.