Steganography text detection technology based on text reconstruction and word order semantic features

Based on the BERT architecture and cosine similarity, combined with text reconstruction and word order semantic features, a steganographic text detection technology is proposed, which solves the problem that the existing technology is difficult to detect highly concealed steganographic text, and realizes high accuracy and low cost steganographic text detection.

CN120163162APending Publication Date: 2025-06-17BEIJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510234083.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

Existing text steganography analysis techniques are difficult to effectively detect highly concealed steganography text generated using deep learning and large language models. Especially in terms of statistical distribution and semantic features, traditional methods are difficult to cope with the development of high-dimensional steganography detection technology.

Method used

Using BERT architecture and cosine similarity, combined with text reconstruction and word order semantic features, a steganographic text detection technology based on text reconstruction and word order semantic features is proposed. This technology reconstructs the messed text through a large language model, and uses the BERT encoder and feature extractor network to compare the semantic differences between the original text, messed text and reconstructed text to realize the detection of steganographic text.

Benefits of technology

High accuracy detection of highly concealed steganographic text is achieved, with good condition adaptability and interpretability, avoiding the high cost and complexity of deploying and training large-scale large-language models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163162A_ABST
    Figure CN120163162A_ABST
Patent Text Reader

Abstract

The invention discloses a steganographic text detection technology based on text reconstruction and word order semantic features, and belongs to the technical field of information hiding. The method comprises the following steps: selecting an open source English steganography text data set containing three fields of social platform speaking, news manuscripts and film review; words in non-steganographic text sentences in the data set are randomly disorganized and sent to a large language model for reconstruction training, and the training target of reconstruction is an original text; and randomly disorganizing a data set text, and inputting the disorganized data set text into the trained model to generate a reconstructed text. And finally, inputting the original text, the disordered text and the reconstructed text into a steganography detection model, calculating a cosine similarity matrix, extracting a semantic difference feature map through a CNN, flattening, splicing semantic vectors of the original text, and inputting the semantic vectors into a classifier to output a binary detection result. Through continuous training, an error between a classifier output result of the model and a real labeling result is continuously reduced, so that parameters of a feature extractor and a classifier are optimized, and the model steganography detection accuracy is improved. According to the method, semantic information and word order features are utilized for classification, high accuracy, low cost and good interpretability are achieved, steganographic texts can be effectively detected, and information safety is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of information hiding, relates to text steganography, and specifically is a steganographic text detection technology based on text reconstruction and word order semantic features Background Technique

[0002] Steganography is a technology that hides secret information in public media. This technology hides secret information in multimedia such as text or images, generates stego media, and transmits it through a public channel to achieve the purpose of secretly transmitting information. Only authorized personnel can detect whether the media is steganographic and accurately extract the secret. Thanks to the ability of text to be losslessly transmitted in public channels, linguistic steganography technology provides a convenient implementation method for hiding information. Especially after the emergence of artificial intelligence generation technology, the abuse of various steganography technologies has raised security concerns in society, calling for powerful language steganalysis technologies to detect carriers containing steganographic information

[0003] Text steganalysis is a technique for determining whether a text contains confidential information. Traditional text steganalysis techniques are limited to finding distribution differences between steganographic texts and ordinary texts from the perspective of symbol statistics. However, with the improvement of steganography technology, the statistical distribution differences caused by information hiding are getting smaller and smaller, increasing the difficulty of steganalysis. In recent years, more advanced generative language steganography technologies have emerged. It generates high-quality text by training a deep learning language model. During the text generation process, a specific encoding method is used to select tokens according to the secret, thereby generating steganography. The emergence of such methods marks that the generation scheme of steganographic texts has developed from traditional schemes that are easy to detect to generation schemes that can automatically generate high-quality steganographic texts using advanced generative models such as AutoEncoder. The highly concealed steganography generated by such schemes not only has almost no difference in statistical distribution from normal texts, but also is close to normal texts in terms of text style and fluency, bringing greater challenges to the semantic feature extraction of fine-tuned deep learning models. In recent years, with the booming development of large language models (LLMs), the technology of steganography schemes is still constantly evolving, and traditional text steganalysis is difficult to cope with the increasingly developed high-dimensional steganalysis detection technology

[0004] With the booming development of natural language processing technology, in recent years, there have also been many attempts to design steganography detection technologies by using deep learning methods to extract high-dimensional semantic features. At the same time, some researchers have proposed that leveraging the text processing capabilities of large language models similar to humans can help steganalysis tools achieve more accurate detection, thereby overcoming the performance bottlenecks of current language steganalysis methods. However, using the capabilities of large language models for language steganalysis faces significant challenges and costs. On the one hand, it is not easy to activate the capabilities of large language models for specific tasks. Researchers have explored the potential of using ChatGPT (Generative Pretrained Transformer) for steganalysis tasks, but due to the lack of targeted design and training, its performance without fine-tuning is mediocre. On the other hand, due to the huge size of large language models, directly fine-tuning, training, and deploying them require a large amount of time and financial investment. At the same time, using various deep learning and large language model methods to implement steganographic text detection technologies aims to enable the model to analyze the high-dimensional semantic features of text, which lacks interpretability to a certain extent and is not conducive to further analysis of the steganography principle. Therefore, when using deep learning models for steganographic text detection, it is very necessary to consider factors such as small size, low cost, and strong interpretability. Summary of the Invention

[0005] To address the above problems, based on the BERT (Bidirectional Encoder Representations from Transformers) architecture and cosine similarity, considering the semantic differences after shuffling the word order of the text, and taking advantage of the superiority of large language models in text reconstruction, the present invention proposes a steganographic text detection technology based on text reconstruction and word order semantic features to perform text steganography detection on the target text and confirm whether the target text contains secret information.

[0006] A steganographic text detection technology based on text reconstruction and word order semantic features provided by the present invention includes the following steps:

[0007] Step 1: Obtain a steganographic text dataset. The dataset contains multiple text segments intercepted in various scenarios, and each text segment has a corresponding annotation indicating whether it is a steganographic text;

[0008] Step 2: For non-steganographic text entries in the dataset, randomly shuffle the words in its text sentence to obtain an unordered word sequence, and send this word sequence into the large language model, requiring the large language model to reconstruct it, and the training objective of the reconstruction is the original non-steganographic text;

[0009] Step 3: Train the large language model in the training set so that the average error between the text reconstructed by the large language model and the original non-steganographic text reaches the error range threshold, and update the parameters of the large language model;

[0010] Step 4: Randomly shuffle the text entries in the dataset (hereinafter referred to as: original text) to obtain an unordered sequence of words (hereinafter referred to as: shuffled text), and then send it into the trained large language model to obtain the reconstructed text output by the large language model (hereinafter referred to as: reconstructed text);

[0011] Step 5: Input the original text, shuffled text, and reconstructed text into the steganographic text detection model respectively to obtain the binary detection result Y=(0,1) of a single text message (0 represents non-steganographic text, 1 represents steganographic text);

[0012] The steganographic text detection model includes an encoder, a feature extractor, and a classifier; the encoder network consists of a BERT architecture, the feature extractor network includes a cosine feature layer and a convolutional layer, and the classifier includes a linear splicing layer, a linear layer, and a Softmax output layer. After inputting the original text, shuffled text, and reconstructed text into the encoder, the encoder output vectors E o , E D and E B are obtained respectively. Input E o and E D into the cosine feature layer and convolutional layer of the feature extractor network to obtain the word order semantic difference feature encoding map O D of the original text and the shuffled text; input E o and E B into the cosine feature layer and convolutional layer of the feature extractor network to obtain the word order semantic difference feature encoding map O B of the original text and the reconstructed text. Input the output vector E O of the original text and the feature encoding map O D into the classifier, and obtain the detection result (0 or 1, 0 represents non-steganographic text, 1 represents steganographic text) of the text shuffling detection scheme through the linear splicing layer, linear layer, and Softmax output layer; input the output vector E O of the original text and the feature encoding map O B into the classifier, and the detection result of the text reconstruction detection scheme can also be obtained. Both schemes can realize the detection technology of steganographic text.

[0013] Step 6: Train the steganographic text detection model so that the error between the output result of the classifier and the true annotation result in the dataset reaches the error range threshold, and update the parameters of the encoder, feature extractor, and classifier.

[0014] Step 7: Use the trained steganographic text detection model to detect the target text.

[0015] Randomly select a text segment as the input, and after randomly shuffling it, obtain the shuffled text of this text segment. Input the shuffled text into the trained large language model to obtain the reconstructed text output by the large language model. Input the original text, the shuffled text, and the reconstructed text into the trained steganographic text detection model to obtain the detection result of whether this text is a steganographic text.

[0016] In the second step, it is required that the large language model reconstructs it, and the reconstruction training target is the original non-steganographic text. The implementation method is as follows: (1) Design the instruction information input to the large language model, and the instruction information contains the instruction to require the large language model to reconstruct the input text; (2) Construct a communication scenario for the large language model to execute the specified task, and sequentially input the instruction information and the shuffled text to be reconstructed to the large language model; (3) Set the target result reconstructed by the large language model as the original sentence to reduce the error and update the parameters during the training process.

[0017] In the fifth step, the text steganography network model performs the following processing on the input original text, shuffled text, and reconstructed text: (1) After inputting the text into the BERT encoder network, obtain the output encoder vector group E of each sentence of the text o ,E D and E B ; (2) Input the vector group E o and E D into the cosine feature layer of the feature extractor network to obtain the cosine similarity matrix M of the original text and the shuffled text encoding vector group D ; (3) Regard M D as the input image of the convolutional neural network (CNN), and input it into the convolutional layer of the feature extractor network to obtain the word order semantic difference feature encoding map O of the original text and the shuffled text D ; (4) Perform similar processing to (2) and (3) on E o and E B to obtain the word order semantic difference feature encoding map O of the original text and the reconstructed text B ; (5) Flatten O D into a one-dimensional vector, and pass it through the linear splicing layer, linear layer, and Softmax output layer of the classifier together with the first vector E of the output vector group of the original text O1 to obtain the detection result of the text shuffling detection scheme; (6) Perform similar processing to (5) on O B to obtain the detection result of the text reconstruction detection scheme.

[0018] The advantages and positive effects of the present invention are as follows: The present invention is a steganographic text detection technology based on text reconstruction and word order semantic features, which takes into account the serialization characteristics of text, the context connection between texts, the statistical characteristics and semantic characteristics of text carriers, etc., and realizes the high-accuracy detection of steganographic texts. The present invention also has good conditional adaptability. The text scrambling detection scheme of the present invention does not involve the deployment and training of large language models with a large volume, and has the advantages of low cost and small volume. The text reconstruction detection scheme of the present invention is a further optimization of the text scrambling detection scheme, with higher accuracy. At the same time, it also has good interpretability, revealing to a certain extent the semantic information of the word sequence of steganographic texts. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 It is a flowchart of the steganographic text detection technology based on text reconstruction and word order semantic features of the present invention;

[0020] Figure 2 It is a schematic diagram of the training and testing of the large language model of the present invention;

[0021] Figure 3 It is a schematic diagram of the reconstructed text, original text and scrambled text generated by the large language model of the present invention;

[0022] Figure 4 It is a schematic diagram of the training and testing process of the steganographic text detection model network based on text reconstruction and word order semantic features of the present invention;

[0023] Figure 5 It is a schematic diagram of the internal structure of the network model of the present invention;

[0024] Figure 6 It is a schematic diagram of the internal parameters of the network model of the present invention;

[0025] Figure 7 It is a schematic diagram of the performance of the model of the present invention in text steganography detection. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0026] The present invention will be further described in detail below in conjunction with the drawings and embodiments.

[0027] In order to elaborate on the features and advantages of the present invention in detail, the present invention will be described in actual application from the full process from training to application.

[0028] The BERT (Bidirectional Encoder Representations from Transformers) model is an optimized Transformer model based on the encoder architecture. Its pre-trained model has good text analysis, understanding, and embedding encoding capabilities. It can convert sentence sequence text into a vector group that includes in-sentence words and the overall sentence meaning, and can also further train and adjust the encoding logic. Cosine similarity is often used to evaluate the similarity between two vectors, and it can also characterize the semantic similarity of a pair of texts after embedding encoding. Based on the above two theories, considering the semantic differences after scrambling the text word order and taking advantage of the large language model in text reconstruction, the present invention proposes a steganographic text detection technology based on text reconstruction and word order semantic features to detect steganographic text in the target text and confirm whether the target text contains secret information.

[0029] As Figure 1 shown, as a steganographic text detection technology based on text reconstruction and word order semantic features, the present invention is described in the following seven steps.

[0030] Step 1: Obtain a steganographic text dataset. The dataset contains multiple text segments intercepted in various scenarios. Each text segment has a corresponding annotation indicating whether it is steganographic text.

[0031] In the embodiment of the present invention, the steganographic text dataset is an open-source dataset selected from the Internet, in English, and contains text samples in three scenarios (social platform posts, news articles, and movie reviews). To facilitate the training of the text steganographic detection model, it is necessary to ensure that the text and annotation formats in the dataset are consistent; if the above requirements are not met, data preprocessing is required to meet the requirements.

[0032] Step 2: For non-steganographic text entries in the dataset, randomly scramble the words in their text sentences to obtain a group of unordered word sequences, and send this word sequence into the large language model, requiring the large language model to reconstruct it, and the training objective of the reconstruction is the original non-steganographic text.

[0033] Step 3: Train the large language model in the training set so that the average error between the text reconstructed by the large language model and the original non-steganographic text reaches the error range threshold, and update the parameters of the large language model.

[0034] As Figure 2As shown in the figure, the shuffling scheme in the embodiment of the present invention is generated by program random numbers, and the large language model Qwen2 that can perform task-specific fine-tuning is adopted. By setting specific communication scenarios, instruction information, and expected target results for the large language model, the output error during the training process of the large language model is reduced and the parameters are updated. During the training process, the LoRA (Low-Rank Adaptation) algorithm is added to the present invention to accelerate the training speed of the large language model and reduce the training consumption. During the testing process, the parameters of the model are no longer updated.

[0035] Step 4: Randomly shuffle the text entries in the data set (hereinafter referred to as: original text) to obtain a set of unordered word sequences (hereinafter referred to as: shuffled text), and then send them into the trained large language model to obtain the reconstructed text output by the large language model (hereinafter referred to as: reconstructed text).

[0036] Select a sample from the steganographic text data set, use the trained large language model, set the same communication scenario and instruction information as when training and testing the large language model, and send the randomly shuffled shuffled text into the trained large language model to obtain the reconstructed text output by the large language model. Examples of the original text, shuffled text, and reconstructed text are as Figure 3 shown.

[0037] Step 5: Input the original text, shuffled text, and reconstructed text into the steganographic text detection model respectively to obtain the binary detection result Y=(0,1) (0 or 1, 0 represents non-steganographic text, 1 represents steganographic text) of a single text message;

[0038] As Figure 4 and Figure 5 shown, the text steganography detection model network of the present invention can be called the CoSim-R model network, including an encoder, a feature extractor, and a classifier. During the training process, each text data in the data set is processed by Step 4 to obtain the original text - shuffled text pair and the original text - reconstructed text pair. Inputting the text pair into the text steganography detection model network of the present invention, the steganography detection analysis result of the original text can be obtained. By comparing the difference between the output result of the model network and the annotation result, the cross-entropy loss function is used to calculate the loss value and then backpropagate the gradient, and continuous training is performed to reduce the difference between the two. During the testing process, except that the gradient is not backpropagated, other steps are the same as those in the training process, and finally the loss value is obtained through the loss function calculation.

[0039] As Figure 5As shown, the encoder network of the text steganography detection model network of the present invention is a fine-tuned BERT model network. The feature extractor network includes a cosine feature layer and a convolutional layer. The classifier network includes a linear splicing layer, a linear layer, and a Softmax output layer. The innovation of the present invention lies in comparing the embedding encoding results of the original text, scrambled text, and reconstructed text to obtain a cosine similarity matrix, and at the same time treating the cosine similarity matrix as an image to be processed by a convolutional neural network (CNN) to obtain a semantic difference feature encoding map. More innovatively, the flattened result of the encoding map is spliced with the semantic vector of the original text and then classified, enhancing the representation ability of the model network for the semantic features and word order features of steganographic texts.

[0040] As Figure 6 shown, the network parameters of the network model of the present invention are as follows.

[0041] Step 5.1, after inputting the text into the BERT encoder network, obtain the encoder vector group output E of each sentence of the text o ,E D and E B ;

[0042] The structure of the vector group output E of the text is as follows:

[0043]

[0044] Among them, H represents the vector dimension preset during embedding encoding, L represents the fixed length preset during embedding encoding, N represents the number of words contained in this segment of text in the encoding end. If N < L - 1, then the subsequent L - N + 1 positions are filled with 0; if N ≥ L - 1, then take the first L - 1 words. represents the H-dimensional vector obtained after embedding encoding of the i-th word in the original sentence sequence. In particular, represents the semantic vector of the entire sentence sequence of the text.

[0045] The method of the present invention performs embedding encoding processing on the input original text, scrambled text, and reconstructed text respectively to obtain the encoding vector groups corresponding to the text carriers.

[0046] Step 5.2, input the vector groups E o and E D into the cosine feature layer of the feature extractor network to obtain the cosine similarity matrix M of the encoding vector groups of the original text and the scrambled text D ;

[0047] The process of obtaining the cosine similarity matrix M D is as follows:

[0048] First, the structure of the cosine similarity matrix M D is M L×L, L represents the fixed length preset during the embedding coding in step 5.1.

[0049] Then, calculate the cosine similarity matrix M D in M ij The formula is as follows:

[0050]

[0051] Among them, the superscripts O and D respectively represent the vector groups of the original text and the scrambled text, and the subscripts i and j represent the vector positions in the vector group. W i and W j represent the i-th and j-th vectors in the vector group.

[0052] Step 5.3, regard M D as the input image of the convolutional neural network (CNN), and input it into the convolutional layer of the feature extractor network to obtain the word order semantic difference feature coding map O D ;

[0053] O D The obtaining process is as follows:

[0054] First, obtain the calculation formula of single-layer convolution according to the structure of the convolutional neural network (CNN):

[0055] I L = σ(W · I L-1 + B)

[0056] Among them, I L represents the output result of the L-th convolutional layer, W is the weight, B is the bias, and σ is the activation function. In the embodiment, σ uses the ReLU activation function.

[0057] Then, regard M D as an L×L image, where the value of each pixel M ij is normalized from [-1, 1] to [0, 1], and pass M D through a multi-layer convolutional neural network:

[0058] Among them, F CNN represents single-layer convolution, and the superscript N represents N times of convolution.

[0059] Step 5.4, perform similar processing on E o and E B as in steps 5.2 and 5.3 to obtain the word order semantic difference feature coding map O B :

[0060] First, the structure of the cosine similarity matrix M B is ML×L , L represents the fixed length preset during the embedding encoding in Step 5.1.

[0061] Then, calculate the cosine similarity matrix M B In M ij The formula is as follows:

[0062]

[0063] Among them, the superscripts O and B respectively represent the vector groups of the original text and the reconstructed text, and the subscripts i and j represent the vector positions in the vector groups. W i And W j Represent the i-th and j-th vectors in the vector groups.

[0064] Finally, regard M B As an L×L image, where each pixel M ij The value of is normalized from [-1, 1] to [0, 1], and M B Pass through a multi-layer convolutional neural network:

[0065] Among them, F CNN Represents a single-layer convolution, and the superscript N represents N convolutions.

[0066] Step 5.5, flatten O D Into a one-dimensional vector L D , and the first vector E O1 Of the output vector group of the original text pass through the linear splicing layer, linear layer, and softmax output layer of the classifier to obtain the detection result of the text scrambling detection scheme;

[0067] The process of obtaining the detection result is as follows:

[0068] First, flatten O D Into a one-dimensional vector L D , The structure of O D Is l×l, where l is the output dimension set by the last convolutional layer. The flattening algorithm is: L D ((x - 1)×l + y) = O D (x, y). Where x and y represent the position coordinates of the original data in the two-dimensional vector.

[0069] Then, splice L D And the first vector E O1 Of the output vector group linearly to obtain the combined vector C D .

[0070] Finally, the combined vector C DAfter passing through multiple linear layers and the softmax activation function, the binary detection result Y=(0,1) of the text scrambling detection scheme is obtained (0 represents non-steganographic text, and 1 represents steganographic text).

[0071]

[0072] Among them, F CNN represents a single-layer convolution, the superscript N represents N convolutions, W is the weight, B is the bias, and σ is the activation function. In the embodiment, σ uses the ReLU activation function in the linear layer and the softmax activation function in the softmax output layer. The final Y selects the closest result in (0,1) as the final Y value.

[0073] Step 5.6, perform a similar process on O B to obtain the detection result of the text reconstruction detection scheme:

[0074] First, flatten O B into a one-dimensional vector L B , and the structure of O B is l×l, where l is the output dimension set by the last convolutional layer. The flattening algorithm is: L B ((x - 1)×l + y)=O B (x,y). Where x and y represent the position coordinates of the original data in the two-dimensional vector.

[0075] Then, linearly concatenate L B with the first vector E O1 of the output vector group to obtain the combined vector C B .

[0076] Finally, after passing the combined vector C B through multiple linear layers and the softmax activation function, the binary detection result Y=(0,1) of the text scrambling detection scheme is obtained (0 represents non-steganographic text, and 1 represents steganographic text).

[0077]

[0078] Among them, F CNN represents a single-layer convolution, the superscript N represents N convolutions, W is the weight, B is the bias, and σ is the activation function. In the embodiment, σ uses the ReLU activation function in the linear layer and the softmax activation function in the softmax output layer. The final Y selects the closest result in (0,1) as the final Y value.

[0079] Step 6: Train the steganographic text detection model so that the output result of the classifier reaches the error range threshold with the true annotation result in the dataset, and update the parameters of the encoder, feature extractor, and classifier.

[0080] The loss function used for training is the cross-entropy loss function:

[0081] The cross-entropy loss function for a single sample is:

[0082]

[0083] where Y is the true label, is the predicted label.

[0084] For all text samples in each batch of the dataset, after passing them through the model network in sequence, the loss value of each single sample is calculated and then summed up to obtain the total loss value.

[0085] Step 7: Use the trained steganographic text detection model to detect the target text.

[0086] Randomly select a text segment as the input, obtain the scrambled text of this text segment after random scrambling, and input the scrambled text into the trained large language model to obtain the reconstructed text output by the large language model. Input the original text, scrambled text, and reconstructed text into the trained steganographic text detection model to obtain the detection result of whether this text is steganographic text.

[0087] Example:

[0088] In the experiment of the present invention, an open-source dataset selected from the network was used, and the language was English. This dataset contains text samples in three scenarios (social platform speeches, news manuscripts, and movie reviews), and the steganographic texts in it were encrypted using two steganographic algorithms, HC and AC. For each algorithm in each scenario, this dataset contains 10,000 non-steganographic texts and 10,000 steganographic texts, totaling 120,000 texts. All texts were divided into training set and test set text data according to a ratio of 4:1.

[0089] The training, testing, and inference processes of this experiment were based on the Python implementation of Pytorch. At the same time, the training process was accelerated using GeForce RTX4090 GPU and CUDA 11.8. The large language model selected was Qwen2-1.5B. The fixed length L = 128 was preset for the encoder of the steganographic text detection model network, and the preset vector dimension H = 768. The ReLU was used as the non-linear activation function σ inside the convolutional neural network and the linear layer. When training the model, to strengthen regularization and avoid overfitting, a dropout mechanism and the Adam optimization method were added.

[0090] Training of large language model: Select all non-steganographic texts from the training set, randomly shuffle the words in the text sentences to obtain a set of unordered word sequences, and feed this word sequence into the large language model. Require the large language model to reconstruct it, and the training objective of the reconstruction is the original non-steganographic text. Update the parameter values of the large language model through the LoRA mechanism, and finally obtain the trained large language model.

[0091] One round of training process: Randomly select 1000 steganographic texts and 1000 non-steganographic texts of each algorithm in each scenario from the training set. First, shuffle the original text to obtain the shuffled text, then input the shuffled text into the trained large language model to obtain the reconstructed text. Input the original text, shuffled text, and reconstructed text of the text sample into the steganographic text detection model network to obtain the detection results predicted by the model network.

[0092] After the detection is completed, for each text, use the cross-entropy loss function to calculate the error between the output detection result and the true annotation result in the dataset.

[0093] Finally, calculate the parameter gradients of the encoder network, feature extractor, and classifier network according to the error, and update the parameter values according to the Adam optimizer and learning rate. Similarly, during each round of training, the model is tested with the test set, and during the test, the network parameter gradients are not updated.

[0094] Train for multiple epochs until the loss value of the test set no longer decreases, then the training process is completed, and the model is exported for steganographic text detection operations.

[0095] As Figure 7 shown, it is the comparison effect between the steganographic text detection method of the present invention and the F-CNN method. The steganographic text detection technology based on text reconstruction and word order semantic features of the present invention is significantly superior to the performance of the F-CNN method in both the text shuffling detection scheme and the text reconstruction detection scheme. It can be seen that the present invention has a certain degree of innovation and has certain practical value.

[0096] Except for the technical features described in the specification, they are all well-known technologies to those skilled in the art. The present invention omits the description of well-known components and well-known technologies to avoid redundancy and unnecessary limitation of the present invention. The described implementation manners in the above embodiments do not represent all implementation manners consistent with the present application. Based on the technical solution of the present invention, various modifications or deformations that can be made by those skilled in the art without creative labor are still within the protection scope of the present invention.

Claims

1. A steganographic text detection technology based on text reconstruction and word order semantic features, characterized in that: The steps include: Step 1: Obtain a steganographic text dataset, which contains multiple text fragments captured in various scenarios. Each text fragment has a corresponding annotation indicating whether it is a steganographic text; Step 2: For the non-steganographic text items in the data set, randomly shuffle the words in the text sentences to obtain a set of disordered word sequences, and send this word sequence to the large language model, requiring the large language model to reconstruct it, and the reconstructed training target is the original non-steganographic text; Step 3: Train the large language model in the training set so that the average error between the text reconstructed by the large language model and the original non-steganographic text reaches the error range threshold, and update the parameters of the large language model; Step 4: randomly shuffle the text items in the data set (hereinafter referred to as the original text) to obtain a set of disordered word sequences (hereinafter referred to as the shuffled text), and then send them to the trained large language model to obtain the reconstructed text output by the large language model (hereinafter referred to as the reconstructed text); Step 5: Input the original text, the scrambled text and the reconstructed text into the stegographic text detection model respectively, and obtain the binary detection result Y=(0,1) of a single text message (0 represents non-steganographic text, 1 represents stegographic text); The text steganalysis detection model includes an encoder, a feature extractor, and a classifier; the encoder network is composed of the BERT architecture, the feature extractor network contains a cosine feature layer and a convolution layer, and the classifier contains a linear concatenation layer, a linear layer, and a Softmax output layer. After the original text, the shuffled text, and the reconstructed text are input into the encoder, the encoder output vector E is obtained respectively. o , E D and E B . o and E D The cosine feature layer and convolution layer of the input feature extractor network are used to obtain the word order semantic difference feature encoding map O of the original text and the scrambled text. D ; E o and E B The cosine feature layer and convolution layer of the input feature extractor network are used to obtain the word order semantic difference feature encoding map O of the original text and the reconstructed text. B , the output vector E of the original text O With feature encoding graph O D Input the classifier, and get the detection result of the text scrambling detection scheme through the linear concatenation layer, linear layer and Softmax output layer (0 or 1, 0 represents non-steganographic text, 1 represents steganographic text); the output vector E of the original text O With feature encoding graph O B By inputting the classifier, the detection results of the text reconstruction detection scheme can also be obtained. Both schemes can realize the detection technology of steganographic text. Step 6: Train the steganographic text detection model so that the error between the output result of the classifier and the actual annotation result in the data set reaches the error range threshold, and update the parameters of the encoder, feature extractor and classifier. Step 7: Use the trained steganographic text detection model to detect the target text. A text segment is randomly selected as input, and the scrambled text of the text segment is obtained after random scrambling. The scrambled text is input into the trained large language model to obtain the reconstructed text output by the large language model. The original text, scrambled text and reconstructed text are input into the trained stegotext detection model to obtain the detection result of whether the text is stegotext.

2. The method according to claim 1, characterized in that: In the above steps 2 and 3, for the non-steganographic text items in the data set, the words in the text sentences are randomly shuffled to obtain a set of disordered word sequences, and the word sequences are sent to the large language model, requiring the large language model to reconstruct them, and the reconstructed training target is the original non-steganographic text. The implementation method is as follows: (1) designing instruction information input to the large language model, the instruction information includes instructions requiring the large language model to reconstruct the input text. (2) Construct a communication scenario that enables the large language model to perform a specified task, and input instruction information and the scrambled text to be reconstructed into the large language model in turn. (3) The target result of the large language model reconstruction is set as the original sentence to achieve error reduction and parameter update during the training process.

3. The method according to claim 1, characterized in that: In step 4, the trained large language model is used to process the original text as follows: Select a sample from the steganographic text dataset, use the trained large language model, set the same communication scenario and instruction information as when training and testing the large language model, and send the randomly scrambled text into the trained large language model to obtain the reconstructed text output by the large language model.

4. The method according to claim 1, characterized in that: In step 5, the original text, the shuffled text, and the reconstructed text are input into the BERT encoder network to obtain the encoder vector group output E for each sentence of text. o ,E D and E B , the vector group E o and E D Input the cosine feature layer of the feature extractor network to obtain the cosine similarity matrix M of the original text and the scrambled text encoding vector group D ; Group vector E o and E B Input the cosine feature layer of the feature extractor network to obtain the cosine similarity matrix M of the encoding vector group of the original text and the reconstructed text B The process includes the following: (1) Cosine similarity matrix M D The structure is M L×L ,L represents the fixed length preset by the BERT encoder network. (2) Calculate the cosine similarity matrix M D Medium ij The formula is as follows: Among them, the superscripts O and D represent the vector groups of the original text and the scrambled text respectively, the subscripts i and j represent the vector positions of the vector groups, and W i With W j Represents the i-th and j-th vectors in the vector group. (3) Cosine similarity matrix M B The structure is M L×L ,L represents the fixed length preset by the BERT encoder network. (4) Calculate the cosine similarity matrix M B Medium ij The formula is as follows: Among them, the superscripts O and B represent the vector groups of the original text and the reconstructed text respectively, the subscripts i and j represent the vector positions of the vector groups, and W i With W j Represents the i-th and j-th vectors in the vector group.

5. The method according to claim 1 or 4, characterized in that: In step 5, the cosine similarity matrix M D The image is regarded as the input image of the convolutional neural network (CNN), and the convolution layer of the input feature extractor network is used to obtain the word order semantic difference feature encoding map O of the original text and the scrambled text. D ; The cosine similarity matrix M B The image is regarded as the input image of the convolutional neural network (CNN), and the convolution layer of the input feature extractor network is used to obtain the word order semantic difference feature encoding map O of the original text and the reconstructed text. B The process includes the following: First, according to the construction of convolutional neural network (CNN), the calculation formula of single-layer convolution is obtained: AND L =σ(W I L-1 +B) Among them, I L represents the output result of the L-th convolutional layer, W is the weight, B is the bias, σ is the activation function, and in the embodiment, σ uses the ReLU activation function. Then, M D is considered as an L×L image, where each pixel M ij The value of is normalized from [-1,1] to [0,1], and M D Through multi-layer convolutional neural network: Among them, F CNN Represents a single layer of convolution, the superscript N represents N convolutions, and batch normalization and maximum pooling are performed after each layer of convolution. Similarly, M B Through multi-layer convolutional neural network: Among them, F CNN Represents a single layer of convolution, the superscript N represents N convolutions, and batch normalization and maximum pooling are performed after each layer of convolution.

6. The method according to claim 1 or 5, characterized in that: In the step 5, O D Flattened into a one-dimensional vector L D , and the first vector E of the output vector group of the original text O1 The detection result of the text shuffle detection scheme is obtained through the linear concatenation layer, linear layer and softmax output layer of the classifier; B Flattened into a one-dimensional vector L B , and the first vector E of the output vector group of the original text O1 The detection results of the text reconstruction detection scheme are obtained through the linear concatenation layer, linear layer and softmax output layer of the classifier.

7. The method according to claim 1, characterized in that In the step 6, the stegotext detection model is trained, and the loss function used in the training is the cross entropy loss function. For all text samples in each batch of the data set, it is necessary to pass through the model network in turn, calculate the loss value of each single sample, and then perform superposition and summation to obtain the total loss value. The error between the output result of the classifier and the actual annotation result in the data set reaches the error range threshold, and the parameters of the encoder, feature extractor and classifier are updated.

8. The method according to claim 1, characterized in that In the step 7, a text segment is randomly selected as input, and after random shuffling, a scrambled text of the text segment is obtained, and the scrambled text is input into the trained large language model to obtain the reconstructed text output by the large language model. The original text, the scrambled text, and the reconstructed text are input into the trained stegotext detection model to obtain a detection result of whether the text is a stegotext.

Citation Information

Cited By

  • Social text site identification method and system based on claim generation and consistency analysis

    CN120780844A