The invention discloses a steganographic
text detection technology based on text reconstruction and
word order semantic features, and belongs to the technical field of
information hiding. The method comprises the following steps: selecting an
open source English
steganography text
data set containing three fields of
social platform speaking, news manuscripts and film review; words in non-steganographic text sentences in the
data set are randomly disorganized and sent to a large
language model for reconstruction training, and the training target of reconstruction is an original text; and randomly disorganizing a
data set text, and inputting the disorganized data set text into the trained model to generate a reconstructed text. And finally, inputting the original text, the disordered text and the reconstructed text into a
steganography detection model, calculating a
cosine similarity matrix, extracting a
semantic difference feature map through a CNN, flattening, splicing semantic vectors of the original text, and inputting the semantic vectors into a classifier to output a binary detection result. Through
continuous training, an error between a classifier output result of the model and a real labeling result is continuously reduced, so that parameters of a feature extractor and a classifier are optimized, and the model
steganography detection accuracy is improved. According to the method,
semantic information and
word order features are utilized for classification, high accuracy, low cost and good
interpretability are achieved, steganographic texts can be effectively detected, and information safety is improved.