Sensitive information detection technology and system based on multi-mode and steganography detection

Through multimodal and steganography detection technology, combined with DistilBERT-LSTM and cross-modal inference mechanism, the problem of accurate detection of sensitive information in multimodal data is solved, efficient and accurate detection of sensitive information is achieved, and data security is improved.

CN120277678AActive Publication Date: 2025-07-08JINAN UNIVERSITY

Patent Information

Application Number
CN202510537308.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-07-08
Estimated Expiration
2045-04-27

AI Technical Summary

Technical Problem

现有敏感信息检测技术难以有效整合多模态数据,无法识别复杂隐写技术,且处理效率低下,无法实时准确识别潜在敏感信息。

Method used

Multimodal and steganography detection technology is adopted to build a multimodal neural network through anti-file format tampering and bypass detection, sensitive information detection based on DistilBERT-LSTM, and joint content and file steganography detection, and combine cross-modal inference mechanism to achieve accurate detection of multimodal data.

Benefits of technology

It significantly improves the comprehensiveness and accuracy of sensitive information detection, improves the operating efficiency of the system, can process large-scale data in real time, prevents complex steganography methods, and provides comprehensive data security guarantees.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277678A_ABST
    Figure CN120277678A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of information security, in particular to a sensitive information detection technology and system based on multi-mode and steganography detection.The sensitive information detection technology and system based on multi-mode and steganography detection.The method comprises the steps that firstly, through a format tampering prevention module, file header information and binary data analysis are utilized, the real type of a file is verified, and format camouflage attacks are prevented; the sensitive data detection module adopts DistilBERT and LSTM structures to preprocess a text, extract text embedding, mine text features through a bidirectional LSTM and Attention mechanism, and finally accurately judge whether the text contains sensitive information or not; the joint content and file steganography detection module fuses a convolutional neural network and multi-modal learning, adopts corresponding models for different data types to carry out steganography information detection, integrates results through a multi-modal neural network, and applies multi-modal feedback and a cross-modal reasoning mechanism to improve the detection accuracy; according to the method, the accuracy, comprehensiveness and system adaptability of sensitive information detection are effectively enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information security technology, and particularly to a sensitive information detection technology and system based on multi-modal and steganography detection, which is used to accurately detect sensitive information and potential steganographic content in multi-modal data, and effectively ensure data security. Background Art

[0002] In today's information age, the storage and transmission methods of data are becoming increasingly diverse and complex, and the security threats faced by sensitive information are also increasing day by day. There are many deficiencies in traditional sensitive information detection methods: First, from the perspective of data modality, most methods are limited to detecting data of a single modality, such as only focusing on text data or image data. However, in actual application scenarios, sensitive information often exists in the form of a combination of multiple modalities. For example, a file may contain both text descriptions and image content at the same time. Existing technologies are difficult to effectively integrate and comprehensively analyze these different modality information.

[0003] Secondly, with the continuous development of steganography technology, the means of hiding sensitive information are becoming more and more sophisticated. Lawbreakers can make sensitive information almost indistinguishable from ordinary data by modifying the least significant bits of files, using redundant spaces in files, or adopting complex encryption and hiding algorithms. Traditional detection methods are difficult to identify these hidden sensitive contents due to the lack of in-depth analysis capabilities for complex steganography technologies.

[0004] In addition, when facing large-scale data, current sensitive information detection systems usually have the problem of low processing efficiency and are unable to accurately identify potential sensitive information from massive data in real time, which poses a severe challenge to data security.

[0005] Although technologies such as deep learning, multi-modal analysis, and steganography are all making continuous progress, there is currently a lack of a sensitive information detection technology and system that can organically integrate these advanced technologies, which can not only effectively integrate multi-modal information, but also have strong steganography detection capabilities, and at the same time can be efficiently deployed. Therefore, developing a sensitive information detection technology and system based on multi-modal and steganography detection is of great significance for improving the accuracy and comprehensiveness of sensitive information detection and ensuring data security. Summary of the Invention

[0006] In view of the deficiencies of existing sensitive information detection technologies, the present invention proposes a sensitive information detection technology and system based on multi-modal and steganography detection. By organically integrating technologies for preventing file format tampering, sensitive information detection, and steganography detection, the method and system achieve accurate detection of sensitive information in multi-modal data, significantly improving the accuracy and comprehensiveness of detection, as well as the operating efficiency and adaptability of the system.

[0007] The object of the present invention is to provide a sensitive information detection technology and system based on multimodality and steganography detection. This method and system can effectively prevent file format tampering from bypassing detection, accurately identify sensitive information in text, comprehensively discover steganographic content hidden in different modality data, and provide all-round protection for data security.

[0008] The present invention proposes a sensitive information detection technology based on multimodality and steganography detection, including:

[0009] Receiving a file to be detected;

[0010] Performing anti-file format tampering bypass detection, including: analyzing file header information, extracting binary structure features and performing format verification to determine the true type of the file;

[0011] Based on the true type, performing sensitive information detection, including: after preprocessing the text, using the DistilBERT model to extract text embedding vectors, capturing long-term dependence features of the text through a bidirectional LSTM and an Attention mechanism and focusing on sensitive regions, and using a fully connected layer to determine whether the text contains sensitive information;

[0012] Performing joint content and file steganography detection, including: respectively using Transformer, LSTM, and CNN models for steganographic feature extraction for text, file binary structure, and images, constructing a multimodal neural network to integrate detection results of different modalities, and applying a cross-modal reasoning mechanism to cross-validate different modality features to improve detection accuracy;

[0013] Outputting comprehensive detection results, including file format risk rating, sensitive information distribution and categories, and steganographic information location and content.

[0014] Preferably, the anti-file format tampering bypass detection specifically includes:

[0015] Identifying the true type of the file based on the magic number in the file header;

[0016] Using a sliding window hash check algorithm to extract and verify header sample information;

[0017] Analyzing the file binary data structure, and calculating byte frequency features and data block entropy value features;

[0018] Comparing the extracted feature vectors with the feature library for similarity to determine the most likely true type of the file;

[0019] Outputting the evaluation result of the file format tampering risk level.

[0020] Preferably, the similarity comparison between the feature vectors and the feature library includes:

[0021] Calculate the cosine similarity between the extracted feature vector and the standard vector in the feature library;

[0022] Based on the similarity calculation result, determine the true type of the file;

[0023] Record the inconsistency between the declared type and the true type of the file to form a risk assessment report.

[0024] Preferably, the sensitive information detection specifically includes:

[0025] Perform a cleaning operation on the text to remove HTML tags and special symbols;

[0026] Apply a multi-level sensitive word dictionary for matching, including basic sensitive words, synonyms, fuzzy words, and domain-specific sensitive words;

[0027] Select an appropriate word segmentation tool according to the text language type, and perform word segmentation and lemmatization;

[0028] Load the pre-trained DistilBERT model to extract context-aware text embedding vectors;

[0029] Construct a bidirectional LSTM network to process the embedding vectors, capturing the long-term dependencies of the text through forward and backward directions;

[0030] Introduce the Attention mechanism to calculate the attention weights, highlighting the key regions of sensitive information;

[0031] Use a fully connected layer and the Sigmoid function to determine whether the text contains sensitive information and generate a detailed report.

[0032] Preferably, the joint content and file steganography detection specifically includes:

[0033] For the text part, use the Transformer model to analyze the document context and detect steganographic information hidden by tiny character changes;

[0034] For the binary structure of the file, use the LSTM model to analyze the characteristics of the file binary data sequence and detect embedded data; for the image part, use the CNN model to analyze the pixel-level changes of the image and identify hidden information; fuse the detection results of different modalities through a multi-modal neural network for comprehensive analysis; apply a cross-modal inference mechanism to improve the detection accuracy and robustness, and output the steganography detection results. Preferably, the multi-modal neural network includes:

[0035] Design a Transformer encoder input layer for the text modality;

[0036] Design an LSTM input layer for the binary structure modality;

[0037] Design a CNN feature extraction layer for the image modality;

[0038] Construct a feature fusion network based on the attention mechanism to automatically learn the correlation and weight between different modalities;

[0039] Output the fused multi-modal feature representation for subsequent steganographic information detection and judgment. Preferably, the cross-modal inference mechanism includes:

[0040] Establish three groups of association networks: text-image, text-binary, and image-binary; construct the mapping relationship between different modality features;

[0041] Based on the detection results and feature clues of different modalities, conduct cross-validation inference; solve the conflict of detection results between modalities, improve the confidence of steganography detection; and form the final steganographic information detection conclusion.

[0042] Preferably, in the sensitive information detection:

[0043] The output of the bidirectional LSTM at time t is where and are the hidden states of the forward and backward LSTMs respectively;

[0044] The Attention mechanism calculates the attention weight where e t =W a h t +b a , and obtains the context vector

[0045] The context vector c passes through a fully connected layer to obtain the probability of sensitive information determination.

[0046] Preferably, it further includes the following steps:

[0047] Generate a comprehensive security assessment report based on the detection results;

[0048] Provide a risk level assessment for the discovered sensitive information and steganographic content;

[0049] Mark and isolate the files confirmed to contain threats;

[0050] Feed the detection results back to the security management system for subsequent threat analysis and defense strategy update.

[0051] A sensitive information detection system based on multi-modal and steganography detection includes:

[0052] A file format tampering bypass detection prevention module, which is used to analyze the file header information, extract binary structure features and perform format verification to determine the true type of the file;

[0053] A sensitive information detection module, which is used to preprocess the text, extract text embedding vectors using the DistilBERT model, capture text features and focus on sensitive areas through a bidirectional LSTM and an Attention mechanism, and use a fully connected layer to determine sensitive information;

[0054] A combined content and file steganography detection module, which is used to perform steganography detection on text, file binary structures, and images using Transformer, LSTM, and CNN models respectively, and construct a multi-modal neural network and a cross-modal inference mechanism to integrate and analyze the detection results;

[0055] A result output module, which is used to generate and output comprehensive detection results, including file format risk ratings, sensitive information distributions and categories, and steganography information locations and contents.

[0056] The beneficial effects of the present invention are:

[0057] 1. By organically integrating technologies for preventing file format tampering, sensitive information detection, and steganography detection, a complete protection chain from file authentication to content detection and then to steganography identification is constructed, significantly improving the comprehensiveness and security of sensitive information detection.

[0058] 2. Using a deep learning architecture combining DistilBERT and bidirectional LSTM for sensitive information detection can effectively identify deformed sensitive words and context-sensitive information in the text, improving the accuracy and efficiency of detection.

[0059] 3. Innovatively introducing a cross-modal inference mechanism, through cross-validation and inference of different modal information, further improves the accuracy and robustness of steganography detection and effectively guards against complex steganography means.

[0060] 4. The system architecture is reasonably designed, the processing flow is clear, it is suitable for real-time detection of large-scale data, and has broad application prospects. Description of the Drawings

[0061] Figure 1 It is a schematic diagram of the overall framework of a sensitive information detection technology and system based on multi-modal and steganography detection of the present invention;

[0062] Figure 2 It is a schematic diagram of the anti-file format tampering bypass detection module of the present invention;

[0063] Figure 3 It is a schematic diagram of the sensitive information detection module based on the DistilBERT-LSTM structure of the present invention;

[0064] Figure 4This is a schematic diagram of the combined content and file steganography detection module of the present invention;

[0065] Figure 5 This is a schematic diagram of the bidirectional LSTM model of the present invention. Specific implementation manners

[0066] Please refer to the attached Figures 1-5 figures. The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be noted that the implementation of the present invention is not limited thereto. Under the inspiration of the present invention, those skilled in the art can make various equivalent transformations or modifications without departing from the spirit of the present invention, and these equivalent transformations or modifications all fall within the protection scope of the present invention.

[0067] Embodiment 1: Overall process of sensitive information detection technology based on multi-modal and steganography detection

[0068] As Figure 1 shown, a sensitive information detection technology based on multi-modal and steganography detection provided by the present invention mainly includes the following steps:

[0069] First, receive the file to be detected. The system receives the file that needs to be detected for sensitive information passed in by the user or other systems, and this file can be various formats of documents, images, compressed packages, etc.

[0070] Second, perform anti-file format tampering bypass detection. This step aims to prevent attackers from disguising the file type by modifying the file extension or tampering with binary data (such as disguising an executable file as a picture file) to bypass security detection. By analyzing the file header information, extracting binary structure features and performing format verification, determine the true type of the file, such as JPEG, PDF, etc., classify by type and record the format tampering risk level.

[0071] Then, based on the determined true type, perform sensitive information detection. This step combines the DistilBERT and LSTM structures. After preprocessing the text, use the DistilBERT model to extract text embedding vectors, capture the long-term dependence features of the text and focus on the sensitive areas through the bidirectional LSTM and Attention mechanisms, use the fully connected layer to determine whether the text contains sensitive information, and finally generate a detailed report in combination with the sensitive dictionary matching results to clarify the category and location of the sensitive information.

[0072] Next, perform combined content and file steganography detection. This step combines convolutional neural networks and multi-modal learning, uses the Transformer, LSTM and CNN models to detect steganographic information for text, file binary structures and images respectively, then uses a multi-modal neural network for comprehensive analysis, and uses a cross-modal reasoning mechanism to improve the detection accuracy and robustness, and outputs the steganography detection results and related information.

[0073] Finally, output the comprehensive detection results. The system integrates the detection results of the above-mentioned steps to generate a comprehensive report including the risk rating of the file format, the distribution and categories of sensitive information, and the location and content of steganographic information, providing a basis for subsequent decision-making and processing.

[0074] The above steps form a complete detection process. There are clear logical associations between the steps. The output of the previous step serves as the input of the next step, jointly achieving a comprehensive detection of sensitive information in multi-modal data.

[0075] Example 2: Detection of bypassing file format tampering

[0076] As Figure 2 shown, this example details the specific steps for detecting bypassing file format tampering.

[0077] First, identify the true file type based on the magic number in the file header. The specific hexadecimal byte sequence in the file header, i.e., the magic number, is the key basis for identifying the true file type and does not rely on the easily tampered file extension. For example, the magic number of a GIF file is [474946383761], and that of a JPEG file is [FFD8FF]. The system uses an efficient byte reading algorithm to locate the starting position of the file, extract an appropriate amount of byte data at the beginning, and precisely compare it with the pre-stored standard magic numbers of various file types. This process strictly follows the byte order and numerical matching principle. Once the match is successful, the true file format can be firmly determined, effectively resisting the attacker's behavior of disguising the file by changing the extension.

[0078] Second, use the sliding window hash verification algorithm to extract and verify the header sample information. Based on the statistical characteristics and information entropy of the file type, this algorithm determines an appropriate sliding window length (128 - 512 bytes for common files), moves the window from the beginning of the file at a fixed step size, extracts byte data as the header sample, calculates its SHA-256 hash value, and compares it with the set of pre-stored standard template hash values of various file headers. If the match is successful, then perform a byte-by-byte precise comparison to confirm the integrity and accuracy of the header information, thereby enhancing the basis for judging the file authenticity.

[0079] Next, analyze the binary data structure of the file and calculate the byte frequency feature and data block entropy value feature. Use professional tools to analyze the binary structure of the file, identify key features such as the directory structure and data block distribution of the file. Calculate the frequency distribution of each byte in the file to form a feature vector; at the same time, divide the file into multiple data blocks and calculate the entropy value of each block to form an entropy value vector. These features are unique for different types of files and can be used as important bases for judging the true file type.

[0080] Then, the extracted feature vectors are compared with the feature library for similarity. The system calculates the cosine similarity between the extracted feature vectors and the standard vectors in the feature library, and determines the most likely true type of the file based on the magnitude of the similarity. The cosine similarity calculation formula is:

[0081]

[0082] Where, and represent the extracted feature vectors and the standard vectors in the feature library respectively, A i and B i represent the i-th element in the vector, and n represents the dimension of the vector. The closer the similarity value is to 1, the more similar the two vectors are, and the more likely the true type of the file is consistent with the type corresponding to the standard vector.

[0083] Finally, the evaluation result of the file format tampering risk level is output. Based on the results of the foregoing detection steps, the system records the inconsistency between the declared type of the file (represented by the extension) and the actual detected type, and classifies the risk rating into three levels: low, medium, and high according to the degree and nature of the inconsistency. At the same time, a detailed evaluation report is generated, including the true type, declared type, risk level, and potential security threats of the file, providing reliable file type information for subsequent sensitive information detection.

[0084] Through the above steps, the system can effectively prevent attackers from bypassing detection through format tampering, ensure that subsequent sensitive information detection and steganography detection are based on the true type of the file, and improve the security and accuracy of the entire detection process.

[0085] Embodiment 3: Sensitive Information Detection Based on DistilBERT-LSTM Structure

[0086] As Figure 3 shown, this embodiment details the specific steps of sensitive information detection based on the DistilBERT-LSTM structure.

[0087] First, perform a cleaning operation on the text to remove HTML tags and special symbols. The system uses regular expressions to match and delete HTML tags, special characters, and redundant whitespace characters in the text, and retains the pure text content. This step ensures the purity of the text to be analyzed, reduces noise interference, and provides a good basis for subsequent sensitive information analysis.

[0088] Secondly, a multi-level sensitive word dictionary is applied for matching. The system constructs a multi-level sensitive word dictionary that includes basic sensitive words, synonyms, fuzzy words, and domain-specific sensitive words. The cleaned text is matched word by word. When sensitive words in the dictionary are found in the text, the system directly marks these words and records their position information in the text. This preliminary screening can quickly identify obvious sensitive content and provide key areas of focus for subsequent in-depth analysis.

[0089] Then, an appropriate word segmentation tool is selected according to the text language type to perform word segmentation and lemmatization. For Chinese texts, the system uses the jieba word segmentation tool; for English texts, the NLTK toolkit is used. After word segmentation, the system also performs lemmatization on each word. For example, "running" in English is restored to "run", and "mice" is restored to "mouse", etc., enabling the model to more accurately understand and process the semantic information in the text.

[0090] Next, a pre-trained DistilBERT model is loaded to extract context-aware text embedding vectors. DistilBERT is a lightweight version of BERT, retaining most of BERT's performance, but with the model size reduced by approximately 40% and the inference speed increased by approximately 60%. The system inputs the text after word segmentation and lemmatization into the DistilBERT model. Through the Transformer encoder in the model, the semantic relationships in the text are deeply understood and encoded, and the embedding vector representation of the text is extracted. These embedding vectors can effectively capture the context information in the text, including semantic associations between words, sentence structures, and grammatical relationships, etc.

[0091] Subsequently, a bidirectional LSTM network is constructed to process the embedding vectors, capturing the long-term dependencies of the text through forward and backward directions. As Figure 5 shown, the bidirectional LSTM consists of a forward LSTM and a backward LSTM, capable of processing the front and back context information of the text simultaneously. Let the sequence of text embedding vectors output by the DistilBERT model be X = [x1, x2, …, x n , where x i is the embedding vector at the i-th time step, with a dimension of d. The output of the bidirectional LSTM at time t is:

[0092]

[0093] where, and are the hidden states of the forward and backward LSTMs respectively, and the final output vector is formed by concatenation, with a dimension of 2d. This structure can effectively capture the long-distance dependencies in the text and provide a richer feature representation for the accurate identification of sensitive information.

[0094] Then, the Attention mechanism is introduced to calculate the attention weights and highlight the key regions of sensitive information. Based on the output H = [h1, h2, …, h n of the bidirectional LSTM, the system introduces the Attention mechanism to automatically learn the importance weights at different time steps. First, calculate the attention scores:

[0095] e t = W a h t + b a ,

[0096] where W a (with dimension 2d×d a ) and b a (with dimension d a ) are learnable parameters. Then, normalize the scores to weights through the softmax function:

[0097]

[0098] Finally, use these weights to perform a weighted sum on the LSTM output to obtain the context vector:

[0099]

[0100] This context vector c can highlight the key parts related to sensitive information in the text, enabling the system to more accurately identify and locate sensitive content.

[0101] Finally, use the fully connected layer and the Sigmoid function to determine whether the text contains sensitive information and generate a detailed report. The system inputs the context vector c into the fully connected layer, and maps the output to a probability value between 0 and 1 through the Sigmoid activation function, indicating the likelihood that the text contains sensitive information:

[0102] p = σ(W f c + b f ),

[0103] where W f and b f are the parameters of the fully connected layer, and σ represents the Sigmoid function. If the probability value exceeds a preset threshold (usually 0.5), it is determined that the text contains sensitive information. The system will combine the aforementioned sensitive dictionary matching results to generate a detailed sensitive information detection report, including the category of sensitive information (such as trade secrets, personal privacy, political sensitivity, etc.), location (start and end positions in the text), and credibility score and other information, providing important basis for subsequent security processing.

[0104] Through the above steps, the sensitive information detection based on the DistilBERT-LSTM structure can effectively identify explicit and implicit sensitive content in the text, with high accuracy and efficiency.

[0105] Example 4: Joint Content and File Steganography Detection

[0106] As Figure 4 shown, this example details the specific steps of joint content and file steganography detection.

[0107] First, for the text part, the Transformer model is used to analyze the document context to detect steganographic information hidden by minute character changes. The system first constructs and trains a Transformer model to enable it to learn the difference between normal text and text containing steganographic information. The training data includes a large number of normal text samples and steganographic text samples generated by various text steganography techniques (such as adding and deleting spaces, replacing special characters, etc.). After training, the text to be detected is input into the model, and the model extracts feature vectors related to steganographic information by deeply analyzing features such as the context relationship, vocabulary distribution, and syntactic structure of the text. These feature vectors can reflect possible abnormal patterns in the text, such as the unnaturalness of vocabulary usage and the abnormality of semantic coherence. Based on these features, the system uses a classifier to determine whether the text contains steganographic information and attempts to extract potential hidden data.

[0108] Second, for the file binary structure, the LSTM model is used to analyze the file binary data sequence features to detect embedded data. The system constructs an LSTM model and trains it using a large amount of binary data containing steganographic data and normal files, enabling the model to learn to identify steganographic patterns in binary data. The model focuses on the binary structure features of the file, such as the arrangement of data blocks and the occurrence frequency of specific byte sequences. For the file to be detected, the system inputs its binary data into the trained LSTM model, and the model extracts feature representations related to steganographic information through the analysis of the data sequence. These features can reveal abnormal patterns in the file binary data, such as the embedding location of hidden data and data anomalies. Based on these features, the system determines whether there is binary steganography in the file and locates the possible positions of hidden data.

[0109] Then, for the image part, a CNN model is used to analyze pixel-level changes in the image and identify hidden information. The system constructs a CNN model and trains the model with a large number of images with and without steganographic information. The CNN model focuses on analyzing pixel-level features of the image and can effectively identify the tiny pixel changes caused by steganography. For the image to be detected, the system inputs it into the trained CNN model. The model extracts the feature maps of the image through convolution and pooling operations. These feature maps reflect the possible steganographic traces in the image, such as abnormal distributions of pixel values and abnormal changes in color channels. Based on these features, the system uses a classifier to determine whether the image contains steganographic information.

[0110] Next, a multi-modal neural network is constructed to fuse feature information of different modalities for comprehensive analysis. The system designs a multi-modal neural network based on the attention mechanism and takes the detection results of the aforementioned three modalities (text, binary, image) as inputs. The network designs dedicated input layers for each modality: the text modality uses a Transformer encoder, the binary structure modality uses an LSTM layer, and the image modality uses a CNN feature extraction layer. After passing through their respective processing layers, the features of these different modalities are fed into the multi-modal fusion network. This network automatically learns the correlations and importance weights between different modalities through the attention mechanism to achieve effective feature fusion. The fused feature representation can comprehensively reflect the possible steganographic information in the file and provide a more comprehensive basis for detection.

[0111] Finally, a cross-modal reasoning mechanism is applied to improve the detection accuracy and robustness and output the steganalysis results. The system constructs three groups of associated networks: text-image, text-binary, and image-binary, to establish the mapping relationship between feature information of different modalities. Through these associated networks, the system can perform cross-modal reasoning. For example, when suspicious steganographic features are found in text detection, the system will check whether there are also abnormalities in the relevant images or binary data to cross-verify the existence of steganographic information. This cross-modal reasoning mechanism can effectively solve the biases and uncertainties that may occur in single-modal detection and improve the overall detection accuracy and robustness of the system. Finally, the system outputs comprehensive steganalysis results, including whether the file contains steganographic information, the type of steganographic information, the possible location, and the extracted hidden content, etc., providing comprehensive information support for subsequent security processing.

[0112] Through the above steps, the combined content and file steganalysis can comprehensively discover steganographic content hidden in different modality data and effectively prevent data leakage and security threats.

[0113] Example 5: Implementation of the Cross-Modal Reasoning Mechanism

[0114] This example details the specific implementation method of the cross-modal reasoning mechanism.

[0115] The cross-modal reasoning mechanism is one of the core innovations of the present invention. By establishing the correlation relationships between different-modal data, it realizes the cross-verification and comprehensive reasoning of information, improving the accuracy and robustness of steganography detection. The implementation of this mechanism mainly includes the following steps:

[0116] Firstly, establish an information correlation model. The system designs and trains three correlation models: text-image correlation model, text-binary correlation model, and image-binary correlation model. These models adopt deep neural network structures and can learn the semantic mapping relationships between different-modal data. For example, the text-image correlation model can learn the corresponding relationships between the visual content in an image and the text description; the text-binary correlation model can learn the correlation between the text content and the binary structural features; and the image-binary correlation model learns the mapping between the image content and its binary representation. These correlation models learn the corresponding relationships that different-modal data should have under normal circumstances through a large amount of training data, providing a benchmark for subsequent anomaly detection.

[0117] Secondly, perform cross-modal feature mapping. For the file to be detected, the system uses the above-mentioned correlation models to map the features of one modality to the feature space of another modality to form expected features. For example, the system can predict the visual features that a relevant image should have based on the text content, or infer the keywords that the corresponding text description should contain based on the image content. These expected features reflect the characteristics that different-modal data should exhibit under normal circumstances, and the deviation from the actual features may imply the existence of steganographic information.

[0118] Then, conduct cross-verification and anomaly detection. The system compares the expected features with the actual features and calculates the difference degree between the two. A significant difference may mean that the data has been tampered with or contains steganographic information. For example, if there are obvious differences between the image features predicted by the text content and the actual image features, or if the text keywords inferred from the image content are missing in the actual text, steganography may exist. The system also analyzes the consistency of the detection results of different modalities. When multiple modalities report anomalies simultaneously, the possibility of steganography is higher; when there are conflicts in the detection results, the system will further analyze the possible reasons.

[0119] Finally, conduct comprehensive reasoning and decision-making. The system uses reasoning models such as Bayesian networks or decision trees to integrate the detection results from each modality and the information from cross-verification, conduct comprehensive reasoning, and give the final steganography detection conclusion. The reasoning model considers various factors, including the detection confidence of each modality, the results of cross-modal verification, historical detection data, etc., and obtains a more accurate judgment by weighted fusion of this information. At the same time, the system also provides an explanation of the reasoning process, including which evidences support the final conclusion and the probability distribution of various possibilities, to help users understand the detection results and make appropriate security decisions.

[0120] Through the above cross-modal reasoning mechanism, the system can overcome the limitations of single-modal detection, achieve more comprehensive and accurate steganography detection, effectively prevent complex steganography means, and improve the data security guarantee ability.

[0121] Embodiment 6: Sensitive Information Detection System Based on Multi-modal and Steganography Detection

[0122] As Figure 1 shown, this embodiment describes the overall architecture and each functional module of the sensitive information detection system based on multi-modal and steganography detection.

[0123] The sensitive information detection system based on multi-modal and steganography detection provided by the present invention mainly includes four core modules: anti-file format tampering bypass detection module, sensitive information detection module, joint content and file steganography detection module, and result output module. These four modules work together to form a complete detection system, which can effectively identify sensitive information and steganographic content in multi-modal data.

[0124] The anti-file format tampering bypass detection module is mainly responsible for verifying the authenticity and integrity of the file. This module analyzes the file header information, extracts binary structure features, and performs format verification to determine the true type of the file. It includes components such as a file header magic number recognition unit, a sliding window hash check unit, a binary structure analysis unit, and a format verification unit. These components use various technical means to prevent attackers from disguising the file type by modifying the file extension or tampering with binary data, ensuring that the subsequent detection process is based on the true type of the file. The output of this module includes information such as the true type judgment of the file and the format tampering risk rating, providing an important basis for subsequent sensitive information detection.

[0125] The sensitive information detection module is based on the DistilBERT-LSTM structure to detect sensitive information in the text content. This module includes components such as a text preprocessing unit, a sensitive word matching unit, a DistilBERT encoding unit, a bidirectional LSTM processing unit, an Attention mechanism unit, and a classification judgment unit. The text preprocessing unit is responsible for cleaning the text and removing irrelevant information; the sensitive word matching unit uses a multi-level sensitive word dictionary for preliminary screening; the DistilBERT encoding unit extracts the context embedding vector of the text; the bidirectional LSTM processing unit captures the long-term dependencies of the text; the Attention mechanism unit focuses on sensitive areas; the classification judgment unit finally determines whether the text contains sensitive information. The output of this module is a detailed sensitive information detection report, including the category, location, and credibility score of the sensitive information.

[0126] The combined content and file steganography detection module integrates multiple models to comprehensively detect various steganographic information. This module includes components such as a text steganography detection unit, a file binary steganography detection unit, an image steganography detection unit, a multimodal neural network unit, and a cross-modal reasoning unit. The text steganography detection unit uses a Transformer model to analyze the document context; the file binary steganography detection unit uses an LSTM model to analyze the binary sequence features; the image steganography detection unit uses a CNN model to detect pixel-level changes; the multimodal neural network unit integrates the detection results of different modalities; the cross-modal reasoning unit improves the detection accuracy through cross-validation. The output of this module is a comprehensive steganography detection result, including information such as whether steganography exists, the type of steganography, and its location.

[0127] The result output module is responsible for integrating the detection results of the foregoing modules to generate a final detection report. This module includes components such as a result fusion unit, a risk assessment unit, a report generation unit, and an interface adaptation unit. The result fusion unit uniformly processes the detection results of each module; the risk assessment unit evaluates the overall security risk of the file; the report generation unit creates a detailed detection report; the interface adaptation unit ensures that the system can be seamlessly docked with other security management systems. The output of this module is a comprehensive security detection report, containing detailed information such as the risk rating of the file format, the distribution and categories of sensitive information, and the location and content of steganographic information.

[0128] There are clear data flow and control flow relationships among these four modules. The output of the anti-file format tampering bypass detection module serves as the input to the sensitive information detection module and the combined content and file steganography detection module; the sensitive information detection module and the combined content and file steganography detection module work in parallel, and their respective outputs are passed to the result output module; the result output module integrates all the information to generate a final report. This modular design makes the system have good scalability and maintainability, while ensuring the integrity and accuracy of the detection process.

[0129] Example 7: Application of the Sensitive Information Detection System Based on Multimodal and Steganography Detection

[0130] This example describes the application of the sensitive information detection system based on multimodal and steganography detection in an actual scenario.

[0131] The sensitive information detection system based on multimodal and steganography detection provided by the present invention can be widely applied to multiple fields such as government data security, enterprise trade secret protection, financial information security, and personal privacy protection. Taking the security audit of enterprise documents as an example, the application process and effects of the system are described below.

[0132] First, the system receives document files from the company's internal document management platform. These documents may be reports, contracts, technical documents, etc. uploaded by employees, and security audits are required to prevent the leakage of sensitive information.

[0133] Secondly, the system performs anti-file format tampering bypass detection. For received document files, the system analyzes their header information and binary structure to confirm the true type of the file. In some cases, the system successfully identified that the seemingly ordinary PPT files were actually malicious files embedded with executable code, effectively preventing security risks.

[0134] The system then performs sensitive information detection based on the true type of the file. For text content, the system uses the DistilBERT-LSTM model for in-depth analysis, which can not only identify obvious sensitive words containing corporate secrets, but also discover sensitive information expressed in disguise through contextual understanding. For example, the system can identify sensitive content involving corporate strategy, such as the third quarter launch plan for the next generation of products, even if restricted words are not explicitly used.

[0135] Next, the system performs joint content and file steganalysis detection. In one detection, the system discovered that a chart in a seemingly ordinary market report contained sensitive information such as a company's customer list. Through multimodal detection and cross-modal reasoning, the system not only discovered the existence of steganalysis, but also successfully extracted the hidden content, providing important support for enterprises to prevent internal information leakage.

[0136] Finally, the system generates a comprehensive detection report, including detailed information such as file format risk rating, sensitive information distribution and category, and steganographic information location and content. At the same time, according to the enterprise's preset security policy, files confirmed to contain sensitive information or steganographic content are marked and isolated, and the detection results are fed back to the enterprise's security management system for subsequent threat analysis and defense strategy updates.

[0137] After a period of application, the system has shown significant advantages in enterprise document security audits: the detection accuracy has increased by about 25%, and the false alarm rate has decreased by about 40%, especially in identifying complex steganography and context-sensitive information; the processing efficiency has increased by about 30%, which can meet the real-time detection needs of large-scale data; the scalability and adaptability of the system have also been verified, and it can be flexibly adjusted according to the ever-changing security needs of the enterprise.

[0138] These practical application results prove that the sensitive information detection system based on multimodal and steganography detection provided by the present invention can effectively solve the limitations of traditional detection methods, provide comprehensive and reliable protection for data security, and has broad application prospects.

[0139] The sensitive information detection technology and system based on multimodality and steganography detection provided by the present invention have significant industrial practicability and can be widely applied in multiple fields:

[0140] 1. In the field of government data security, the system can help government departments review whether the documents contain national secrets or sensitive information, preventing information leakage and security incidents.

[0141] 2. In terms of protecting enterprise trade secrets, the system can detect whether the internal documents and externally sent materials contain sensitive information such as core technologies, business strategies, customer information, etc., protecting the commercial interests of enterprises.

[0142] 3. In the field of financial information security, the system can identify sensitive financial data and personal information in financial documents, ensuring compliance and preventing data leakage.

[0143] 4. In terms of personal privacy protection, the system can detect and filter personal sensitive information such as ID numbers, bank accounts, medical records, etc., preventing privacy violations.

[0144] 5. In the field of network security protection, the system can be used as a deep detection component and integrated into a larger security architecture to provide comprehensive data security protection.

[0145] The present invention adopts a modular design, which is easy to implement and deploy. It can be flexibly configured and optimized according to the requirements of different application scenarios, and has strong practicability and adaptability. At the same time, the high efficiency and high accuracy of the system enable it to meet the requirements of real-time detection of large-scale data in actual applications, and have significant industrialization value.

[0146] The present invention provides a sensitive information detection technology and system based on multimodality and steganography detection. Through three core steps: bypassing detection by preventing file format tampering, sensitive information detection based on the DistilBERT-LSTM structure, and joint content and file steganography detection, it realizes the comprehensive detection of sensitive information in multimodal data. The system has significant advantages in aspects such as file authenticity verification, sensitive content recognition, and steganographic information discovery, and can effectively improve the accuracy, comprehensiveness, and system operation efficiency of sensitive information detection, providing a strong guarantee for data security.

[0147] The innovative points of the present invention are mainly reflected in: introducing an anti-file format tampering bypass detection mechanism to ensure that the detection is based on the true type of the file; using the DistilBERT-LSTM structure for sensitive information detection, which improves the accuracy and efficiency of text sensitive information recognition; innovatively introducing a cross-modal reasoning mechanism, through cross-verification and reasoning of different modal information, further improving the accuracy and robustness of steganography detection. These innovations give the present invention significant advantages in dealing with complex and changing data security threats and meet the urgent needs of modern data security protection.

[0148] In summary, a sensitive information detection technology and system based on multi-modal and steganography detection provided by the present invention constructs a comprehensive, efficient and accurate sensitive information detection solution by integrating a variety of advanced technologies, which is of great significance for ensuring data security and preventing information leakage, and has broad application prospects and social value.

[0149] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A sensitive information detection technology based on multimodality and steganography detection, characterized in that, include: Receive files to be tested; Perform anti-file format tampering bypass detection, including: analyzing file header information, extracting binary structure features and performing format verification to determine the true type of the file; Based on the true type, sensitive information detection is performed, including: after preprocessing the text, using the DistilBERT model to extract text embedding vectors, using the bidirectional LSTM and Attention mechanism to capture the long-term dependency features of the text and focus on sensitive areas, and using the fully connected layer to determine whether the text contains sensitive information; Perform joint content and file steganalysis detection, including: extracting steganalysis features using Transformer, LSTM and CNN models for text, file binary structure and image, respectively, building a multimodal neural network to integrate detection results of different modalities, and applying a cross-modal reasoning mechanism to cross-validate features of different modalities to improve detection accuracy; Output comprehensive detection results, including file format risk rating, sensitive information distribution and category, and steganographic information location and content.

2. The method according to claim 1, wherein The anti-file format tampering bypass detection specifically includes: Identify the true type of the file based on the magic number in the file header; Use a sliding window hash check algorithm to extract and verify header sample information; Analyze the binary data structure of the file and calculate the byte frequency characteristics and data block entropy value characteristics; Compare the extracted feature vector with the feature library for similarity to determine the most likely true type of the file; Output file format tampering risk level assessment results.

3. The method according to claim 2, wherein The similarity comparison between the feature vector and the feature library includes: Calculate the cosine similarity between the extracted feature vector and the standard vector in the feature library; Based on the similarity calculation results, determine the true type of the file; Record the inconsistency between the declared type and the actual type of the file and form a risk assessment report.

4. The method according to claim 1, wherein The sensitive information detection specifically includes: Perform text cleaning operations to remove HTML tags and special symbols; Use a multi-level sensitive word dictionary for matching, including basic sensitive words, synonyms, fuzzy words, and domain-specific sensitive words; Select an appropriate word segmentation tool according to the text language type to perform word segmentation and word form restoration; Load the pre-trained DistilBERT model and extract context-aware text embedding vectors; Build a bidirectional LSTM network to process the embedding vectors and capture the long-term dependencies of the text through forward and backward passes; Introduce the Attention mechanism to calculate the attention weight and highlight the key areas of sensitive information; Use the fully connected layer and Sigmoid function to determine whether the text contains sensitive information and generate a detailed report.

5. The method according to claim 1, characterized in that The combined content and file steganalysis detection specifically includes: For the text part, the Transformer model is used to analyze the document context and detect the hidden steganographic information in small character changes. According to the binary structure of the file, the LSTM model is used to analyze the characteristics of the binary data sequence of the file and detect the embedded data; For the image part, the CNN model is used to analyze the pixel-level changes of the image and identify hidden information; The detection results of different modalities are integrated through multimodal neural networks for comprehensive analysis; Apply a cross-modal reasoning mechanism to improve detection accuracy and robustness, and output steganography detection results.

6. The method according to claim 5, wherein The multi-modal neural network includes: Design a Transformer encoder input layer for the text modality; Design an LSTM input layer for the binary structure modality; Design a CNN feature extraction layer for the image modality; Construct a feature fusion network based on the attention mechanism to automatically learn the correlation and weights between different modalities; Output the fused multi-modal feature representation for subsequent steganography information detection and judgment.

7. The method according to claim 5, wherein The cross-modal reasoning mechanism includes: Establish three groups of association networks: text-image, text-binary, and image-binary; Construct the mapping relationship between different modality features; Based on the detection results and feature clues of different modalities, perform cross-validation reasoning; Resolve the conflicts between the detection results of different modalities and improve the steganography detection confidence; Form the final steganography information detection conclusion.

8. The method according to claim 1, characterized in that In the sensitive information detection: The output of the bidirectional LSTM at time t is where and are the hidden states of the forward and backward LSTMs respectively; The Attention mechanism calculates the attention weights where e t = W a h t + b a , and obtains the context vector The context vector c passes through a fully connected layer to obtain the sensitive information determination probability.

9. The method according to claim 1, characterized in that It also includes the following steps: Generate a comprehensive security assessment report based on the detection results; Provide risk level ratings for the discovered sensitive information and steganographic content; Mark and isolate the files confirmed to contain threats; Feed the detection results back to the security management system for subsequent threat analysis and defense strategy update.

10. A sensitive information detection system based on multi-modal and steganography detection for implementing the method according to any one of claims 1-9, characterized in that, It includes: A file format tampering bypass detection module for analyzing the file header information, extracting binary structure features, and performing format verification to determine the true type of the file; A sensitive information detection module for preprocessing the text, using the DistilBERT model to extract text embedding vectors, capturing text features and focusing on sensitive areas through a bidirectional LSTM and an Attention mechanism, and using a fully connected layer to determine sensitive information; A joint content and file steganography detection module for performing steganography detection on text, file binary structures, and images respectively using Transformer, LSTM, and CNN models, and constructing a multi-modal neural network and a cross-modal reasoning mechanism to integrate and analyze the detection results; A result output module for generating and outputting comprehensive detection results, including file format risk ratings, sensitive information distribution and categories, and steganography information locations and contents.

Citation Information

Patent Citations

  • A CNN and RNN fusion model-based network heterogeneous concurrent steganography channel detection method

    CN109729070A

  • File classification method and device, medium and equipment

    CN113704184A

  • Aspect-level sentiment analysis method fusing multi-modal data

    CN114936623A

  • Sensitive information extraction method based on natural semantic processing and deep learning

    CN115718792A

  • Cross-modal sensitive information identification method

    CN117668292A

Cited By

  • Splicing-scale-adaptive structured sensitive data detection method and system

    CN120804635A

  • Code-level sensitive information detection method, electronic equipment and storage medium

    CN121579326A