Hidden file extraction and detection method, system and equipment
By performing file size classification and logistic regression model analysis on outbound emails, combined with semantic association detection of graph neural network and BERT model, a keyword fragmented pattern library is built, which accurately identify and warning the key data in the emails is achieved, solving the problem of false alarms and missed detection in the existing technology, and improving information security protection capabilities.
Patent Information
- Application Number
- CN202510631621.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2045-05-16
AI Technical Summary
When the prior art performs hidden file extraction and detection of external emails, it is difficult to fully and accurately identify key data in the emails, especially when facing emails of different sizes and encrypted content, there is a risk of false alarms and missed detection, and it is impossible to effectively protect information security.
By monitoring email transmission, detecting file size and classification, extracting email text features, calculating the first probability value using logistic regression model, comparing the attachment content, analyzing semantic associations with graph neural network and BERT model, building a keyword fragmented mode library, and performing dual probability judgment to generate early warnings.
It significantly improves the detection accuracy of key email information, reduces the risks of false alarms and missed detection, supports automated response, provides efficient and reliable data leakage prevention solutions, and adapts to the security needs of different business scenarios.
Smart Images

Figure CN120455096A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of security management and control technology, and in particular to a method, system and device for extracting and detecting hidden files. Background Art
[0002] Steganography, a technique for concealing secret information within other media, is widely used in the field of information security. Existing technologies for extracting and detecting hidden files from outgoing emails primarily rely on data hiding and detection techniques. These methods analyze the statistical characteristics and pattern variations of multimedia data (such as images, audio, and video) or text files within emails to detect and extract potentially hidden information. With the advancement of technology, steganography has been combined with encryption techniques and deep learning, improving the stealth and robustness of information hiding. However, this has also brought new challenges to the extraction and detection of hidden files.
[0003] Chinese invention patent application number 202010086751.5 discloses a sensitive information detection method: intercepting an outgoing email and extracting first text data; obtaining a preset monitoring field and identifying a first monitoring field value corresponding to the preset monitoring field from the first text data; combining the first and second text data to generate a first combined feature, which is then input into a sensitive data detection model to obtain a first sensitivity probability; extracting the attachment from the outgoing email when the first sensitivity probability is less than or equal to a preset value; performing anti-hiding analysis on the file in the attachment and determining whether the parsed file data has changed; if the parsed file data has changed, determining that the outgoing email has a data leak; extracting the changed data from the parsed file data and generating a first warning message; and sending the extracted data and the first warning message to a management terminal. This method can improve the accuracy of email detection.
[0004] Traditional email monitoring methods are often simplistic and lack comprehensive and accurate identification of critical data within emails. Emails of different sizes have varying probabilities of containing critical information, and critical data can be hidden within both email text and attachments. Relying solely on a single criterion or simple detection is ineffective. Therefore, a more complex and sophisticated email monitoring and critical data detection mechanism is required. This mechanism comprehensively considers multiple factors, including email size, text content, and attachment content, to accurately identify and issue alerts on emails containing critical data, effectively protecting the information security of enterprises or institutions. Summary of the Invention
[0005] This application provides a hidden file extraction and detection method and system to accurately identify and warn emails containing key data, thereby improving the accuracy of detecting key information in emails.
[0006] The present application provides a method for extracting and detecting hidden files, the method comprising: S1, monitors email transmission and intercepts outgoing emails; S2, detects the size of outgoing email files and classifies and marks them; S3, extract email text features and set the probability threshold of the key data detection model according to the size; Pre-analyze the email title to identify the subject and type; based on the pre-analysis results, check the word order to identify encrypted emails; decrypt the encrypted email and extract feature data to execute S4; analyze the semantic behavior correlation factors of the email text and mark suspicious encrypted emails if they exceed the standard to execute S5; S4, inputting feature data into the model and calculating a first probability value; S5, compare low-probability emails and analyze new attachment content; S6, evaluating the second probability value, determining whether the email contains the keyword, and sending an early warning to the terminal.
[0007] Preferably, the probability threshold of the key data detection model includes: cleaning the title and body of the email and removing irrelevant characters; then splitting the text with the jieba word segmentation tool, extracting and marking feature data vectors based on the key data; collecting marked email data, and using regression analysis to statistically analyze the relationship between file size and the probability of occurrence of key information; setting the probability threshold of each email category based on the statistical results, configuring the grading parameters to the logistic regression model, and setting the corresponding threshold by category.
[0008] Preferably, the step S4, calculating the first probability value, comprises: converting the feature data vector extracted in step S3 into Input into the logistic regression model; calculate the first probability value of the email containing key information; the calculation formula of the first probability value is:
[0009] is the intercept, Features The weight coefficient of is the eigenvector The value of the i-th dimension of .
[0010] Preferably, the checking of word order to identify encrypted emails includes: extracting the body of the email based on the subject and type of the email text; building a word order logic rule library, matching the body with the rules one by one, and determining whether there are word order logic errors; if there are errors, marking the email text as an encrypted email and calculating the correlation factor.
[0011] Preferably, the semantic behavior of the email text includes: building a domain keyword fragmentation pattern library, cutting historical email texts into character sequences using a statistical word segmentation algorithm; setting a window parameter range based on the domain pattern, scanning and recording relationship paths after probabilistic parameter selection; improving the initial set of keyword potential relationship paths and recording positions; identifying path jump points and recording relevant information; during contextual semantic analysis, extracting the character sequence after detecting the first keyword, checking for similar patterns, and if any, comparing the jump points with the path, and calculating the matching degree using the edit distance.
[0012] Preferably, the matching degree includes: in the text scanning process, detecting the first keyword of the keyword, extracting the character sequence starting with the keyword from the keyword potential relationship path ,in Represents the i-th character in the sequence. If there is a jump point, its position index is i; the extracted character sequence S is compared with the domain keyword fragmentation pattern library Match and check whether there is a similar fragmentation pattern. If there is a similar pattern, further compare the jump points and relationship paths in S with Whether they are consistent; use the edit distance algorithm to calculate the matching degree; edit distance Indicates converting the character sequence S into a pattern The minimum number of editing operations required; let the length of S be n, The length of is m, then the similarity score sim for: , matching degree Equal to the similarity score , at this time the matching degree has been normalized to the [0,1] interval.
[0013] Preferably, the evaluating the second probability value comprises: S61, detecting keyword region blocks according to a keyword fragmentation pattern library; S62, mapping the detected area blocks into nodes, and automatically completing the logical edges between the nodes based on domain-related semantic rules to form a complete semantic chain; S63, uses graph neural network to calculate the transition probability between nodes; S64, setting a transition probability threshold. When the path probability exceeds the threshold, it is determined to be a valid association, and these validly associated area blocks are determined as the identified keyword set; S65: Calculate a second probability value based on the keyword set.
[0014] Preferably, the calculating of the second probability value includes: the second probability value calculation formula is:
[0015] This can be obtained by analyzing the frequency of occurrence of the keyword in a large number of normal and suspicious attachments, and expressing each keyword as corresponding to a suspicious probability calculated based on its semantic association and occurrence. ; Represented as a keyword set, each keyword The corresponding weight is .
[0016] The present application also provides a hidden file extraction and detection system, the system comprising: The acquisition module is used to monitor the email transmission between the intranet and the extranet, intercept all emails sent from the intranet to the extranet; and perform file size detection on the intercepted emails, and classify and mark the emails according to the preset file size classification standards; A detection module is configured to extract key data features from the email text, set classification parameters in a key data detection model, and set a probability threshold based on file size; input the extracted feature text data into the key data detection model to calculate a first probability value that the email contains key information; A parsing module, configured to compare the email title and body content with the attachment content, and parse out the new content in the attachment; A probability determination module is configured to input the parsed attachment's newly added content into a key data detection model to evaluate a second probability value, and determine whether the email contains key data by combining the first probability value and the second probability value; The early warning module is used to extract the specific key data content from emails determined to contain key data, generate early warning information and send it to the management terminal.
[0017] A device includes a processor, a memory, a network interface and a database connected by a system bus, and is applied to the aforementioned method for extracting and detecting hidden files.
[0018] One or more technical solutions provided in this application have at least the following technical effects or advantages: Through a hierarchical detection mechanism and dynamic threshold settings, the system achieves precise identification and early warning of key email information. The system first categorizes files by size and extracts key features. It then uses a logistic regression model to calculate key probability values. It then compares and analyzes the body text and attachments to identify hidden information. Finally, it combines dual probabilistic analysis to generate structured early warnings. This approach significantly improves detection accuracy, effectively reduces the risk of false positives and missed detections, and supports automated response, providing an efficient and reliable solution for data leakage prevention.
[0019] Through title pre-analysis, word order logic checking, and metadata association analysis, the system achieves accurate identification and processing of encrypted emails. The system first uses text classification and rule matching to detect anomalies in the title subject and word order, and then uses public key decryption to obtain the plaintext content. For emails that cannot be decrypted, BERT semantic modeling and user behavior analysis are used to calculate contextual association factors to effectively identify suspicious encryption behavior. While ensuring the privacy of encrypted communications, this solution significantly improves the coverage and accuracy of key information detection, addressing the pain point of traditional methods that are ineffective for encrypted content, and adapting to the security needs of different business scenarios through dynamic thresholds.
[0020] By building a keyword fragmentation pattern library and a probabilistic sliding window scanning mechanism, the system achieves accurate identification of fragmented key information within hidden files. The system first establishes a fragmentation feature library encompassing various patterns, such as character insertion and homophone replacement. Using dynamic window parameter settings, the system performs multi-granular scanning of file content, recording potential relationship paths and analyzing jump point features. This system then combines the edit distance algorithm to calculate matching scores, effectively addressing the difficulty traditional methods have in detecting hidden information that has been deliberately segmented and deformed.
[0021] Leveraging a keyword fragmentation pattern library, we accurately detect regional blocks and construct semantic chains based on domain semantic rules, effectively mining potential connections between keywords. Graph neural networks calculate node transition probabilities, combined with thresholds to determine valid connections, enabling accurate identification of keyword clusters. Furthermore, a secondary probability value is calculated based on the frequency of keyword occurrence in both normal and suspicious attachments, quantifying the degree of suspicion for each keyword. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 Schematic diagram of the process of extracting and detecting hidden files according to an embodiment of the present invention; Figure 2 2 is a structural block diagram of a hidden file extraction and detection system according to an embodiment of the present invention. DETAILED DESCRIPTION
[0023] To facilitate understanding of the present invention, the present application will be described more comprehensively below with reference to the relevant drawings; the drawings show preferred embodiments of the present invention, but the present invention can be implemented in many different forms and is not limited to the embodiments described herein; on the contrary, the purpose of providing these embodiments is to enable a more thorough and comprehensive understanding of the disclosed content of the present invention.
[0024] It should be noted that the terms “vertical”, “horizontal”, “up”, “down”, “left”, “right” and similar expressions used in this document are for illustrative purposes only and do not represent the only implementation method.
[0025] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this invention pertains; the terms used herein in the specification of the present invention are for the purpose of describing specific embodiments only and are not intended to limit the present invention; the term "and / or" used herein includes any and all combinations of one or more of the associated listed items.
[0026] Example 1: Figure 1 The figure is a flow chart of a hidden file extraction and detection method according to an embodiment of the present invention.
[0027] like Figure 1 As shown, a hidden file extraction and detection method includes the following steps: S1, the server monitors the email transmission between the intranet and the extranet, and intercepts all emails sent from the intranet to the extranet.
[0028] S2, detect the file size of the intercepted emails and classify and mark them according to the file size.
[0029] Specifically, the files are classified and marked according to their size, including: The server extracts the file size information of the intercepted email.
[0030] Emails are categorized and marked according to preset file size classification standards (such as small emails, medium emails, and large emails).
[0031] Among them, the file size information is the total size of the email (including attachments). The classification standard can be set according to the file size of emails sent daily. For reference, set small file emails: emails with a size not exceeding 10MB, usually with text content; medium file emails: emails with a size between 10MB and 50MB, which may contain a small number of attachments; large file emails: emails with a size exceeding 50MB, usually with large files attached (such as compressed files, high-definition pictures).
[0032] S3, extracts key data features from the email text, sets classification parameters in the key data detection model, and sets probability thresholds based on file size.
[0033] Among them, email text refers to the title and body of the email; key data is the pre-set text information fields that need to be monitored and tested, such as "estimated time", "bid amount", "invoice information", etc.
[0034] Specifically, key data features are extracted from email text, including: Clean the email title and body, remove irrelevant characters (such as HTML tags, special symbols, etc.), and convert the text into a format suitable for analysis.
[0035] Use the Chinese word segmentation tool Jieba to split the text into individual words or phrases.
[0036] The corresponding feature data vector is extracted from the email text based on the key data.
[0037] Among them, the feature data vector is expressed as , Expressed as keyword weight (e.g. "contract" = 0.8, "amount" = 0.9), Named entity recognition results (e.g., "Person Name: Zhang San" = 1, "Organization: XX Company" = 1).
[0038] Mark the feature data as key data features for subsequent operations.
[0039] Specifically, the classification parameters are set in the key data detection model, including: Collect a large amount of labeled email data, including file size, content, and whether it contains key information.
[0040] Regression analysis statistical method is used to analyze the relationship between file size and the probability of key information appearing.
[0041] Based on the statistical results, one or more probability thresholds are set for each email category.
[0042] Configure the classification parameters to the key data detection model (logistic regression model) and set the corresponding probability threshold for the model based on the email category.
[0043] It should be noted that in actual applications, when emails are classified and input into the key data detection model, the model will apply the corresponding probability threshold to make judgments based on the category of the email.
[0044] S4, inputting the extracted characteristic text data into a key data detection model to calculate a first probability value that the email contains key information.
[0045] Specifically, calculating the first probability value of the email containing key information includes: The feature data vector extracted from S3 Input into the logistic regression model.
[0046] Calculate the first probability value of the email containing key information.
[0047] The calculation formula for the first probability value is:
[0048] is the intercept, Features The weight coefficient of (obtained through training), is the i-th dimensional value of the eigenvector .
[0049] For example, assume that the coefficients after model training are = -2.5, = 1.8, = 2.0, and the input features are = [0.8, 0.9, 1, 1] (corresponding to "contract", "amount", "person's name", "organization"), then:
[0050] If the email is classified as a medium file (threshold 0.5), since 0.56 > 0.5, it is marked as critical.
[0051] S5. Compare the email title and body content with the attachment content where the first probability value is lower than the threshold, and parse out the new content existing in the attachment.
[0052] Specifically, parse out the new content existing in the attachment, including: Only process emails where the first probability value P1 is lower than the classification threshold (such as for small files, P1 < 0.7).
[0053] Use the Apache Tika tool to convert the attachment (PDF / DOCX / Excel) into plain text (remove noises such as headers, footers, table borders in the attachment, and retain the core content).
[0054] Segment the words in the email body and attachment text respectively, and obtain the word sets and .
[0055] Remove stop words (such as "de", "shi") and words that have already appeared in the body, and obtain the new word set of the attachment:
[0056] Retain named entities (person's name, account number) and key data (such as "invoice", "key").
[0057] For example, the email body: "Please refer to the attachment for the financial summary." The attachment content: "Key for 2024: X5Y9Z2, internal account number: 622848...". Parsing result: Δ = {("X5Y9Z2", key), ("622848...", bank account number)}, it is determined that there is new content, and S7 secondary evaluation is triggered.
[0058] S6. Input the new content in the attachment into the key data detection model to evaluate the second probability value and determine whether the email contains key data.
[0059] S7: For emails determined to contain key data, extract the specific key data content, generate warning information, and send it to the management terminal.
[0060] The technical solutions in the above embodiments of the present application have at least the following technical effects or advantages: Through a hierarchical detection mechanism and dynamic threshold settings, the system achieves precise identification and early warning of key email information. The system first categorizes files by size and extracts key features. It then uses a logistic regression model to calculate key probability values. It then compares and analyzes the body text and attachments to identify hidden information. Finally, it combines dual probabilistic analysis to generate structured early warnings. This approach significantly improves detection accuracy, effectively reduces the risk of false positives and missed detections, and supports automated response, providing an efficient and reliable solution for data leakage prevention.
[0061] Example 2: In the email key information detection solution of Example 1, when the email content is encrypted, the system faces the dilemma of being unable to extract key data features. Traditional methods can only process plaintext emails and are helpless against encrypted content, resulting in significant vulnerabilities in the detection system. In particular, when emails utilize asymmetric or hybrid encryption, conventional decryption methods are impossible to obtain the content, and effective analysis of encryption behavior characteristics is lacking. This limitation forces the system to rely entirely on manual review for encrypted emails, which is not only inefficient but also prone to missed detections.
[0062] Therefore, the embodiments of the present application are optimized based on the above embodiments.
[0063] In some embodiments, in step S3, extracting key data features from the email text further includes: S31, pre-analyze the title of the email to identify the subject and type of the email text.
[0064] Specifically, the email title is pre-analyzed, including: Collect a large number of email headers from email servers, email clients or related databases. These headers should cover a variety of possible topics and types.
[0065] Clean the collected titles to remove noise data, then store the cleaned email titles in a database or file to build an email title corpus (which can be classified and stored according to certain rules, such as by time, sender, etc.).
[0066] Split the email title into individual words. For example, for the title "Project Progress Report and Next Week's Plan," the word segmentation results are "project," "progress," "report," "and," "next week," and "plan."
[0067] After word segmentation, mark each word with its part of speech, such as noun, verb, adjective, etc. For example, "project" is marked as a noun, and "report" is marked as a verb.
[0068] Using text classification algorithms, email titles are classified into different subject and type categories, such as business emails, financial emails, contract emails, etc., based on features such as keywords and phrases in the titles.
[0069] S32, based on the pre-analysis result of the email, perform a word order logic check on the email text to identify encrypted emails.
[0070] Specifically, word order logic checks include: Extract the email body text based on the subject and type of the email text.
[0071] Build a word order logic rule library to define the rules for sentence component integrity, word collocation rationality, and sentence logical relationship coherence.
[0072] Among them, the sentence component completeness rule is to check whether the sentence has basic components such as subject, predicate, and object. For example, in a simple sentence "I eat", "I" is the subject, "eat" is the predicate, and "meal" is the object. If the sentence lacks these basic components, there may be a word order logic error. The word collocation rationality rule is to determine common word collocation patterns, such as the collocation of verbs and nouns, the collocation of adjectives and nouns, etc. For example, "playing basketball" is a reasonable collocation, while "playing football" does not conform to common collocation habits. The sentence logical relationship coherence rule is to analyze the logical relationship between sentences, such as cause and effect, transitional relationship, parallel relationship, etc. For example, "Because it rained, we canceled the outdoor activities." There is a cause and effect relationship here. If the logical relationship is confusing, there may be a problem.
[0073] Match the email body text with the rules in the word order logic rule base one by one.
[0074] Based on rule matching, determine whether there are any word order logic errors in the email text.
[0075] If there is a word order logic error in the email text, the email text with the word order logic error will be marked as an encrypted email.
[0076] S33, use the public key to decrypt the encrypted email. If the decryption is successful, the plaintext content of the email is obtained and key data features are extracted; if the decryption fails, execute step S34.
[0077] S34, extracting email metadata, analyzing and calculating the correlation factor between email text context semantics and user behavior.
[0078] Metadata refers to non-encrypted data such as sender / recipient, timestamp, and attachment type.
[0079] Specifically, the correlation factors include: Specify the email metadata that needs to be extracted, including sender, recipient, timestamp, attachment type, etc.
[0080] Parse and extract these metadata from the email data, and then use the BERT pre-trained language model to capture the contextual semantic information in the text.
[0081] Preprocess the email text, including word segmentation, stop word removal, stemming, and other operations, to convert the text into a format suitable for model input.
[0082] The preprocessed email text is input into the BERT model to obtain the vector representation of the text.
[0083] Collect the user's historical email data, build a user behavior model, and convert the user behavior model into a vector representation. For example, you can combine the parameters of the time behavior model, frequency behavior model, and sender behavior model into a vector as the vector representation of the user behavior model.
[0084] Email data includes information such as the time, frequency, and sender of emails received. By analyzing the distribution of when users typically receive emails, we can use a probability distribution model (Gaussian distribution) to fit the timing patterns of email receipts. We can also calculate the frequency of email receipts, including the average frequency and the range of frequency fluctuations. We can also analyze the range of senders from whom users receive emails, build sender whitelists or blacklists, and calculate the probability of different senders sending emails.
[0085] Calculate the similarity between the email text context semantic vector and the user behavior model vector.
[0086] Assuming that the email text context semantic vector is A and the user behavior model vector is B, the similarity calculation formula is:
[0087] and are the i-th element of vectors A and B respectively, and n is the dimension of the vector.
[0088] S35, setting a threshold value of the correlation factor. When the correlation factor exceeds the threshold value, the email is marked as a suspicious encrypted email, and step S5 is executed.
[0089] Among them, the threshold value of the correlation factor is set based on historical data and experience.
[0090] The technical solutions in the above embodiments of the present application have at least the following technical effects or advantages: Through title pre-analysis, word order logic checking, and metadata association analysis, the system achieves accurate identification and processing of encrypted emails. The system first uses text classification and rule matching to detect anomalies in the title subject and word order, and then uses public key decryption to obtain the plaintext content. For emails that cannot be decrypted, BERT semantic modeling and user behavior analysis are used to calculate contextual association factors to effectively identify suspicious encryption behavior. While ensuring the privacy of encrypted communications, this solution significantly improves the coverage and accuracy of key information detection, addressing the pain point of traditional methods that are ineffective for encrypted content, and adapting to the security needs of different business scenarios through dynamic thresholds.
[0091] Example 3: In Example 2, when email content uses conventional encryption, the system can determine risk through metadata analysis and behavioral feature detection. However, when email content undergoes keyword fragmentation (e.g., by inserting spaces or replacing homophones), traditional semantic analysis and keyword matching methods struggle to effectively identify this intentionally hidden critical information. Due to the diverse and subtle nature of fragmentation, using a unified detection model can lead to both increased false positives and missed detections. To improve the ability to identify intentionally hidden information, a multi-layered fragmentation feature analysis system is necessary.
[0092] Therefore, the embodiments of the present application are optimized based on the above embodiments.
[0093] In some embodiments, in step S34, the email text context semantics further includes: S341, builds a domain keyword fragmentation pattern library, segments historical email texts based on a statistical word segmentation algorithm, and divides the text into character sequences.
[0094] Specifically, the keyword fragmentation pattern library includes: A large number of email samples (including normal and suspicious emails) were analyzed in detail to manually identify common keyword fragmentation processing patterns, such as interval insertion, homophone replacement, and symbol separation.
[0095] The collected patterns are divided into sub-libraries according to the different fragmentation types, such as character insertion, character replacement, symbol usage, etc. For example, the character insertion sub-library contains all instances of the space insertion pattern.
[0096] Write a detailed description for each pattern, explaining its characteristics and identification methods, and provide specific examples to facilitate subsequent queries and use. For example, for the "interval insertion" pattern, the description could be "interpolate the keyword characters into other characters" and the example could be "pay 1, pay 3, and succeed 9."
[0097] All sub-libraries are aggregated and organized into a keyword fragmentation pattern library.
[0098] S342, setting the window parameter range according to the domain model, performing sliding window scanning after probabilistically selecting parameters, and recording them as relationship paths.
[0099] Domain patterns refer to the application fields in which keywords exist, such as finance and law. Keywords in the financial field might be monetary policy and market, while keywords in the legal field might be suspected and legal provisions. The combination of these keywords forms a common pattern of character sequences.
[0100] Specifically, after probabilistically selecting parameters, a sliding window scan is performed, including: Determine the window size based on common patterns in the domain keyword fragmentation pattern library. For example, for interval insertion mode, set the window size to 3-7 characters; for symbol separation mode, set the window size to 2-5 characters. Keep the step size between 1-3 characters.
[0101] Set probabilities for different window sizes. For example, let the window size W be 3, 4, 5, 6, and 7, with corresponding probabilities P(W=3)=0.25, P(W=4)=0.4, P(W=5)=0.25, P(W=6)=0.05, and P(W=7)=0.05.
[0102] Assign probabilities for different step sizes. Suppose step sizes S are 1, 2, and 3, with corresponding probabilities P(S=1)=0.25, P(S=2)=0.5, and P(S=3)=0.25.
[0103] Starting from the beginning of the email text, the window size W and moving step size S are selected based on the probability. After each move, the character sequence within the window is recorded. These records initially constitute the potential relationship path of the keywords.
[0104] S344, improve the initial set of potential relationship paths of keywords and record the location information.
[0105] Specifically, the initial set of potential relationship paths includes: When scanning the email text in the sliding window, the character sequences in each window are completely recorded. These sequences constitute the initial set of potential relationship paths of keywords.
[0106] After initially obtaining the character sequence records through sliding window scanning of the email text, further refine the initial set of potential relationship paths for keywords. Specifically, completely record the character sequences within each window. For each recorded character sequence, determine and record its starting position and ending position in the email text. For example, for the character sequence "abc", if it starts from the 5th character and ends at the 7th character in the email text, the starting position is 5 and the ending position is 7.
[0107] Analyze the recorded character sequences to explore the potential relationships between characters. For example, count the number of times characters repeat, calculate the interval distances between characters and find patterns. Based on the analysis results, screen and sort the character sequences, removing those sequences that clearly do not have the characteristics of keywords, such as sequences with too short lengths (e.g., only 1 character) or sequences with irregular character repetitions.
[0108] S345, Identify the jump points in the path, record their positions, types, and lengths, and associate them with the sequences.
[0109] Among them, the jump point is the position point of the keyword where discontinuous, spaced, and other special situations occur. For example, in "支1付3成9功", the positions of the numbers 1, 3, and 9 are jump points; the associated sequence is to associate the recorded jump point information with the corresponding character sequence for subsequent comparative analysis.
[0110] S346, When performing context semantic analysis of the email text, after detecting the first keyword, extract the character sequence starting with this keyword and check whether there is a similar fragmented pattern.
[0111] S347, If there is a similar pattern, then compare the jump points and the path, and use the edit distance to calculate the similarity and normalize it to obtain the matching degree.
[0112] Specifically, the matching degree includes: During the text scanning process, once the first keyword of the keyword is detected, extract the character sequence starting with this keyword from the potential relationship path of the keyword , where represents the i-th character in the sequence. If there is a jump point, its position index is i. For example, if the keyword is "支付", when "支" is scanned and meets the characteristics of the first keyword, extract the subsequent relevant character sequence.
[0113] Match the extracted character sequence S with the domain keyword fragmented pattern library to check whether there is a similar fragmented pattern. If there is a similar pattern, then further compare the jump points and the relationship path in S with whether they are consistent.
[0114] The matching degree is calculated using the edit distance algorithm. Indicates converting the character sequence S into a pattern The minimum number of editing operations (insertion, deletion, replacement) required. Let the length of S be n, The length of is m, then the similarity score sim for:
[0115] Match is equal to the similarity score , at this time the matching degree has been normalized to the [0,1] interval.
[0116] Set a matching threshold T (for example, 0.8). If the character sequence S matches a pattern in the pattern library Matching degree If the matching degree exceeds the threshold T, the character sequence S is determined to be a true keyword; if the matching degree does not exceed the threshold, the text is scanned and the matching operation is continued until a qualified keyword is found or the scanning is completed.
[0117] For example, in an outgoing email, the text content is "Please transfer 1 funds to the designated account as soon as possible, otherwise your WeChat account will be frozen." First, construct a fragmented pattern library for domain keywords. By analyzing a large number of fraudulent and normal emails, common keyword fragmentation processing patterns such as "interval insertion", "homophone replacement", and "symbol separation" are manually identified. For example, "transfer 1 funds" is an interval insertion pattern, "WeChat" is a homophone replacement pattern, and "Alipay@" is a symbol separation pattern. Based on this, sub-libraries for character insertion, character replacement, and symbol usage are established, and each pattern is also described and examples are given. Then, use a statistical-based word segmentation algorithm to segment the email text, obtaining character sequences such as "Please", "as soon as", "transfer 1 funds", "to", "the", "designated", "account", "otherwise", "your", "WeChat", "account", etc. Next, set the window parameter range according to the domain pattern. For the "interval insertion" pattern, the window size is set to 3 - 7 characters, and the step size is between 1 - 3 characters. The probabilities of the window size and step size are also set, such as P(W = 3) = 0.25, etc. Then, perform a sliding window scan according to the probabilistically selected parameters, starting from the beginning of the text. For example, if W = 4 and S = 2 are randomly selected for the first time, record character sequences within the window such as "Please transfer as soon as", initially forming a potential relationship path for keywords. Then, complete the initial set of potential relationship paths for keywords and record the position information. For example, the starting position of "transfer 1 funds" is 4, and the ending position is 6; the starting position of "WeChat" is 16, and the ending position is 18. At the same time, analyze and filter to remove sequences with too short length or irregular repetition. Then, identify the jump points in the path. For example, the position index of the number 1 in "transfer 1 funds" is 5, and the position index of the number 1 in "WeChat" is 17. Record their positions, types, and lengths and associate them with the sequences.
[0118] When performing context semantic analysis on the email text, it is detected that "transfer" and "WeChat" may be the initial keywords. Extract the character sequences "transfer 1 funds" and "WeChat" starting with them, and match them with the fragmented pattern library for domain keywords. It is found that "transfer 1 funds" conforms to the "interval insertion" pattern, and "WeChat" conforms to the "homophone replacement" pattern. Then, extract the character sequences S1 = "transfer 1 funds" and S2 = "WeChat", and match them with the corresponding patterns in the pattern library respectively. Calculate the similarity through the edit distance algorithm. Assume the length of S1 is 3, and the length of the corresponding standard pattern "transfer" is 2. The edit distance d(S1, m1) = 1, and the similarity score sim(S1, m1) ≈ 0.67; the length of S2 is 3, and the length of the corresponding standard pattern "WeChat" is 2. The edit distance d(S2, m2) = 1, and the similarity score sim(S2, m2) ≈ 0.67.
[0119] Set the matching threshold T to 0.8. Since sim(S1,m1)≈0.67<0.8 and sim(S2,m2)≈0.67<0.8, the matching degree does not exceed the threshold. Continue scanning the text and performing matching operations (in this example, there are no other suspicious sequences in the email text segment; in practice, scanning will continue). If the matching degree exceeds the threshold after subsequent scanning or parameter adjustment, the character sequence is determined to be a true keyword.
[0120] It should be noted that a feedback mechanism needs to be established in the system. When the system encounters a new fragmentation pattern in actual application, it should be added to the domain keyword fragmentation pattern library in a timely manner.
[0121] The technical solutions in the above embodiments of the present application have at least the following technical effects or advantages: By building a keyword fragmentation pattern library and a probabilistic sliding window scanning mechanism, the system achieves accurate identification of fragmented key information within hidden files. The system first establishes a fragmentation feature library encompassing various patterns, such as character insertion and homophone replacement. Using dynamic window parameter settings, the system performs multi-granular scanning of file content, recording potential relationship paths and analyzing jump point features. This system then combines the edit distance algorithm to calculate matching scores, effectively addressing the difficulty traditional methods have in detecting hidden information that has been deliberately segmented and deformed.
[0122] Example 4: The method of Example 3 may be difficult to fully and accurately capture the potential semantic connections between keywords, resulting in low keyword recognition accuracy and an inability to effectively distinguish between normal and suspicious keywords. Since the frequency of occurrence and semantic association patterns of keywords in different texts are different, the use of a unified recognition rule will inevitably lead to recognition errors and weak generalization capabilities. In view of the significant differences in the fragmentation patterns and semantic associations of keywords in different texts in history, the required indicators for accurate recognition of keywords are bound to be different. In order to more accurately identify keywords, especially to distinguish between normal and suspicious keywords, it is necessary to comprehensively consider the fragmentation patterns, semantic associations, and appearances of keywords in different types of texts.
[0123] Therefore, the embodiments of the present application are optimized based on the above embodiments.
[0124] In some embodiments, in step S6, evaluating the second probability value includes: S61: Detect keyword region blocks according to a keyword fragmentation pattern library.
[0125] Keyword region blocks are generated by scanning the input text character by character or word by word, matching text segments with patterns in the fragmented pattern library. When a segment matching a pattern is found, it is marked as a potential keyword region block.
[0126] S62, maps the detected area blocks into nodes, and automatically completes the logical edges between the nodes based on domain-related semantic rules to form a complete semantic chain.
[0127] Specifically, each detected keyword region block is mapped to a node in a graph. Each node represents a potential keyword or part of a keyword. Based on the knowledge and semantic relationships in a specific field, a set of semantic rules are defined. For example, in the medical field, if two keywords are both related to a certain disease, and one keyword is a symptom or treatment method of the other keyword, then there may be a semantic association between them. When judging the semantic association, a variety of factors can be considered, such as the similarity of word meanings, the closeness of the association in domain knowledge, etc. All nodes are traversed, and whether there is a logical association between the nodes is determined according to the semantic rules. If there is an association, a logical edge is added between the corresponding nodes, and the weight of the edge can be initialized according to the strength of the association (for example, a strong association is given a higher initial weight, and a weak association is given a lower initial weight).
[0128] S63, uses graph neural network to calculate the transition probability between nodes.
[0129] Specifically, a graph neural network (GNN) is constructed using the graph structure constructed in step S62. A feature vector is initialized for each node. The feature vector can be encoded based on the text content, part of speech, and other information of the node. For example, word embedding technology (such as Word2Vec and GloVe) is used to convert the text content of the node into a vector representation. When initializing the feature vector, other attribute information of the node, such as the node's position in the text and frequency of occurrence, can be further combined to more comprehensively represent the characteristics of the node. Message passing and aggregation operations are performed in the GNN. Each node updates its own feature vector by interacting with messages from its neighboring nodes. After multiple layers of message passing and aggregation, the final feature vector of each node is obtained. Then, the transition probability between nodes is obtained by calculating the similarity between the node feature vectors. The similarity calculation can use methods such as cosine similarity, and the specific method can be selected according to the actual situation.
[0130] S64, setting a transition probability threshold. When the path probability exceeds the threshold, it is determined to be a valid association, and these validly associated area blocks are determined as the identified keyword set.
[0131] Among them, for the path between any two nodes in the graph, the path probability is calculated. The path probability can be obtained by multiplying the transition probabilities of all edges on the path. When calculating the path probability, if the path contains multiple nodes, the transition probabilities between adjacent nodes need to be multiplied in the order of the path. The keyword set is represented as , each keyword The corresponding weight is .
[0132] Specifically, a transition probability threshold θ is set. For all paths with a probability exceeding θ, the corresponding nodes are considered to have a valid connection. The regional blocks corresponding to all nodes with a valid connection are then identified as the set of identified keywords. When determining the keyword set, if multiple paths involve the same regional block, a comprehensive assessment of the importance of the regional block based on the path probabilities can be performed to more rationally determine the keyword set.
[0133] S65: Calculate a second probability value based on the keyword set.
[0134] The second probability value calculation formula is:
[0135] This can be obtained by analyzing the frequency of occurrence of the keyword in a large number of normal and suspicious attachments, and expressing each keyword as corresponding to a suspicious probability calculated based on its semantic association and occurrence. .
[0136] The technical solutions in the above embodiments of the present application have at least the following technical effects or advantages: Leveraging a keyword fragmentation pattern library, we accurately detect regional blocks and construct semantic chains based on domain semantic rules, effectively mining potential connections between keywords. Graph neural networks calculate node transition probabilities, combined with thresholds to determine valid connections, enabling accurate identification of keyword clusters. Furthermore, a secondary probability value is calculated based on the frequency of keyword occurrence in both normal and suspicious attachments, quantifying the degree of suspicion for each keyword.
[0137] Furthermore, an embodiment of the present invention also provides a hidden file extraction and detection system.
[0138] Figure 2 It is a structural diagram of a hidden file extraction and detection system according to an embodiment of the present invention.
[0139] like Figure 2 As shown, a hidden file extraction and detection system includes: an acquisition module, a detection module, a parsing module, a probability determination module and an early warning module.
[0140] The acquisition module is used to monitor the email transmission between the intranet and the extranet, intercept all emails sent from the intranet to the extranet; and perform file size detection on the intercepted emails, and classify and mark the emails according to the preset file size classification standards.
[0141] The detection module is used to extract key data features from the email text, set classification parameters in the key data detection model, and set the probability threshold according to the file size; input the extracted feature text data into the key data detection model to calculate the first probability value that the email contains key information.
[0142] The parsing module is used to compare the email title and body content with the attachment content, and parse out the new content in the attachment.
[0143] The probability determination module is used to input the parsed new content of the attachment into the key data detection model to evaluate the second probability value, and combine the first probability value and the second probability value to determine whether the email contains key data.
[0144] The early warning module is used to extract the specific key data content from emails determined to contain key data, generate early warning information and send it to the management terminal.
[0145] It should be noted that other specific implementation contents of the hidden file extraction and detection system according to the embodiment of the present invention may refer to the above-mentioned hidden file extraction and detection method.
[0146] Furthermore, an embodiment of the present invention also provides a device.
[0147] The device includes a processor, memory, a network interface, and a database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operating system and computer program stored in the non-volatile storage medium. The database of the device is used to store email data. The network interface of the device is used to communicate with external terminals via a network connection.
[0148] It should be noted that other specific implementation contents of the hidden file extraction and detection device according to the embodiment of the present invention may refer to the above-mentioned hidden file extraction and detection method.
[0149] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Various modifications and variations are readily apparent to those skilled in the art. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.
Claims
1. A method for extracting and detecting hidden files, characterized in that: The method comprises: S1, monitors email transmission and intercepts outgoing emails; S2, detects the size of outgoing email files and classifies and marks them; S3, extract email text features and set the probability threshold of the key data detection model according to the size; Pre-analyze the email title to identify the subject and type; based on the pre-analysis results, check the word order to identify encrypted emails; decrypt the encrypted email and extract feature data to execute S4; analyze the semantic behavior correlation factors of the email text and mark suspicious encrypted emails if they exceed the standard to execute S5; S4, inputting feature data into the model and calculating a first probability value; S5, compare low-probability emails and analyze new attachment content; S6, evaluating the second probability value, determining whether the email contains the keyword, and sending an early warning to the terminal.
2. The method for extracting and detecting hidden files according to claim 1, wherein: The probability threshold of the key data detection model includes: cleaning the title and body of the email and removing irrelevant characters; then splitting the text using the Jieba word segmentation tool, extracting and marking feature data vectors based on the key data; collecting marked email data, and using regression analysis to statistically analyze the relationship between file size and the probability of key information appearing; setting probability thresholds for each email category based on the statistical results, configuring the grading parameters to the logistic regression model, and setting corresponding thresholds by category.
3. The method for extracting and detecting hidden files according to claim 1, wherein: The step S4, calculating the first probability value, includes: converting the feature data vector extracted in step S3 into Input into the logistic regression model; calculate the first probability value of the email containing key information; the calculation formula of the first probability value is: is the intercept, Features The weight coefficient of is the eigenvector The value of the i-th dimension of .
4. The method for extracting and detecting hidden files according to claim 1, wherein: The method of checking word order to identify encrypted emails includes: extracting the body of the email based on the subject and type of the email text; building a word order logic rule library, matching the body of the email with the rules one by one, and determining whether there is a word order logic error; if there is an error, marking the email text as an encrypted email and calculating a correlation factor.
5. The method for extracting and detecting hidden files according to claim 1, wherein: The semantic behavior of the email text includes: building a domain keyword fragmentation pattern library, using a statistical word segmentation algorithm to cut historical email text into character sequences; setting a window parameter range based on the domain pattern, scanning and recording the relationship path after probabilistic parameter selection; improving the initial set of potential relationship paths of keywords and recording the positions; identifying path jump points and recording relevant information; during contextual semantic analysis, extracting the character sequence after detecting the first keyword, checking for similar patterns, and if any, comparing the jump points and paths, and calculating the matching degree using the edit distance.
6. The method for extracting and detecting hidden files according to claim 5, wherein: The matching degree includes: in the text scanning process, the first keyword of the keyword is detected, and the character sequence starting with the keyword is extracted from the keyword potential relationship path. ,in Represents the i-th character in the sequence. If there is a jump point, its position index is i; the extracted character sequence S is compared with the domain keyword fragmentation pattern library Match and check whether there is a similar fragmentation pattern. If there is a similar pattern, further compare the jump points and relationship paths in S with Whether they are consistent; use the edit distance algorithm to calculate the matching degree; edit distance Indicates converting the character sequence S into a pattern The minimum number of editing operations required; let the length of S be n, The length of is m, then the similarity score sim for: , matching degree Equal to the similarity score , at this time the matching degree has been normalized to the [0,1] interval.
7. The method for extracting and detecting hidden files according to claim 1, wherein: The evaluating the second probability value comprises: S61, detecting keyword region blocks according to a keyword fragmentation pattern library; S62, mapping the detected area blocks into nodes, and automatically completing the logical edges between the nodes based on domain-related semantic rules to form a complete semantic chain; S63, uses graph neural network to calculate the transition probability between nodes; S64, setting a transition probability threshold. When the path probability exceeds the threshold, it is determined to be a valid association, and these validly associated area blocks are determined as the identified keyword set; S65: Calculate a second probability value based on the keyword set.
8. The method for extracting and detecting hidden files according to claim 7, wherein: The calculating of the second probability value includes: the second probability value calculation formula is: This can be obtained by analyzing the frequency of occurrence of the keyword in a large number of normal and suspicious attachments, and expressing each keyword as corresponding to a suspicious probability calculated based on its semantic association and occurrence. ; Represented as a set of keywords, each keyword The corresponding weight is .
9. A hidden file extraction and detection system, applied to a hidden file extraction and detection method according to any one of claims 1 to 8, characterized in that: The system comprises: The acquisition module is used to monitor the email transmission between the intranet and the extranet, intercept all emails sent from the intranet to the extranet; and perform file size detection on the intercepted emails, and classify and mark the emails according to the preset file size classification standards; A detection module is configured to extract key data features from the email text, set classification parameters in a key data detection model, and set a probability threshold based on file size; input the extracted feature text data into the key data detection model to calculate a first probability value that the email contains key information; A parsing module, configured to compare the email title and body content with the attachment content, and parse out the new content in the attachment; A probability determination module is configured to input the parsed attachment's newly added content into a key data detection model to evaluate a second probability value, and determine whether the email contains key data by combining the first probability value and the second probability value; The early warning module is used to extract the specific key data content from emails determined to contain key data, generate early warning information and send it to the management terminal.
10. A device comprising: A processor, a memory, a network interface and a database connected by a system bus, characterized in that a device is applied to a hidden file extraction and detection method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Sensitive information detection method, device, computer equipment and storage medium
CN111310205B
Error sample recognition method and device
CN107291774A
Method and system for detecting mail sensitive information in express industry, device and storage medium
CN108876233A
Method and device for searching keywords according to semantics
CN110209765A
Sensitive information detection method and device, computer equipment and storage medium
CN111310205A