APT information acquisition and classification system and method based on multi-source information

By combining a multimodal crawler module and a word frequency classifier, dynamic UA spoofing and IOC tagging, the problem of centralized analysis of OSINT data is solved, enabling efficient acquisition and classification of APT intelligence and improving the intelligent analysis capabilities of APT defense.

CN120910256APending Publication Date: 2025-11-07SUN YAT SEN UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510931560.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-07
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing technologies lack the ability to prevent APT attacks in advance. OSINT data is scattered and difficult to analyze centrally. The problem of data silos from multiple sources and heterogeneous data is prominent. The ability to automatically collect data and analyze it is weak, which affects the accurate prediction and defense against APT attacks.

Method used

It employs a multimodal crawler module for dynamic user agent spoofing, a word frequency classifier module for cleaning and IOC tagging, and combines an expert-annotated hot word list and an APT organization list to achieve a multi-dimensional scoring algorithm for document classification and output threat intelligence.

Benefits of technology

By overcoming anti-crawling blockades, eliminating non-semantic interference, and resolving feature drift, the system improves the automated collection and intelligent analysis capabilities of APT intelligence, thereby enhancing the ability to detect and defend against APT attacks in advance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120910256A_ABST
    Figure CN120910256A_ABST
Patent Text Reader

Abstract

The invention discloses an APT information acquisition and classification system and method based on multi-source information. The system comprises a multi-mode crawler module and a word frequency classifier module. The multi-mode crawler module is used for receiving two preset input sources: initiating a request and downloading resources, and avoiding an anti-crawling strategy by dynamically assembling a browser UA head and keep-alive connection; performing exception processing, decoding and URLs extraction on the response content; the word frequency classifier module is used for cleaning a text output by a crawler, removing irrelevant information and replacing IOCs with a standardized tag; based on the hot word list and the APT organization list labeled by the expert, classifying the documents through a multi-dimensional scoring algorithm; and outputting the threat intelligence document and discarding the irrelevant content.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of cyberspace security, in particular, to an APT intelligence acquisition and classification system and method based on multi-source information. BACKGROUND

[0002] Advanced Persistent Threat (APT) is a long-term and hidden computer intrusion process targeting specific targets, which is a serious network threat.

[0003] To effectively combat APT, existing technical means mainly focus on "in-process response" and "post-tracing". Based on host system logs, a tracing graph is constructed, and related innovative algorithms on the graph are developed to achieve rapid identification and blocking when an attack occurs; based on communication traffic, innovative identification methods are used to identify APT communication and block data theft; based on honeypot, APT software is captured, and techniques such as sandbox and disassembly are used to profile APT organizations to complete attack attribution. These methods have good defense capability in certain scenarios, but also face challenges: too strong dependence on "knownness", i.e. unable to achieve "prevention". Attackers use zero-day exploits and supply chain pollution to achieve immediate attack and long-term concealment, bypassing most existing defense systems and posing a serious challenge to traditional detection techniques.

[0004] In recent years, open-source threat intelligence (OSINT) has gradually become an important part of the APT defense system. Global security vendors (FireEye, Kaspersky, 360, and Qianxin, etc.) regularly publish APT OSINT; various hacker forums, CVE vulnerability libraries, and threat intelligence platforms also provide a large amount of OSINT regularly. By analyzing and utilizing these OSINT, one can quickly grasp the latest attack group dynamics, attack tactics, and IOC (Indicators of Compromises), etc., significantly improving the ability to advance the perception, automatic defense, and rapid response to APT attacks. However, applying OSINT to complete APT defense still faces the following challenges: (1) OSINT is scattered in various network security forums and security company websites, making it difficult to collect and analyze centrally.

[0005] (2) Most OSINT does not include national background and political information, while APT is often closely related to geopolitical conflicts or regional hot events, hindering accurate prediction and analysis of APT.

[0006] (3) There is a lack of multi-source heterogeneous OSINT extraction and analysis systems, with a prominent data silo problem, weak automated collection and intelligent analysis capabilities, making it difficult to support large-scale APT defense tracing system construction.

[0007] And the application number is 202411768791.2 Chinese invention application discloses "APT attack technology identification and matching method and system based on threat intelligence", including: obtaining multiple threat intelligence sentences;According to the training of the initial classification model of multiple threat intelligence sentences, the final classification model is obtained, and the detection sentence is input into the final classification model in turn, and the classification result of each detection sample is obtained;Each word in the detection sentence with the classification result is converted into a high-dimensional context embedding vector representation, and the high-dimensional context embedding vector representation is extracted, and the structured relationship triple corresponding to each high-dimensional context embedding vector representation is obtained;Match the structured relationship triple with any ATT&CK technology information, and obtain the target ATT&CK technology information corresponding to the detection sentence according to the matching result. SUMMARY

[0008] To solve the technical problems in the above background art, the present application provides a kind of APT intelligence acquisition and classification system and method based on multi-source information, and the technical scheme adopted by the present application is: The first aspect of the present application provides a kind of APT intelligence acquisition and classification system based on multi-source information, the system includes multimodal crawler module and word frequency classifier module; The multimodal crawler module is used to: accept two types of input sources: initiate request and download resources, avoid anti-crawling strategy by dynamically assembling browser UA header and keep-alive connection; Abnormal processing, decoding and URL extraction are performed on response content; The word frequency classifier module is used to: clean the text output by crawler, remove irrelevant information and replace IOCs with standardized labels; Based on expert annotated hot word list and APT organization list, classify documents through multi-dimensional scoring algorithm; Output threat intelligence documents and discard irrelevant content.

[0009] As a preferred scheme, the two types of input sources include: OSINT input: user-configured threat intelligence URLs; Geopolitical news input: URLs generated according to geopolitical hot keywords.

[0010] As a preferred scheme, the multimodal crawler module includes input scheduling unit, request proxy unit, exception control unit, automatic identification module and resource parsing unit; The input scheduling unit is used to switch the two types of URL input sources of OSINT and geopolitical news. The request agent unit is configured to randomly implant a UA header of a preset browser in a request message; The exception control unit is configured to take corresponding response measures for different error status codes; The automatic identification module is configured to automatically identify resource types in a timestamp+extension format and record file hash values; The resource analysis unit is configured to decode bytecode and extract text and nested URLs in HTML / PDF.

[0011] As a preferred solution, the resource analysis unit comprises an HTML processing unit and a PDF processing unit; The HTML processing unit is configured to extract URLs and text of HTML resources through BeautifulSoup respectively, and remove duplicate sentences through a sent_tokenize method in nltk.tokenize; The PDF processing unit is configured to extract text page by page through PyPDF2 and splice.

[0012] As a preferred solution, the word frequency classifier module comprises a regular cleaning unit, an IOC tagging unit, and an expert knowledge base unit; The regular cleaning unit is configured to remove irrelevant information that may interfere with the normal work of the classifier by using a preset regular expression; The IOC tagging unit is configured to perform tagging processing on URLs, Hash, IP addresses, paths, and email addresses respectively; The expert knowledge base unit is configured to store OSINT distinguishing hot words and APT organization names.

[0013] As a preferred solution, the multi-dimensional scoring algorithm comprises: A hot word scoring unit configured to calculate a basic score value according to the following formula:

[0014] Wherein, hotword_score is the current document hot word score, doc is the current document, wi is the i-th word in the current document, H and N are the hot word table and the length of the hot word table respectively, and RANK(wi) is used to obtain the ranking of wi in H; A feature tagging unit configured to perform the following scoring operations: When an APT organization name or alias is hit, the corresponding score is increased apt_score ; When a CVE vulnerability is hit, the corresponding score is increased cve_score ; A tag weighting unit configured to perform differential scoring on IOC tags: When a label corresponding to a hit Hash or path is hit, the corresponding score is increased hp_score ; When a label corresponding to a hit email address is hit, the corresponding score is increased eml_ score ; When a label corresponding to a hit URL is hit, the corresponding score is increased url_ score ; A total score calculation unit aggregates the scores to obtain a total score according to the following formula: .

[0015] As a preferred solution, the word frequency classifier module further comprises a decision unit for: calculating a document score length ratio:

[0016] wherein length is the document length; When score_length_ratio > 15, it is determined as an OSINT document, otherwise it is discarded.

[0017] The second aspect of the present application provides an APT intelligence acquisition and classification method based on multi-source information, the method comprising: inputting URLs of OSINT and geopolitical news; dynamically assembling a UA header to initiate a request and download resources, and performing retry or abortion on abnormal responses; parsing HTML / PDF resources, extracting a body and nested URLs and de-duplicating; performing regular cleaning and IOC labeling to generate a purified text; calculating a document score based on a hot word table, an APT list and a CVE feature; outputting a threat intelligence document by a preset threshold.

[0018] The third aspect of the present application provides a computer readable storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to implement the steps of the foregoing APT intelligence acquisition and classification method based on multi-source information.

[0019] The fourth aspect of the present application provides a computer device comprising a storage medium, a processor and a computer program stored in the storage medium and executable by the processor, wherein the computer program is executed by the processor to implement the steps of the foregoing APT intelligence acquisition and classification method based on multi-source information.

[0020] Compared with the prior art, the present application has the beneficial effects that: The application breaks through the anti-crawling blockage through the multi-modal crawler module for dynamic UA camouflage (4 browser rotation); the application carries out 7 types of regular cleaning through the word frequency classifier module to remove page codes / addresses / random codes and other non-semantic interference; the application carries out IOC tagging through the word frequency classifier module to replace URL / Hash / IP with semantic tags, solving feature drift. BRIEF DESCRIPTION OF DRAWINGS

[0021] Figure 1 A multi-source information-based APT intelligence acquisition and classification system architecture diagram is provided for the embodiment; Figure 2 A multi-source information-based APT intelligence acquisition and classification method flowchart is provided for the embodiment. DETAILED DESCRIPTION The accompanying drawings are only used for illustrative description and cannot be understood as a limitation to the application; It should be clear that the described embodiments are only part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present application.

[0022] The terms used in the present application are only for the purpose of describing specific embodiments, and are not intended to limit the present application. The singular forms "a", "said" and "the" used in the present application and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein means and includes any or all possible combinations of one or more associated listed items.

[0023] The following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present application. Instead, they are merely examples of apparatuses and methods consistent with some aspects of the present application, as detailed in the appended claims. In the description of the present application, it should be understood that the terms "first", "second", "third", etc. are only used to distinguish similar objects, and do not necessarily describe a specific order or sequence, nor can they be understood as indicating or implying relative importance. For those of ordinary skill in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances.

[0024] Further, in the description of the present application, "a plurality of" means two or more, unless otherwise specified. The association relationship of "and / or" describes the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B can represent the three cases of A existing alone, A and B existing together, and B existing alone. The character " / " generally represents an "or" relationship between the associated objects before and after it. The present application is further described below in conjunction with the accompanying drawings and embodiments.

[0025] The present application is further described below in conjunction with the accompanying drawings and embodiments.

[0026] Embodiment 1 Please refer to Figure 1 The embodiment provides an APT intelligence acquisition and classification system based on multi-source information, which comprises a multi-modal crawler module and a word frequency classifier module. The multi-modal crawler module is configured to: accept two types of input sources: initiate a request and download resources, avoid anti-crawling strategies by dynamically assembling browser UA headers and keep-alive connection; perform exception handling, decoding and URL extraction on the response content; The word frequency classifier module is configured to: clean the text output by the crawler, remove irrelevant information and replace IOCs with standardized labels; classify documents based on an expert-labeled hot word list and an APT organization list through a multi-dimensional scoring algorithm; output threat intelligence documents and discard irrelevant content.

[0027] In a specific embodiment, the two types of input sources include: OSINT input: user-configured threat intelligence URLs; Geopolitical news input: URLs generated according to geopolitical hot keywords.

[0028] In a specific embodiment, the multi-modal crawler module comprises an input scheduling unit, a request proxy unit, an exception control unit, an automatic identification module and a resource parsing unit. The input scheduling unit is configured to switch between OSINT and geopolitical news URL input sources; The request proxy unit is configured to randomly implant the UA header of a preset browser in the request message; The exception control unit is configured to take corresponding response measures for different error status codes; The automatic identification module is configured to automatically identify resource types in timestamp+extension format and record file hash values; The resource parsing unit is used to decode bytecode and extract the main text and nested URLs from HTML / PDF.

[0029] In one specific embodiment, the resource parsing unit includes an HTML processing unit and a PDF processing unit; The HTML processing unit is used to extract the URLs and body of HTML resources using BeautifulSoup, and to segment sentences using the send_tokenize method in nltk.tokenize to remove duplicate sentences. The PDF processing unit is used to extract and concatenate text page by page using PyPDF2.

[0030] In one specific embodiment, the word frequency classifier module includes a regular expression cleaning unit, an IOC tagging unit, and an expert knowledge base unit; The regular expression cleaning unit is used to remove irrelevant information in the corpus that may interfere with the normal operation of the classifier using a preset regular expression. The IOC tagging unit is used to tag URLs, hashes, IP addresses, paths, and email addresses respectively. The expert knowledge base unit is used to store OSINT distinguishing hot words and APT organization names.

[0031] In one specific embodiment, the multi-dimensional scoring algorithm includes: The hot word scoring unit is used to calculate the base score according to the following formula:

[0032] Where hotword_score is the hotword score of the current document, doc is the current document, wi is the i-th word in the current document, H and N are the hotword list and the length of the hotword list, respectively, and RANK(wi) is used to get the ranking of wi in H; The feature labeling unit performs the following scoring operation: When an APT organization's name or alias is matched, the corresponding score is increased. apt_score ; When a CVE vulnerability is detected, the corresponding score is increased. cve_score ; The tag weighting unit applies differentiated scoring to IOC tags: When a tag corresponding to a hash or path is matched, the corresponding score is increased. hp_score ; When the tag corresponding to the email address is matched, the corresponding score is increased. eml_ score ; When the tag corresponding to the URL is hit, the corresponding score is increased. url_ score ; The total score is calculated by aggregating the scores according to the following formula: .

[0033] In one specific embodiment, the word frequency classifier module further comprises a decision unit for: Calculating the document score length ratio:

[0034] Where length is the document length; When score_length_ratio > 15, it is determined to be an OSINT document, otherwise it is discarded.

[0035] Embodiment 2 Please refer to Figure 1 This embodiment can be regarded as an improved or extended embodiment based on embodiment 1, specifically: an APT intelligence acquisition and classification system based on multi-source information, comprising a multi-modal crawler module and a word frequency classifier module; The multi-modal crawler module is configured to: Accept two types of input sources: Initiate requests and download resources, and evade anti-crawling strategies by dynamically assembling browser UA headers and keep-alive connection; Perform exception handling, decoding and URL extraction on the response content.

[0036] Specifically, the multi-modal crawler receives URLs as input, which are independent of the content pointed to by the URLs, and can be configured for OSINT and geopolitical news. Geopolitical news uses the GNews API to parse geopolitical hot keywords and returns geopolitical news URLs to the crawler module. The OSINT input requires the user to specify URLs that are definitely pointing to OSINT to improve efficiency and ensure quality, and is also used for subsequent word frequency classifier construction. In actual work, the user can choose MITRE Cti-Master data or MispGalaxy threat-actor URLs, which all point to OSINT of various security vendors and forums, or manually configure URLs pointing to trusted OSINT; The crawler termination condition can be set to obtain the URLs in the pages pointed to by these URLs, that is, to obtain the resources pointed to by the manually input URLs, and then parse the URLs contained in these resources and collect the resources pointed to by them, and no further parsing is performed.

[0037] The multi-modal crawler is obtained by the get() method of the request library in Python, the headers field is randomly assembled to evade the anti-crawler mechanism of the website, the request sent by four browsers of Firefox, Safari, Chrome and IE is assembled, the type of the allowed receiving field is set to 'application / pdf, application / xhtml+xml, image / png, image / jpeg, image / gif, image / webp, * / *', a keep-alive connection is required, and an exception handling is set to capture different types of exceptions: for 404 exceptions, write to the log and return; for 500, the returned data is incomplete, then reassemble the header, wait for 10s and request the resource again, the number of requests for the same resource is not more than 5 times. The different types of resources requested are automatically identified, saved in the form of "timestamp.file type extension", and the file hash value is recorded, and the same hash value file is not requested and saved repeatedly.

[0038] The OSINT is a natural language document in different languages, and the get() method returns and saves in binary form. In order to ensure the accuracy of the parsed code, the chardet method is used to automatically detect the document code (chardet.detect), and the file is read with the parsed code. For HTML, it often contains navigation, rendering and other irrelevant information, in order to ensure that only URLs and text are extracted, the present application uses the BeautifulSoup library in bs4 to parse, extracts a tags with href attributes, extracts all URLs, and uses the get_text() method to obtain pure text information in HTML. Since there are multiple nested or <section>Some content can be different in different or <section>The same content is extracted multiple times due to multiple nesting in the label, so the sent_tokenize method in nltk.tokenize is introduced to divide sentences and remove duplicate sentences. For PDF, PyPDF2 parses the file, and the.get_text() method is used to extract text page by page and concatenate. The extracted document is saved as a corpus.

[0039] The word frequency classifier module is used to: Clean the text output by the crawler, remove irrelevant information and replace IOCs with standardized labels; Based on the expert annotated hotword list and APT organization list, classify the document through multi-dimensional scoring algorithm; Output threat intelligence documents and discard irrelevant content.

[0040] Specifically, considering that the crawler may capture OSINT irrelevant web pages or even advertising web pages, policy and legal statements, etc., the classifier is needed to classify the web pages into OSINT and non-OSINT for subsequent processing. The present application further proposes a word frequency-based classifier. Since MITRE Cti-Master data or MispGalaxy threat-actor URLs point to various security vendors, forum OSINT, these corpora can be used to depict the OSINT word frequency distribution. First, a set of regular expressions is developed to remove irrelevant information that may interfere with the normal operation of the classifier, including the following contents: (1) Figure or table number (e.g. Table 1, Table 1, Figure 1 , Figure 1, Fig. 1, etc.) r'\b(?:Figure|Table|Fig|Fig\.)\s*\d+[a-zA-Z]?\b' (2) Page number information (e.g. Page 1) r'^\s*(Page)?\s*\d+\s*( / [\d]+)?\s*$' (3) Address information, often the address information of the organization or forum publishing OSINT r'\b(?:Inc|Ltd|LLC|Co|Company|Corporation)\b|(?:Street|St\.|Avenue|Ave\.|Road|Blvd|Suite|Zip|No\.)\b' (4) Directory, often with continuous dot + page number combination r'([\.·•_\-—]\s*){5,}' (5) Meaningless long combination of letters and numbers, which may be garbled characters obtained by the.get_text() method r'\b[a-zA-Z0-9]{20,}\b' (6) Isolated punctuation r'([,.:!?。!?,:;])\1{1,}' (7) Isolated characters r'^\s*(\d+[.\-\)]*\s*)+' There are a large number of IOCs in OSINT, the specific content of which is not important to the classifier, but the position of the IOCs needs to be preserved, so a matching and replacing statement is designed to replace the following content: (1) Replace the specific URL with <url>,sub r'hxxps?: / / |hxxp?: / / |https?: / / \S+|http?: / / \S+|www\.\S+|\b[a-zA-Z0-9.-]+\.(com|net|org|cn|info|biz|io|top)\b' to" <url>" (2) Replace the specific Hash with <hash>, typical hash values have MD5 (32-bit alphanumeric) subr'\b[a-fA-F0-9]{32}\b', to ' <hash>';SHA1 (40 digits of letters and numbers) sub r'\b[a-fA-F0-9]{40}\b' to ' <hash>' and SHA256 (64-bit alphanumeric combination) sub r'\b[a-fA-F0-9]{64}\b'to ' <hash> (3) replace the specific IP address with <ip>sub r'\b\d{1,3}(?:\.\d{1,3}){3}\b' to ' <ip> (4) Replace Linux / Unix style paths with <path>sub r'\b( / [a-zA-Z0-9._\-]+)+\b' to' <path>replacing Windows-style paths with <path>, consider Windows paths may be single backslash \ and double backslash \\ sub r'\b[a-zA-Z]:\\(?:[^\\ / :*?"<>|\r\n]+\\)*[^\\ / :*?"<>|\r\n]*' to <path> (5) Replace the email address with <email>sub r'\b[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}\b' to ' <email> The cleaned corpus does not contain irrelevant content such as headers, page numbers, and table of contents, and all specific URLs, Hash values, IP addresses, paths, and EMAILs are replaced by tags.

[0041] Then, the frequency of each word in the corpus is counted, and all sentences in the corpus are tokenized using the word_tokenize method in nltk.tokenize, converted to all lowercase, and counted. The top 2000 words with the highest frequency of occurrence are output. Among these words, there are words unique to OSINT and frequently used, such as null, malware, malicious, group, etc., and words that are also used in non-OSINT documents, such as one, new, data, use, etc. This invention invited three experts who have been engaged in OSINT application and APT threat detection for a long time to discuss and label these 2000 words according to whether the word can distinguish OSINT with high probability. Among them, the number of words that can distinguish OSINT with high probability is 621, and the remaining 1379 words are considered to be not representative, and their occurrence cannot determine whether the document is OSINT. These 621 hot words are ranked in descending order of frequency of occurrence and saved in hotwords.txt for use by the classifier.

[0042] OSINT often also involves an introduction to APT organizations, so the name of the APT organization is also an important indicator to distinguish OSINT from non-OSINT. This invention also collects 177 common APT organization names and aliases and saves them in apt_groups.txt for use by the classifier.

[0043] The document scoring algorithm is the core of the classifier. For all OSINT to be classified, extract the text content and clean the tokens as described above. For each word, if it hits a word in the hot word list, accumulate different scores according to the position of the hit word in the list. The higher the frequency of the hit word (i.e., the more frequently occurring word in known OSINT) contributes more to the document score. The scoring function is shown in the following formula:

[0044] where, hotword_score is the current document hot word score, doc is the current document, w i is the i th word in the current document, H and N are the hot word table and the length of the hot word table ( N = 621), respectively, and RANK( w i ) gets w i In​ H ordering in the.

[0045] If the APT organization name and alias are hit, apt_score = 300, otherwise 0. If the CVE vulnerability is hit, i.e., matches r'CVE-\d{4}-\d{4,7}', then cve_score = 300, otherwise 0. Considering that hashes and file paths do not often appear in non-OSINT, if the <hash>or <path>, hp_score = 300, otherwise 0; in comparison, the probability of non-OSINT appearing URL and mailbox is higher, if hit <email>, eml_score = 200, else 0, if hit <url>, url_score = 100, otherwise 0. The total score of the document is shown in the following formula: To remove the influence of the length of the document on the total score, introduce the score-length ratio:

[0046] wherein, length is the length of the document.

[0047] In the present application, the classifier takes score_length_ratio>15 as OSINT, otherwise discarded.

[0048] Embodiment 3 Please refer to Figure 2 The embodiment provides an APT intelligence acquisition and classification method based on multi-source information, which comprises the following steps: S1: inputting URLs of OSINT and geopolitical news; S2: dynamically assembling a UA header to initiate a request and download resources, and performing retry or abortion on abnormal response; S3: parsing HTML / PDF resources, extracting a body and nested URLs and removing duplicates; S4: performing regular cleaning and IOC labeling to generate a purified text; S5: calculating a document score based on a hot word table, an APT list and a CVE feature; S6: determining and outputting a threat intelligence document through a preset threshold.

[0049] Embodiment 4 A computer readable storage medium, which stores a computer program, wherein the computer program is executed by a processor to realize the steps of the APT intelligence acquisition and classification method based on multi-source information in the embodiment 3.

[0050] Embodiment 5 A computer device, comprising a storage medium, a processor and a computer program stored in the storage medium and executable by the processor, wherein the computer program is executed by the processor to realize the steps of the APT intelligence acquisition and classification method based on multi-source information in the embodiment 3.

[0051] Obviously, the above embodiments of the present application are merely exemplary and are not intended to limit the embodiments of the present application. Based on the above description, other different forms of changes or modifications can be made by those skilled in the art. Here, it is not necessary or possible to exhaust all the embodiments. Any modification, equivalent replacement and improvement made within the spirit and principle of the present application should be included in the protection scope of the claims of the present application.< / url> < / email> < / path> < / hash> < / email> < / email> < / path> < / path> < / path> < / path> ​< / ip> < / ip> ​< / hash> < / hash> < / hash> < / hash> < / url> < / url> < / section> < / section>

Claims

1. A multi-source information-based APT intelligence acquisition and classification system, characterized in that, The system comprises a multi-modal crawler module and a word frequency classifier module; The multi-modal crawler module is used for: accepting two types of preset input sources: initiating a request and downloading resources, avoiding anti-crawling strategies by dynamically assembling browser UA headers and keep-alive connection; performing exception handling, decoding and URL extraction on response content; The word frequency classifier module is used for: cleaning the text output by the crawler, removing irrelevant information and replacing IOCs with standardized labels; based on the expert-labeled hot word list and APT organization list, classifying the document through a multi-dimensional scoring algorithm; outputting threat intelligence documents and discarding irrelevant content.

2. The APT intelligence acquisition and classification system based on multi-source information according to claim 1, characterized in that, The two types of input sources include: OSINT input: user-configured threat intelligence URLs; geopolitical news input: URLs generated according to geopolitical hot keywords.

3. The APT intelligence acquisition and classification system based on multi-source information according to claim 2, characterized in that, The multi-modal crawler module comprises an input scheduling unit, a request proxy unit, an exception control unit, an automatic identification module and a resource parsing unit; The input scheduling unit is used to switch between OSINT and geopolitical news URL input sources; The request proxy unit is used to randomly implant the UA header of a preset browser in the request message; The exception control unit is used to take corresponding response measures for different error status codes; The automatic identification module is used to automatically identify resource types in timestamp+extension format and record file hash values; The resource parsing unit is used to decode bytecodes and extract text and nested URLs in HTML / PDF.

4. The APT information acquisition and classification system based on multi-source information according to claim 3, characterized in that, The resource parsing unit comprises an HTML processing unit and a PDF processing unit; The HTML processing unit is used to extract URLs and text of HTML resources through BeautifulSoup respectively, and to remove duplicate sentences by using the sent_tokenize method in nltk.tokenize; The PDF processing unit is used to extract text page by page through PyPDF2 and splice them.

5. The APT intelligence acquisition and classification system based on multi-source information according to claim 1, characterized in that, The word frequency classifier module comprises a regular cleaning unit, an IOC labeling unit and an expert knowledge base unit; The regular cleaning unit is used to remove irrelevant information that may interfere with the normal work of the classifier using a preset regular expression; The IOC labeling unit is used to label URLs, Hash, IP addresses, paths and email addresses respectively; The expert knowledge base unit is used to store OSINT distinctive hot words and APT organization names.

6. The APT intelligence acquisition and classification system based on multi-source information according to claim 1, characterized in that, The multi-dimensional scoring algorithm includes: A hot word scoring unit for calculating the base score according to the following formula: where hotword_score is the current document hot word score, doc is the current document, wi is the ith word in the current document, H and N are the hot word list and hot word list length respectively, and RANK(wi) is used to obtain the ranking of wi in H; A feature labeling unit for performing the following scoring operations: When a hit is made on an APT organization name or alias, increase the corresponding score apt_score ; When hitting CVE vulnerabilities, increase the corresponding score cve_score ; A label weighting unit for performing differential scoring on IOC labels: When a hit Hash or path corresponds to a label, increase the corresponding score hp_score ; When a label corresponding to a hit mail address is hit, a corresponding score is increased eml_ score ; When a tag corresponding to the hit URL is hit, the corresponding score is increased url_ score ; A total score calculation unit for aggregating scores to obtain the total score according to the following formula: 。 7. The APT information acquisition and classification system based on multi-source information according to claim 6, characterized in that, The word frequency classifier module further comprises a decision unit for: calculating a document score length ratio: wherein length is the document length; When score_length_ratio > 15, it is determined to be an OSINT document, otherwise it is discarded.

8. An APT intelligence acquisition and classification method based on multi-source information, characterized in that, The method comprises: inputting OSINT, geonews URLs; dynamically assembling UA headers to initiate requests and download resources, performing retries or aborts on abnormal responses; parsing HTML / PDF resources, extracting the main text and nested URLs and removing duplicates; performing regular cleaning and IOC tagging to generate clean text; calculating a document score based on a hot word table, an APT list, and a CVE feature; outputting threat intelligence documents by a preset threshold.

9. A computer readable storage medium having stored thereon a computer program, characterized in that: The computer program, when executed by a processor, implements the steps of the APT intelligence acquisition and classification method based on multiple sources of information according to claim 8.

10. A computer device, comprising: The computer program, when executed by a processor, implements the steps of the APT intelligence acquisition and classification method based on multiple sources of information according to claim 8.

Citation Information

Patent Citations

  • APT attack technology identification and matching method and system based on threat intelligence

    CN119760705A