An analysis method based on hybrid threat intelligence data
By using multiple types of data storage middleware and semantic analysis models, threat intelligence data is stored and analyzed in digital warehouses, which solves the problem of inaccurate access storage and information extraction of multi-source heterogeneous data, and realizes the establishment of correlation rules and multi-dimensional display of threat intelligence data, improving threat prevention and attack detection capabilities.
Patent Information
- Application Number
- CN202410763470.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-13
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2044-06-13
AI Technical Summary
The prior art is difficult to effectively solve the problem of the access storage of multi-source heterogeneous threat intelligence data, inaccurate data extraction, and the establishment and application display form of hybrid threat intelligence data that cannot form a standardized system.
Multiple types of data storage middleware (database), semantic analysis models and knowledge graph display middleware are used to store data warehouses, extract and analyze, establish relationships and display threat intelligence data in multiple dimensions. Specific steps include data traction and storage, data preprocessing, information extraction and transformation, comparison and analysis to establish association relationships and graph display.
It realizes the rational storage and analysis of hybrid threat intelligence data, improves the capabilities of threat prevention, attack detection and response, and attack tracing, and forms a multi-dimensional, drill-down analysis knowledge graph.
Smart Images

Figure CN118627499B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of network security, and particularly relates to an analysis method based on hybrid threat intelligence data. Background Art
[0002] At present, threat intelligence data not only comes from multiple different sources, but also includes hybrid data (including structured and unstructured) and discrete data (data distributed in different systems or platforms). From the perspective of data structure, hybrid threat intelligence data includes structured data, semi-structured data, and unstructured data. Structured data includes external data interfaces, threat indicators, and attack indicators. Semi-structured data includes attack traffic packets and abnormal logs. Unstructured data includes meeting codes, sample files, analysis reports, etc. There are the following several problems in analyzing the relationships between these data in daily life:
[0003] 1. The access and storage of multi-source heterogeneous data cannot be stored in a reasonable data warehouse according to the data characteristics and application scenarios, which is convenient for the query and application of subsequent application system development;
[0004] 2. The extraction of data information is inaccurate. When extracting keywords or summary information, the training and optimization of the analysis model cannot be carried out according to actual needs;
[0005] 3. The establishment of association rules for hybrid threat intelligence data and the application display form cannot form a standardized system.
[0006] The appearance of the present invention solves the above several problems. By using various types of data storage middleware (databases), semantic analysis models, and knowledge graph display middleware, threat intelligence data is stored in a data warehouse, extracted and analyzed, relationships are established, and multi-dimensional display is performed. Its specific advantages are as follows:
[0007] 1. According to the formats and content characteristics of various types of data, a comprehensive data warehouse suitable for hybrid threat intelligence data is established. By configuring multiple data source drivers, relational databases, cache databases, full-text retrieval databases, and graph databases are accessed. At the same time, flexible storage expansion can be performed in combination with the transformation of the user's data types and usage scenarios, realizing the reasonable storage of hybrid data information;
[0008] 2. It can extract and analyze knowledge from hybrid threat intelligence data. Using the NER word segmentation technology in natural language processing (NLP), after the general model is retrained with sample data, the automatic extraction of threat information entities for structured data and unstructured data is realized; then, part-of-speech tagging and dependency relationship analysis are performed on these entity information to obtain information knowledge and information association rules.
[0009] 3. Store information and knowledge in the graph database according to the association rules. At the same time, it is possible to compare, transform, and enhance the written information data with other result data to obtain more context information, thereby improving the content displayed in the knowledge graph. Summary of the Invention
[0010] (1) Technical Problems to be Solved
[0011] The technical problem to be solved by the present invention is how to provide an analysis method based on hybrid threat intelligence data to solve problems in aspects such as access storage of multi-source heterogeneous data, inaccurate extraction of data information, establishment of association rules for hybrid threat intelligence data, and inability to form a standardized system for application display forms.
[0012] (2) Technical Solutions
[0013] To solve the above technical problems, the present invention proposes an analysis method based on hybrid threat intelligence data. The method includes the following steps:
[0014] The first step, data introduction and storage: According to the data warehouse design, introduce hybrid threat intelligence data into the data warehouse to achieve classified storage of multi-source information;
[0015] The second step, data preprocessing: Perform standardized preprocessing on the introduced original information data, including data deduplication, establishment of data unique identifiers, data format checking, and data classification;
[0016] The third step, information extraction and transformation: Extract information points from the intelligence data stored in the data warehouse, and extract threat information entities and establish their association relationships through built-in association rule conditions and natural language processing technologies;
[0017] The fourth step, comparative analysis and establishment of association relationships: Establish a graph database, establish association relationships for all knowledge point data in the graph database by creating association rules, and store the relationships in the graph database for subsequent display and application of the threat situation knowledge graph;
[0018] The fifth step, graph display: Use the front-end graph display component to organize the corresponding data structure using the breadth-first search algorithm BFS for knowledge graph display.
[0019] (3) Beneficial Effects
[0020] The present invention proposes an analysis method based on hybrid threat intelligence data. Through in-depth mining and analysis of various threat intelligence data, the present invention assists security management personnel in making reasonable judgments on threat trends, thereby helping security management personnel improve their capabilities in threat prevention, attack detection and response, and attack tracing.
[0021] The technology of the present invention is simple to operate and has a perfect process. It extracts information knowledge points using the keyword weight method, and utilizes the efficient recursive logic of the front-end graph display component and the graph database to form a multi-dimensional and drill-down analysis knowledge graph for result display and application. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 It is a flowchart of the analysis method based on hybrid threat intelligence data of the present invention;
[0023] Figure 2 It is a flowchart of the part-of-speech tagging of a piece of text by the threat information entity extraction model;
[0024] Figure 3 It is a flowchart of establishing association relationships for comparative analysis;
[0025] Figure 4 It is a schematic diagram of the knowledge graph display and search. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0026] To make the objectives, contents and advantages of the present invention clearer, the following further describes in detail the specific embodiments of the present invention with reference to the drawings and embodiments.
[0027] The present invention takes threat intelligence data as the core and is applied to the field of network security.
[0028] The core of the analysis of hybrid threat intelligence data by the present invention lies in the following aspects:
[0029] First, the original information data accessed is preprocessed in a standardized manner, including data deduplication, establishing unique identifiers, data format checking, etc. Through the above means, a large amount of irrelevant information or incorrect information is filtered out, avoiding incorrect analysis or loss of key information;
[0030] Second, natural language processing (NLP) technology is used to construct a threat information entity extraction model applicable to the field of threat intelligence and capable of training and optimization. By extracting information points with higher keyword weights for analysis, knowledge points of threat intelligence data are established;
[0031] Third, the relationship between incremental data and stock data is mined, and then a relationship graph is generated according to the knowledge point association rules, where nodes represent threat information entities and edges represent the relationships between threat information entities, and the graph database is used to store the graph for subsequent analysis and utilization;
[0032] Fourth, using the efficient recursive logic of the front-end graph display component and the graph database, the threat intelligence knowledge points in the library are formed into a multi-dimensional and drill-down analysis knowledge graph page for result display and application.
[0033] How to effectively integrate these hybrid heterogeneous intelligence data to form a unified analysis platform and data warehouse is a current technical challenge; there may be a large amount of invalid information in the collected raw data, and there may be false alarms and missed reports. Without manual participation in data annotation training to optimize the analysis model, inaccurate and missing information points will be extracted; if the knowledge point data is not reasonably stored in the graph database, the knowledge front-end graph display component cannot efficiently obtain message information from the background interface service for efficient rendering.
[0034] The technical application fields based on the analysis of hybrid threat intelligence data are very extensive, including information technology, enterprises, universities, etc. Figure 1 It is a schematic flow diagram of the analysis method based on hybrid threat intelligence data provided by the present invention. As Figure 1 shown, the method includes the following steps:
[0035] The first step, data ingestion and storage: According to the data warehouse design, the hybrid threat intelligence data is ingested into the data warehouse to achieve multi-source information classification storage. Among them, it is divided into offline data and online data according to the data source; and into structured data and unstructured data according to the data type.
[0036] The data types of multi-source information include:
[0037] 1. External data interface data: Customize and develop interface functions to complete the interface calls to multiple external threat intelligence platforms, and store the interface return data in a relational database;
[0038] 2. Threat attack indicators: Structured JSON data, directly analyzed, converted, and stored;
[0039] 3. Semi-structured data: Attack traffic packets and abnormal logs, and custom-developed parsing services can be used to read the content according to their specific structures;
[0040] 4. Unstructured data: mainly includes conference codes, sample files, and analysis report data. The Hash unique identifiers are extracted from the conference codes and sample files, and the existing comprehensive analysis reports, technical analysis reports, and internal summary reports are uploaded through the import function for the analysis report data. After uploading, the interface service is responsible for reading the full text content of the report and storing it in the full-text retrieval database. The information extracted can be associated and compared with the interface data, message data, and report data.
[0041] The second step, data preprocessing: Normalize and preprocess the ingested raw information data, including data deduplication, establishing data unique identifiers, data format checking, data classification, etc.
[0042] S21. After specifying one or more fields according to the rule configuration, set the conditions to be met (the relationship of AND or OR when there are multiple conditions), and use this as the standard for data deduplication.
[0043] S22. Establish a unique data identifier to facilitate association with other types of data.
[0044] S23. According to the national standard requirements of the "Information Security Technology - Network Security Threat Information Format Specification", perform data format specification checks on the content of the information fields to be stored.
[0045] S24. Classify and label the data to support the classification and filtering of data on the retrieval display page.
[0046] Step 3. Information point extraction and conversion: Extract information points from the intelligence data stored in the data warehouse, and extract threat information entities and establish their association relationships through built-in association rule conditions and natural language processing technologies.
[0047] 1. Extract from external data interfaces: Extract information such as intelligence type, credibility, severity level, intelligence source name, and alert host.
[0048] 2. Extract threat indicators: Extract information such as threat type, IOC identifier, status, severity level, and attack group.
[0049] 3. Extract attack indicators: Extract information such as type, IOA identifier, action, rule, number of times, attack method, and attack behavior.
[0050] 4. Extract data from attack traffic packets: Extract the five-tuple information (network transmission protocol, source IP, destination IP, source port, destination port) within the traffic packet.
[0051] 5. Extract data from abnormal log data: Extract information such as event ID and time from the abnormal log data.
[0052] 6. Extract data from conference code data: Extract Hash value information, including MD5, SHA1, and SHA256.
[0053] 7. Extract unstructured data from analysis reports: Through natural language processing (NLP) technology, construct a threat information entity extraction model applicable to the field of threat intelligence and capable of training and optimization. Figure 2 For the threat information entity extraction model, through the part-of-speech tagging process of a piece of text, extract keywords and summaries after extracting the content and performing part-of-speech tagging:
[0054] S301. First, establish a feature function f ja set, and each feature function takes the following parameters as input: ① sentence s; ② the position i of a certain word in the sentence; ③ the label l of the current word i ; ④ the label l of the previous word i-1 .
[0055] S302. Next, assign a weight λ j to each feature function f j . After inputting a sentence s, by summing all the weighted feature functions, the score when the current sentence s is labeled with a certain label sequence l is obtained: where m is the number of feature functions and n is the number of words in the sentence.
[0056] S303. Finally, taking the power and normalizing can convert the score into a probability p(l|s): The summation over l' in the denominator represents considering all possible labelings of the sentence s. According to the above weighted part-of-speech tagging results, it can be used to extract information knowledge points in the analysis report, such as: attacking organization, attack tool, technical and tactical means, meeting code name, etc.
[0057] Step 4. Comparative analysis to establish association relationships: such as Figure 3 , establish a graph database. By creating association rules, establish association relationships for all the knowledge point data in the graph database and store the relationships in the graph database for subsequent display applications of the threat situation knowledge graph.
[0058] Establish associations using the relationships in the graph database. The extracted information has many attributes according to different classifications. Establish equal, inclusion, or bidirectional inclusion association relationships according to the attributes of different information and store them in the graph database for subsequent data association display of the knowledge graph.
[0059] 1. Added classification management to the relationship rules based on the graph database, facilitating users to classify and query system data and perform data drilling; the data display effect is clearer and more understandable when sorted by classification;
[0060] 2. For the same type of data, drill-down displays from different sources support different display presentation results. For example, for a keyword, if it is associated from a report, it can be displayed as the keyword name, and if it is associated from an information point, it can be displayed as the report name;
[0061] 3. The relationship rules have added multiple relationship matching rules. It is not just equality. Data associations can be established through conditions such as the source containing the target or bidirectional inclusion, making the data association ability more powerful.
[0062] The graph database uses Cypher statements, a database query language similar to Structured Query Language (SQL), to perform operations such as adding, deleting, modifying, and querying the stored data. For example:
[0063] ① Create an APT organization named "APT32": CREATE(n:Organization{name:'APT32'}) RETURN n
[0064] CREATE is the create operation, APT is the label representing the type of the node. The curly braces {} represent the attributes of the node. The meaning of this statement is to create a node with the label APT, which has a name attribute with the value APT32.
[0065] ② Create a country / region node: CREATE(n:Location{country:'Vietnam'})
[0066] Similar to the above statement, a node with the label Location and the attribute Vietnam is created.
[0067] ③ Establish an association relationship between the above two nodes through the statement: MERGE(a)-[:Country / Region]->(b).
[0068] ④ Query all nodes related to a: MATCH(a)--() RETURN a.
[0069] Step 5: Graph display: Use the front-end graph display component to organize the corresponding data structure for knowledge graph display, as shown in Figure 4. The front-end graph display component uses the Breadth-First-Search (BFS) algorithm in the graph database. The algorithm logic is as follows:
[0070] S51. Access the starting vertex, which is the data element that triggers the retrieval. It can be an APT organization, a traffic packet, or an executable file, etc.;
[0071] S52. Starting from the vertex, sequentially access each unvisited adjacent node W 1, W 2...... W n , and then continue to sequentially access all unvisited adjacent nodes related to W 1, W 2...... W n ;
[0072] S54. Then, starting from these visited nodes, access all their unvisited adjacent nodes, and so on until all nodes in the graph have been visited.
[0073] Here, the breadth-first search algorithm is specifically illustrated by an example. Given Figure 4 As follows. Assume starting from node a for access, and a is enqueued first. At this time, the queue is not empty, and the head element a is taken out. Since b and c are adjacent to a and have not been visited, b and c are visited in sequence and enqueued in sequence. The queue is not empty, the element b is taken out, and the nodes d and e that are adjacent to b and have not been visited are visited in sequence and enqueued. Here, a is also adjacent to b, but a has been marked as visited, so it is not visited repeatedly. At this time, the queue is not empty, the element c is taken out, and the nodes f and g that are adjacent to c and have not been visited are visited and enqueued. At this time, the element d is taken out, but there are no nodes adjacent to d that have not been visited, so no operation is performed. Continue to take out the head element e, and h is enqueued. Finally, after taking out the queue element h, the queue is empty, and the loop automatically jumps out.
[0074] The key points of the present invention are as follows:
[0075] 1. The introduction, storage, analysis, and display of hybrid threat intelligence data are integrated in the same software, simplifying the operation process and reducing the technical threshold for users;
[0076] 2. The constructed hybrid multi-source threat intelligence data warehouse supports storage in multiple types of databases.
[0077] 3. A threat entity information analysis and extraction model constructed for the threat intelligence field, and it can be trained and optimized;
[0078] 4. A customized threat intelligence knowledge graph front-end display application component.
[0079] Through in-depth mining and analysis of various types of threat intelligence data, the present invention assists security management personnel in making reasonable judgments on threat trends, thereby helping security management personnel improve their capabilities in threat prevention, attack detection and response, and attack traceability.
[0080] The technology of the present invention is simple to operate and has a perfect process. It extracts information knowledge points using the keyword weight method, and forms a multi-dimensional and drill-down analysis knowledge graph by utilizing the efficient recursive logic of the front-end graph display component and the graph database for result display and application.
[0081] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and deformations can be made, and these improvements and deformations should also be regarded as the protection scope of the present invention.
Claims
1. An analysis method based on hybrid threat intelligence data, characterized in that: The method comprises the following steps: Step 1: Data ingestion and storage: Based on the data warehouse design, the hybrid threat intelligence data is ingested into the data warehouse to achieve classified storage of multi-source information. Step 2: Data preprocessing: Standardize and preprocess the imported raw information data, including data deduplication, data unique identification, data format checking and data classification; Step 3: Information extraction and conversion: Extract information points from intelligence data stored in the data warehouse, extract threat information entities and establish their association relationships through built-in association rule conditions and natural language processing technology; Step 4: Comparative analysis to establish association relationships: Establish a graph database, create association rules, establish association relationships between the data of all knowledge points in the graph database, and store the relationships in the graph database for subsequent threat situation knowledge graph display applications; Step 5: Graph display: Use the front-end graph display component to organize the corresponding data structure using the breadth-first search algorithm BFS to display the knowledge graph; in, For unstructured data such as analysis reports, a threat information entity extraction model that is applicable to the field of threat intelligence and can be trained and optimized is constructed through natural language processing (NLP) technology; the threat information entity extraction model extracts keywords and abstracts through the part-of-speech tagging process of a piece of text, extracts content, and tags the part-of-speech; the part-of-speech tagging process of the threat information entity extraction model includes: S301, first establish a characteristic function f j Each feature function takes the following parameters as input: sentence s, the position i of a word in the sentence, and the label l of the current word. i , the label of the previous word l i-1 ; S302, next give each characteristic function f j Assign weight λ j , after inputting a sentence s, we sum all weighted feature functions to get the score when a certain label sequence l is attached to s: Among them, m is the number of feature functions, n is the number of words in the sentence; S303, finally exponentiation and normalization are performed to convert the score into probability p(l|s): The sum of l' in the denominator means considering all possible tags for sentence s, and extracting information points in the analysis report based on the above weighted part-of-speech tagging results.
2. The analysis method based on hybrid threat intelligence data according to claim 1, characterized in that: The data types of the multi-source information in the first step include: external data interface data, threat attack indicators, semi-structured data and unstructured data; External data interface data: Customize and develop interface functions to complete interface calls to multiple external threat intelligence platforms, and store interface return data in a relational database; Threat attack indicators: structured JSON data, direct analysis, conversion, and storage; Semi-structured data: attack traffic packets and abnormal logs, which can be read by customizing and developing parsing services based on their specific structures; Unstructured data: including conference codes, sample files, and analysis report data. Conference codes and sample files are extracted from their unique hash identifiers. Analysis report data uploads existing comprehensive analysis reports, technical analysis reports, and internal summary reports through the import function. After uploading, the interface service is responsible for reading the full text of the report and storing it in the full-text retrieval database.
3. The analysis method based on hybrid threat intelligence data according to claim 1, characterized in that: The second step of preprocessing includes: S21. After the specified fields are configured according to the rules, the conditions are set to perform data deduplication; S22. Establish a unique identifier for data to facilitate association with other types of data; S23. According to the national standard requirements, check the data format of the information field to be stored; S24. Classify and annotate the data, thereby supporting the classification filtering of the data on the search display page.
4. The analysis method based on hybrid threat intelligence data according to claim 1, characterized in that: The information point extraction in the third step includes: Extraction of external data interface: extract intelligence type, credibility, severity level, intelligence source name, alarm host information; Extract threat indicators: extract threat type, IOC identification, status, severity level, and attack group information; Attack indicator extraction: extraction type, IOA identification, action, rule, number of times, attack method, and attack behavior information; Extract attack traffic packet data: extract the five-tuple information in the traffic packet; Extract abnormal log data: extract event ID and time information from abnormal log data; Extract conference code data: extract hash value information, including MD5, SHA1, and SHA256.
5. The analysis method based on hybrid threat intelligence data according to claim 1, characterized in that: In the step S303, the extracted information points include: attack organization, attack tools, technical and tactical means, and conference code name.
6. The analysis method based on hybrid threat intelligence data according to claim 1, characterized in that: The fourth step uses the relationships in the graph database to establish associations. The extracted information has many attributes depending on the classification. Equal, inclusive or bidirectional inclusive associations are established based on the attributes of different information and stored in the graph database for subsequent data association display in the knowledge graph.
7. The analysis method based on hybrid threat intelligence data according to claim 6, characterized in that: In the fourth step, classification management is added to the relationship rules based on the graph database, which is used for classified query and data drilling of system data; for the same type of data, drill-down display from different sources supports different display presentation results. If a keyword is associated from a report, it is displayed as the keyword name; if it is associated from an information point, it is displayed as the report name; the relationship rules include a variety of relationship matching rules, which establish associations between data through equality, source inclusion target, and bidirectional inclusion conditions, making the data association capability more powerful.
8. The analysis method based on hybrid threat intelligence data according to claim 6, characterized in that: Graph databases use the database query language Cypher statements to add, delete, modify and query stored data.
9. The analysis method based on hybrid threat intelligence data according to claim 6, characterized in that: The fifth step of the breadth-first search algorithm BFS includes: S51, access the starting vertex, that is, the data element that triggers the retrieval, including: APT organization, traffic package, or executable file; S52, starting from the vertex, sequentially visit each unvisited adjacent node W associated with the vertex 1, W 2...... W n , and then continue to visit the W 1, W 2...... W n All unvisited adjacent nodes with associated relationships; S54, starting from these visited nodes, visit all their unvisited adjacent nodes, and so on until all nodes in the graph have been visited.
Citation Information
Patent Citations
Threat intelligence attribution method based on graph attention mechanism
CN116467438A
Visual development system for knowledge graph
CN118093895A