A multi-source cross-modal threat intelligence processing method, device, equipment and medium

By constructing a multi-source cross-modal threat intelligence processing method, the problem of difficulty in integrating multi-source cross-modal data in traditional methods is solved, efficient fusion of multi-source cross-modal data and the improvement of knowledge graphs is achieved, and the accuracy and reliability of intelligence analysis are improved.

CN120296688BActive Publication Date: 2025-08-15QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510788310.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-08-15
Estimated Expiration
2045-06-13

AI Technical Summary

Technical Problem

Traditional threat intelligence processing methods are difficult to effectively integrate multi-source cross-modal data, resulting in imperfect construction of knowledge graphs and the inability to fully utilize the rich information of multi-source cross-modal data. It is difficult to integrate cross-modal features, affecting the accuracy and reliability of intelligence analysis.

Method used

Multi-source cross-modal threat intelligence processing methods are adopted, and data collection, preprocessing and feature extraction are carried out by building a multi-source cross-modal four-dimensional acquisition architecture, and cross-modal feature alignment and fusion technology are used, combined with cross-modal semantic alignment technology of graph structures, knowledge graphs are built to realize unified representation and efficient utilization of natural language, images, malicious code and system logs.

Benefits of technology

It realizes comprehensive and accurate processing of multi-source cross-modal data, enriches the structure and semantic expression of the knowledge graph, improves the accuracy and reliability of intelligence analysis, and provides stronger support for subsequent threat analysis and decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296688B_ABST
    Figure CN120296688B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of electronic digital data processing, and in particular relates to a multi-source cross-modal threat intelligence processing method, device, equipment and medium. Through the "two-level progressive fusion" technical route, the innovation of the threat intelligence processing paradigm is achieved: the cross-modal fusion layer focuses on the deep coupling of heterogeneous features of natural language CTI and image CTI, and adopts the three-level processing mechanism of "intra-modal self-attention, cross-modal cross-attention, and LSTM temporal fusion" to achieve the mapping of natural language semantics and visual information; the multi-source alignment layer jointly models the cross-modal fused CTI feature vector with malicious code and log data, constructs a three-dimensional association network of "CTI knowledge graph, malicious code graph, log timing graph", and realizes cross-source semantic space unification through the heterogeneous graph attention network. It solves the problems of traditional threat intelligence processing methods in the face of multi-source cross-modal data, such as difficulty in collection and integration, insufficient preprocessing, difficulty in cross-modal feature fusion, and imperfect knowledge graph construction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of electronic digital data processing, and in particular relates to a multi-source cross-modal threat intelligence processing method, device, equipment and medium. Background Art

[0002] In recent years, cyberattacks have exhibited multi-stage, multi-technique combinations. Defenders need to integrate multi-source, cross-modal data, such as natural language reports, malicious code, and logs, to build a correlated threat awareness system. Traditional threat intelligence processing methods primarily target single-source, single-modality data, making it difficult to effectively integrate multi-source, cross-modal threat intelligence. Furthermore, when constructing knowledge graphs, traditional threat intelligence processing methods often fail to fully utilize the rich information from multi-source, cross-modal data. The construction of knowledge graphs requires in-depth semantic understanding and relationship mining of data. Traditional methods struggle to achieve comprehensive, accurate modeling and efficient utilization of multi-source, cross-modal threat intelligence, resulting in incomplete knowledge graph structure and semantic representation. Furthermore, threat intelligence processing methods also face difficulties in fusing features from different modalities. Data from different modalities have different feature spaces and semantic information. How to effectively fuse them together to form a unified feature representation that captures the correlations and complementary information between modalities is a problem that needs to be addressed.

[0003] Chinese patent document CN119829942A discloses an intelligent cloud simulation multimodal intelligence processing method, apparatus, device, and storage medium. An intelligence data processing model is established based on the BERT model. A feature multi-source intelligence dataset is input into the intelligence data processing model for identification. A cloud simulation platform is established, and intelligence identification content is input into the cloud simulation platform to establish a virtual intelligence space. Real-time feedback data is obtained based on the intelligence evolution plan. The intelligence identification content is adjusted based on a reinforcement learning algorithm to obtain target intelligence information. However, its data processing is not targeted enough. Although it can process multimodal data, it does not conduct in-depth mining and analysis of the characteristics of different modal data when processing threat intelligence. The extraction and fusion of natural language, image, and audio features are mainly achieved through NLP and CNN models. This does not fully consider the particularity and complexity of threat intelligence in the field of cyberspace security, and lacks in-depth mining and analysis of key information such as potential attack patterns and malicious code in threat intelligence. Secondly, the fusion effect is limited. It fuses different modal features by sequentially connecting them end to end. This simple connection method may not fully capture the complex relationships and potential associations between different modalities. In practical applications, this fusion method may result in inaccurate and incomplete feature representation after fusion, thus affecting the accuracy and reliability of intelligence analysis.

[0004] Chinese patent document CN116150392A uses a multi-head attention layer to aggregate relational features of an initial threat intelligence knowledge graph, aggregating different semantic features to obtain attention weights for multiple preset relational types. Taking into account the characteristics of each entity based on the preset relational type, the node attention layer in the graph attention network aggregates the different neighbor features of each node to obtain the aggregated neighbor features of each node based on the multiple preset relational types. The attention weights of the multiple preset relational types and the aggregated neighbor features of each node based on the multiple preset relational types are then concatenated and fused to obtain the final target features of each node. Finally, the target features are embedded into the semantic information of the initial threat intelligence knowledge graph, making the semantic information of the threat intelligence knowledge graph more comprehensive and effective. However, this method focuses on processing the semantic information of the initial threat intelligence knowledge graph, mainly aggregating relational features and node semantic information through the graph attention network. It lacks the comprehensive collection, preprocessing, and feature extraction of multi-source cross-modal data, and lacks an effective multi-source data fusion and alignment algorithm, making it difficult to address the large differences in feature space and semantic information of data from different modalities. The data type coverage is insufficient, and there is no mention of how to integrate various heterogeneous data such as natural language, images, malicious code, logs, etc., which makes it impossible to fully utilize the rich information in multi-source cross-modal data to build a complete knowledge graph, resulting in the knowledge graph being lacking in semantic expression and structural integrity. Summary of the Invention

[0005] The present invention aims to provide a multi-source cross-modal threat intelligence processing method to comprehensively and accurately process and integrate threat intelligence from different sources and modalities, so as to solve the problems of traditional threat intelligence processing methods when facing multi-source cross-modal data, such as difficulty in collection and integration, insufficient preprocessing, difficulty in cross-modal feature fusion, and imperfect knowledge graph construction.

[0006] The present invention also discloses a device loaded with a multi-source cross-modal threat intelligence processing method.

[0007] The invention also discloses an electronic device for implementing the method.

[0008] The present invention also discloses a machine-readable storage medium for implementing the above method.

[0009] In order to solve the above technical problems, the present invention provides a multi-source cross-modal threat intelligence processing method, which includes:

[0010] S1. Data Collection: Build a multi-source, cross-modal, four-dimensional acquisition architecture to collect multi-source, cross-modal threat intelligence data including natural language, images, malicious code, and system logs;

[0011] S2. Data preprocessing: Clean, transform and extract features of the multi-source cross-modal threat intelligence data collected in step S1 to obtain natural language feature vectors , image feature vector , malicious code feature vector , system log feature vector ;

[0012] S3. Cross-modal feature alignment and fusion:

[0013] S31, the natural language feature vector obtained in step S2 and image feature vector Perform feature alignment to obtain the aligned natural language feature vector and image feature vector ;

[0014] S32. Aligned natural language feature vector and image feature vector Perform intra-modal self-attention calculation to obtain the natural language feature vectors after intra-modal self-attention calculation and image feature vector ;Will and Perform cross-modal attention fusion to obtain multi-head feature vectors ;

[0015] S33, the multi-head feature vector obtained in step S32 After splicing, the final cross-modal fusion vector is obtained ;

[0016] S4, multi-source data fusion alignment: using graph-based cross-modal semantic alignment technology, the code feature vector obtained after pre-processing in step S2 is aligned and the log eigenvector And the cross-modal fusion vector obtained after step S3 Align in a unified semantic space to obtain a multi-source heterogeneous graph G;

[0017] S5. Knowledge graph construction: Based on the threat intelligence analysis requirements and the multi-source heterogeneous graph G fused and aligned in step S4, a knowledge graph is constructed through graph model definition and design, knowledge extraction and fusion, storage and query optimization.

[0018] Preferably, step S1 is specifically as follows: natural language threat intelligence data is collected using a distributed crawler cluster; image threat intelligence data is collected using an automated screenshot system; malicious code threat intelligence data is collected using a sandbox behavior capture device; system log threat intelligence data is collected using a time series log processing pipeline.

[0019] Preferably, in step S2, the natural language threat intelligence data T is subjected to text segmentation and word vector representation extraction to obtain a natural language feature vector ;

[0020] By analyzing image threat intelligence data Perform image size unification and feature extraction to obtain image feature vector ;

[0021] Through malicious code threat intelligence data Perform static and dynamic feature analysis to obtain malicious code feature vectors ;

[0022] By extracting key features and converting numerical values of system log threat intelligence data L, we can obtain the system log feature vector .

[0023] Preferably, in step S31, the natural language feature vector obtained in step S2 is and image feature vector Feature alignment is to transform the feature vectors of different dimensions into and Mapping to the target dimension is achieved as follows:

[0024] The natural language feature vector Align to obtain the aligned natural language feature vector ,as follows:

[0025] (1)

[0026] in, is the aligned natural language feature vector, is the natural language feature vector, is the weight matrix, is the bias vector, Used to transform natural language feature vectors Map to target dimension;

[0027] For the image feature vector , transform the dimension to the target dimension through a fully connected layer:

[0028] (2)

[0029] in, is the aligned image feature vector, represents the fully connected layer, is the image feature vector.

[0030] Preferably, step S32 and Perform cross-modal attention fusion to obtain multi-head feature vectors , specifically:

[0031] After the self-attention calculation within the modality, and Perform linear transformation to obtain query Q, key and numerical values vector; calculate the attention score using the query vector Q and the key vector The inner product of , and the attention weight is calculated by the softmax function; the attention score and the value vector Multiply them together to get the weighted representation of each feature vector under different attention heads; connect the weighted representations of each attention head and transform them linearly Generate the final fusion representation vector , the formula is as follows:

[0032] Will and Transformed into query vector Query, key-value vector Key and value vector Value, where the query vector comes from natural language and the key-value vector comes from the image:

[0033] (3)

[0034] (4)

[0035] (5)

[0036] in, is a learnable parameter matrix, Q refers to the query vector Query, Refers to the key-value vector Key, is the value vector Value, It refers to the natural language feature vector obtained after the self-attention calculation within the modality, It refers to the image feature vector obtained after the intra-modal self-attention calculation;

[0037] The weighted representation of each attention head is calculated as:

[0038] (6)

[0039] in, refers to the output of the i-th attention head, refers to the attention calculation function, is the query weight matrix of the i-th attention head, refers to the key weight matrix of the i-th attention head, refers to the value weight matrix of the i-th attention head, It refers to the natural language feature vector obtained after the self-attention calculation within the modality, It refers to the image feature vector obtained after the intra-modal self-attention calculation;

[0040] Long-headed eigenvector The calculation is:

[0041] (7)

[0042] in, is the multi-head feature vector, h is the number of attention heads, is the final linear transformation matrix, is the multi-head attention mechanism function, refers to the output of each independent attention head, It refers to the natural language feature vector obtained after the self-attention calculation within the modality, It refers to the image feature vector obtained after intra-modal self-attention calculation.

[0043] Preferably, in step S33, the multi-head feature vector obtained in step S32 is After splicing, the final cross-modal fusion vector is obtained , specifically:

[0044] Long-headed eigenvector and the natural language feature vector calculated by intra-modal self-attention Spliced together to form a new feature vector ; The concatenated feature vector Input into the LSTM model and obtain the intermediate vector representation through weight adjustment ; Then the middle vector and the image feature vector calculated by intra-modal self-attention Spliced together to form a new feature vector ; The concatenated feature vector Input into the LSTM model and obtain the final cross-modal fusion vector through weight adjustment ,in,

[0045] (8)

[0046] (9)

[0047] (10)

[0048] (11)

[0049] is the concatenated feature vector, is the multi-headed eigenvector, is the natural language feature vector calculated by intra-modal self-attention, is the middle vector, represents a long memory network, is the middle vector and the image feature vector calculated by intra-modal self-attention The new feature vector formed by splicing, is the final cross-modal fusion vector.

[0050] Further preferably, the weight adjustment is achieved through a back-propagation algorithm and an optimizer.

[0051] More preferably, the optimizer SGD is used to update the weights according to the gradient calculated by the back propagation algorithm, and the cross-modal fusion feature vector Z is finally obtained.

[0052] Preferably, step S4 is specifically as follows: based on the malicious code feature vector obtained in step S2 and system log feature vector Build malicious code call graphs separately and system log timing diagram , based on the cross-modal fusion vector obtained in step S3 Constructing a cross-modal fusion network threat intelligence CTI knowledge graph , use the graph neural network HAN to generate node embedding for each graph, solve the optimal matching matrix through the gradient descent optimization algorithm, realize the fusion alignment of multi-source data, and form a multi-source heterogeneous graph G.

[0053] Further preferably, the generating of node embedding for each graph using the graph neural network HAN is specifically as follows:

[0054] Using the heterogeneous graph attention network HAN, meta-paths are used to generate hierarchical graph node embedding vectors that retain semantic features:

[0055] (12)

[0056] in, It means the Nodes in the layer The embedding vector of is the meta-path, is the meta-path level attention coefficient, It represents the node under the meta-path P The result of the transformation of the embedding vector is Is the activation function;

[0057] Through the meta-path level attention mechanism, hierarchical feature extraction is implemented for each graph structure, and the node embeddings under different meta-paths are weighted fused to obtain the final node embedding matrix:

[0058] (13)

[0059] (14)

[0060] in, It is a graph structure that is analyzed through graph neural network The node embedding matrix obtained after processing is It is a graph structure that is analyzed through graph neural network The node embedding matrix obtained after processing is It is a graph structure that is analyzed through graph neural network The node embedding matrix obtained after processing is It is a graph neural network model. is a matrix Dimensions, It is a graph structure The number of nodes, It is a graph structure The number of nodes, It is a graph structure The number of nodes, It is a graph structure The node set of It is a graph structure The node set of It is a graph structure The node set of , d is the dimension of the embedding vector.

[0061] Further preferably, the optimal matching matrix is solved by the gradient descent optimization algorithm to achieve the fusion alignment of multi-source data, specifically:

[0062] calculate and The similarity matrix between:

[0063] (15)

[0064] in: yes and The similarity matrix between Yes The embedding vector of node j in the graph, Yes The embedding vector of node k in the graph, sim is the similarity calculation function;

[0065] Define the alignment cost function: (16)

[0066] in: yes and The matching matrix between them, j is node j, k is node k, yes and The similarity matrix between It is a graph structure The number of nodes, It is a graph structure The number of nodes, yes The alignment cost function between ; Solve the optimal matching matrix through gradient descent optimization algorithm ; According to the obtained optimal matching matrix ,Will and The fused embedding is :

[0067] (17)

[0068] in: It is a graph structure and The node embedding matrix obtained after fusion, represents element-wise multiplication, yes and The fused optimal matching matrix is used to indicate which nodes have been aligned. It is a graph structure that is analyzed through graph neural network The node embedding matrix obtained after processing is It is a graph structure that is analyzed through graph neural network The node embedding matrix obtained after processing;

[0069] calculate and after fusion The similarity matrix between:

[0070] (18)

[0071] in: yes and The similarity matrix between Yes Nodes in the graph The embedding vector of yes The fused node j is embedded; sim is the similarity calculation function;

[0072] Alignment optimization, define the alignment cost function:

[0073] (19)

[0074] in: yes and The alignment cost function between It is a graph structure The number of nodes, It is a graph structure The number of nodes, It is a graph structure The number of nodes, yes and The similarity matrix between yes and The matching matrix between them is optimized to minimize , solve the optimal matching matrix through the gradient descent optimization algorithm , the optimal matching matrix guide and The alignment relationship between nodes.

[0075] Preferably, the graph model definition and design in step S5 specifically include: determining the ontology structure of the knowledge graph, including entity types, relationship types, and attribute definitions of entities and relationships; the knowledge extraction and fusion include extracting corresponding entities, relationships, and attribute information from the multi-source heterogeneous graph G after fusion and alignment in S4 according to the graph model definition; the storage specifically includes: selecting the graph database Neo4j to efficiently store the constructed knowledge graph.

[0076] In another aspect of the present invention, a multi-source cross-modal threat intelligence processing device is provided, the device comprising:

[0077] The multi-source cross-modal data collection module is used to collect multi-source intelligence data sets from various channels, including data in various forms such as natural language, images, logs, and codes to obtain the initial multi-source intelligence data set. The data collection module collects multi-source cross-modal threat intelligence data from various data sources according to preset strategies and transmits it to the data preprocessing module;

[0078] The data preprocessing module is used to preprocess the multi-source cross-modal intelligence data set and is responsible for cleaning, converting and extracting features from the collected multi-source cross-modal threat intelligence data;

[0079] The cross-modal feature fusion module is used to combine feature vectors from different modalities to form a unified feature representation. It uses intra-modal self-attention, cross-modal cross-attention mechanisms, and feature splicing and transformation techniques to achieve deep coupling of heterogeneous features between natural language and image CTI, generate a cross-modal fusion vector, and then pass it to the multi-source data alignment module.

[0080] The multi-source data alignment module adopts cross-modal semantic alignment technology based on graph structure to align multi-source cross-modal threat intelligence data, malicious code analysis data and system log data in a unified semantic space. It solves the optimal matching matrix through optimization algorithm, realizes the fusion alignment of multi-source data, constructs a three-dimensional association network, and finally forms a multi-source heterogeneous graph G, providing an integrated data foundation for the knowledge graph construction module.

[0081] The knowledge graph construction module is responsible for graph model definition and design. It selects the graph database Neo4j to efficiently store the constructed knowledge graph, extracts knowledge from the multi-source heterogeneous graph G and resolves conflicts, constructs a complete knowledge graph, and stores it in the graph database for subsequent query and analysis, thereby achieving comprehensive and accurate representation and efficient utilization of network threat intelligence.

[0082] In another aspect of the present invention, an electronic device is provided, comprising:

[0083] at least one processor; and,

[0084] A memory storing instructions, which, when executed by the at least one processor, causes the at least one processor to execute the multi-source cross-modal threat intelligence processing method as described above.

[0085] In another aspect of the present invention, a machine-readable storage medium is provided, which stores executable instructions. When the instructions are executed, the machine executes the multi-source cross-modal threat intelligence processing method as described above.

[0086] Compared with the prior art, the present invention has the following beneficial effects:

[0087] (1) The present invention provides a "two-level progressive fusion" technology route, which solves the traditional fusion dilemma and defects of traditional methods. Traditional methods directly perform global fusion on four types of heterogeneous data - natural language / images / codes / logs, which will lead to an explosion of feature space dimensions and serious interference between modalities, such as code features drowning out natural language semantics.

[0088] (2) This invention achieves an innovation in the threat intelligence processing paradigm through a step-by-step processing architecture of "cross-modal feature fusion → multi-source data alignment". The first level, the cross-modal fusion layer, focuses on the deep coupling of heterogeneous features of natural language CTI and image CTI, and adopts a three-level processing mechanism of "intra-modal self-attention → cross-modal cross-attention → LSTM temporal fusion" to first realize the mapping of natural language semantics and visual information; the second level, the multi-source alignment layer, jointly models the cross-modal fused CTI feature vector with malicious code and log data, constructs a three-dimensional association network of "CTI knowledge graph-malicious code graph-log temporal graph", and realizes cross-source semantic space unification through the heterogeneous graph attention network.

[0089] (3) The present invention can enrich the structure and semantic expression of the knowledge graph. Traditional methods often fail to fully utilize the rich information of multi-source cross-modal data when constructing knowledge graphs, resulting in imperfect structure and semantic expression of the knowledge graph. The present invention can fully utilize the rich information of multi-source cross-modal data through cross-modal feature fusion and multi-source data alignment, making the structure of the knowledge graph more reasonable and the semantic expression richer. Through the application of hierarchical graph embedding representation and heterogeneous graph attention network HAN, the complex relationship between nodes can be captured more accurately, the accuracy and reliability of the knowledge graph can be improved, and more powerful support can be provided for subsequent threat analysis and decision-making. BRIEF DESCRIPTION OF THE DRAWINGS

[0090] Figure 1 This is a general processing flow chart of Example 1 of the present invention;

[0091] Figure 2 This is a flowchart of the multi-source cross-modality CTI two-level progressive fusion process in Example 1 of the present invention;

[0092] Figure 3 This is a diagram of the cross-modal natural language and image CTI feature extraction and fusion model in Example 1 of the present invention;

[0093] Figure 4 This is a flowchart of multi-source data graph alignment in Example 1 of the present invention. DETAILED DESCRIPTION

[0094] Example 1

[0095] The present invention provides a multi-source cross-modal threat intelligence processing method, which combines Figure 1 As shown, the architecture model combines Figure 2 As shown:

[0096] S1. Data collection: Build a multi-source, cross-modal, four-dimensional acquisition architecture to collect multi-source, cross-modal threat intelligence data including natural language, images, malicious code, and system logs.

[0097] Natural language CTI, or natural language threat intelligence data, is collected using a distributed crawler cluster; image CTI is collected using an automated screenshot system; malicious code threat intelligence data is collected using a sandbox behavior capture device; and system log threat intelligence data is collected using a time series log processing pipeline.

[0098] Collect natural language threat intelligence, including detailed descriptions of attack techniques from sources such as MITRE ATT&CK, security blogs, and news websites. Specifically, configure the ATT&CK Navigator API to batch retrieve TTPs metadata from the MITRE ATT&CK knowledge base. Deploy a Scrapy-Redis distributed crawler cluster to capture natural language technical analysis reports published by security vendors such as Palo Alto Unit 42 and FireEye. Use the standardized interface of the MISP threat intelligence platform to extract natural language descriptions of related events and indicators of compromise (IoCs).

[0099] Collect threat intelligence in the form of images, searching for relevant images such as malware interface screenshots or attack flow charts from security reports, blog posts, or technical forums. Specifically, build a PDFBox-JSON parsing pipeline to extract embedded attack flow charts from PDF documents such as Kaspersky APT reports and Mandiant Red Team manuals; develop a Selenium automated capture system to take regional screenshots of visual dashboards from threat intelligence platforms such as Rapid7 and Recorded Future; and deploy Tesseract 5.0, an OCR natural language recognition engine, to convert technical descriptions in screenshots into structured natural language.

[0100] Collect malicious code samples from malware sample repositories or analysis reports published by security companies. Specifically, obtain sandbox behavior logs and sample metadata through the VirusTotal Intelligence API; configure the CAPE sandbox dynamic analysis system to capture API call sequences and process tree information; and perform CodeQL scanning on GitHub code repositories to identify potentially malicious code snippets.

[0101] Collect log information, obtaining log examples related to attack behavior from analysis reports published by security companies, open source log datasets, or internal system logs. Specifically, establish an Elastic Stack log processing pipeline to collect raw logs from firewalls and IDS devices; forward Windows system security logs via the Syslog-NG protocol, focusing on capturing Event ID 4688 process creation and 4624 login events; and extract historical log archive files from the HDFS distributed storage cluster, with a time span of at least six months.

[0102] To increase automation, we built a data pipeline based on Apache NiFi, integrating the PDFBox parser for PDF vector graphics and the Tika natural language extractor to automate the collection of open source intelligence. We used tools such as the ELK Stack to collect and analyze log data, employing the Grok pattern matching engine to convert unstructured logs into fields. We used the CAPESandbox tool to analyze malicious code and generate reports, and deployed the BinaryNinja plugin to extract the AST structure of deobfuscated code.

[0103] To ensure the correspondence between data, unified identifiers are used when collecting and organizing data. In terms of semantic association, a mapping relationship between MITRE ATT&CK technical identifiers and Cyber Kill Chain attack stages is established. In terms of spatiotemporal association, the created_time_ref field of the STIX2.1 specification is extended in the JSON data structure to record the ISO 8601 standard timestamp of each data source. The IEEE 1588 Precision Time Protocol (PTP) is used to achieve cross-device clock synchronization, with a time deviation of ≤1ms.

[0104] Organize the collected natural language descriptions, images, logs, and malicious codes into a unified JSON file format.

[0105] S2. Data preprocessing: Clean, transform and extract features of the multi-source cross-modal threat intelligence data collected in step S1 to obtain natural language feature vectors , image feature vector , malicious code feature vector , system log feature vector ;

[0106] During the data preprocessing phase, the multi-source, cross-modal threat intelligence data collected by S1 undergoes systematic cleaning, conversion, and feature extraction. This aims to remove noise and redundant information, unify data formats, and extract feature representations for various data types, laying the foundation for subsequent feature fusion and knowledge graph construction. This includes text segmentation and word vector representation extraction for natural language threat intelligence, image resizing and feature extraction for image threat intelligence, static and dynamic feature analysis of malicious code samples, and key feature extraction and numerical conversion of system logs. These operations ensure data accuracy and usability, improving the efficiency and quality of subsequent analysis and fusion.

[0107] Specifically, we use the WordPunctTokenizer from the NLTK library to segment natural language into words or phrases, remove stop words and punctuation, and obtain the word vector representation of each word. We load the BERT-base pre-trained model and input the natural language sequence after word segmentation. , generate dynamic word vectors , perform mean pooling operation on the sequence word vector to generate natural language feature vector .

[0108] Image threat intelligence data Preprocessing is performed to convert the original image into a format suitable for model input, while removing noise and standardizing image data. The bilinear interpolation algorithm is used to unify the image size to 224×224 pixels. , perform channel normalization. Load the ResNet-50 pre-trained model, remove the last fully connected layer, extract the second-to-last layer features, and first transform the input image Divided into multiple areas Then each area Input to the ResNet model. For each region , after being processed by the ResNet-50 model, the feature representation of the region is obtained The feature vector of the entire image is obtained by averaging and normalizing the features of each region. .

[0109] Then, the malicious code sample Use the IDA Pro decompiler to parse the PE file, extract static features, extract the function call graph and API import table, perform abstract syntax tree (AST) analysis on the code, extract code structure information, and count the frequency of specific code patterns. Use the code embedding technology CodeBERT to analyze dynamic behavior, convert code snippets into embedding vectors, and extract key functions and API calls as features. Use One-Hot Encoding to convert the extracted features into numerical representations to obtain malicious code feature vectors. .

[0110] Parsing log data , use the log parsing tool ELK Stack to extract key features from the log, configure the Grok pattern matching engine, define a regular template to parse unstructured logs: %{TIMESTAMP_ISO8601:timestamp} %{LOGLEVEL:level} - %{DATA:event}, use One-Hot Encoding to convert the extracted information into a numerical representation to obtain the system log feature vector .

[0111] S3. Cross-modal feature alignment and fusion

[0112] Cross-modal feature fusion combines feature vectors from different modalities, such as natural language CTI and image CTI, to form a unified feature representation that better captures the associations and complementary information between the modalities. For each modality, we first need to extract a feature vector that represents its semantic information.

[0113] S31, the natural language feature vector obtained in step S2 and image feature vector Perform feature alignment to obtain the aligned natural language feature vector and image feature vector ;

[0114] After S2 preprocessing, the feature vectors of natural language CTI and image CTI are obtained. 、 Perform feature alignment to align features of different modalities to the same feature space so that fusion operations can be performed directly and feature vectors of different dimensions are mapped to the target dimension through linear transformation. Figure 3 shown.

[0115] The feature vectors of different dimensions are mapped to the target dimension through linear transformation as follows:

[0116] The natural language feature vector Align to obtain aligned natural language features :

[0117] (1)

[0118] in, is the aligned natural language feature vector, is the weight matrix, is the natural language feature vector, Is the bias vector used to convert the natural language feature vector Mapped to the target dimension. To ensure that the image feature vector and natural language feature vectors The dimension of the image feature vector is the same as that of the image feature vector. , transform the dimension to the target dimension through a fully connected layer:

[0119] (2)

[0120] in, is the aligned image feature vector, represents the fully connected layer, is the image feature vector.

[0121] S32. Aligned natural language feature vector and image feature vector Perform intra-modal self-attention calculation to obtain the natural language feature vectors after intra-modal self-attention calculation and image feature vector ;Will and Perform cross-modal attention fusion to obtain multi-head feature vectors ;

[0122] In the embodiment of the present invention, the aligned natural language feature vector and image feature vector First, we perform intra-modal self-attention calculation and implement a multi-head self-attention mechanism on the natural language features. The number of heads h=8 is determined by neural architecture search (NAS). The image features are processed using the patch-level attention mechanism of Vision Transformer, which divides the feature map into a 16×16 patch sequence. The natural language feature vector after intra-modal self-attention calculation is obtained. and image feature vector ;

[0123] Will and Perform cross-modal attention fusion to obtain multi-head feature vectors Specifically:

[0124] After the self-attention calculation within the modality, and Through the multi-head attention mechanism, each head can learn different semantic information, enabling the model to pay weighted attention to these features separately, thereby better capturing the complex relationships between features.

[0125] right and Perform linear transformation to obtain query Q, key and numerical values Vector. Calculate the attention score using the query vector Q and the key vector The inner product of , and the attention weight is calculated by the softmax function. Multiply them together to get the weighted representation of each feature vector under different attention heads. Connect the weighted representations of each attention head and transform them linearly. Generate the final fusion representation vector .

[0126] Will and Transformed into query vector Query, key-value vector Key and value vector Value, where the query vector comes from natural language and the key-value vector comes from the image:

[0127] (3)

[0128] (4)

[0129] (5)

[0130] in, is a learnable parameter matrix, Q refers to the query vector Query, Refers to the key-value vector Key, is the value vector Value, It refers to the natural language feature vector obtained after the self-attention calculation within the modality, It refers to the image feature vector obtained after the self-attention calculation within the modality. The role of the weight matrix is to convert the input feature vector into the query vector Q and perform a linear transformation on the output of the multi-head attention mechanism to generate the final fusion representation. During the training process of the model, The parameters of the model are updated through the back-propagation algorithm to maximize the performance indicators of the model.

[0131] Calculate the attention score. The output of the multi-head attention mechanism is composed of the results of each attention head, which are linearly transformed and concatenated to form a comprehensive representation. For each attention head, the attention weight needs to be calculated. This is achieved by calculating the inner product of the query vector Q and the key vector V, and then applying the softmax function. The outputs of multiple attention heads are concatenated and linearly transformed to obtain the final multi-head feature vector :

[0132] (6)

[0133] Where h is the number of attention heads, is the final linear transformation matrix, is the long-headed eigenvector, is the multi-head attention mechanism function, refers to the output of each independent attention head, It refers to the natural language feature vector obtained after the self-attention calculation within the modality, It refers to the image feature vector obtained after intra-modal self-attention calculation.

[0134] The weighted representation of each attention head is calculated as:

[0135] (7)

[0136] in, refers to the output of the i-th attention head, refers to the attention calculation function, is the query weight matrix of the i-th attention head, refers to the key weight matrix of the i-th attention head, refers to the value weight matrix of the i-th attention head, It refers to the natural language feature vector obtained after the self-attention calculation within the modality, It refers to the image feature vector obtained after intra-modal self-attention calculation.

[0137] S33, the multi-head feature vector obtained in step S32 After splicing, that is, LSTM time series fusion, the final cross-modal fusion vector is obtained , specifically;

[0138] The multi-headed feature vector and the natural language feature vector calculated by intra-modal self-attention Spliced together to form a new feature vector ; The concatenated feature vector Input into the LSTM model and obtain the intermediate vector representation through weight adjustment , the LSTM model can capture the long-term dependencies in sequence data; then the intermediate vector and the image feature vector calculated by intra-modal self-attention Spliced together to form a new feature vector ; The concatenated feature vector Input into the LSTM model and obtain the final cross-modal fusion vector through weight adjustment ,in,

[0139] (8)

[0140] (9)

[0141] (10)

[0142] (11)

[0143] is the concatenated feature vector, is the multi-headed eigenvector, is the natural language feature vector calculated by intra-modal self-attention, is the middle vector, represents a long memory network, is the middle vector and the image feature vector calculated by intra-modal self-attention The new feature vector formed by splicing, is the final cross-modal fusion vector.

[0144] In LSTM processing, weight adjustment is automatically performed during the training process. The LSTM weights primarily include the weight matrix and bias vector for the input gate, forget gate, output gate, and cell state. Weight adjustment is achieved using the backpropagation algorithm and optimizers such as Adam and SGD. The specific steps are as follows:

[0145] During the forward propagation process, LSTM calculates the output and the new state based on the input data and the current state:

[0146] (12)

[0147] (13)

[0148] (14)

[0149] (15)

[0150] (16)

[0151] in, is the input gate, the output value at time step t; σ is the sigmoid function, is the weight matrix of the input gate, is the input, is the hidden state weight matrix of the input gate; is the hidden state at the previous moment, is the bias term of the input gate; is the output value of the forget gate at time step t, is the weight matrix of the forget gate, is the hidden state weight matrix of the forget gate, Bias vector of forget gate; is a memory unit, the value at time step t, storing long-term state information; ⊙ is element-by-element multiplication, It is the memory state of the previous moment; is the hyperbolic tangent function, an activation function; is the input weight matrix of the candidate values of the memory cell, is the hidden state weight matrix of the candidate values of the memory cell, is the bias vector of the candidate value of the memory cell; is the output gate, the output value at time step t, is the input weight matrix of the output gate, is the hidden state weight matrix of the output gate, is the bias vector of the output gate; It is the hidden state, and its value at time step t contains the output information of the current time step and will be passed to the next time step.

[0152] According to the output of the model and the true label Calculate the loss function L = Loss( , ), calculate the gradient of loss with respect to weights through the back-propagation algorithm:

[0153] (17)

[0154] in, represents the loss function, is the weight matrix of the input gate. is the weight matrix of the forget gate. is the input weight matrix of the output gate. is the input weight matrix of the candidate values of the memory cell.

[0155] Use the optimizer SGD to update the weights according to the calculated gradient, and the final cross-modal fusion feature vector is Z for subsequent use.

[0156] S4, multi-source data fusion alignment: using graph-based cross-modal semantic alignment technology, the code feature vector obtained after pre-processing in step S2 is aligned and the log eigenvector And the cross-modal fusion vector obtained after S3 processing Align in a unified semantic space to obtain a multi-source heterogeneous graph G, and combine the processing methods Figure 4 shown.

[0157] Multi-source graph structure modeling first converts the feature representations of different data sources into graph data structures and constructs a three-dimensional association network.

[0158] Specifically, based on the malicious code feature vector obtained in step S2 and system log feature vector Build malicious code call graphs separately and system log sequence diagram , based on the cross-modal fusion vector obtained in step S3 Constructing a cross-modal fusion CTI knowledge graph , use the graph neural network HAN to generate node embedding for each graph, solve the optimal matching matrix through the gradient descent optimization algorithm, realize the fusion alignment of multi-source data, and form a multi-source heterogeneous graph G.

[0159] Cross-modal fusion CTI knowledge graph :With the attack entity as the core, the threat intelligence after cross-modal fusion is converted into an attribute graph structure based on the final cross-modal fusion vector Z. Define three types of nodes: attacker Actor, attack technique Technique, and tactical stage Tactic. Node attributes include natural language description feature vectors, image region feature vectors, and cross-modal fusion vectors. Nodes are connected by edges such as "adopted technology" and "belonging to stage". Edge attributes include MITRE ATT&CK technology identifiers and confidence weights. Malicious code call graph : Code features obtained based on static analysis and dynamic sandbox results , build a function-level call relationship graph. Define PE file nodes to include hash values and compilation timestamps, API call nodes to include call sequence hashes, and process behavior nodes to include sandbox behavior labels. Use directed edges to represent function call paths, and edge weights are calculated based on call frequency and context relevance. System log time sequence diagram : Log features obtained based on the timestamp and process chain of log events Build a time-series association graph. Define the host node IP address, asset type, event node EventID, event level, process node PID, and execution path. Connect them through time-series edges such as "Generate Event" and "Call Process." Edge attributes include time intervals and event triggering frequencies.

[0160] Specifically, we use the heterogeneous graph attention network HAN and utilize meta-paths to generate node embeddings for each graph that retain the semantic and structural information of the nodes. Using HAN, we use meta-paths to generate embeddings:

[0161] (18)

[0162] in, It means the Nodes in the layer The embedding vector of is the meta-path, is the meta-path level attention coefficient, is the result of the transformation of the embedding vector of node i under the meta-path P, Is the activation function.

[0163] The meta path definition, for example, is Design the "attacker-technique-stage" meta-path to capture the attack implementation logic chain; Define the "file-API-process" meta-path to reveal the malicious code execution mode; A "host-event-process" meta-path is constructed to restore the trajectory of attack behavior. Node-level attention calculation, guided by each meta-path, aggregates neighborhood information through a multi-head attention mechanism, performing a weighted combination of node representations from different meta-paths to generate the final node embedding.

[0164] Through the meta-path level attention mechanism, hierarchical feature extraction is implemented for each graph structure, and the node embeddings under different meta-paths are weighted fused to obtain the final node embedding matrix:

[0165] (19)

[0166] (20)

[0167] in, It is a graph structure that is analyzed through graph neural network The node embedding matrix obtained after processing is It is a graph structure that is analyzed through graph neural network The node embedding matrix obtained after processing is It is a graph structure that is analyzed through graph neural network The node embedding matrix obtained after processing is It is a graph neural network model. is a matrix Dimensions, It is a graph structure The number of nodes, It is a graph structure The number of nodes, It is a graph structure The number of nodes, It is a graph structure The node set of It is a graph structure The node set of It is a graph structure The node set of , d is the dimension of the embedding vector.

[0168] Specifically, the optimal matching matrix is solved by the gradient descent optimization algorithm to achieve the fusion alignment of multi-source data, specifically:

[0169] calculate and The similarity matrix between:

[0170] (twenty one)

[0171] in: yes and The similarity matrix between Yes The embedding vector of node j in the graph, Yes The embedding vector of node k in the graph, sim is the similarity calculation function.

[0172] Define the alignment cost function: (twenty two)

[0173] in: yes and The matching matrix between them, j is node j, k is node k, yes and The similarity matrix between It is a graph structure The number of nodes, It is a graph structure The number of nodes, yes The optimization goal is to minimize . Solve the optimal matching matrix through the gradient descent optimization algorithm At the same time, according to the alignment result of the first step, and The fused embedding is :

[0174] (twenty three)

[0175] in: It is a graph structure and The node embedding matrix obtained after fusion, Represents element-wise multiplication. yes and The fused optimal matching matrix is used to indicate which nodes have been aligned. The graph structure is analyzed through the graph neural network The node embedding matrix obtained after processing is The graph structure is analyzed through the graph neural network The node embedding matrix obtained after processing.

[0176] calculate and after fusion The similarity matrix between:

[0177] (twenty four)

[0178] in: yes and The similarity matrix between Yes The embedding vector of node i in the graph, yes The fused node j is embedded; sim is the similarity calculation function.

[0179] Alignment optimization, define the alignment cost function:

[0180] (25)

[0181] in: yes and The alignment cost function between It is a graph structure The number of nodes, It is a graph structure The number of nodes, It is a graph structure The number of nodes, yes and The similarity matrix between yes and The optimization goal is to minimize . Solve the optimal matching matrix through the gradient descent optimization algorithm , the optimal matching matrix guide and The alignment relationship between nodes.

[0182] pass , we can determine which nodes are corresponding in different graph structures, thus achieving semantic and structural alignment of nodes. Specifically, and The node through Align and merge. The aligned nodes exist as unified nodes in G, preserving the semantic and structural information from different graph structures. The edges in G include not only the edges in the original graph structure, but also the edges obtained through The newly created edges after alignment represent the association relationship between different data sources.

[0183] Through S4's multi-source data fusion and alignment, a unified graph structure is ultimately obtained, integrating cross-modal threat intelligence, malicious code sample analysis, and system log information. Nodes from each data source are aligned, forming a semantically rich, clearly linked, and unified multi-source heterogeneous graph G. This graph preserves the detailed features of the original data and, through cross-source alignment, establishes the inherent connections between different data, providing a solid data foundation and structural framework for the subsequent construction of a knowledge graph. Based on the multi-source heterogeneous graph G, further knowledge extraction, integration, and semantic representation can be performed to achieve comprehensive and accurate modeling and efficient utilization of network threat intelligence, thereby further constructing a knowledge graph.

[0184] S5. Knowledge graph construction: Based on the threat intelligence analysis requirements and the multi-source heterogeneous graph G fused and aligned in step S4, a knowledge graph is constructed through graph model definition and design, knowledge extraction and fusion, storage and query optimization.

[0185] After completing the fusion and alignment of multi-source data, the knowledge graph construction phase begins. This phase aims to further integrate the fused and aligned multi-source data to form a structured and semantic knowledge graph, enabling comprehensive and accurate representation and efficient utilization of cyber threat intelligence.

[0186] First, the graph model is defined and designed. Based on the integrated data characteristics and threat intelligence analysis requirements, the ontology structure of the knowledge graph is determined, including entity types, relationship types, and attribute definitions of entities and relationships.

[0187] For example, an entity type might be: Attacker: The attacker entity might be "Alpha Hacking Team," with attributes such as name "Alpha Hacking Team," organization "Independent Hacker Group," and attack motive "Financial Gain." Attack Target: The attack target entity might be "E-commerce Platform," with attributes such as name "E-commerce Platform," type "Online Shopping," and sensitive information such as "User Order Information." Malicious Code: The malicious code entity might be "Omega Stealer," with attributes such as name "Omega Stealer," type "Information Stealing Trojan," and transmission method "Phishing Email Attachment." System Event: The system event entity might be "Unauthorized File Access Attempt," with attributes such as name "Unauthorized File Access Attempt," time "2023-11-2022:15:00," and affected file path " / var / www / html / data."

[0188] Relationship type: Attackers use attack techniques: For example, the relationship between "Alpha Hacking Team" and "SQL injection attack technique" has attributes such as exploitation method "Gaining database access by injecting malicious SQL code" and impact level "medium." Malicious code triggers system events: For example, the relationship between "Omega Stealer" and "Unauthorized file access attempt" has attributes such as trigger time "2023-11-20 22:15:00" and trigger frequency "3 times."

[0189] Next, knowledge extraction and fusion are performed. From the multi-source data after S4 fusion and alignment, the corresponding entities, relationships, and attribute information are extracted according to the graph schema definition. Entities such as attackers and attack targets in the CTI graph are directly mapped to the corresponding entity types in the knowledge graph, and their attribute information in the CTI graph is retained. Functions and API nodes in the malicious code call graph are converted into attributes or behavioral descriptions of malicious code-related entities in the knowledge graph based on their semantic meaning in malicious code analysis. Event nodes in the log sequence graph are positioned and associated in the knowledge graph based on characteristics such as event type and log source.

[0190] During the knowledge fusion process, we prioritize resolving potential entity alignment and attribute conflicts between different data sources. We leverage the cross-source node alignment relationships established in S4 to ensure that information from different data sources describing the same entity or event is accurately integrated into the corresponding nodes of the knowledge graph. For example, CTI data describes a "fileless attack" technique, whose attributes include the attack method "executing malicious code from memory" and the common attack target "Windows systems." Simultaneously, during analysis of malicious code samples, we discovered a code snippet characterized by "using PowerShell commands to execute malicious scripts in memory." Using the cross-source node alignment relationships established in S4, we identify this code snippet as describing the same attack behavior as the "fileless attack" technique in CTI. We then link the two in the knowledge graph, establish a direct edge, and integrate their attribute information. The resulting knowledge graph node contains not only the description of the "fileless attack," but also the specific code feature "PowerShell command execution" and the common attack target of both attacks, "Windows systems." This enriches the knowledge graph and improves the understanding and utilization of threat intelligence.

[0191] Then, we optimize the storage and query of the knowledge graph and choose the graph database Neo4j to efficiently store the constructed knowledge graph.

[0192] Example 2

[0193] This embodiment provides a multi-source cross-modal threat intelligence processing device, the device comprising:

[0194] The multi-source cross-modal data collection module is used to collect multi-source intelligence data sets from various channels, including natural language, images, logs, codes and other forms of data to obtain the initial multi-source intelligence data set. The data collection module collects multi-source cross-modal threat intelligence data from various data sources according to the preset strategy and transmits it to the data preprocessing module;

[0195] The data preprocessing module is used to preprocess the multi-source cross-modal intelligence data set and is responsible for cleaning, converting and extracting features from the collected multi-source cross-modal threat intelligence data;

[0196] The cross-modal feature fusion module is used to combine feature vectors from different modalities to form a unified feature representation. It uses intra-modal self-attention, cross-modal cross-attention mechanisms, and feature splicing and transformation techniques to achieve deep coupling of heterogeneous features between natural language and image CTI, generate a cross-modal fusion vector, and then pass it to the multi-source data alignment module.

[0197] The multi-source data alignment module uses graph-based cross-modal semantic alignment technology to align multi-source cross-modal threat intelligence data, malicious code analysis data, and system log data in a unified semantic space. An optimization algorithm is used to solve the optimal matching matrix, achieving fusion alignment of multi-source data. A three-dimensional association network is constructed, ultimately forming a multi-source heterogeneous graph G, which provides an integrated data foundation for the knowledge graph construction module.

[0198] The knowledge graph construction module is responsible for graph model definition and design. It selects the graph database Neo4j to efficiently store the constructed knowledge graph, extracts knowledge from the multi-source heterogeneous graph G and resolves conflicts, constructs a complete knowledge graph, and stores it in the graph database for subsequent query and analysis, thereby achieving comprehensive and accurate representation and efficient utilization of network threat intelligence.

[0199] Example 3

[0200] This embodiment further provides an electronic device, including:

[0201] at least one processor; and,

[0202] A memory storing instructions, which, when executed by the at least one processor, causes the at least one processor to execute the multi-source cross-modal threat intelligence processing method as described above.

[0203] In this embodiment, electronic devices may include, but are not limited to: personal computers, server computers, workstations, desktop computers, laptop computers, notebook computers, mobile computing devices, smart phones, tablet computers, cellular phones, personal digital assistants (PDAs), handheld devices, messaging devices, wearable computing devices, consumer electronic devices, and the like.

[0204] Example 4

[0205] This embodiment also provides a machine-readable storage medium storing executable instructions, which, when executed, enable the machine to execute the multi-source cross-modal threat intelligence processing method described above.

[0206] Specifically, a system or device equipped with a readable storage medium can be provided, on which software program codes that implement the functions of any of the above-mentioned embodiments are stored, and a computer or processor of the system or device can read and execute instructions stored in the readable storage medium.

[0207] In this case, the program code itself read from the machine-readable medium can implement the functions of any one of the above embodiments, and thus the machine-readable code and the machine-readable storage medium storing the machine-readable code constitute part of this specification.

[0208] Examples of readable storage media include floppy disks, hard disks, magneto-optical disks, optical disks (e.g., CD-ROMs, CD-Rs, CD-RWs, DVD-ROMs, DVD-RAMs, DVD-RWs, DVD-RWs), magnetic tapes, non-volatile memory cards, and ROMs. Alternatively, the program code may be downloaded from a server computer or a cloud via a communication network.

[0209] Obviously, the above embodiments of the present invention are merely examples for the purpose of clearly illustrating the technical solutions of the present invention, and are not intended to limit the specific implementation methods of the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the claims of the present invention shall be included within the scope of protection of the claims of the present invention.

Claims

1. A multi-source cross-modal threat intelligence processing method, characterized in that: The method comprises: S1. Data Collection: Build a multi-source, cross-modal, four-dimensional acquisition architecture to collect multi-source, cross-modal threat intelligence data including natural language, images, malicious code, and system logs; S2. Data preprocessing: Clean, transform and extract features of the multi-source cross-modal threat intelligence data collected in step S1 to obtain natural language feature vectors , image feature vector , malicious code feature vector , system log feature vector ; S3. Cross-modal feature alignment and fusion: S31, the natural language feature vector obtained in step S2 and image feature vector Perform feature alignment to obtain the aligned natural language feature vector and image feature vector ; S32. Aligned natural language feature vector and image feature vector Perform intra-modal self-attention calculation to obtain the natural language feature vectors after intra-modal self-attention calculation and image feature vector ;Will and Perform cross-modal attention fusion to obtain multi-head feature vectors ; S33, the multi-head feature vector obtained in step S32 After splicing, the final cross-modal fusion vector is obtained ; S4, multi-source data fusion alignment: using graph-based cross-modal semantic alignment technology, the code feature vector obtained after pre-processing in step S2 is aligned and the log eigenvector And the cross-modal fusion vector obtained after step S3 Align in a unified semantic space to obtain a multi-source heterogeneous graph G; S5 Knowledge Graph Construction: Based on the threat intelligence analysis requirements and the multi-source heterogeneous graph G fused and aligned in step S4, a knowledge graph is constructed through graph model definition and design, knowledge extraction and fusion, storage and query optimization.

2. The multi-source cross-modal threat intelligence processing method according to claim 1, characterized in that: The step S1 specifically includes: collecting natural language threat intelligence data using a distributed crawler cluster; collecting image threat intelligence data using an automated screenshot system; collecting malicious code threat intelligence data using a sandbox behavior capture device; and collecting system log threat intelligence data using a time series log processing pipeline.

3. The multi-source cross-modal threat intelligence processing method according to claim 1, characterized in that: Step S31: The natural language feature vector obtained in step S2 is and image feature vector Perform feature alignment, By linear transformation, the feature vectors of different dimensions are and Mapping to the target dimension is achieved as follows: The natural language feature vector Align to obtain the aligned natural language feature vector ,as follows: (1) in, is the aligned natural language feature vector, is the natural language feature vector, is the weight matrix, is the bias vector, Used to transform natural language feature vectors Map to target dimension; For the image feature vector , transform the dimension to the target dimension through a fully connected layer: (2) in, is the aligned image feature vector, represents the fully connected layer, is the image feature vector; Step S32 and Perform cross-modal attention fusion to obtain multi-head feature vectors , specifically: After the self-attention calculation within the modality, and Perform linear transformation to obtain query Q, key and numerical values vector; calculate the attention score using the query vector Q and the key vector The inner product of , and the attention weight is calculated by the softmax function; the attention score and the value vector Multiply them together to get the weighted representation of each feature vector under different attention heads; connect the weighted representations of each attention head and transform them linearly Generate the final fusion representation vector , the formula is as follows: Will and Transformed into query vector Query, key-value vector Key and value vector Value, where the query vector comes from natural language and the key-value vector comes from the image: (3) (4) (5) in, is a learnable parameter matrix, Q refers to the query vector Query, Refers to the key-value vector Key, is the value vector Value, It refers to the natural language feature vector obtained after the self-attention calculation within the modality, It refers to the image feature vector obtained after the intra-modal self-attention calculation; The weighted representation of each attention head is calculated as: (6) in, refers to the output of the i-th attention head, refers to the attention calculation function, is the query weight matrix of the i-th attention head, refers to the key weight matrix of the i-th attention head, refers to the value weight matrix of the i-th attention head, It refers to the natural language feature vector obtained after the self-attention calculation within the modality, It refers to the image feature vector obtained after the intra-modal self-attention calculation; Long-headed eigenvector The calculation is: (7) Where h is the number of attention heads, is the final linear transformation matrix, is the long-headed eigenvector, is the multi-head attention mechanism function, refers to the output of each independent attention head, It refers to the natural language feature vector obtained after the self-attention calculation within the modality, It refers to the image feature vector obtained after the intra-modal self-attention calculation; Step S33: The multi-head feature vector obtained in step S32 is After splicing, the final cross-modal fusion vector is obtained , specifically: Long-headed eigenvector and the natural language feature vector calculated by intra-modal self-attention Spliced together to form a new feature vector ; The concatenated feature vector Input into the LSTM model and obtain the intermediate vector representation through weight adjustment ; Then the middle vector and the image feature vector calculated by intra-modal self-attention Spliced together to form a new feature vector ; The concatenated feature vector Input into the LSTM model and obtain the final cross-modal fusion vector through weight adjustment ,in, (8) (9) (10) (11) is the concatenated feature vector, is the multi-headed eigenvector, is the natural language feature vector calculated by intra-modal self-attention, is the middle vector, represents a long memory network, is the middle vector and the image feature vector calculated by intra-modal self-attention The new feature vector formed by splicing, is the final cross-modal fusion vector.

4. The multi-source cross-modal threat intelligence processing method according to claim 3, characterized in that: The weight adjustment is achieved through a back-propagation algorithm and an optimizer.

5. The multi-source cross-modal threat intelligence processing method according to claim 1, characterized in that: Step S4 is specifically as follows: based on the malicious code feature vector obtained in step S2 and system log feature vector Build malicious code call graphs separately and system log timing diagram , based on the cross-modal fusion vector obtained in step S3 Constructing a cross-modal fusion network threat intelligence knowledge graph , use the graph neural network HAN to generate node embedding for each graph, solve the optimal matching matrix through the gradient descent optimization algorithm, realize the fusion alignment of multi-source data, and form a multi-source heterogeneous graph G.

6. The multi-source cross-modal threat intelligence processing method according to claim 5, characterized in that: The use of the graph neural network HAN to generate node embeddings for each graph is specifically as follows: Using the heterogeneous graph attention network HAN, meta-paths are used to generate hierarchical graph node embedding vectors that retain semantic features: (12) in, It means the Nodes in the layer The embedding vector of is the meta-path, is the meta-path level attention coefficient, It represents the node under the meta-path P The result of the transformation of the embedding vector is Is the activation function; Through the meta-path level attention mechanism, hierarchical feature extraction is implemented for each graph structure, and the node embeddings under different meta-paths are weighted fused to obtain the final node embedding matrix: (13) (14) in, It is a graph structure that is analyzed through graph neural network The node embedding matrix obtained after processing is It is a graph structure that is analyzed through graph neural network The node embedding matrix obtained after processing is It is a graph structure that is analyzed through graph neural network The node embedding matrix obtained after processing is It is a graph neural network model. is a matrix Dimensions, It is a graph structure The number of nodes, It is a graph structure The number of nodes, It is a graph structure The number of nodes, It is a graph structure The node set of It is a graph structure The node set of It is a graph structure The node set of , d is the dimension of the embedding vector; The optimal matching matrix is solved by the gradient descent optimization algorithm to achieve the fusion alignment of multi-source data, specifically: calculate and The similarity matrix between: (15) in: yes and The similarity matrix between Yes The embedding vector of node j in the graph, Yes The embedding vector of node k in the graph, sim is the similarity calculation function; Define the alignment cost function: (16) in: yes and The matching matrix between them, j is node j, k is node k, yes and The similarity matrix between It is a graph structure The number of nodes, It is a graph structure The number of nodes, yes The alignment cost function between them is optimized to minimize , solve the optimal matching matrix through the gradient descent optimization algorithm , according to the obtained optimal matching matrix ,Will and The fused embedding is : (17) in: It is a graph structure and The node embedding matrix obtained after fusion, represents element-wise multiplication, yes and The fused optimal matching matrix is used to indicate which nodes have been aligned. It is a graph structure that is analyzed through graph neural network The node embedding matrix obtained after processing is It is a graph structure that is analyzed through graph neural network The node embedding matrix obtained after processing; calculate and after fusion The similarity matrix between: (18) in: yes and The similarity matrix between Yes Nodes in the graph The embedding vector of yes The fused node j is embedded; sim is the similarity calculation function; Alignment optimization, define the alignment cost function: (19) in: yes and The alignment cost function between It is a graph structure The number of nodes, It is a graph structure The number of nodes, It is a graph structure The number of nodes, yes and The similarity matrix between yes and The matching matrix between them is optimized to minimize , solve the optimal matching matrix through the gradient descent optimization algorithm , the optimal matching matrix guide and The alignment relationship between nodes.

7. The multi-source cross-modal threat intelligence processing method according to claim 1, characterized in that: The graph model definition and design described in step S5 specifically include: determining the ontology structure of the knowledge graph, including entity types, relationship types, and attribute definitions of entities and relationships; the knowledge extraction and fusion includes extracting corresponding entities, relationships, and attribute information from the multi-source heterogeneous graph G after fusion and alignment in S4 according to the graph model definition; the storage specifically includes: selecting the graph database Neo4j to efficiently store the constructed knowledge graph.

8. A device for a multi-source cross-modal threat intelligence processing method, characterized in that: The device comprises: The multi-source cross-modal data collection module is used to collect multi-source intelligence data sets from various channels, including data in various forms such as natural language, images, logs, and codes to obtain the initial multi-source intelligence data set. The data collection module collects multi-source cross-modal threat intelligence data from various data sources according to preset strategies and transmits it to the data preprocessing module; The data preprocessing module is used to preprocess the multi-source cross-modal intelligence data set and is responsible for cleaning, converting and extracting features from the collected multi-source cross-modal threat intelligence data; The cross-modal feature fusion module combines feature vectors from different modalities to form a unified feature representation. It uses intra-modal self-attention, cross-modal cross-attention mechanisms, and feature splicing and transformation techniques to achieve deep coupling of heterogeneous features between natural language and image CTI, generating a cross-modal fusion vector that is then passed to the multi-source data alignment module. The multi-source data alignment module uses graph-based cross-modal semantic alignment technology to align multi-source cross-modal threat intelligence data, malicious code analysis data, and system log data in a unified semantic space. It uses an optimization algorithm to solve the optimal matching matrix, achieves fusion alignment of multi-source data, constructs a three-dimensional association network, and ultimately forms a multi-source heterogeneous graph G, providing an integrated data foundation for the knowledge graph construction module. The knowledge graph construction module is responsible for graph model definition and design. It selects the graph database Neo4j to efficiently store the constructed knowledge graph, extracts knowledge from the multi-source heterogeneous graph G and resolves conflicts, constructs a complete knowledge graph, and stores it in the graph database for subsequent query and analysis, thereby achieving comprehensive and accurate representation and efficient utilization of network threat intelligence.

9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and, A memory storing instructions, which, when executed by the at least one processor, causes the at least one processor to execute the multi-source cross-modal threat intelligence processing method according to any one of claims 1 to 7.

10. A machine-readable storage medium storing executable instructions, characterized in that: When the instruction is executed, the machine executes the multi-source cross-modal threat intelligence processing method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Threat intelligence knowledge graph processing method and device, equipment and storage medium

    CN116150392A

  • Intelligent cloud simulation multi-modal intelligence processing method, device and equipment and storage medium

    CN119829942A

  • Threat intelligence data processing method and computer readable storage medium

    CN117668244A

  • Method and device for constructing threat intelligence knowledge spectrogram and electronic equipment

    CN118410178A