A network intrusion detection method and system based on multimodal deep learning
Patent Information
- Application Number
- CN202611038472.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-13
- Publication Date
- 2026-09-11
AI Technical Summary
[0006]本发明的目的在于:针对目前存在的现有网络入侵检测过程中,其所构建的语义空间或知识图谱在模型训练完成后即被固定,无法随新型攻击的出现而在线扩展与更新,当遭遇训练阶段未曾见过的攻击类型时,模型因缺乏对应的语义表征而难以有效识别;新型攻击的发现、语义建模、知识更新与模型优化等环节彼此割裂,未能形成完整的闭环反馈机制,检测到无法识别的未知流量后,系统无法自动完成从新型攻击识别到检测能力更新的完整流程的问题
在本申请的方案中:
Smart Images

Figure CN122741201A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of network intrusion detection, and more specifically, to a network intrusion detection method and system based on multimodal deep learning. Background Technology
[0002] Network intrusion detection is a core technology in network security defense systems, identifying malicious activities by monitoring network traffic or host behavior in real time. As network attack methods become increasingly complex and diverse, detection methods relying solely on a single data source (such as analyzing only network traffic characteristics or detecting only host system call sequences) are insufficient to effectively address increasingly sophisticated intrusion behaviors. To improve detection accuracy and generalization capabilities, researchers in the field of network security have begun exploring technical approaches that integrate multiple heterogeneous data sources for joint analysis. This involves simultaneously utilizing data from multiple sources, such as network traffic, system logs, and user behavior, to uncover correlations and complementary features across data sources, aiming to construct a more comprehensive intrusion behavior representation model.
[0003] In existing technologies, joint analysis by fusing multiple heterogeneous data sources has become an important technical direction for improving intrusion detection performance. A common approach to cross-data source fusion is to use attention mechanisms to dynamically weight and aggregate features from different data sources to model the dependencies between cross-source features, thereby improving the expressive power of the fused features. Building on this, to further address the semantic gap between features from different data sources, some research has proposed using contrastive learning to map multi-source features to a unified semantic space for alignment, bringing related features from different data sources closer together in this space to enhance the model's ability to identify abnormal behavior. In addition, other solutions construct intrusion scenario knowledge graphs, mapping samples in the feature space to the semantic space, and using the known attack type information stored in the graph to achieve reasoning and identification of unseen intrusion categories.
[0004] However, the existing methods mentioned above still have the following shortcomings: First, the semantic space or knowledge graph they construct is fixed after the model is trained and cannot be expanded and updated online with the emergence of new attacks. When encountering attack types that have not been seen during the training phase, the model is unable to effectively identify them due to the lack of corresponding semantic representations. Second, the discovery of new attacks, semantic modeling, knowledge updating and model optimization are isolated from each other and fail to form a complete closed-loop feedback mechanism. After detecting unidentifiable unknown traffic, the system cannot automatically complete the complete process from identifying new attacks to updating detection capabilities.
[0005] Therefore, we have made improvements to this by proposing a network intrusion detection method and system based on multimodal deep learning. Summary of the Invention
[0006] The purpose of this invention is to address the problem that in existing network intrusion detection processes, the semantic space or knowledge graph constructed is fixed after model training and cannot be expanded and updated online with the emergence of new attacks. When encountering attack types not seen during the training phase, the model is unable to effectively identify them due to the lack of corresponding semantic representations. Furthermore, the discovery of new attacks, semantic modeling, knowledge updating, and model optimization are disconnected from each other, failing to form a complete closed-loop feedback mechanism. After detecting unidentifiable unknown traffic, the system cannot automatically complete the entire process from identifying new attacks to updating detection capabilities.
[0007] To achieve the above-mentioned objectives, this invention provides a network intrusion detection method and system based on multimodal deep learning to solve the aforementioned problems.
[0008] The application is as follows: A network intrusion detection method based on multimodal deep learning includes the following steps: Obtain network traffic data and extract the graph structure features and time-series statistical features of the network traffic; The graph structure features and the temporal statistical features are mapped to a shared semantic space, and matched with the semantic description features of multiple attack types stored in the shared semantic space. The attack type is determined based on the matching results. When unidentified unknown traffic is detected or an external trigger command is received, a dynamic evolution operation is performed: Generate a structured semantic description of the unknown traffic and encode it as a new semantic description vector; The newly added semantic description vector is subjected to consistency verification. After the verification is passed, it and its associated information are injected into the knowledge graph. Incremental parameter updates are performed on the semantic encoder to complete the online evolution of the detection model.
[0009] As a preferred technical solution of this application, before mapping the graph structure features and the temporal statistical features to the shared semantic space, the method further includes: Extract the text semantic view of the network traffic, which contains behavioral descriptions of various attacks; Multimodal contrastive learning is used to align the graph structure features, the temporal statistical features, and the text semantic view to the shared semantic space; The optimization objectives of the multimodal contrastive learning include: minimizing the distance between different modal features of the same sample, maximizing the distance between corresponding modal features of different samples, and ensuring that the trend of modal feature changes of the same flow remains consistent across different time windows.
[0010] As a preferred technical solution of this application, the constraint that the modal characteristic change trend of the same flow rate remains consistent under different time windows specifically includes: Calculate the first change of the graph structure features under different time windows, and the second change of the time series statistical features under the corresponding time windows; The optimization objective is to minimize the difference between the first change and the second change.
[0011] As a preferred technical solution of this application, the matching with the semantic description features of multiple types of attacks already stored in the shared semantic space includes: The similarity between the graph structure features and the temporal statistical features and the semantic description features of various attacks is calculated respectively, and the two similarities are weighted and fused to obtain a comprehensive similarity. The attack type is determined based on the relationship between the maximum value of the comprehensive similarity and the adaptive threshold; The adaptive threshold is dynamically determined based on the similarity distribution of confirmed normal traffic and attack traffic.
[0012] As a preferred technical solution of this application, the generation of the structured semantic description of the unknown traffic includes: Using the graph structure features and temporal statistical features of the unknown traffic as input conditions, the existing attack descriptions that are closest to the unknown traffic are retrieved from the knowledge graph as reference templates. Based on the reference template, a structured semantic description containing attack type labels, core behavioral features, and typical communication patterns is generated.
[0013] As a preferred technical solution of this application, the consistency verification of the newly added semantic description vector includes: Calculate the similarity distribution between the newly added semantic description vector and the existing attack semantic description vectors in the shared semantic space; Based on the similarity distribution, determine whether the newly added semantic description vector satisfies the local consistency condition and the global consistency condition; If the conditions are met, the newly added semantic description vector will be directly injected into the knowledge graph. If the condition is not met, the main positions of the existing vectors in the shared semantic space are fixed, and the positions of the newly added semantic description vector and its neighboring vectors are adjusted through iterative optimization until the consistency condition is met.
[0014] As a preferred technical solution of this application, the incremental parameter update of the semantic encoder includes: Only the encoder responsible for generating semantic description vectors is fine-tuned, while the encoder parameters used to encode graph structure features and temporal statistical features remain unchanged; The semantic encoder is updated incrementally by performing gradient updates for a limited number of rounds using labeled samples of the unknown traffic.
[0015] As a preferred technical solution of this application, the extraction of the graph structure features of the network traffic includes: The network traffic is parsed into heterogeneous graph data, where nodes in the heterogeneous graph data represent network communication entities and edges represent communication relationships between entities. The heterogeneous graph data is encoded using a graph neural network, and graph-level feature representations are obtained after node-level information aggregation and global pooling.
[0016] As a preferred technical solution of this application, the extraction of the time-series statistical features of the network traffic includes: Extract the serialized statistical indicators of the network traffic within a preset time window. The serialized statistical indicators include at least the packet size sequence and the packet arrival time interval sequence. The serialized statistical indicators are encoded by a time encoder to obtain a time-series feature representation.
[0017] A network intrusion detection system based on multimodal deep learning includes: The feature extraction module is used to acquire network traffic data and extract the graph structure features and time-series statistical features of the network traffic. The zero-shot inference module is used to map the graph structure features and the temporal statistical features to a shared semantic space, and match them with the semantic description features of multiple types of attacks stored in the shared semantic space, and determine the attack type based on the matching result. The dynamic evolution module is used to generate a structured semantic description of the unknown traffic and encode it into a new semantic description vector when an unidentifiable unknown traffic is detected or an external trigger command is received. The module performs a consistency check on the new semantic description vector, and after the check passes, it injects the vector and its associated information into the knowledge graph. The module also performs incremental parameter updates on the semantic encoder to complete the online evolution of the detection model.
[0018] Compared with the prior art, the beneficial effects of the present invention are as follows: In the scheme of this application: 1. When this method detects unidentifiable unknown traffic or receives an external trigger command, it performs a dynamic evolution operation: generating a structured semantic description of the unknown traffic and encoding it as a new semantic description vector, which is then injected into the knowledge graph after consistency verification, and the semantic encoder is updated incrementally. This operation allows the shared semantic space to be supplemented online with semantic representations of new attacks after model deployment, integrating the originally fragmented discovery, modeling, updating and optimization steps described in the background technology into continuous steps in the same process. This enables the model to acquire the ability to identify attack types that have not been seen during the training phase without full retraining.
[0019] 2. In the process of multimodal feature alignment, multimodal contrastive learning is used to align graph structure features, temporal statistical features, and text semantic views to a shared semantic space. The constraint is to minimize the difference between the first change of graph structure features under different time windows and the second change of temporal statistical features under the corresponding time windows. This constraint forces the model to establish a coupling relationship between the two modalities in temporal evolution during the optimization process, so that the feature offsets of the two modalities remain coordinated when network traffic behavior changes dynamically. This alleviates the problem of cross-modal feature misalignment caused by single-modal temporal fluctuations, thereby improving the reliability of subsequent matching judgments.
[0020] 3. In the attack type determination stage, the similarity between graph structure features and temporal statistical features and the semantic description features of various attack types is calculated separately. The two similarities are then weighted and fused, and an adaptive threshold is dynamically determined based on the similarity distribution of confirmed normal traffic and attack traffic to determine the attack type. Weighted fusion utilizes the differentiated representation capabilities of different modalities of attack behavior, enabling the matching results to comprehensively reflect information from both structured semantics and temporal behavioral patterns. The adaptive threshold is dynamically adjusted based on the measured values of similarity distribution in the actual network environment, allowing the determination boundary to adaptively shift with changes in traffic distribution. Compared to a fixed threshold, this reduces the probability of misjudgment caused by natural traffic fluctuations or attack variant distribution shifts.
[0021] 4. When performing consistency verification on newly added semantic description vectors, the similarity distribution between the vector and existing attack semantic description vectors in the shared semantic space is calculated. Based on this, it is determined whether both local and global consistency conditions are met. Direct injection is allowed only when both conditions are met; otherwise, the positions of the new vector and its neighboring vectors are adjusted through iterative optimization until the conditions are met. This verification mechanism provides quantitative constraints for maintaining the clarity of local inter-class boundaries and the stability of the global topological structure in the semantic space when injecting new knowledge online. This ensures that the expansion process of the knowledge graph maintains a quantifiable compatibility relationship with the existing semantic distribution, reducing the risk of destroying the existing class distinguishability due to the introduction of new knowledge.
[0022] 5. After dynamic evolution is completed, only the parameters of the semantic encoder are fine-tuned, while the encoder parameters used to encode graph structure features and temporal statistical features remain fixed. Gradient updates are performed in a limited number of rounds using labeled samples with unknown traffic. This update strategy limits the parameter adjustment range to the semantic encoder part, while the parameters of the underlying multimodal feature extraction module remain unchanged during the update process. This approach maintains the detection performance against known attacks while introducing new attack recognition capabilities. Furthermore, because the number of updated parameters is less than that of a full update, the required computational resources and labeled sample size are reduced, enabling online evolution to be completed under limited resource conditions. Attached Figure Description
[0023] Figure 1 A flowchart illustrating the network intrusion detection method based on multimodal deep learning provided in this application; Figure 2 A schematic diagram of the module architecture of the network intrusion detection system based on multimodal deep learning provided in this application. Detailed Implementation
[0024] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0025] The present invention will now be described in further detail with reference to the accompanying drawings.
[0026] Example: Refer to Figure 1 This invention provides a network intrusion detection method based on multimodal deep learning. Its core lies in constructing an online scalable shared semantic space, mapping the graph structure features and temporal statistical features of network traffic to this space, and matching them with the semantic description features of various attacks to determine the attack type. When unknown traffic is detected, a complete closed loop of structured semantic description generation, consistency verification, knowledge graph injection, and incremental update of semantic encoder is automatically triggered, thereby enabling the detection model to continuously evolve online without the need for full retraining.
[0027] Specifically, the overall process of this method is as follows: First, network traffic data is acquired and multimodal features are extracted; then, multimodal contrastive learning is used to align the features of each modality to a shared semantic space; then, zero-shot inference is performed to identify known attacks or detect unknown traffic; when unknown traffic is detected, a dynamic evolution operation is triggered to generate its structured semantic description and encode it as a new semantic description vector, which is then injected into the knowledge graph after consistency verification; finally, incremental parameter updates are performed on the semantic encoder to complete the online evolution of the model; each stage is described in detail below.
[0028] (I) Multimodal Feature Extraction First, network traffic data is acquired: raw network data packets can be captured in real time using traffic acquisition tools deployed at key network nodes (such as core switches or gateways). After data cleaning and flow reconstruction, traffic samples to be detected are generated. The data cleaning includes removing invalid data packets, deduplication, and protocol parsing preprocessing; flow reconstruction includes aggregating data packets into a complete session flow based on the five-tuple (source IP address, destination IP address, source port number, destination port number, and protocol type). The above acquisition and preprocessing operations are well-known technologies in this field, and their specific implementation methods will not be elaborated further.
[0029] Next, the graph structure features of the network traffic are extracted.
[0030] Preferably, network traffic is parsed into heterogeneous graph data; the nodes of the heterogeneous graph data represent network communication entities (including but not limited to IP addresses, port numbers, and protocol types), and the edges represent the communication relationships between entities (including but not limited to TCP connections or UDP sessions). Specifically, a time window (preferably 10 to 120 seconds, which can be dynamically adjusted according to network traffic density) can be set. Within this window, an edge is established between each pair of communicating entities. The edge attributes can include one or more of the following statistics: number of communications, total number of bytes, average packet length, etc. Based on this, the heterogeneous graph data is encoded using a graph neural network. After node-level information aggregation and global pooling, a graph-level feature representation is obtained. Those skilled in the art can choose specific architectures such as graph attention networks, graph convolutional networks, or GraphSAGE to implement graph neural network encoding according to actual needs. The number of layers in the graph neural network is preferably 2 to 4, and the hidden layer dimension is preferably 128 to 512 dimensions. The specific values can be configured according to the scale of the training data and computing resources. This embodiment does not impose a unique limitation.
[0031] Then, the time-series statistical features of the network traffic are extracted.
[0032] Preferably, serialized statistical indicators of network traffic within a preset time window are extracted. These serialized statistical indicators include at least a packet size sequence and a packet arrival time interval sequence. Specifically, for each time window (which can be consistent with the window for the aforementioned graph structure feature extraction or can be configured independently), the size (number of bytes) of each packet within the window and the arrival time interval (milliseconds) between adjacent packets are recorded, forming two time-series sequences. Subsequently, the serialized statistical indicators are encoded using a time-series encoder to obtain a time-series feature representation. Those skilled in the art can select a Long Short-Term Memory network, a gated recurrent unit, or a time-series convolutional network as the time-series encoder according to actual needs. The time-series encoder encodes the variable-length input sequence into a fixed-dimensional vector by capturing long-range dependencies in the sequence. The hidden layer dimension of the time-series encoder is preferably consistent with the graph-level feature vector dimension to facilitate subsequent alignment operations.
[0033] Furthermore, before mapping the graph structure features and the temporal statistical features to the shared semantic space, a textual semantic view of the network traffic is extracted. This textual semantic view contains behavioral descriptions of various attacks, the sources of which include, but are not limited to, publicly available cybersecurity threat intelligence databases (such as CVE vulnerability descriptions, ATT&CK attack tactical descriptions) and text log records of historical attack events. Preferably, a pre-trained language model (such as BERT or its lightweight variant) can be used to encode the aforementioned textual information and extract textual semantic feature vectors. The extraction of this textual semantic view is performed during the initial training phase and the knowledge graph construction phase. During the online inference phase, it can be selectively executed or omitted depending on computing resources; this embodiment does not impose such limitations.
[0034] (ii) Multimodal contrastive learning alignment Before mapping the features of each modality to the shared semantic space, multimodal contrastive learning is first used to align the graph structure features, the temporal statistical features, and the text semantic view to the shared semantic space. Specifically, the features of the above three modalities are projected to a shared semantic space with a uniform dimension through their respective linear or nonlinear mapping networks (the projection dimension is preferably 128 to 512 dimensions, consistent with the dimension of each modality feature).
[0035] The optimization objectives of the multimodal contrastive learning include the following three aspects: First, minimize the distance between different modal features of the same sample; for the same traffic sample, its graph structure features, temporal statistical features and text semantic views should be close to each other after being projected onto a shared semantic space, so that the cross-modal representations of the same sample are clustered in the space; this goal aims to eliminate the semantic gap between different data sources and make multi-source features comparable in a unified space.
[0036] Second, maximize the distance between modal features corresponding to different samples; for different traffic samples, features of the same modality should be spaced out in the shared semantic space, so that the representations of the same modality of different samples are dispersed in the space. This goal aims to ensure that the semantic space has sufficient discriminative power and avoid confusion between samples of different categories in the space.
[0037] Third, it constrains the modal characteristics of the same flow rate to maintain a consistent trend across different time windows. Specifically, let the same flow rate be in the first... The graphical feature projection vectors for each time window are: In the The graphical feature projection vectors for each time window are: The flow rate was in the first... The temporal feature projection vector of each time window is In the The temporal feature projection vector of each time window is The first change in the graph's structural features within adjacent time windows is defined as: The second change of a time series statistical feature under corresponding adjacent time windows is defined as: With the first change Compared with the second change Minimizing the difference between them is the optimization objective. Even if the evolution trends of the two modes remain synchronized over time, the optimization objective function is expressed as: This constraint ensures that the evolutionary trend of graph structure and the evolutionary trend of time-series statistics remain coordinated in the shared semantic space, thereby enhancing the model's consistency in modeling the dynamic changes of traffic behavior over time.
[0038] The three optimization objectives mentioned above can be integrated into a joint optimization problem through a weighted combination. The weight coefficients of each objective are all positive real numbers, and their specific values can be determined by searching for the optimal combination on the validation set. In this embodiment, they are not limited to any single fixed value. After the multimodal contrastive learning is trained, both graph structure features and temporal statistical features can be mapped to the shared semantic space through their respective projection networks. The mapped vectors are adjacent to the text semantic view of the same sample in the space, laying the foundation for subsequent zero-shot inference.
[0039] Through the above multimodal contrastive learning alignment operation, graph structure features, temporal statistical features, and text semantic view are established with a unified representation benchmark in a shared semantic space. In particular, the temporal consistency constraint forces the model to establish a coupling relationship between the two modalities in temporal evolution during the optimization process, so that when network traffic behavior changes dynamically, the feature offsets of the two modalities remain coordinated, thereby alleviating the problem of cross-modal feature misalignment caused by single-modal temporal fluctuations, and thus improving the reliability of subsequent matching judgments.
[0040] (III) Zero-sample inference judgment After completing the multimodal contrastive learning alignment, the graph structure features and the temporal statistical features are mapped to the shared semantic space and matched with the semantic description features of multiple attack types stored in the shared semantic space. The attack type is determined based on the matching results.
[0041] Those skilled in the art should understand that the "semantic description features of multiple attack types stored in the shared semantic space" originate from the vector set stored after encoding the text descriptions of known attack types through a text semantic view encoder during the initialization phase. This initialization process can employ the same pre-trained language model as the aforementioned text semantic view extraction to encode the publicly available text descriptions of various known attacks (such as DDoS attacks, port scanning, brute-force attacks, SQL injection, and other common attack types), and store the encoding results as a semantic description feature library. This feature library is constructed during the initial deployment of the system and can be continuously expanded during subsequent dynamic evolution.
[0042] The matching process specifically includes: The first similarity between the graph structure features and the semantic description features of various attacks, and the second similarity between the temporal statistical features and the semantic description features of various attacks are calculated separately. The two similarities are then weighted and fused to obtain a comprehensive similarity. The similarity can be measured using well-known vector similarity measurement methods in the art, such as cosine similarity or the negative value of Euclidean distance.
[0043] Suppose that the shared semantic space stores Semantic description feature vector of class attack , , ..., For the graph feature projection vector of the current traffic sample and temporal feature projection vector , and the first Semantic description feature vector of class attack The overall similarity between them is determined by the following formula: in, and These are the fusion weights for graph structure features and time series statistical features, respectively, all of which are non-negative real numbers and satisfy the following conditions: + = 1. Preferably, the fusion weights can be configured with equal weights (i.e., = = 0.5); or dynamically allocated based on the independent detection performance of each modality on the historical validation set, for example, using the normalized value of the independent detection accuracy of each modality as the fusion weight; those skilled in the art can choose an appropriate weight determination method according to the actual application scenario, and this embodiment does not limit it to the only value method.
[0044] Subsequently, the attack type is determined based on the relationship between the maximum value of the comprehensive similarity and the adaptive threshold. Let the maximum value of the comprehensive similarity be: like If so, the traffic is identified as a known attack type, specifically: The corresponding attack type, i.e. ;like If so, the traffic is classified as unidentifiable unknown traffic.
[0045] Wherein, the adaptive threshold The similarity distribution of confirmed normal traffic and attack traffic is dynamically determined. Specifically, the system maintains a fixed-length sliding window (preferably with 1000 to 5000 samples, adjustable based on storage resources and traffic change frequency), recording the comprehensive similarity value and category label of confirmed samples within the window; let the mean comprehensive similarity of normal samples within the window be... Standard deviation is The average overall similarity of the known attack samples is Standard deviation is Then the adaptive threshold Determined by the following formula: in, This is an adjustable offset coefficient, preferably ranging from -1.0) to 1.0); when When = 0, the threshold is the arithmetic mean of the similarity values of the two classes of samples; when When the threshold is > 0, it shifts towards the normal side, making the judgment criteria more stringent and helping to reduce the false alarm rate; when When the threshold is less than 0, it shifts towards the attack side, making the judgment conditions more lenient and helping to reduce the false negative rate. The specific value can be determined by optimizing the validation set based on the tolerance of false positive rate and false negative rate in the actual application scenario; the sliding window adopts a first-in-first-out update strategy to ensure that the threshold can be adaptively adjusted with the dynamic changes of the network environment.
[0046] The aforementioned synergistic mechanism of weighted fusion and adaptive threshold determination has clear technical benefits: weighted fusion utilizes the differentiated representation capabilities of different modalities for attack behavior, enabling the matching results to comprehensively reflect information from both graph structure semantics and temporal behavior patterns; the adaptive threshold is dynamically adjusted based on the measured values of similarity distribution in the actual network environment, allowing the determination boundary to adaptively move with changes in traffic distribution; compared to a fixed threshold, the above mechanism reduces the probability of misjudgment caused by natural traffic fluctuations or attack variant distribution shifts, thus improving the robustness of detection.
[0047] (iv) Dynamic evolution of unknown flow When unidentified unknown traffic is detected or an external trigger command is received, the following dynamic evolution operation is performed: Step 1: Generate semantic description and encode it Generate a structured semantic description of the unknown traffic and encode it as a new semantic description vector.
[0048] Preferably, the graph structure features and temporal statistical features of the unknown traffic are used as input conditions, and the existing attack description closest to the unknown traffic is retrieved from the knowledge graph as a reference template. Specifically, the graph feature vector of the unknown traffic is projected... With time series feature projection vector By concatenating or averaging, the query vector is obtained. Then, search for related terms in the semantic vector index of the knowledge graph. The one or more closest existing attack description vectors are used as reference templates. The similarity measurement method used in the retrieval is consistent with that used in the aforementioned matching stage.
[0049] Based on the reference template, a structured semantic description containing attack type labels, core behavioral characteristics, and typical communication patterns is generated. The structured semantic description preferably adopts the standard knowledge graph triple format, i.e., (head entity, relation, tail entity), such as ("Unknown attack type 1", "has characteristics", "high-frequency short packet communication"), ("Unknown attack type 1", "target port", "445"), etc. Specifically, the generation method is as follows: using the triple structure of the retrieved reference template as the skeleton, the actual statistical characteristics of the unknown traffic (including but not limited to average packet length, packet rate, port distribution, uplink / downlink traffic ratio, etc.) are filled into the corresponding positions of the triples to form a new set of triples; those skilled in the art can also use other equivalent structured data formats (such as JSON structure or attribute graphs), as long as the description content includes the three core elements of attack type identifier, behavioral characteristics, and communication pattern.
[0050] Subsequently, the generated set of triples is encoded into a new semantic description vector using a semantic encoder. The dimension of this vector is the same as the semantic description vectors already existing in the shared semantic space. The dimensions remain consistent.
[0051] The aforementioned dynamic evolution operation is initiated when unidentifiable unknown traffic is detected or an external trigger command is received. Its execution includes a complete sequence from the generation and encoding of structured semantic descriptions to consistency verification, knowledge graph injection, and incremental updates of the semantic encoder. This operation enables the shared semantic space to be supplemented online with semantic representations of new attacks after model deployment. It integrates the originally fragmented "discovery-modeling-update-optimization" links in the background technology into continuous steps in the same process, so that when the model encounters an attack type not seen during the training phase, it can obtain the ability to identify the attack type through the above operation without full retraining.
[0052] Step 2: Consistency Verification For the newly added semantic description vector Consistency checks are performed to ensure that new knowledge can coexist harmoniously with the existing knowledge system, avoiding the introduction of noise or disruption of the topological structure of the existing semantic space.
[0053] Specifically, the newly added semantic description vector is calculated. With the existing set of attack semantic description vectors in the shared semantic space The similarity distribution, i.e., calculating the similarity distribution separately. For all ( = 1,2, ..., The value of ).
[0054] Based on the similarity distribution, it is determined whether the newly added semantic description vector satisfies the local consistency condition and the global consistency condition. The present invention provides the following preferred method for determining the above two consistency conditions: Local consistency condition: The average similarity between a newly added vector and its nearest neighboring existing vectors in the shared semantic space should not be lower than a set lower limit. Let... express In the existing vector set The set consisting of the nearest neighbor vectors, The integer is preferably between 5 and 20. The local consistency condition is then expressed as: in, This is the lower limit threshold for local similarity, preferably a real value between 0.3 and 0.6, with the specific value depending on the similarity distribution among existing vectors in the training set. The physical meaning of this condition is that new knowledge should not deviate from the existing semantic manifold in isolation, but should maintain semantic coherence with surrounding knowledge.
[0055] Global consistency condition: The introduction of a new vector should not disrupt the relative distance relationships between existing vectors in the shared semantic space. The specific determination method is as follows: when introducing... Previously, the cosine similarity values between all existing vectors were calculated to form a similarity matrix. In the introduction After that, Similarity values with all existing vectors are appended to Then, an augmented similarity matrix is formed. Then calculate the original Between the vectors and Spearman rank correlation coefficient corresponding to the similarity values The global consistency condition is expressed as: in, The threshold for global consistency is preferably a real value between 0.85 and 0.95. The physical meaning of this condition is that the injection of new knowledge should not distort the relative semantic relationships between existing knowledge, that is, it should not disrupt the topological ordering structure of the original semantic space, thereby maintaining the overall stability of the semantic space.
[0056] If both the local consistency condition and the global consistency condition are satisfied, then the newly added semantic description vector will be... The knowledge graph is directly injected.
[0057] If the consistency condition is not met, the positions of the existing vectors in the shared semantic space are fixed, and the positions of the newly added semantic description vector and its neighboring vectors are adjusted through iterative optimization until the consistency condition is met. Specifically, this iterative optimization, under the premise of satisfying the consistency constraint, aims to minimize the total displacement of the newly added vector and its neighboring vectors, that is, to move the vector positions as little as possible to achieve the consistency requirement. Preferably, this optimization can be solved using the gradient descent method, with the number of iterations preferably between 10 and 100, and an upper limit is set on the step size to ensure that the computational cost is controllable and does not compromise the spatial stability.
[0058] After the verification is passed (whether the conditions are met directly or after iterative adjustments), the newly added semantic description vector will be... The knowledge graph is then injected with its associated information. This associated information includes at least: a new attack type label (which can be automatically generated according to a preset naming rule, such as "unknown attack type X"), a generation timestamp, a set of triples representing structured semantic descriptions, and a confidence score (this score can be determined based on the average similarity between the new vector and its nearest neighboring existing vectors, i.e., ...). .
[0059] The aforementioned consistency verification mechanism provides quantitative constraints for maintaining the clarity of local inter-class boundaries and the stability of the global topological structure in the semantic space when injecting new knowledge online. This ensures that the expansion process of the knowledge graph maintains a quantifiable compatibility with the existing semantic distribution, reducing the risk of destroying the existing class distinguishability due to the introduction of new knowledge.
[0060] (v) Online evolution of the model Incremental parameter updates are performed on the semantic encoder to complete the online evolution of the detection model. The semantic encoder refers to a neural network module that encodes structured semantic descriptions (set of triples) into semantic description vectors. Its specific architecture can be a Transformer encoder or an encoder based on a multilayer perceptron, but this embodiment does not limit it to a single architecture.
[0061] This phase employs a "freeze-fine-tuning" strategy: only the encoder responsible for generating semantic description vectors (i.e., the semantic encoder) undergoes parameter fine-tuning, while the encoder parameters used to encode graph structure features and temporal statistical features remain unchanged. This strategy aims to protect the learned general traffic awareness capabilities and prevent catastrophic forgetting during the learning of new knowledge.
[0062] Specifically, a limited number of gradient updates are performed using labeled samples of the unknown traffic to complete the incremental parameter update of the semantic encoder. The labeled samples refer to data marked after expert confirmation of the unknown traffic that has triggered the generation of structured semantic descriptions. Preferably, labels confirmed by human experts are used as supervision signals; a semi-automatic method can also be used (e.g., security analysts confirm or correct the semantic descriptions generated by the system on the system interface before submitting). The optimization objective of the incremental update is to make the encoding result of the semantic encoder for new samples as close as possible to its target vector in the shared semantic space. The target vector can be determined by the optimal similarity matching result confirmed by human verification or directly specified by experts. This optimization process is implemented using gradient descent algorithms and their variants (such as the Adam optimizer), which are well-known in the art. The specific parameter update rules are also well-known technologies and will not be elaborated here.
[0063] The "limited number of rounds" is preferably 1 to 5 training rounds, or the stopping condition is dynamically determined based on the loss changes on the monitoring set: the update is terminated early when the loss on the monitoring set no longer decreases within several consecutive evaluation steps. During the update process, the learning rate is set to a fraction of the initial training learning rate. to This is to ensure the smoothness of parameter updates and avoid drastic disturbances to the learned knowledge.
[0064] The incremental update strategy described above limits parameter adjustment to the semantic encoder portion, while the parameters of the underlying multimodal feature extraction module remain unchanged during the update process. This maintains the detection performance against known attacks while introducing new attack recognition capabilities. Furthermore, because the number of updated parameters is much smaller than that of a full update, it requires fewer computational resources and labeled samples, enabling online evolution to be completed under resource-constrained conditions.
[0065] See Figure 2 As shown, the present invention also provides a network intrusion detection system based on multimodal deep learning, which includes a feature extraction module, a zero-shot inference module, and a dynamic evolution module.
[0066] The feature extraction module is used to acquire network traffic data and extract the graph structure features and temporal statistical features of the network traffic. Its specific operation is consistent with the multimodal feature extraction stage described earlier: this module contains a graph feature extraction submodule and a temporal feature extraction submodule, responsible for constructing heterogeneous graph data and encoding it using a graph neural network, and extracting and encoding the temporal sequence using a temporal encoder, respectively. This module outputs fixed-dimensional graph-level feature vectors and temporal feature vectors for use by subsequent modules.
[0067] The zero-shot inference module maps the graph structure features and the temporal statistical features to a shared semantic space, and matches them with the semantic description features of multiple attack types stored in the shared semantic space to determine the attack type based on the matching results. This module internally includes a projection network unit, a similarity calculation unit, an adaptive threshold calculation unit, and a determination unit. This module receives the feature vector output by the feature extraction module and outputs the determination result (a known attack type label or an "unknown traffic" identifier) and the corresponding comprehensive similarity value.
[0068] The dynamic evolution module is used to generate a structured semantic description of the unknown traffic and encode it into a new semantic description vector when an unidentifiable unknown traffic is detected or an external trigger command is received. The module performs a consistency check on the new semantic description vector, and if the check passes, injects it and its associated information into the knowledge graph. It also performs incremental parameter updates on the semantic encoder to complete the online evolution of the detection model. This module internally includes a semantic description generation subunit (including reference template retrieval and triple generation functions), a semantic encoder, a consistency check subunit (including similarity distribution calculation and local / global condition determination functions), a knowledge graph management subunit (including vector injection interface and adjacency update functions), and an incremental update subunit (including gradient calculation and parameter freezing control functions). This module receives the "unknown traffic" identifier and corresponding features output by the zero-shot inference module, completes the evolution, and sends the updated knowledge graph index back to the zero-shot inference module, enabling subsequent inference to identify the newly injected attack type.
[0069] Preferably, the above modules can be deployed using a loosely coupled software architecture, and the modules interact with each other through standardized data interfaces. Those skilled in the art can flexibly choose the deployment method according to the actual network scale and detection latency requirements, for example, it can be deployed on a general server, cloud computing platform or edge computing device, and this embodiment is not limited thereto.
[0070] In this invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," "linking," and "fixing," etc., should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral part; they can refer to a mechanical connection, an electrical connection, or a connection that allows communication between them; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication of two components or the interaction between two components, unless otherwise explicitly limited. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.
[0071] Obviously, the embodiments described above are merely some embodiments of the present invention, not all embodiments. The accompanying drawings show preferred embodiments of the present invention, but do not limit the scope of the patent. The present invention can be implemented in many different forms; rather, these embodiments are provided to provide a more thorough and complete understanding of the disclosure of the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing specific embodiments, or make equivalent substitutions for some of the technical features. Any equivalent structures made using the content of this specification and drawings, directly or indirectly applied to other related technical fields, are similarly within the scope of patent protection of this invention.
Claims
1. A network intrusion detection method based on multimodal deep learning, characterized in that, include: Obtain network traffic data and extract the graph structure features and time-series statistical features of the network traffic; The graph structure features and the temporal statistical features are mapped to a shared semantic space, and matched with the semantic description features of multiple attack types stored in the shared semantic space. The attack type is determined based on the matching results. When unidentified unknown traffic is detected or an external trigger command is received, a dynamic evolution operation is performed: Generate a structured semantic description of the unknown traffic and encode it as a new semantic description vector; The newly added semantic description vector is subjected to consistency verification. After the verification is passed, it and its associated information are injected into the knowledge graph. Incremental parameter updates are performed on the semantic encoder to complete the online evolution of the detection model.
2. The network intrusion detection method based on multimodal deep learning according to claim 1, characterized in that, Before mapping the graph structure features and the temporal statistical features to the shared semantic space, the method further includes: Extract the text semantic view of the network traffic, which contains behavioral descriptions of various attacks; Multimodal contrastive learning is used to align the graph structure features, the temporal statistical features, and the text semantic view to the shared semantic space; The optimization objectives of the multimodal contrastive learning include: minimizing the distance between different modal features of the same sample, maximizing the distance between corresponding modal features of different samples, and ensuring that the trend of modal feature changes of the same flow remains consistent across different time windows.
3. The network intrusion detection method based on multimodal deep learning according to claim 2, characterized in that, The requirement to maintain consistent modal characteristic trends for the same flow rate across different time windows specifically includes: Calculate the first change of the graph structure features under different time windows, and the second change of the time series statistical features under the corresponding time windows; The optimization objective is to minimize the difference between the first change and the second change.
4. The network intrusion detection method based on multimodal deep learning according to claim 3, characterized in that, The matching with the semantic description features of multiple attack types already stored in the shared semantic space includes: The similarity between the graph structure features and the temporal statistical features and the semantic description features of various attacks is calculated respectively, and the two similarities are weighted and fused to obtain a comprehensive similarity. The attack type is determined based on the relationship between the maximum value of the comprehensive similarity and the adaptive threshold; The adaptive threshold is dynamically determined based on the similarity distribution of confirmed normal traffic and attack traffic.
5. The network intrusion detection method based on multimodal deep learning according to claim 4, characterized in that, The structured semantic description for generating the unknown traffic includes: Using the graph structure features and temporal statistical features of the unknown traffic as input conditions, the existing attack descriptions that are closest to the unknown traffic are retrieved from the knowledge graph as reference templates. Based on the reference template, a structured semantic description containing attack type labels, core behavioral features, and typical communication patterns is generated.
6. The network intrusion detection method based on multimodal deep learning according to claim 5, characterized in that, The consistency check of the newly added semantic description vector includes: Calculate the similarity distribution between the newly added semantic description vector and the existing attack semantic description vectors in the shared semantic space; Based on the similarity distribution, determine whether the newly added semantic description vector satisfies the local consistency condition and the global consistency condition; If the conditions are met, the newly added semantic description vector will be directly injected into the knowledge graph. If the condition is not met, the main positions of the existing vectors in the shared semantic space are fixed, and the positions of the newly added semantic description vector and its neighboring vectors are adjusted through iterative optimization until the consistency condition is met.
7. The network intrusion detection method based on multimodal deep learning according to claim 6, characterized in that, The incremental parameter update of the semantic encoder includes: Only the encoder responsible for generating semantic description vectors is fine-tuned, while the encoder parameters used to encode graph structure features and temporal statistical features remain unchanged; The semantic encoder is updated incrementally by performing gradient updates for a limited number of rounds using labeled samples of the unknown traffic.
8. The network intrusion detection method based on multimodal deep learning according to claim 7, characterized in that, The extraction of graph structure features from the network traffic includes: The network traffic is parsed into heterogeneous graph data, where nodes in the heterogeneous graph data represent network communication entities and edges represent communication relationships between entities. The heterogeneous graph data is encoded using a graph neural network, and graph-level feature representations are obtained after node-level information aggregation and global pooling.
9. The network intrusion detection method based on multimodal deep learning according to claim 8, characterized in that, The extraction of the time-series statistical features of the network traffic includes: Extract the serialized statistical indicators of the network traffic within a preset time window. The serialized statistical indicators include at least the packet size sequence and the packet arrival time interval sequence. The serialized statistical indicators are encoded by a time encoder to obtain a time-series feature representation.
10. A network intrusion detection system based on multimodal deep learning, used to implement the network intrusion detection method based on multimodal deep learning as described in any one of claims 1-9, characterized in that, include: The feature extraction module is used to acquire network traffic data and extract the graph structure features and time-series statistical features of the network traffic. The zero-shot inference module is used to map the graph structure features and the temporal statistical features to a shared semantic space, and match them with the semantic description features of multiple types of attacks stored in the shared semantic space, and determine the attack type based on the matching result. The dynamic evolution module is used to generate a structured semantic description of the unknown traffic and encode it into a new semantic description vector when an unidentifiable unknown traffic is detected or an external trigger command is received. The module performs a consistency check on the new semantic description vector, and after the check passes, it injects the vector and its associated information into the knowledge graph. The module also performs incremental parameter updates on the semantic encoder to complete the online evolution of the detection model.