Network intrusion detection method and system based on multi-modal deep learning

By constructing a time-continuous multi-track dynamic sensing track set and a modal attention allocation mechanism, the problem of weakened modal information and wasted resources in existing network intrusion detection is solved, achieving efficient detection and improved robustness of network intrusion behavior.

CN120498883BActive Publication Date: 2026-02-27聊城大学东昌学院 +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510909760.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-02
Publication Date
2026-02-27
Estimated Expiration
2045-07-02

AI Technical Summary

Technical Problem

Existing network intrusion detection methods lack the ability to dynamically adjust the weights of perception intensity and state changes of different modalities within different time windows, resulting in the weakening or ignoring of key modal information, affecting the accuracy and robustness of detection. Furthermore, multimodal deep learning models struggle to fully exploit cross-modal correlations, leading to wasted computational resources and performance bottlenecks.

Method used

By constructing a time-continuous multi-track dynamic sensing track set, integrating a modal attention allocation mechanism, dynamically allocating modal parameters, establishing cross-modal dependencies, adopting a cross-modal residual connection mechanism, constructing a multimodal behavior aggregation map, extracting abnormal behavior map structure, and performing modal source traceability labeling, the data is finally input into a deep joint discriminator for detection.

Benefits of technology

It significantly improves the accuracy of modal recognition and judgment, adapts to complex dynamic environments, solves the problem of uncaptured modal linkage, improves the availability and scalability of the system, and achieves efficient detection of network intrusion behavior.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120498883B_ABST
    Figure CN120498883B_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of Internet security, and discloses a network intrusion detection method and system based on multi-modal deep learning, which comprises collecting and modal attribution of multi-source data from different channels, constructing a time-continuous multi-track dynamic perception track set, integrating a modal attention distribution mechanism, dynamically distributing modal parameters according to the modal energy compound of each type of modal under different perception tracks, and outputting an initial index tensor after multi-track modal fusion; receiving the initial index tensor, establishing a cross-modal dependency relationship by constructing a bidirectional response tensor between the modal and the track; constructing a semantic trigger matrix based on the semantic labels in the cross-modal dependency relationship, adopting a cross-modal residual connection mechanism, preserving low-order modal coupling features through a parallel residual path, and outputting a modal linkage node; the recognition ability of the detection on complex intrusion behaviors is improved, and more comprehensive and accurate detection on complex network attack behaviors is realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of Internet security, more particularly, the present application relates to a network intrusion detection method and system based on multi-modal deep learning. BACKGROUND

[0002] The patent with the patent publication number CN116886393A discloses a network security intrusion detection method and system, comprising the following steps: step S1, collecting internet traffic data; step S2, inputting the internet traffic data into an intrusion detection model to obtain a detection result output by the intrusion detection model; the intrusion detection model is obtained by training based on an optimized NN-LSTM algorithm and is used for detecting the internet traffic data. The technical scheme of the present application can realize the function of network intrusion active defense.

[0003] The existing network intrusion detection method and system have the following main problems:

[0004] The existing network intrusion detection method usually takes different modal data as static and equal-weight input channels, lacks dynamic weight adjustment capability for the perception intensity and state change of each modal in different time windows, and thus some key modal information is weakened or ignored, affecting the accuracy and robustness of detection. Network intrusion detection involves communication load, protocol triggering, system call, user behavior and other heterogeneous modal features, and the existing method is difficult to uniformly process the dimension difference and expression form of different modalities, making it difficult for the multi-modal deep learning model to fully mine cross-modal correlations and limiting the accurate identification of intrusion behaviors.

[0005] The existing system uses a static fusion strategy in multi-modal information fusion, which cannot dynamically adjust the contribution of each modal with the time window, and is prone to cause strong modal redundancy or weak modal being ignored. The existing fusion strategy ignores the potential correlation or redundant dependence between modalities, resulting in unreasonable attention allocation or loose fusion structure; the existing system does not perform differential processing on modal features, does not enhance the extraction of important modalities, and does not effectively reduce the dimension of weak modalities, causing waste of computing resources and model performance bottleneck, and restricting the application of multi-modal deep learning in real-time network intrusion detection.

[0006] In view of this, the present application proposes a network intrusion detection method and system based on multi-modal deep learning to solve the above problems. SUMMARY

[0007] In order to overcome the above-mentioned defects of the prior art and achieve the above-mentioned purposes, the present application provides the following technical scheme: a network intrusion detection method based on multi-modal deep learning, comprising:

[0008] S1, collect and modal attribution of multi-source data from different channels, construct a time-continuous multi-track dynamic perception track set, integrate a modal attention allocation mechanism, dynamically allocate modal parameters according to the modal energy of each modal under different perception tracks, and output the initial index tensor after multi-track modal fusion;

[0009] S2, receive the initial index tensor, establish a cross-modal dependency relationship by constructing a bidirectional response tensor between modalities and tracks; construct a semantic trigger matrix based on the semantic labels in the cross-modal dependency relationship; adopt a cross-modal residual connection mechanism to preserve low-order modal coupling features through parallel residual paths, and output modal linkage nodes;

[0010] S3, according to the modal linkage nodes output by the semantic trigger matrix, construct a multi-modal behavior aggregation graph, extract highly aggregated subgraphs and frequent path patterns appearing in the multi-modal behavior aggregation graph, and generate an abnormal behavior graph structure body labeled with abnormal confidence;

[0011] S4, based on the abnormal behavior graph structure body, extract attack sub-paths with time constraints and behavior logic causal chains through graph structure backtracking and pattern inversion analysis; adopt an abnormal behavior causal chain backtracking mechanism to track the smallest traceable trigger unit of abnormal behavior, and generate an attack graph differential structure in combination with a pre-set normal behavior template graph;

[0012] S5, modal source traceability labeling is performed on the attack graph differential structure, the variation of modal parameter structure is detected, a multi-modal time sequence perception vector group is constructed, and input into a pre-constructed deep joint discriminator for network intrusion behavior detection; when modal drift or track structure anomaly is detected, trigger an intrusion response mechanism, and push the detection result to a network security response terminal.

[0013] Preferably, the method for constructing a time-continuous multi-track dynamic perception track set comprises:

[0014] Collect multi-source data from different channels, including network communication data, host call data, user operation behavior data and external device access data; real-time monitoring is performed through a pre-set data collection agent and a listening port, and the collected multi-source data is encapsulated into a standardized data stream format with a timestamp;

[0015] A preset modal energy function is provided, which comprehensively considers data type, data structure and behavior characteristics; the collected multi-source data is classified and classified based on the modal attribution determination of the modal energy function, and is divided into network modal, system modal, user modal and context modal;

[0016] According to the timestamp information of the multi-source data, the multi-source data after modal attribution will be constructed into different modal track units within a preset time window according to the timestamp, the modal track units are connected in time sequence to form a time-continuous multi-track dynamic perception track set, and the multi-track dynamic perception track set includes a network perception track, a system perception track, a user perception track and a context perception track.

[0017] Preferably, the method for obtaining the initial index tensor comprises:

[0018] For each type of modal corresponding perception track, a set of modal energy factors of the modal is extracted in each time window as a state description index of the modal in the time period, and the state description index includes a communication load of a network modal, an abnormal protocol trigger rate, an entropy value of a calling sequence of a system modal, a resource access frequency, an operation jump frequency of a user modal, an environmental variable stability of a context modal and a device transformation rate.

[0019] The set of modal energy factors is sequentially aggregated and normalized to form a modal energy complex for perception, and a weighted scoring function is used to score the modal energy complex of each perception track in a preset time window to obtain a modal attention score;

[0020] An attention allocation mechanism is introduced, the modal attention score is normalized according to the modal energy complex of each perception track, the attention weight in the current multi-track dynamic perception track set is calculated, the cross-dependence probability between the modes is considered, the attention weight is optimized, the optimized attention weight is mapped to a modal parameter weight factor, and attention allocation is performed.

[0021] Based on the attention allocation result, each perception track is given a corresponding modal parameter weight factor, a preset modal attention score threshold is set, a feature extraction enhancement process is performed on the modal greater than the preset modal attention score threshold, and a dimension reduction operation is performed on the modal less than the preset modal attention score threshold.

[0022] After the optimization and allocation of each perception track are completed, the modal energy complex of the multi-track dynamic perception track is weighted and fused based on the modal parameter weight factor, and the fused modal energy complex is packaged into an initial index tensor in a unified format.

[0023] Preferably, the method for establishing a cross-modal dependence relationship comprises:

[0024] Based on the initial index tensor, a bidirectional response tensor between the modal and the perception track is constructed, the bidirectional response tensor including a response matrix of the track to the modal and a response matrix of the modal to the track; an orbit feature vector corresponding to each perception track in each time window is obtained, the orbit feature vector being generated by feature aggregation of different modal energy composite sub-features carried by the track;

[0025] By using the cosine similarity to measure the matching degree between the orbit feature vector and the modal energy composite sub, a response weight of the track to the modal is formed; according to the energy composite sub of different modal of each perception track in the time window, the response matrix of the track to the modal is obtained;

[0026] By normalizing each dimension feature value in the modal energy composite sub, multiplying the normalized feature value by the corresponding modal parameter weight factor and summing, a feature contribution degree vector is obtained; the response matrix of the modal to the track is obtained by combining the modal attention score of the modal in each time window and the feature contribution degree vector;

[0027] The track to modal response matrix and the modal to track response matrix are fused into a bidirectional response tensor, which represents the interaction response relationship between the modal and the perception track; based on the bidirectional response tensor, the similarity between the modal energy composite sub feature vectors is calculated by using the cosine similarity function, a dependence strength matrix between the modal is generated, a cross-modal dependence tensor is constructed, and the dependence strength between different modal in the time window is reflected through the cross-modal dependence tensor.

[0028] Preferably, the construction method of the modal linkage node comprises:

[0029] Based on the dependence strength in the cross-modal dependence tensor, the semantic labels of each modal are extracted, the semantic labels including the semantic categories, context variables and mutual relationships between the modal; a set of semantic trigger rules is preset, and it is judged whether the semantic response of modal B is triggered if the semantic label of modal A appears a state in any time window;

[0030] A semantic trigger matrix is constructed, which is a two-dimensional numerical matrix, each row in the matrix represents a modal as a trigger source, and each column represents a modal as a response target; if it is judged that the semantic response of modal B is triggered, a value representing the trigger strength is filled in the position A, B in the semantic trigger matrix, and 0 is filled in the position not triggered;

[0031] After obtaining the semantic trigger matrix, a cross-modal residual connection mechanism is introduced to quantify the coupling degree between the modal, the similarity between the energy composite sub feature vectors of any two modal is calculated by using the cosine similarity function, the modal coupling strength score is obtained, and a modal coupling strength score threshold is preset;

[0032] The modal coupling feature with a modal coupling strength score greater than a preset modal coupling strength score threshold is divided into a high-order modal coupling feature; the modal coupling feature with a modal coupling strength score less than or equal to the preset modal coupling strength score threshold is divided into a low-order modal coupling feature; and a parallel residual connection structure is adopted to reserve the low-order modal coupling feature in a parallel residual path and output a modal linkage node.

[0033] Preferably, the method for obtaining the abnormal behavior graph structure comprises:

[0034] A multi-modal behavior aggregation graph is constructed with the modal linkage node as a vertex, the multi-modal behavior aggregation graph is a directed graph structure, the edges of the graph represent semantic trigger relationships between modes, and the edge weights represent trigger strengths; the node density of a subgraph in the multi-modal behavior aggregation graph is obtained by dividing the actual number of edges of the subgraph by a preset maximum number of edges;

[0035] A preset subgraph node density threshold is used to define a highly aggregated subgraph greater than the preset subgraph node density threshold; the highly aggregated subgraph and a frequent path pattern are extracted, the frequent path pattern including a time sequence of modal triggering and an occurrence frequency thereof in different time windows;

[0036] The modal linkage structure of the current time window is compared with the historical highly aggregated subgraph and the frequent path pattern, and a structure deviation degree is calculated, and when the structure deviation degree is greater than a preset structure deviation threshold, it is marked as an abnormal behavior candidate; a weighted combination function is constructed based on the structure deviation degree and the time sequence of the abnormal behavior, and an abnormal confidence score is output;

[0037] According to the abnormal behavior candidate and the abnormal confidence score, an abnormal behavior graph structure body labeled with an abnormal confidence is generated; the abnormal behavior graph structure body includes an abnormal path sequence, a related modal identifier, time window information, a trigger edge attribute, and an abnormal confidence value.

[0038] Preferably, the method for obtaining the attack sub-path comprises:

[0039] The modal nodes and trigger edges in the abnormal behavior graph structure body are time-stamped to construct a behavior graph with time sequence attributes; a graph structure backtracking analysis is performed from the abnormal path endpoint in the behavior graph, with a maximum time span and a maximum path depth being limited, and a backtracking path satisfying a semantic trigger logic and a time sequence is extracted; the semantic trigger logic includes the existence of a trigger edge and the trigger strength of the edge being greater than a preset trigger strength threshold;

[0040] The backtracking path is structurally and semantically matched with a preset attack behavior pattern library to identify a candidate attack sub-path; the candidate attack sub-path is determined for time consistency, causal logic coherence, and abnormal confidence aggregation, and an attack sub-path satisfying the conditions is extracted.

[0041] Preferably, the attack graph microstructure acquisition method comprises:

[0042] Based on the extracted attack sub-path, continue to backtrack in the abnormal behavior graph structure, limit the maximum backtracking depth and time span threshold, recursively track all reachable abnormal confidence nodes and trigger relationship edges, and form an abnormal behavior causal chain graph; in the abnormal behavior causal chain graph, identify the smallest connected subgraph structure that affects any node in the attack sub-path as the smallest traceable trigger unit;

[0043] Define the constraint conditions met by the smallest traceable trigger unit, which include that the smallest connected subgraph composed of nodes and edges in the smallest traceable trigger unit must be time-sequentially reachable to the preset target attack node; the trigger strength of each edge is greater than the preset trigger strength threshold; the abnormal confidence score is greater than the preset abnormal confidence score threshold; and the deletion of any node or edge in the unit will cause the attack sub-path to be unreconstructable;

[0044] Compare the abnormal behavior causal chain graph with the preset normal behavior template graph for structural difference, and generate an attack graph microstructure; the attack graph microstructure includes a node difference set, an edge difference set, a trigger strength difference set, and an abnormal confidence score set.

[0045] Preferably, the network intrusion behavior detection method comprises:

[0046] Mark the modal source corresponding to each node or edge in the node difference set and the edge difference set in the attack graph microstructure for modality source traceability, and retain the trigger link information in the attack sub-path;

[0047] Extract the modality parameter structure of each traceable marked node, and compare it with the modality parameter structure of the corresponding node in the preset normal behavior template graph to determine whether there is a modality parameter structure variation; when the modality parameter structure difference is greater than the preset modality parameter structure difference threshold, it is marked as having a modality parameter structure variation;

[0048] Combine the time stamp, modality identifier, trigger strength, and abnormal confidence score of the nodes in the attack graph microstructure to construct a corresponding multi-modality time sequence perception vector group, input the multi-modality time sequence perception vector group into a pre-constructed deep joint discriminant model, perform network intrusion behavior detection, and output the recognition result of the network intrusion behavior.

[0049] The network intrusion detection system based on multi-modality deep learning comprises:

[0050] A multi-track dynamic construction module collects and mode attribution multi-source data from different channels, constructs a time-continuous multi-track dynamic perception track set, integrates a mode attention allocation mechanism, dynamically allocates mode parameters according to the mode energy complex of each mode under different perception tracks, and outputs an initial index tensor after multi-track mode fusion;

[0051] A mode track intersection module receives the initial index tensor, establishes a cross-modal dependency relationship by constructing a bidirectional response tensor between modes and tracks, constructs a semantic trigger matrix based on the semantic labels in the cross-modal dependency relationship, adopts a cross-modal residual connection mechanism, preserves low-order mode coupling features through a parallel residual path, and outputs a mode linkage node;

[0052] A linkage graph mining module constructs a multi-modal behavior aggregation graph according to the mode linkage node output by the semantic trigger matrix, extracts highly aggregated subgraphs and frequent path patterns appearing in the multi-modal behavior aggregation graph, and generates an abnormal behavior graph structure body labeled with an abnormal confidence;

[0053] An attack trace backtracking module extracts attack sub-paths with time constraints and behavior logic causal chains through graph structure backtracking and pattern inversion analysis based on the abnormal behavior graph structure body; an abnormal behavior causal chain backtracking mechanism is adopted to track the smallest traceable trigger unit of the abnormal behavior, and an attack graph differential structure is generated in combination with a preset normal behavior template graph;

[0054] An intrusion behavior detection module performs mode source traceability labeling on the attack graph differential structure, detects mode parameter structure variation, constructs a multi-modal time sequence perception vector group, and inputs it into a pre-constructed deep joint discriminator for network intrusion behavior detection; when mode drift or track structure anomaly is detected, an intrusion response mechanism is triggered, and the detection result is pushed to a network security response terminal.

[0055] Compared with the prior art, the present application has the following beneficial effects:

[0056] The present application introduces a set of mode energy factors as a perception index system, measures the dynamic state of different modes in each time window, and combines different indicators such as communication load, call sequence entropy, and environmental variable stability to truly reflect the behavior change trend of each mode in the time sequence dimension. By aggregating and normalizing the mode energy factors to form a mode energy complex, the mode state is uniformly expressed, the accuracy of mode recognition and judgment is significantly improved, and the ability to adapt to complex dynamic environments is achieved;

[0057] A scoring function based on dot product attention mechanism is adopted to dynamically calculate the importance score of each modal trajectory within the current time window, thereby strengthening the guiding role of key modal information and improving the overall perception quality. A cross-dependency strength parameter between modalities is introduced. After normalizing the modal attention weights, the initial attention weights are optimized based on the degree of inter-modal correlation, effectively solving the problem of uncaptured linkage between different modalities in existing multimodal systems. The standardized encapsulation method overcomes the problems of inconsistent multimodal output structures and cumbersome processing in existing technologies, improving the system's usability and scalability. Attached Figure Description

[0058] Figure 1 This is a schematic diagram of the network intrusion detection method based on multimodal deep learning of the present invention;

[0059] Figure 2 This is a schematic diagram of the network intrusion detection system based on multimodal deep learning according to the present invention;

[0060] Figure 3 This is a schematic diagram of the method for obtaining attack sub-paths provided by the present invention. Detailed Implementation

[0061] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0062] Example 1

[0063] Please see Figure 1 and Figure 3 As shown, Embodiment 1 further illustrates the network intrusion detection method based on multimodal deep learning proposed in this invention, including:

[0064] With the increasing complexity of network environments and the continuous evolution of attack methods, network intrusion detection has become a core component of network security defense systems. Existing network intrusion detection methods primarily rely on the analysis of multi-source information such as network traffic, system logs, protocol behavior, and user operations to identify potential malicious behaviors or abnormal patterns. In recent years, with the development of deep learning technology, more and more research has incorporated multimodal features into intrusion detection models, attempting to improve detection performance by fusing different types of data modalities (such as communication load, system call sequences, user behavior trajectories, and protocol triggering events).

[0065] However, the existing multi-modal network intrusion detection method still has the following key technical bottlenecks:

[0066] The existing method generally processes various types of modal features as static and equal-weight inputs, lacking dynamic modeling and adjustment ability for the change of perception intensity, the difference in expression ability and the relevance to attack events of different modalities in different time windows. This problem leads to the weakening of some key modalities in a specific attack stage, making it difficult to fully exploit important features and affecting the detection accuracy and robustness of the model.

[0067] The multi-modal data involved in network intrusion detection generally has heterogeneity and high dimensionality, such as structured sequences of protocol data, semi-structured text of system logs, and time series trajectories of user behavior, with large dimension span and significant differences in expression. In the process of modal unification, the existing method often uses a fixed structure of deep learning network, which is difficult to adapt to the heterogeneity and dynamics between modalities, resulting in insufficient cross-modal feature correlation learning ability and limiting accurate identification of complex intrusion behaviors.

[0068] The existing system generally adopts a static fusion strategy in the multi-modal fusion stage, lacking a mechanism to dynamically adjust the contribution of each modality according to time evolution. This strategy can easily lead to problems such as strong modal information redundancy and weak modal feature loss, and cannot accurately focus on modal features related to intrusion behavior. In addition, the existing system does not fully consider the potential dependency and complementarity between modalities in the fusion structure, resulting in unreasonable attention mechanism allocation or loose fusion structure, further weakening the model performance.

[0069] In the process of modal feature processing, the existing intrusion detection system generally uses a one-size-fits-all approach for different modalities, lacking enhancement mechanisms for key modalities and compression mechanisms for redundant modalities, which not only wastes computational resources but also increases training difficulty and reduces inference efficiency, restricting the practical application of multi-modal deep learning methods in network intrusion detection scenarios with high real-time requirements.

[0070] To effectively solve the above problems, the present application proposes a network intrusion detection method based on multi-modal deep learning, comprising:

[0071] S1, collect and attribute multi-source data from different channels, construct a set of time-continuous multi-track dynamic perception tracks, integrate a modal attention allocation mechanism, dynamically allocate modal parameters according to the modal energy complex of each modality in different perception tracks, and output an initial index tensor after multi-track modal fusion;

[0072] S2, receive an initial index tensor, establish a cross-modal dependency relationship by constructing a bidirectional response tensor between modes and tracks; construct a semantic trigger matrix based on the semantic labels in the cross-modal dependency relationship, use a cross-modal residual connection mechanism, preserve low-order modal coupling features through parallel residual paths, and output modal linkage nodes;

[0073] S3, construct a multi-modal behavior aggregation graph based on the modal linkage nodes output by the semantic trigger matrix, extract highly aggregated subgraphs and frequent path patterns appearing in the multi-modal behavior aggregation graph, and generate an abnormal behavior graph structure body labeled with abnormal confidence;

[0074] S4, based on the abnormal behavior graph structure body, extract attack sub-paths with time constraints and behavior logic causal chains through graph structure backtracking and pattern inversion analysis; use an abnormal behavior causal chain backtracking mechanism to track the smallest traceable trigger unit of abnormal behavior, and generate an attack graph differential structure in combination with a pre-set normal behavior template graph;

[0075] S5, label the attack graph differential structure for modal source traceability, detect modal parameter structure variation, construct a multi-modal time sequence perception vector group, and input it into a pre-constructed deep joint discriminator for network intrusion behavior detection; when modal drift or track structure anomaly is detected, trigger an intrusion response mechanism, and push the detection result to a network security response terminal.

[0076] The method for constructing a time-continuous multi-track dynamic perception track set comprises:

[0077] Collecting multi-source data from different channels, the multi-source data including network communication data, host call data, user operation behavior data, and external device access data; performing real-time monitoring through a pre-set data collection agent and a listening port, and encapsulating the collected multi-source data into a standardized data stream format with a timestamp;

[0078] The network communication data includes TCP / IP traffic, DNS resolution records, HTTP request information, network connection session start time, and duration; the host call data includes process creation, termination, file access log, kernel log, and security audit log;

[0079] The user operation behavior data includes login and logout records (time, method, source address), user command history (such as Shell / Bash commands), and terminal remote access records; the external device access data includes USB storage device plug-in events, external device transmission file records (transmission direction, file type, timestamp), storage device mounting and unmounting records;

[0080] The preset modality energy function is a function that comprehensively considers data types, data structures and behavior characteristics; the modality attribution determination and classification of the collected multi-source data are performed based on the modality energy function, and the multi-source data are classified into network modality, system modality, user modality and context modality;

[0081] According to the timestamp information of the multi-source data, the multi-source data after modality attribution will construct different modality track units in a preset time window according to the timestamp, and the modality track units are connected in time sequence to form a time-continuous multi-track dynamic perception track set, which includes a network perception track, a system perception track, a user perception track and a context perception track.

[0082] The method for obtaining the initial index tensor comprises:

[0083] For the perception track corresponding to each type of modality, a set of modality energy factors of the modality is extracted in each time window as a state description index of the modality in the time period, and the state description index includes a communication load of the network modality, an abnormal protocol trigger rate, an entropy value of a calling sequence of the system modality, a resource access frequency, an operation jump frequency of the user modality, an environmental variable stability and a device transformation rate of the context modality.

[0084] The set of modality energy factors is sequentially aggregated and normalized to form a modality energy complex used for perception, which represents the perception intensity of the modality in the current perception track window; a weighted scoring function is used to score the modality energy complex of each perception track in a preset time window to obtain a modality attention score;

[0085] The weighted scoring function is: ; wherein, represents the modality in the time window ; represents a global query vector in the time window , which is obtained by summing all the energy complex feature vectors of the perception tracks in the time window ; represents the modality energy complex feature vector of the modality in the time window ; represents the dimension of the modality energy complex feature vector, which is used to prevent the value after the vector inner product from being too large; it should be noted that the weighted scoring function is based on the dot product attention mechanism and is used to calculate the attention score of each modality track in a specified time window;

[0086] A attention allocation mechanism is introduced, and the modal attention score is normalized according to the modal energy complexor of each perception track, and the attention weight in the current multi-track dynamic perception track set is calculated; the cross-dependence probability between modes is considered, the attention weight is optimized, the optimized attention weight is mapped to the modal parameter weight factor, and attention allocation is performed;

[0087] The optimized attention weight is ; wherein, represents the optimized attention weight of the modal in the time window . represents the original attention weight of the modal in the time window . represents the weight adjustment coefficient, which is used to balance the proportion of the original attention weight and the dependent modal contribution, and is determined according to expert experience, The value range of is 0 to 1. represents the cross-dependence intensity between the modal and the modal . represents the original attention weight of the modal in the time window . and represent the index of the modal;

[0088] Based on the attention allocation result, each perception track is given a corresponding modal parameter weight factor, a preset modal attention score threshold is set, feature extraction enhancement processing is performed on the modal greater than the preset modal attention score threshold, and dimensionality reduction operation is performed on the modal less than the preset modal attention score threshold.

[0089] After the optimization and allocation of each perception track are completed, the modal energy complexor of the multi-track dynamic perception track is weighted and fused based on the modal parameter weight factor, and the fused modal energy complexor is packaged into an initial index tensor in a unified format.

[0090] The problems existing in the prior art are solved, and the problems are as follows: in the prior art, different modalities are usually regarded as parallel input channels, only simple splicing or average fusion is performed, and fine modeling and dynamic differentiation of the perception intensity and state performance of each modality in different time periods are lacked; the index dimensions and expression methods of different modalities are inconsistent, and it is difficult to support unified analysis across modalities; the existing system uses a static fusion strategy in multi-modal information fusion, and cannot dynamically adjust the contribution of each modality with a time window, which easily leads to strong modality redundancy or weak modality being ignored. The existing fusion strategy ignores the potential correlation or redundant dependence between modalities, leading to unreasonable attention allocation or loose fusion structure; the features of weak modalities cannot be compressed, and the features of strong modalities cannot be extracted, causing performance bottleneck and resource waste of the model; and a standardized and high-universal output expression structure is lacked.

[0091] The beneficial effects of the prior art are as follows: by introducing a set of modality emotion factors as a perception index system, the dynamic state of different modalities in each time window is measured, and combined with differential indexes such as communication load, call sequence entropy and environmental variable stability, the behavior change trend of each modality in the time dimension can be truly reflected. By aggregating and normalizing the modality emotion factors, a modality emotion complex is formed, which realizes the unified expression of the modality state, significantly improves the accuracy of modality recognition and judgment, and has the ability to adapt to complex dynamic environments.

[0092] The scoring function based on the dot product attention mechanism can dynamically calculate the importance score of each modality track in the current time window; the guiding role of key modality information is strengthened, and the overall perception quality is improved. The cross-dependence strength parameter between modalities is introduced, and after normalizing the modality attention weight, the initial attention weight is adjusted according to the correlation between modalities, effectively solving the problem that the linkage between different modalities in the existing multi-modal system is not captured; the standardized packaging method overcomes the inconsistency of the multi-modal output structure in the prior art and the cumbersome processing problem, improving the usability and expansibility of the system.

[0093] The method for establishing the cross-modality dependence relationship includes:

[0094] Based on the initial index tensor, a bidirectional response tensor between the modality and the perception track is constructed, and the bidirectional response tensor includes a response matrix from the track to the modality and a response matrix from the modality to the track; the track feature vector corresponding to each perception track in each time window is obtained, and the track feature vector is generated by feature aggregation of different modality emotion complex features carried by the track;

[0095] The matching degree between the track feature vector and the modal affective complex is measured by using cosine similarity to form the response weight of the track to the modal; the response matrix of the track to the modal is obtained according to the affective complex of each perception track to different modes in the time window, which represents the response degree of the track set to each mode;

[0096] By normalizing each dimension feature value in the modal affective complex, multiplying the normalized feature value by the corresponding modal parameter weight factor and summing, the feature contribution degree vector is obtained; the response matrix of the modal to the track is obtained by combining the modal attention score of the modal in each time window and the feature contribution degree vector, which represents the response preference of the modal to each perception track;

[0097] The track-to-modal response matrix and the modal-to-track response matrix are fused into a bidirectional response tensor to represent the interaction response relationship between the modal and the perception track; based on the bidirectional response tensor, the similarity between the modal affective complex feature vectors is calculated by using the cosine similarity function to generate the dependence strength matrix between the modes, and a cross-modal dependence tensor is constructed to reflect the dependence strength between different modes in the time window.

[0098] The construction method of the modal linkage node includes:

[0099] Based on the dependence strength in the cross-modal dependence tensor, the semantic labels of each modal are extracted, including the semantic categories of each modal (such as voice, image, text, etc.), the context variables (such as "high noise", "night", "high interaction density", etc.), and the mutual relationship between modes (such as "main- auxiliary coupling", "multi-directional feedback", etc.); a set of semantic trigger rules are preset to determine whether the semantic response of modal B will be triggered if the semantic label of modal A appears a state in any time window; for example: if the image modal detects "moving object", it may trigger the voice recognition modal to enter the "activated" state; if the temperature modal is marked as "high temperature alarm", the text modal is triggered to generate a prompt content;

[0100] A semantic trigger matrix is constructed, which is a two-dimensional numerical matrix, each row of which represents a modal as a trigger source, and each column represents a modal as a response target; if it is determined that the semantic response of modal B is triggered, a value representing the trigger strength is filled in the position A, B in the semantic trigger matrix, and 0 is filled in the position not triggered;

[0101] After obtaining the semantic trigger matrix, a cross-modal residual connection mechanism is introduced to quantify the coupling degree between modes, the similarity between the affective complex feature vectors of any two modes is calculated by using the cosine similarity function, the modal coupling strength score is obtained, and the modal coupling strength score threshold is preset.

[0102] The modal coupling feature with a modal coupling strength score greater than a preset modal coupling strength score threshold is divided into a high-order modal coupling feature; the modal coupling feature with a modal coupling strength score less than or equal to the preset modal coupling strength score threshold is divided into a low-order modal coupling feature; and a parallel residual connection structure is used to retain the low-order modal coupling feature in a parallel residual path and output a modal linkage node.

[0103] The method for obtaining the abnormal behavior graph structure body includes:

[0104] A multi-modal behavior aggregation graph is constructed with the modal linkage node as a vertex, the multi-modal behavior aggregation graph is a directed graph structure, the edges of the graph represent semantic trigger relationships between modes, and the edge weights represent trigger strengths; the node density of a subgraph in the multi-modal behavior aggregation graph is obtained by dividing the actual number of edges of the subgraph by a preset maximum number of edges;

[0105] A preset subgraph node density threshold is used to define a highly aggregated subgraph as a subgraph greater than the preset subgraph node density threshold; the highly aggregated subgraph and a frequent path pattern are extracted, the frequent path pattern including a time sequence of modal triggering and an occurrence frequency thereof in different time windows;

[0106] The modal linkage structure of the current time window is compared with the historical highly aggregated subgraph and the frequent path pattern, and a structure deviation degree is calculated, and when the structure deviation degree is greater than a preset structure deviation threshold, it is marked as an abnormal behavior candidate; a weighted combination function is constructed based on the structure deviation degree and the time sequence of the abnormal behavior, and an abnormal confidence score is output;

[0107] According to the abnormal behavior candidate and the abnormal confidence score, an abnormal behavior graph structure body labeled with an abnormal confidence is generated; the abnormal behavior graph structure body includes an abnormal path sequence, a related modal identifier, time window information, a trigger edge attribute, and an abnormal confidence value.

[0108] The method for obtaining the attack sub-path includes:

[0109] The modal nodes and trigger edges in the abnormal behavior graph structure body are time-stamped, and a behavior graph with time sequence attributes is constructed; a graph structure backtracking analysis is performed from the abnormal path end point of the behavior graph, a maximum time span and a maximum path depth are limited, and a backtracking path that satisfies semantic trigger logic and a time sequence order is extracted; the semantic trigger logic includes the existence of a trigger edge and the trigger strength of the edge being greater than a preset trigger strength threshold;

[0110] The backtracking path is structurally and semantically matched with a preset attack behavior pattern library, and a candidate attack sub-path is identified; the candidate attack sub-path is determined for time consistency, causal logic coherence, and abnormal confidence aggregation, and an attack sub-path that satisfies the conditions is extracted.

[0111] It should be noted that the time consistency represents that the time difference of all adjacent nodes on the candidate attack sub-path is less than the preset time difference threshold; the causal logical consistency represents that the connection between events in the candidate attack sub-path must exist in the semantic trigger matrix, and the corresponding trigger strength is higher than the preset trigger strength threshold; the abnormal confidence aggregation represents that the average abnormal confidence score of the candidate attack sub-path is greater than the preset average abnormal confidence score threshold.

[0112] The method for obtaining the attack graph microstructure comprises:

[0113] Based on the extracted attack sub-path, the abnormal behavior graph structure is continued to be traced back, the maximum backtracking depth and the time span threshold are limited, all reachable abnormal confidence nodes and trigger relationship edges are recursively tracked, and an abnormal behavior causal chain graph is formed; in the abnormal behavior causal chain graph, a minimum connected subgraph structure affecting any node in the attack sub-path is identified as a minimum traceable trigger unit;

[0114] The constraint conditions satisfied by the minimum traceable trigger unit are defined, and the constraint conditions include that the minimum connected subgraph composed of the nodes and edges in the minimum traceable trigger unit must be time-sequentially reachable to the preset target attack node; the trigger strength of each edge is greater than the preset trigger strength threshold; the abnormal confidence score is greater than the preset abnormal confidence score threshold; and the deletion of any node or edge in the unit will cause the attack sub-path to be unreconstructable;

[0115] The abnormal behavior causal chain graph and the preset normal behavior template graph are compared in terms of structural differences to generate an attack graph microstructure; the attack graph microstructure includes a node difference set, an edge difference set, a trigger strength difference set, and an abnormal confidence score set.

[0116] The node difference set represents a node that exists in the abnormal behavior causal chain graph but does not exist in the preset normal behavior template graph; the edge difference set represents a trigger relationship that exists in the abnormal behavior causal chain graph but does not exist in the preset normal behavior template graph; the strength difference set represents that the trigger strength difference between the abnormal edge and the corresponding edge in the preset normal behavior template graph exceeds the preset edge trigger strength difference threshold; and the abnormal confidence score set is an aggregated subgraph formed by a continuous high-confidence region in the attack path.

[0117] The method for detecting network intrusion behavior comprises:

[0118] The node difference set and the edge difference set in the attack graph microstructure are marked for modality source traceability, the modality source corresponding to each node or edge is identified, and the trigger link information in the attack sub-path is retained;

[0119] Extract the modal parameter structure of each node marked by traceability, and compare it with the modal parameter structure of the corresponding node in the preset normal behavior template graph to determine whether there is a modal parameter structure variation. When the modal parameter structure difference is greater than the preset modal parameter structure difference threshold, it is marked as having a modal parameter structure variation;

[0120] In combination with the timestamps, modal identifiers, trigger strengths and abnormal confidence scores of the nodes in the attack graph differential structure, a corresponding multi-modal time sequence perception vector group is constructed. The multi-modal time sequence perception vector group is input into a pre-constructed deep joint discriminant model for network intrusion behavior detection, and an identification result of the network intrusion behavior is output.

[0121] The preset modal attention score threshold is set by the staff. By collecting different modal attention scores, the average of multiple modal attention scores is taken as the preset modal attention score threshold. Similarly, the preset modal coupling strength score threshold, the preset subgraph node density threshold, the preset structure deviation threshold, the preset trigger strength threshold, the time span threshold, the preset abnormal confidence score threshold and the preset modal parameter structure difference threshold are set.

[0122] In this embodiment, by introducing a set of modal emotion factors as a perception index system, the dynamic state of different modalities in each time window is measured. In combination with different indicators such as communication load, call sequence entropy and environmental variable stability, the behavior change trend of each modality in the time sequence dimension can be truly reflected. By aggregating and normalizing the modal emotion factors, a modal emotion complex is formed, which realizes the unified expression of the modal state, significantly improves the accuracy of modal recognition and judgment, and has the ability to adapt to complex dynamic environments.

[0123] The scoring function based on the dot product attention mechanism can dynamically calculate the importance score of each modal track in the current time window. The guiding role of key modal information is strengthened, and the overall perception quality is improved. The cross-dependence strength parameter between modalities is introduced. After normalizing the modal attention weight, the initial attention weight is optimized by combining the degree of mutual association between modalities, effectively solving the problem that the linkage between different modalities in the existing multi-modal system is not captured. The standardized packaging method overcomes the inconsistency of the multi-modal output structure in the prior art, and the processing is cumbersome, improving the usability and expandability of the system.

[0124] Embodiment 2

[0125] Please refer to Figure 2 The network intrusion detection system based on multi-modal deep learning of the embodiment shown in the figure comprises:

[0126] The multi-track dynamic construction module collects and mode attribution of multi-source data from different channels, constructs a time-continuous multi-track dynamic perception track set, integrates a mode attention allocation mechanism, dynamically allocates mode parameters according to the mode energy of each mode under different perception tracks, and outputs an initial index tensor after multi-track mode fusion.

[0127] The mode track intersection module receives the initial index tensor, establishes a cross-modal dependency relationship by constructing a bidirectional response tensor between modes and tracks, constructs a semantic trigger matrix based on semantic labels in the cross-modal dependency relationship, adopts a cross-modal residual connection mechanism, preserves low-order mode coupling features through a parallel residual path, and outputs a mode linkage node.

[0128] The linkage graph mining module constructs a multi-modal behavior aggregation graph according to the mode linkage node output by the semantic trigger matrix, extracts highly aggregated subgraphs and frequent path patterns appearing in the multi-modal behavior aggregation graph, and generates an abnormal behavior graph structure body labeled with an abnormal confidence;

[0129] The attack trace backtracking module extracts attack sub-paths with time constraints and behavior logic causal chains through graph structure backtracking and pattern inversion analysis based on the abnormal behavior graph structure body; adopts an abnormal behavior causal chain backtracking mechanism to track the smallest traceable trigger unit of the abnormal behavior, and generates an attack graph differential structure in combination with a preset normal behavior template graph.

[0130] The intrusion behavior detection module labels the attack graph differential structure with a mode source traceability, detects mode parameter structure variation, constructs a multi-modal time sequence perception vector group, and inputs it into a pre-constructed deep joint discriminator for network intrusion behavior detection; when mode drift or track structure anomaly is detected, an intrusion response mechanism is triggered, and the detection result is pushed to a network security response terminal.

[0131] Since the electronic device introduced in the embodiment is the electronic device used to implement the network intrusion detection method and system based on multi-modal deep learning in the embodiment, the specific implementation of the electronic device and its various forms can be understood by those skilled in the art based on the network intrusion detection method and system based on multi-modal deep learning introduced in the embodiment, so the implementation of the electronic device in the method of the embodiment will not be described in detail. As long as the electronic device used to implement the network intrusion detection method and system based on multi-modal deep learning in the embodiment is implemented by those skilled in the art, it belongs to the scope of protection of the present application.

[0132] The above formulas are all dimensionless numerical calculations, the formulas are obtained by collecting a large amount of data to simulate a formula of the most recent real situation, and preset parameters and threshold values in the formulas are set by a person skilled in the art according to actual conditions.

[0133] The above only describes the preferred embodiments of the present application, and the protection scope of the present application is not limited to the above-described embodiments. Any technical solution falling within the concept of the present application belongs to the protection scope of the present application. It should be noted that, for ordinary technical users in the technical field, some improvements and refinements without departing from the principles of the present application are also considered to be within the protection scope of the present application.

Claims

1. A network intrusion detection method based on multi-modal deep learning, characterized in that, The method comprises the following steps: S1, collecting and modal attribution of multi-source data from different channels, constructing a time-continuous multi-track dynamic perception track set, integrating a modal attention allocation mechanism, dynamically allocating modal parameters according to the modal energy complex of each type of modal under different perception tracks, and outputting an initial index tensor after multi-track modal fusion; The method for obtaining the initial index tensor comprises: For each type of modal corresponding to the perception track, a set of modal energy factors of the modal is extracted in each time window as a state description index of the modal in the time period. The state description index includes the communication load of the network modal, the abnormal protocol trigger rate, the calling sequence entropy value of the system modal, the resource access frequency, the operation jump frequency of the user modal, the environment variable stability of the context modal, and the device transformation rate; The modal emotion factor set is time-aggregated and normalized to form a modal emotion complex for perception; a weighted scoring function is used to score the modal emotion complex of each perception track within a preset time window to obtain a modal attention score; An attention allocation mechanism is introduced. According to the modal energy complex of each perception track, the modal attention score is normalized, and the attention weight in the current multi-track dynamic perception track set is calculated. Considering the cross-dependence probability between modes, the attention weight is optimized, the optimized attention weight is mapped to the modal parameter weight factor, and attention allocation is performed; Based on the attention allocation result, each perception track is given a corresponding modal parameter weight factor. A preset modal attention score threshold is set. Feature extraction enhancement processing is performed on the modal greater than the preset modal attention score threshold, and dimensionality reduction operation is performed on the modal less than the preset modal attention score threshold; After the optimization and allocation of each perception track are completed, the modal energy complex of the multi-track dynamic perception track is weighted and fused based on the modal parameter weight factor, and the fused modal energy complex is packaged into an initial index tensor in a unified format; S2, receiving the initial index tensor, establishing a cross-modal dependence relationship by constructing a bidirectional response tensor between the modal and the track; constructing a semantic trigger matrix based on the semantic labels in the cross-modal dependence relationship; using a cross-modal residual connection mechanism, preserving low-order modal coupling features through parallel residual paths, and outputting a modal linkage node; S3, constructing a multi-modal behavior aggregation graph according to the modal linkage node output by the semantic trigger matrix, extracting highly aggregated subgraphs and frequent path patterns appearing in the multi-modal behavior aggregation graph, and generating an abnormal behavior graph structure body labeled with abnormal confidence; S4, based on the abnormal behavior graph structure body, extracting attack sub-paths with time constraints and behavior logic causal chains through graph structure backtracking and pattern inversion analysis; using an abnormal behavior causal chain backtracking mechanism to track the smallest traceable trigger unit of the abnormal behavior, and combining a preset normal behavior template graph to generate an attack graph differential structure; S5, marking the modal signal source traceability of the attack graph differential structure, detecting the modal parameter structure variation, constructing a multi-modal time sequence perception vector group, and inputting it into a pre-constructed deep joint discriminator for network intrusion behavior detection; when the modal drift or track structure anomaly is detected, the intrusion response mechanism is triggered, and the detection result is pushed to the network security response terminal.

2. The network intrusion detection method based on multi-modal deep learning according to claim 1, characterized in that, The method for constructing a time-continuous multi-track dynamic perception track set comprises: Collect multi-source data of different channels, including network communication data, host call data, user operation behavior data and external device access data; Real-time monitoring is performed through a preset data collection agent and a listening port, and the collected multi-source data is all packaged into a standardized data stream format with a time stamp; A preset modal affective function is provided, which comprehensively considers data types, data structures and behavior characteristics; The collected multi-source data is classified based on modal attribution determination of the modal affective function, and is divided into network modal, system modal, user modal and context modal; According to the time stamp information of the multi-source data, the multi-source data after modal attribution will be constructed into different modal track units in a preset time window according to the time stamp, the modal track units are connected in time sequence, and a time-continuous multi-track dynamic perception track set is formed, which includes network perception track, system perception track, user perception track and context perception track. 3.The multi-modal deep learning based network intrusion detection method of claim 2, wherein, The method for establishing a cross-modal dependency relationship includes: Based on the initial index tensor, a bidirectional response tensor between the modal and the perception track is constructed, the bidirectional response tensor includes a response matrix from the track to the modal and a response matrix from the modal to the track; Obtain the track feature vector corresponding to each perception track in each time window, which is generated by feature aggregation of different modal affective composite features carried by the track; By using cosine similarity to measure the matching degree between the track feature vector and the modal affective composite, the response weight of the track to the modal is formed; According to the affective composite of different modal of each perception track in the time window, the response matrix from the track to the modal is obtained; By normalizing each dimension feature value in the modal affective composite, multiplying the normalized feature value by the corresponding modal parameter weight factor and summing, the feature contribution degree vector is obtained; Combined with the modal attention score and the feature contribution degree vector of the modal in each time window, the response matrix from the modal to the track is obtained; The track-to-modal response matrix and the modal-to-track response matrix are fused into a bidirectional response tensor, which represents the interaction response relationship between the modal and the perception track; Based on the bidirectional response tensor, the similarity between the modal affective composite feature vectors is calculated by using the cosine similarity function, a dependency strength matrix between the modes is generated, a cross-modal dependency tensor is constructed, and the dependency strength between different modes in the time window is reflected through the cross-modal dependency tensor.

4. The network intrusion detection method based on multi-modal deep learning according to claim 3, characterized in that, The construction method of the modal linkage node includes: Based on the dependency strength in the cross-modal dependency tensor, the semantic labels of each modal are extracted, including the semantic categories, context variables and mutual relationships between the modes; A set of semantic trigger rules are preset to determine whether the semantic response of modal B will be triggered if the semantic label of modal A appears a state in any time window; A semantic trigger matrix is constructed, the semantic trigger matrix being a two-dimensional numerical matrix, each row in the matrix representing a modality as a trigger source, and each column representing a modality as a response target; if it is judged that the semantic response of modality B is triggered, a numerical value representing the triggering intensity is filled in the position A, B in the semantic trigger matrix, and 0 is filled in the position not triggered; After obtaining the semantic trigger matrix, a cross-modality residual connection mechanism is introduced to quantify the coupling degree between the modalities, the similarity between the energy complex feature vectors of any two modalities is calculated through a cosine similarity function, and a modality coupling strength score is obtained, and a preset modality coupling strength score threshold is set; The modality coupling features with the modality coupling strength score greater than the preset modality coupling strength score threshold are divided into high-order modality coupling features; the modality coupling features with the modality coupling strength score less than or equal to the preset modality coupling strength score threshold are divided into low-order modality coupling features; a parallel residual connection structure is used, the low-order modality coupling features are retained in the parallel residual path, and a modality linkage node is output.

5. The network intrusion detection method based on multi-modal deep learning according to claim 4, characterized in that, The method for obtaining the abnormal behavior graph structure body includes: A multi-modality behavior aggregation graph is constructed with the modality linkage node as a vertex, the multi-modality behavior aggregation graph being a directed graph structure, the edges of the graph representing the semantic trigger relationship between the modalities, and the edge weights representing the triggering intensity; the actual edge number of a subgraph in the multi-modality behavior aggregation graph is divided by a preset maximum edge number to obtain a subgraph node density; A preset subgraph node density threshold is set, and a subgraph greater than the preset subgraph node density threshold is defined as a highly aggregated subgraph; the highly aggregated subgraph and a frequent path pattern are extracted, the frequent path pattern including a time sequence of modality triggering and an occurrence frequency of the time sequence in different time windows; The modality linkage structure of the current time window is compared with the historical highly aggregated subgraph and the frequent path pattern, a structure deviation degree is calculated, and when the structure deviation degree is greater than a preset structure deviation threshold, the abnormal behavior candidate is marked; a weighted combination function is constructed based on the structure deviation degree and the time sequence of the abnormal behavior, and an abnormal confidence score is output; According to the abnormal behavior candidate and the abnormal confidence score, an abnormal behavior graph structure body labeled with an abnormal confidence is generated; the abnormal behavior graph structure body includes an abnormal path sequence, a modality identifier involved, time window information, a trigger edge attribute, and an abnormal confidence value.

6. The network intrusion detection method based on multi-modal deep learning of claim 5, wherein, The method for obtaining the attack sub-path includes: The modality nodes and the trigger edges in the abnormal behavior graph structure body are time-stamped to construct a behavior graph with time sequence attributes; the behavior graph is analyzed by backtracking from the end point of the abnormal path, the maximum time span and the maximum path depth are limited, and a backtracking path satisfying the semantic trigger logic and the time sequence is extracted; the semantic trigger logic includes the existence of the trigger edge and the trigger intensity of the edge being greater than a preset trigger intensity threshold; The backtracking path is matched with a preset attack behavior pattern library in structure and semantics, and a candidate attack sub-path is identified; the candidate attack sub-path is determined in terms of time consistency, causal logic coherence, and abnormal confidence aggregation, and an attack sub-path satisfying the conditions is extracted.

7. The network intrusion detection method based on multi-modal deep learning of claim 6, wherein, The method for obtaining the attack graph microstructure includes: Based on the extracted attack sub-path, continue to backtrack in the abnormal behavior graph structure, limit the maximum backtracking depth and time span threshold, recursively track all reachable abnormal confidence nodes and trigger relationship edges, form an abnormal behavior causal chain graph; In the abnormal behavior causal chain graph, identify the smallest connected subgraph structure that affects any node in the attack sub-path as the smallest traceable trigger unit; Define the constraint conditions that the smallest traceable trigger unit satisfies, including the smallest connected subgraph composed of the nodes and edges in the smallest traceable trigger unit must be reachable to the preset target attack node in the time sequence; The trigger strength of each edge is greater than the preset trigger strength threshold; The abnormal confidence score is greater than the preset abnormal confidence score threshold; The deletion of any node or edge in the unit will cause the attack sub-path to be unreconstructable; Compare the abnormal behavior causal chain graph with the preset normal behavior template graph for structural difference, and generate an attack graph differential structure; The attack graph differential structure includes a node difference set, an edge difference set, a trigger strength difference set, and an abnormal confidence score set. 8.The multi-modal deep learning based network intrusion detection method of claim 7, wherein, The method for detecting network intrusion behavior comprises: Marking the modal source corresponding to each node or edge in the node difference set and the edge difference set in the attack graph differential structure for modal signal source traceability, and retaining the trigger link information in the attack sub-path; Extract the modal parameter structure of each traceable marked node, and compare it with the modal parameter structure of the corresponding node in the preset normal behavior template graph to determine whether there is a modal parameter structure variation; When the modal parameter structure difference is greater than the preset modal parameter structure difference threshold, it is marked as having a modal parameter structure variation; Combine the time stamp, modal identifier, trigger strength and abnormal confidence score of the node in the attack graph differential structure to construct a corresponding multi-modal time sequence perception vector group, input the multi-modal time sequence perception vector group into the pre-constructed deep joint discriminant model, and perform network intrusion behavior detection, and output the recognition result of the network intrusion behavior.

9. A network intrusion detection system based on multi-modal deep learning for implementing the network intrusion detection method based on multi-modal deep learning according to any one of claims 1 to 8, characterized in that, It comprises: A multi-track dynamic construction module acquires and attributes multiple sources of data from different channels, constructs a time-continuous multi-track dynamic perception track set, integrates a modal attention allocation mechanism, dynamically allocates modal parameters according to the modal energy complex of each type of modal under different perception tracks, and outputs an initial index tensor after multi-track modal fusion; A modal track intersection module receives the initial index tensor, establishes a cross-modal dependency relationship by constructing a bidirectional response tensor between the modal and the track; A semantic trigger matrix is constructed based on the semantic labels in the cross-modal dependency relationship, and a cross-modal residual connection mechanism is adopted to preserve low-order modal coupling features through a parallel residual path and output modal linkage nodes; A linkage graph mining module constructs a multi-modal behavior aggregation graph according to the modal linkage nodes output by the semantic trigger matrix, extracts highly aggregated subgraphs and frequent path patterns appearing in the multi-modal behavior aggregation graph, and generates an abnormal behavior graph structure labeled with abnormal confidence; The attack trace backtracking module extracts attack sub-paths with time constraints and behavior logic causal chains through graph structure backtracking and pattern inversion analysis based on the abnormal behavior graph structure body; an abnormal behavior causal chain backtracking mechanism is adopted to track the minimum traceable trigger unit of abnormal behavior, and an attack graph differential structure is generated in combination with a preset normal behavior template graph; The intrusion behavior detection module marks the modal source traceability of the attack graph differential structure, detects the modal parameter structure variation, constructs a multi-modal time sequence perception vector group, and inputs it into the pre-constructed deep joint discriminator for network intrusion behavior detection; when modal drift or track structure anomaly is detected, the intrusion response mechanism is triggered, and the detection result is pushed to the network security response terminal.

Citation Information

Patent Citations

  • Network security intrusion detection method and system

    CN116886393A

  • Multi-modal emotion recognition method based on situational attention neural network

    CN112348075A

  • Unsupervised cross-modal retrieval method based on attention mechanism enhancement

    CN113971209A