Network intrusion detection method and system based on multi-modal deep learning

By constructing a multi-track dynamically perceived track set and cross-modal dependencies and dynamically allocating modal parameters, the problem of weakening modal information and uncatched linkage in the existing network intrusion detection methods is solved, and a higher precision network intrusion detection is achieved.

CN120498883AActive Publication Date: 2025-08-15聊城大学东昌学院 +1

Patent Information

Application Number
CN202510909760.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-02
Publication Date
2025-08-15
Estimated Expiration
2045-07-02

AI Technical Summary

Technical Problem

Existing network intrusion detection methods lack the ability to adjust dynamic weights to perceive intensity and state changes in different modes within different time windows, resulting in the weakening or neglect of key mode information, affecting the accuracy and robustness of detection, and multimodal deep learning models are difficult to fully explore cross-modal correlations and deal with heterogeneous modal features.

Method used

By constructing a time-continuous multi-track dynamic perception track set, integrating a modal attention allocation mechanism, dynamically allocating modal parameters, establishing a cross-modal dependency relationship, adopting a cross-modal residual connection mechanism, building a multi-modal behavior aggregation map, extracting anomaly behavior map structure, and performing modal source traceability marking, and inputting a depth joint discriminator for detection.

Benefits of technology

It significantly improves the accuracy of modal recognition and judgment, adapts to complex dynamic environments, solves the problem that the linkage between different modals is not captured, and improves the usability and scalability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120498883A_ABST
    Figure CN120498883A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of Internet security, and discloses a multi-modal deep learning-based network intrusion detection method and system, which comprises the steps of carrying out collection and modal attribution on multi-source data from different channels, constructing a time-continuous multi-track dynamic sensing track set, integrating a modal attention distribution mechanism, and carrying out multi-modal deep learning. According to the modal energy sensing complexes of each type of modals under different sensing orbits, dynamically distributing modal parameters, and outputting an initial index tensor after multi-orbit modal fusion; receiving an initial index tensor, and establishing a cross-modal dependency relationship by constructing a bidirectional response tensor between a modal and an orbit; constructing a semantic trigger matrix based on semantic tags in a cross-modal dependency relationship, adopting a cross-modal residual connection mechanism, retaining low-order modal coupling features through a parallel residual path, and outputting modal linkage nodes; according to the invention, the identification capability of detection on complex intrusion behaviors is improved, and more comprehensive and accurate detection on complex network attack behaviors is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of Internet security technology, and more specifically, to a network intrusion detection method and system based on multimodal deep learning. Background Art

[0002] Patent publication number CN116886393A discloses a network security intrusion detection method and system, comprising the following steps: Step S1: Collecting internet traffic data; Step S2: Inputting the internet traffic data into an intrusion detection model to obtain detection results output by the intrusion detection model; the intrusion detection model is trained based on an optimized NN-LSTM algorithm and is used to detect the internet traffic data. This invention's technical solution can implement active network intrusion defense.

[0003] Existing network intrusion detection methods and systems have the following major problems: Existing network intrusion detection methods typically treat data from different modalities as static, equally weighted input channels. These methods lack the ability to dynamically adjust the weights of each modality's sensory strength and state changes within different time windows. This results in the weakening or neglect of certain key modal information, impacting detection accuracy and robustness. Network intrusion detection involves multiple heterogeneous modal features, including communication load, protocol triggers, system calls, and user behavior. Existing methods struggle to uniformly address the dimensional differences and representations of these modalities, making it difficult for multimodal deep learning models to fully exploit cross-modal correlations, limiting the accurate identification of intrusions.

[0004] Existing systems use static fusion strategies for multimodal information fusion, failing to dynamically adjust the contribution of each modality over time. This can easily lead to redundancy of strong modalities or neglect of weak modalities. Existing fusion strategies ignore potential correlations or redundant dependencies between modalities, resulting in irrational attention allocation or loose fusion structures. Existing systems also fail to differentiate modal features, enhance the extraction of important modalities, or effectively reduce the dimensionality of weak modalities. This results in wasted computing resources and model performance bottlenecks, hindering the application of multimodal deep learning in real-time network intrusion detection.

[0005] In view of this, the present invention proposes a network intrusion detection method and system based on multimodal deep learning to solve the above problems. Summary of the Invention

[0006] In order to overcome the above-mentioned defects of the prior art and to achieve the above-mentioned objectives, the present invention provides the following technical solution: a network intrusion detection method based on multimodal deep learning, comprising: S1. Collect and assign modal attributes to multi-source data from different channels, construct a time-continuous multi-track dynamic perception track set, integrate the modal attention allocation mechanism, dynamically allocate modal parameters based on the modal energy complex of each modality under different perception tracks, and output the initial index tensor after multi-track modal fusion; S2: Receive the initial index tensor and establish a cross-modal dependency relationship by constructing a bidirectional response tensor between the modality and the track. Construct a semantic trigger matrix based on the semantic labels in the cross-modal dependency relationship. Adopt a cross-modal residual connection mechanism to preserve low-order modal coupling features through parallel residual paths and output modal linkage nodes. S3. Based on the modal linkage nodes output by the semantic trigger matrix, a multimodal behavior aggregation graph is constructed, highly clustered subgraphs and frequent path patterns appearing in the multimodal behavior aggregation graph are extracted, and an abnormal behavior graph structure annotated with abnormal confidence is generated; S4. Based on the abnormal behavior graph structure, through graph structure backtracking and pattern inversion analysis, we extract attack subpaths with time constraints and behavioral logic causal chains. We use the abnormal behavior causal chain backtracking mechanism to track the smallest traceable trigger unit of abnormal behavior, and generate the differential structure of the attack graph by combining it with the preset normal behavior template graph. S5. Mark the modal source traceability of the differential structure of the attack graph, detect the variation of the modal parameter structure, construct a multimodal time series perception vector group, and input it into the pre-built deep joint discriminator for network intrusion behavior detection; when modal drift or track structure abnormality is detected, trigger the intrusion response mechanism and push the detection results to the network security response terminal.

[0007] Preferably, the method for constructing a temporally continuous multi-track dynamic perception track set includes: Collect multi-source data from different channels, including network communication data, host call data, user operation behavior data, and external device access data; conduct real-time monitoring through preset data collection agents and listening ports, and encapsulate the collected multi-source data into a standardized data stream format with timestamps; A modal energy function is preset, which comprehensively considers data type, data structure and behavioral characteristics; based on the modal energy function, the collected multi-source data is classified into modal attribution, network modality, system modality, user modality and contextual modality; Based on the timestamp information of multi-source data, the multi-source data after modality attribution will construct different modal track units within the preset time window according to its timestamp, and connect the modal track units in chronological order to form a temporally continuous multi-track dynamic perception track set. The multi-track dynamic perception track set includes network perception track, system perception track, user perception track and context perception track.

[0008] Preferably, the method for obtaining the initial index tensor includes: For each type of modality's corresponding perception track, the modal sensitivity factor set of the modality is extracted in each time window as the modality's state characterization indicators in that time period. The state characterization indicators include the communication load of the network modality, the abnormal protocol trigger rate, the call sequence entropy value of the system modality, the resource access density, the operation jump frequency of the user modality, and the environmental variable stability and device conversion rate of the context modality. The modal sensory factor set is aggregated and normalized in time series to form a modal sensory complex for perception; a weighted scoring function is used to evaluate each perception track in a preset time window. The modal sense composite within the model is scored to obtain the modal attention score; An attention allocation mechanism is introduced. Based on the modal energy complex of each perception track, the modal attention score is normalized and its attention weight in the current multi-track dynamic perception track set is calculated. The attention weight is tuned considering the cross-dependency probability between modalities and mapped to the modal parameter weight factor for attention allocation. Based on the attention allocation results, a corresponding modal parameter weight factor is assigned to each perception track, a modal attention score threshold is preset, feature extraction and enhancement processing is performed on the modalities with a modal attention score greater than the preset modal attention score threshold, and dimensionality reduction operation is performed on the modalities with a modal attention score less than the preset modal attention score threshold; After completing the tuning and allocation of each perception track, the modal perception energy complexes of the multi-track dynamic perception track are weightedly fused based on the modal parameter weight factors, and the fused modal perception energy complexes are encapsulated as the initial index tensor in a unified format.

[0009] Preferably, the method for establishing a cross-modal dependency relationship includes: Based on the initial index tensor, a bidirectional response tensor is constructed between the modality and the perception track. The bidirectional response tensor includes the track-to-modality response matrix and the modality-to-track response matrix. The track feature vector corresponding to each perception track in each time window is obtained. The track feature vector is generated by feature aggregation of the different modal sensory energy composite sub-features carried by the track. By using cosine similarity to measure the matching degree between the orbital feature vector and the modal sensory complex, the orbital response weight to the modality is formed; based on the sensory complex of each sensing orbit to different modalities within the time window, the orbital to modal response matrix is obtained; By normalizing the eigenvalues of each dimension in the modal energy complex, multiplying the normalized eigenvalues with the corresponding modal parameter weight factors and summing them, a feature contribution vector is obtained. The modal attention score and feature contribution vector of the modal in each time window are combined to obtain the mode-to-orbit response matrix. The track-to-modal response matrix and the modal-to-track response matrix are fused into a bidirectional response tensor to characterize the interactive response relationship between the modality and the perception track. Based on the bidirectional response tensor, the similarity between the modal sensory energy composite sub-eigenvectors is calculated through the cosine similarity function to generate the inter-modal dependency strength matrix and construct a cross-modal dependency tensor to reflect the dependency strength between different modalities within the time window.

[0010] Preferably, the method for constructing the modal linkage node includes: Based on the dependency strength in the cross-modal dependency tensor, the semantic labels of each modality are extracted. The semantic labels include the semantic categories of each modality, contextual variables, and the relationship between modalities. A set of semantic triggering rules is preset to determine whether, within any time window, if a state of the semantic label of modality A appears, it will trigger a semantic response in modality B. Construct a semantic trigger matrix. The semantic trigger matrix is a two-dimensional numerical matrix. Each row in the matrix represents a modality as a trigger source, and each column represents a modality as a response target. If the semantic response of modality B is triggered, a value representing the trigger strength is filled in at position A or B in the semantic trigger matrix, and 0 is filled in at the untriggered position. After obtaining the semantic trigger matrix, a cross-modal residual connection mechanism is introduced to quantify the degree of coupling between the modalities. The similarity between the sensory-energy composite sub-feature vectors of any two modalities is calculated using the cosine similarity function to obtain the modal coupling strength score, and a modal coupling strength score threshold is preset. The modal coupling features with a modal coupling strength score greater than a preset modal coupling strength score threshold are classified as high-order modal coupling features; the modal coupling features with a modal coupling strength score less than or equal to the preset modal coupling strength score threshold are classified as low-order modal coupling features; a parallel residual connection structure is adopted to retain the low-order modal coupling features in the parallel residual path and output the modal linkage node.

[0011] Preferably, the method for obtaining the abnormal behavior graph structure includes: A multimodal behavior aggregation graph is constructed with modal linkage nodes as vertices. The multimodal behavior aggregation graph is a directed graph structure. The edges of the graph represent the semantic triggering relationship between modalities, and the edge weights represent the triggering strength. The actual number of edges in the subgraph of the multimodal behavior aggregation graph is divided by the preset maximum number of edges to obtain the subgraph node density. A subgraph node density threshold is preset, and subgraphs with a density greater than the preset subgraph node density threshold are defined as highly clustered subgraphs; highly clustered subgraphs and frequent path patterns are extracted. Frequent path patterns include the time series of modal triggering and its occurrence frequency in different time windows; The modal linkage structure of the current time window is compared with the historical highly clustered subgraphs and frequent path patterns to calculate the structural deviation. When the structural deviation is greater than the preset structural deviation threshold, it is marked as an abnormal behavior candidate. A weighted combination function is constructed based on the temporal sequence of the structural deviation and the abnormal behavior to output the abnormality confidence score. Based on the abnormal behavior candidates and abnormal confidence scores, an abnormal behavior graph structure marked with abnormal confidence is generated; the abnormal behavior graph structure contains abnormal path sequence, involved modal identification, time window information, triggering edge attributes and abnormal confidence value.

[0012] Preferably, the method for acquiring the attack sub-path includes: Timestamp the modal nodes and trigger edges in the abnormal behavior graph structure to construct a behavior graph with time-series attributes. Backtrack the graph structure from the end point of the abnormal path in the behavior graph, limit the maximum time span and maximum path depth, and extract the backtracking path that meets the semantic trigger logic and time sequence. The semantic trigger logic includes the existence of the trigger edge and the trigger strength of the edge greater than the preset trigger strength threshold. The backtracking path is structurally and semantically matched with the preset attack behavior pattern library to identify candidate attack sub-paths; the candidate attack sub-paths are judged for temporal consistency, causal logical coherence, and abnormal confidence aggregation, and the attack sub-paths that meet the conditions are extracted.

[0013] Preferably, the method for obtaining the differential structure of the attack graph includes: Based on the extracted attack subpath, the system continues to backtrack in the abnormal behavior graph structure, limiting the maximum backtracking depth and time span threshold, and recursively traces all reachable abnormal confidence nodes and trigger relationship edges to form an abnormal behavior causal chain graph. In the abnormal behavior causal chain graph, the system identifies the minimum connected subgraph structure that affects any node in the attack subpath as the minimum traceable trigger unit. Define the constraints that the minimum traceable trigger unit must satisfy. These constraints include: the minimum connected subgraph consisting of the nodes and edges contained in the minimum traceable trigger unit must be able to reach the preset target attack node in time sequence; the trigger strength of each edge must be greater than the preset trigger strength threshold; the anomaly confidence score must be greater than the preset anomaly confidence score threshold; and deleting any node or edge in the unit will make the attack subpath unreconstructable. The structural differences between the abnormal behavior causal chain graph and the preset normal behavior template graph are compared to generate the attack graph differential structure; the attack graph differential structure includes a node difference set, an edge difference set, a trigger strength difference set, and an anomaly confidence score set.

[0014] Preferably, the method for detecting network intrusion behavior includes: The node difference set and edge difference set in the differential structure of the attack graph are marked with modal source traceability to identify the modal source corresponding to each node or edge, and the trigger link information in the attack sub-path is retained; The modal parameter structure of each traceably marked node is extracted and compared with the modal parameter structure of the corresponding node in the preset normal behavior template diagram to determine whether there is modal parameter structure variation. When the modal parameter structure difference is greater than the preset modal parameter structure difference threshold, it is marked as modal parameter structure variation. Combining the timestamps, modal identifications, trigger strengths, and anomaly confidence scores of the nodes in the differential structure of the attack graph, a corresponding multimodal time-series perception vector group is constructed. The multimodal time-series perception vector group is input into a pre-built deep joint discriminant model to detect network intrusion behavior and output the identification results of the network intrusion behavior.

[0015] Network intrusion detection system based on multimodal deep learning, including: The multi-track dynamic construction module collects and modally attributes multi-source data from different channels, constructs a temporally continuous multi-track dynamic perception track set, integrates the modal attention allocation mechanism, dynamically allocates modal parameters based on the modal energy complex of each modality in different perception tracks, and outputs the initial index tensor after multi-track modal fusion; The modal-track cross module receives the initial index tensor and establishes a cross-modal dependency relationship by constructing a bidirectional response tensor between the modality and the track. It builds a semantic trigger matrix based on the semantic labels in the cross-modal dependency relationship, adopts a cross-modal residual connection mechanism, preserves low-order modal coupling features through parallel residual paths, and outputs modal linkage nodes. The linkage graph mining module constructs a multimodal behavior aggregation graph based on the modal linkage nodes output by the semantic trigger matrix, extracts highly clustered subgraphs and frequent path patterns appearing in the multimodal behavior aggregation graph, and generates an abnormal behavior graph structure annotated with anomaly confidence. The attack trace backtracking module, based on the abnormal behavior graph structure, extracts attack subpaths with time constraints and behavioral logic causal chains through graph structure backtracking and pattern inversion analysis. It uses the abnormal behavior causal chain backtracking mechanism to track the smallest traceable trigger unit of abnormal behavior and generates the differential structure of the attack graph by combining it with the preset normal behavior template graph. The intrusion behavior detection module performs modal source traceability marking on the differential structure of the attack graph, detects the variation of the modal parameter structure, constructs a multimodal time series perception vector group, and inputs it into the pre-built deep joint discriminator for network intrusion behavior detection; when modal drift or track structure abnormality is detected, the intrusion response mechanism is triggered and the detection results are pushed to the network security response terminal.

[0016] Compared with the prior art, the present invention has the following beneficial effects: This invention introduces a set of modal sensing factors as a perception indicator system to measure the dynamic state of different modalities within each time window. Combined with differentiated indicators such as communication load, call sequence entropy, and environmental variable stability, it can truly reflect the behavioral change trends of each modality in the temporal dimension. By aggregating and normalizing the modal sensing factors to form a modal sensing complex, a unified expression of modal states is achieved, significantly improving the accuracy of modal recognition and judgment, and having the ability to adapt to complex dynamic environments. A scoring function based on the dot-product attention mechanism dynamically calculates the importance score of each modal track within the current time window, strengthening the guiding role of key modal information and improving overall perceptual quality. A cross-modal dependency strength parameter is introduced. After normalizing the modal attention weights, the initial attention weights are fine-tuned based on the degree of intermodal correlation, effectively addressing the problem of under-representation of intermodal linkages in existing multimodal systems. A standardized encapsulation approach overcomes the inconsistent multimodal output structures and cumbersome processing issues of existing technologies, improving the system's usability and scalability. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 This is a flow chart of a network intrusion detection method based on multimodal deep learning according to the present invention; Figure 2 This is a schematic diagram of the structure of a network intrusion detection system based on multimodal deep learning of the present invention; Figure 3 This is a flow chart of the method for obtaining attack sub-paths provided by the present invention. DETAILED DESCRIPTION

[0018] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0019] Example 1 See also Figure 1 and Figure 3 As shown, Example 1 further illustrates the network intrusion detection method based on multimodal deep learning proposed by the present invention, including: With the increasing complexity of network environments and the continuous evolution of attack methods, network intrusion detection has become a core component of network security defense systems. Existing network intrusion detection methods primarily rely on analyzing multiple sources of information, such as network traffic, system logs, protocol behavior, and user operations, to identify potential malicious behavior or abnormal patterns. In recent years, with the advancement of deep learning technology, a growing number of studies have incorporated multimodal features into intrusion detection models, attempting to improve detection performance by integrating different data modalities (such as communication load, system call sequences, user behavior trajectories, and protocol triggering events).

[0020] However, existing multimodal network intrusion detection methods still have the following key technical bottlenecks: Existing methods generally treat various modal features as static, equally weighted inputs, lacking the ability to dynamically model and adjust for variations in the perceived strength of different modalities within different time windows, differences in their expressive power, and their correlation with attack events. This issue weakens the processing of certain key modalities during specific attack phases, making it difficult to fully exploit important features, impacting the detection accuracy and robustness of the model.

[0021] The multimodal data involved in network intrusion detection is often heterogeneous and high-dimensional. For example, protocol data is structured sequences, system logs are semi-structured text, and user behavior is time-series trajectories. These data span a wide range of dimensions and exhibit significant differences in expression. Existing methods often employ fixed-structure deep learning networks to unify modalities. These networks struggle to adapt to the heterogeneity and dynamic nature of these modalities, resulting in insufficient cross-modal feature correlation learning capabilities and limiting the accurate identification of complex intrusion behaviors.

[0022] Existing systems generally employ static fusion strategies during the multimodal fusion phase, lacking mechanisms to dynamically adjust the contribution of each modality based on temporal evolution. This strategy can easily lead to issues such as redundant information in strong modalities and missing features in weak modalities, making it impossible to accurately focus on modal features relevant to intrusion behavior. Furthermore, existing systems fail to fully consider the potential dependencies and complementarities between modalities in their fusion architecture, resulting in irrational attention allocation or a loose fusion structure, further weakening model performance.

[0023] In the process of modal feature processing, existing intrusion detection systems usually adopt a one-size-fits-all approach for different modalities. They lack the enhancement mechanism for key modalities and the compression mechanism for redundant modalities. This not only wastes computing resources, but also brings about problems such as increased training difficulty and reduced inference efficiency. This restricts the practical application of multimodal deep learning methods in network intrusion detection scenarios with high real-time requirements.

[0024] In order to effectively solve the above problems, the present invention proposes a network intrusion detection method based on multimodal deep learning, including: S1. Collect and assign modal attributes to multi-source data from different channels, construct a time-continuous multi-track dynamic perception track set, integrate the modal attention allocation mechanism, dynamically allocate modal parameters based on the modal energy complex of each modality under different perception tracks, and output the initial index tensor after multi-track modal fusion; S2: Receive the initial index tensor and establish a cross-modal dependency relationship by constructing a bidirectional response tensor between the modality and the track. Construct a semantic trigger matrix based on the semantic labels in the cross-modal dependency relationship. Adopt a cross-modal residual connection mechanism to preserve low-order modal coupling features through parallel residual paths and output modal linkage nodes. S3. Based on the modal linkage nodes output by the semantic trigger matrix, a multimodal behavior aggregation graph is constructed, highly clustered subgraphs and frequent path patterns appearing in the multimodal behavior aggregation graph are extracted, and an abnormal behavior graph structure annotated with abnormal confidence is generated; S4. Based on the abnormal behavior graph structure, through graph structure backtracking and pattern inversion analysis, we extract attack subpaths with time constraints and behavioral logic causal chains. We use the abnormal behavior causal chain backtracking mechanism to track the smallest traceable trigger unit of abnormal behavior, and generate the differential structure of the attack graph by combining it with the preset normal behavior template graph. S5. Mark the modal source traceability of the differential structure of the attack graph, detect the variation of the modal parameter structure, construct a multimodal time series perception vector group, and input it into the pre-built deep joint discriminator for network intrusion behavior detection; when modal drift or track structure abnormality is detected, trigger the intrusion response mechanism and push the detection results to the network security response terminal.

[0025] Methods for constructing a temporally continuous multi-track dynamic-aware track set include: Collect multi-source data from different channels, including network communication data, host call data, user operation behavior data, and external device access data; conduct real-time monitoring through preset data collection agents and listening ports, and encapsulate the collected multi-source data into a standardized data stream format with timestamps; Network communication data includes TCP / IP traffic, DNS resolution records, HTTP request information, network connection session start time and duration; host call data includes process creation and termination, file access logs, kernel logs, and security audit logs; User operation behavior data includes login and logout records (time, method, source address), user command history (such as Shell / Bash commands), and terminal remote access records; external device access data includes USB storage device plug-in and unplug events, external device file transfer records (transfer direction, file type, timestamp), and storage device mount and unmount records; A modal energy function is preset, which comprehensively considers data type, data structure and behavioral characteristics; based on the modal energy function, the collected multi-source data is classified into modal attribution, network modality, system modality, user modality and contextual modality; Based on the timestamp information of multi-source data, the multi-source data after modality attribution will construct different modal track units within the preset time window according to its timestamp, and connect the modal track units in chronological order to form a temporally continuous multi-track dynamic perception track set. The multi-track dynamic perception track set includes network perception track, system perception track, user perception track and context perception track.

[0026] Methods for obtaining the initial index tensor include: For each type of modality's corresponding perception track, the modal sensitivity factor set of the modality is extracted in each time window as the modality's state characterization indicators in that time period. The state characterization indicators include the communication load of the network modality, the abnormal protocol trigger rate, the call sequence entropy value of the system modality, the resource access density, the operation jump frequency of the user modality, and the environmental variable stability and device conversion rate of the context modality. The modal perception factor set is aggregated and normalized in time series to form a modal perception compound for perception, which represents the perception strength of the modality in the current perception track window; a weighted scoring function is used to evaluate the perception strength of each perception track in the preset time window. The modal sense composite within the model is scored to obtain the modal attention score; The weighted scoring function is: ;in, Indicates modality In the time window modality attention score of ; Indicates that in the time window The global query vector is obtained by dividing the time window The sensory energy composite sub-feature vectors of all sensory tracks in the system are summed dimension by dimension; Indicates modality In the time window The modal sense energy composite sub-eigenvector of ; Represents the dimension of the modal sense energy composite sub-feature vector, which is used to prevent the value after the vector inner product from being too large. It should be noted that the weighted scoring function is based on the dot product attention mechanism, which is used to calculate the attention score of each modal track within a specified time window. An attention allocation mechanism is introduced. Based on the modal energy complex of each perception track, the modal attention score is normalized and its attention weight in the current multi-track dynamic perception track set is calculated. The attention weight is tuned considering the cross-dependency probability between modalities and mapped to the modal parameter weight factor for attention allocation. The tuned attention weight is ;in, Indicates modality In the time window The tuned attention weights; Indicates modality In the time window The original attention weight of Represents the weight adjustment coefficient, which is used to balance the original attention weight and the proportion of dependent modal contribution. According to the expert experience method, The value range is from 0 to 1; Indicates modality With modal The strength of cross-dependence between Indicates modality In the time window The original attention weight of and The index of the modality; Based on the attention allocation results, a corresponding modal parameter weight factor is assigned to each perception track, a modal attention score threshold is preset, feature extraction and enhancement processing is performed on the modalities with a modal attention score greater than the preset modal attention score threshold, and dimensionality reduction operation is performed on the modalities with a modal attention score less than the preset modal attention score threshold; After completing the tuning and allocation of each perception track, the modal perception energy complexes of the multi-track dynamic perception track are weightedly fused based on the modal parameter weight factors, and the fused modal perception energy complexes are encapsulated as the initial index tensor in a unified format.

[0027] The following problems existing in the existing technology are solved: In the existing technology, multimodal perception usually regards different modalities as parallel input channels, and only performs simple splicing or average fusion, lacking detailed modeling and dynamic differentiation of the perception intensity and state performance of each modality in different time periods; the indicator dimensions and expression methods of different modalities are inconsistent, making it difficult to support unified cross-modal analysis; the existing system uses a static fusion strategy in multimodal information fusion, which cannot dynamically adjust the contribution of each modality over time windows, and easily leads to strong modal redundancy or weak modal neglect. The existing fusion strategy ignores the potential correlation or redundant dependence between modalities, resulting in irrational attention allocation or loose fusion structure; it is impossible to perform feature compression on weak modalities and enhance the extraction of strong modalities, resulting in model performance bottlenecks and resource waste; there is a lack of standardized and highly versatile output expression structure; Compared with existing technologies, the advantages are as follows: by introducing a set of modal sensing factors as a perception indicator system, the dynamic state of different modalities in each time window is measured. Combined with differentiated indicators such as communication load, call sequence entropy, and environmental variable stability, it can truly reflect the behavioral change trends of each modality in the time series dimension. By aggregating and normalizing the modal sensing factors to form a modal sensing complex, a unified expression of modal state is achieved, significantly improving the accuracy of modal recognition and judgment, and having the ability to adapt to complex dynamic environments; A scoring function based on the dot-product attention mechanism dynamically calculates the importance score of each modal track within the current time window, strengthening the guiding role of key modal information and improving overall perceptual quality. A cross-modal dependency strength parameter is introduced. After normalizing the modal attention weights, the initial attention weights are fine-tuned based on the degree of intermodal correlation, effectively addressing the problem of under-representation of intermodal linkages in existing multimodal systems. A standardized encapsulation approach overcomes the inconsistent multimodal output structures and cumbersome processing issues of existing technologies, improving the system's usability and scalability.

[0028] Methods for establishing cross-modal dependencies include: Based on the initial index tensor, a bidirectional response tensor is constructed between the modality and the perception track. The bidirectional response tensor includes the track-to-modality response matrix and the modality-to-track response matrix. The track feature vector corresponding to each perception track in each time window is obtained. The track feature vector is generated by feature aggregation of the different modal sensory energy composite sub-features carried by the track. By using cosine similarity to measure the matching degree between the orbital feature vector and the modal sensory complex, the orbital response weight to the modality is formed. Based on the sensory complex of each sensing orbit to different modalities within the time window, the orbital to modal response matrix is obtained, which represents the response degree of the orbital set to each modality. By normalizing the eigenvalues of each dimension in the modal sensory energy complex, multiplying the normalized eigenvalues with the corresponding modal parameter weight factors and summing them, a feature contribution vector is obtained. Combining the modal attention score and feature contribution vector of the modality in each time window, the modal-to-track response matrix is obtained, which represents the response preference of the modality to each perception track. The track-to-modal response matrix and the modal-to-track response matrix are fused into a bidirectional response tensor to characterize the interactive response relationship between the modality and the perception track. Based on the bidirectional response tensor, the similarity between the modal sensory energy composite sub-eigenvectors is calculated through the cosine similarity function to generate the inter-modal dependency strength matrix and construct a cross-modal dependency tensor to reflect the dependency strength between different modalities within the time window.

[0029] The construction methods of modal linkage nodes include: Based on the dependency strength in the cross-modal dependency tensor, the semantic labels of each modality are extracted. The semantic labels include the semantic categories of each modality (such as speech, image, text, etc.), contextual variables (such as "high noise", "nighttime", "high interaction density", etc.), and the relationship between modalities (such as "primary-auxiliary coupling", "multi-directional feedback", etc.). A set of semantic triggering rules is preset to determine whether, within any time window, if the semantic label of modality A appears in a certain state, it will trigger the semantic response of modality B. For example, if the image modality detects a "moving object", it may trigger the speech recognition modality to enter the "activated" state; if the temperature modality is marked as "high temperature alarm", it will trigger the text modality to generate prompt content. Construct a semantic trigger matrix. The semantic trigger matrix is a two-dimensional numerical matrix. Each row in the matrix represents a modality as a trigger source, and each column represents a modality as a response target. If the semantic response of modality B is triggered, a value representing the trigger strength is filled in at position A or B in the semantic trigger matrix, and 0 is filled in at the untriggered position. After obtaining the semantic trigger matrix, a cross-modal residual connection mechanism is introduced to quantify the degree of coupling between the modalities. The similarity between the sensory-energy composite sub-feature vectors of any two modalities is calculated using the cosine similarity function to obtain the modal coupling strength score, and a modal coupling strength score threshold is preset. The modal coupling features with a modal coupling strength score greater than a preset modal coupling strength score threshold are classified as high-order modal coupling features; the modal coupling features with a modal coupling strength score less than or equal to the preset modal coupling strength score threshold are classified as low-order modal coupling features; a parallel residual connection structure is adopted to retain the low-order modal coupling features in the parallel residual path and output the modal linkage node.

[0030] The method for obtaining the abnormal behavior graph structure includes: A multimodal behavior aggregation graph is constructed with modal linkage nodes as vertices. The multimodal behavior aggregation graph is a directed graph structure. The edges of the graph represent the semantic triggering relationship between modalities, and the edge weights represent the triggering strength. The actual number of edges in the subgraph of the multimodal behavior aggregation graph is divided by the preset maximum number of edges to obtain the subgraph node density. A subgraph node density threshold is preset, and subgraphs with a density greater than the preset subgraph node density threshold are defined as highly clustered subgraphs; highly clustered subgraphs and frequent path patterns are extracted. Frequent path patterns include the time series of modal triggering and its occurrence frequency in different time windows; The modal linkage structure of the current time window is compared with the historical highly clustered subgraphs and frequent path patterns to calculate the structural deviation. When the structural deviation is greater than the preset structural deviation threshold, it is marked as an abnormal behavior candidate. A weighted combination function is constructed based on the temporal sequence of the structural deviation and the abnormal behavior to output the abnormality confidence score. Based on the abnormal behavior candidates and abnormal confidence scores, an abnormal behavior graph structure marked with abnormal confidence is generated; the abnormal behavior graph structure contains abnormal path sequence, involved modal identification, time window information, triggering edge attributes and abnormal confidence value.

[0031] Methods for obtaining attack subpaths include: Timestamp the modal nodes and trigger edges in the abnormal behavior graph structure to construct a behavior graph with time-series attributes. Backtrack the graph structure from the end point of the abnormal path in the behavior graph, limit the maximum time span and maximum path depth, and extract the backtracking path that meets the semantic trigger logic and time sequence. The semantic trigger logic includes the existence of the trigger edge and the trigger strength of the edge greater than the preset trigger strength threshold. The backtracking path is structurally and semantically matched with the preset attack behavior pattern library to identify candidate attack sub-paths; the candidate attack sub-paths are judged for temporal consistency, causal logical coherence, and abnormal confidence aggregation, and the attack sub-paths that meet the conditions are extracted.

[0032] It should be noted that temporal consistency means that the time difference between all adjacent nodes on the candidate attack sub-path is less than the preset time difference threshold; causal logical coherence means that the connection between events in the candidate attack sub-path must exist in the semantic trigger matrix, and the corresponding trigger strength must be higher than the preset trigger strength threshold; anomaly confidence aggregation means that the average anomaly confidence score of the candidate attack sub-path is greater than the preset average anomaly confidence score threshold.

[0033] Methods for obtaining the differential structure of the attack graph include: Based on the extracted attack subpath, the system continues to backtrack in the abnormal behavior graph structure, limiting the maximum backtracking depth and time span threshold, and recursively traces all reachable abnormal confidence nodes and trigger relationship edges to form an abnormal behavior causal chain graph. In the abnormal behavior causal chain graph, the system identifies the minimum connected subgraph structure that affects any node in the attack subpath as the minimum traceable trigger unit. Define the constraints that the minimum traceable trigger unit must satisfy. These constraints include: the minimum connected subgraph consisting of the nodes and edges contained in the minimum traceable trigger unit must be able to reach the preset target attack node in time sequence; the trigger strength of each edge must be greater than the preset trigger strength threshold; the anomaly confidence score must be greater than the preset anomaly confidence score threshold; and deleting any node or edge in the unit will make the attack subpath unreconstructable. The structural differences between the abnormal behavior causal chain graph and the preset normal behavior template graph are compared to generate the attack graph differential structure; the attack graph differential structure includes a node difference set, an edge difference set, a trigger strength difference set, and an anomaly confidence score set.

[0034] The node difference set represents the nodes that exist in the abnormal behavior causal chain graph but do not exist in the preset normal behavior template graph; the edge difference set represents the trigger relationship that exists in the abnormal behavior causal chain graph but does not exist in the preset normal behavior template graph; the strength difference set represents the difference in trigger strength between the abnormal edge and the corresponding edge in the preset normal behavior template graph that exceeds the preset edge trigger strength difference threshold; the abnormal confidence score set is the clustered subgraph formed by the continuous high confidence areas in the attack path.

[0035] Methods for network intrusion behavior detection include: The node difference set and edge difference set in the differential structure of the attack graph are marked with modal source traceability to identify the modal source corresponding to each node or edge, and the trigger link information in the attack sub-path is retained; The modal parameter structure of each traceably marked node is extracted and compared with the modal parameter structure of the corresponding node in the preset normal behavior template diagram to determine whether there is modal parameter structure variation. When the modal parameter structure difference is greater than the preset modal parameter structure difference threshold, it is marked as modal parameter structure variation. Combining the timestamps, modal identifications, trigger strengths, and anomaly confidence scores of the nodes in the differential structure of the attack graph, a corresponding multimodal time-series perception vector group is constructed. The multimodal time-series perception vector group is input into a pre-built deep joint discriminant model to detect network intrusion behavior and output the identification results of the network intrusion behavior.

[0036] The preset modal attention score threshold is set by the staff. By collecting different modal attention scores, the average of multiple modal attention scores is taken as the preset modal attention score threshold; similarly, the preset modal coupling strength score threshold, the preset subgraph node density threshold, the preset structure deviation threshold, the preset trigger strength threshold, the time span threshold, the preset anomaly confidence score threshold and the preset modal parameter structure difference threshold are set.

[0037] This embodiment introduces a set of modal sensing factors as a perception indicator system to measure the dynamic state of different modalities within each time window. Combined with differentiated indicators such as communication load, call sequence entropy, and environmental variable stability, it can truly reflect the behavioral change trends of each modality in the time series dimension. By aggregating and normalizing the modal sensing factors to form a modal sensing compound, a unified expression of the modal state is achieved, significantly improving the accuracy of modal recognition and judgment, and having the ability to adapt to complex dynamic environments; A scoring function based on the dot-product attention mechanism dynamically calculates the importance score of each modal track within the current time window, strengthening the guiding role of key modal information and improving overall perceptual quality. A cross-modal dependency strength parameter is introduced. After normalizing the modal attention weights, the initial attention weights are fine-tuned based on the degree of intermodal correlation, effectively addressing the problem of under-representation of intermodal linkages in existing multimodal systems. A standardized encapsulation approach overcomes the inconsistent multimodal output structures and cumbersome processing issues of existing technologies, improving the system's usability and scalability.

[0038] Example 2 See also Figure 2 As shown, the network intrusion detection system based on multimodal deep learning in this embodiment includes: The multi-track dynamic construction module collects and modally attributes multi-source data from different channels, constructs a temporally continuous multi-track dynamic perception track set, integrates the modal attention allocation mechanism, dynamically allocates modal parameters based on the modal energy complex of each modality in different perception tracks, and outputs the initial index tensor after multi-track modal fusion; The modal-track cross module receives the initial index tensor and establishes a cross-modal dependency relationship by constructing a bidirectional response tensor between the modality and the track. It builds a semantic trigger matrix based on the semantic labels in the cross-modal dependency relationship, adopts a cross-modal residual connection mechanism, preserves low-order modal coupling features through parallel residual paths, and outputs modal linkage nodes. The linkage graph mining module constructs a multimodal behavior aggregation graph based on the modal linkage nodes output by the semantic trigger matrix, extracts highly clustered subgraphs and frequent path patterns appearing in the multimodal behavior aggregation graph, and generates an abnormal behavior graph structure annotated with anomaly confidence. The attack trace backtracking module, based on the abnormal behavior graph structure, extracts attack subpaths with time constraints and behavioral logic causal chains through graph structure backtracking and pattern inversion analysis. It uses the abnormal behavior causal chain backtracking mechanism to track the smallest traceable trigger unit of abnormal behavior and generates the differential structure of the attack graph by combining it with the preset normal behavior template graph. The intrusion behavior detection module performs modal source traceability marking on the differential structure of the attack graph, detects the variation of the modal parameter structure, constructs a multimodal time series perception vector group, and inputs it into the pre-built deep joint discriminator for network intrusion behavior detection; when modal drift or track structure abnormality is detected, the intrusion response mechanism is triggered and the detection results are pushed to the network security response terminal.

[0039] Since the electronic device introduced in this embodiment is an electronic device used to implement the network intrusion detection method and system based on multimodal deep learning in the embodiments of this application, based on the network intrusion detection method and system based on multimodal deep learning introduced in the embodiments of this application, those skilled in the art can understand the specific implementation of the electronic device of this embodiment and its various variations, so how the electronic device implements the method in the embodiments of this application will not be described in detail here. As long as those skilled in the art implement the electronic device used by the network intrusion detection method and system based on multimodal deep learning in the embodiments of this application, they all fall within the scope of protection of this application.

[0040] The above formulas are all dimensionless and numerical calculations. The formulas are obtained by collecting a large amount of data and performing software simulation to obtain the most recent real situation. The preset parameters and thresholds in the formulas are set by technicians in this field according to actual conditions.

[0041] The above description is merely a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiment. All technical solutions based on the principles of the present invention are within the scope of protection of the present invention. It should be noted that for users of ordinary skill in the art, various improvements and modifications that do not depart from the principles of the present invention should also be considered within the scope of protection of the present invention.

Claims

1. A network intrusion detection method based on multimodal deep learning, characterized in that: include: S1. Collect and assign modal attributes to multi-source data from different channels, construct a time-continuous multi-track dynamic perception track set, integrate the modal attention allocation mechanism, dynamically allocate modal parameters based on the modal energy complex of each modality under different perception tracks, and output the initial index tensor after multi-track modal fusion; S2, receiving the initial index tensor, establishes cross-modal dependency by constructing a bidirectional response tensor between the modality and the track; A semantic trigger matrix is constructed based on the semantic labels in the cross-modal dependency relationship. A cross-modal residual connection mechanism is adopted to preserve low-order modal coupling features through parallel residual paths and output modal linkage nodes. S3. Based on the modal linkage nodes output by the semantic trigger matrix, a multimodal behavior aggregation graph is constructed, highly clustered subgraphs and frequent path patterns appearing in the multimodal behavior aggregation graph are extracted, and an abnormal behavior graph structure annotated with abnormal confidence is generated; S4. Based on the abnormal behavior graph structure, through graph structure backtracking and pattern inversion analysis, we extract attack subpaths with time constraints and behavioral logic causal chains. We use the abnormal behavior causal chain backtracking mechanism to track the smallest traceable trigger unit of abnormal behavior, and generate the differential structure of the attack graph by combining it with the preset normal behavior template graph. S5. Mark the modal source traceability of the differential structure of the attack graph, detect the variation of the modal parameter structure, construct a multimodal time series perception vector group, and input it into the pre-built deep joint discriminator for network intrusion behavior detection; when modal drift or track structure abnormality is detected, trigger the intrusion response mechanism and push the detection results to the network security response terminal.

2. The network intrusion detection method based on multimodal deep learning according to claim 1 is characterized in that: The method for constructing a temporally continuous multi-track dynamic perception track set includes: Collect multi-source data from different channels, including network communication data, host call data, user operation behavior data, and external device access data; conduct real-time monitoring through preset data collection agents and listening ports, and encapsulate the collected multi-source data into a standardized data stream format with timestamps; A modal energy function is preset, which comprehensively considers data type, data structure and behavioral characteristics; based on the modal energy function, the collected multi-source data is classified into modal attribution, network modality, system modality, user modality and contextual modality; Based on the timestamp information of multi-source data, the multi-source data after modality attribution will construct different modal track units within the preset time window according to its timestamp, and connect the modal track units in chronological order to form a temporally continuous multi-track dynamic perception track set. The multi-track dynamic perception track set includes network perception track, system perception track, user perception track and context perception track.

3. The network intrusion detection method based on multimodal deep learning according to claim 2 is characterized in that: The method for obtaining the initial index tensor includes: For each type of modality's corresponding perception track, the modal sensitivity factor set of the modality is extracted in each time window as the modality's state characterization indicators in that time period. The state characterization indicators include the communication load of the network modality, the abnormal protocol trigger rate, the call sequence entropy value of the system modality, the resource access density, the operation jump frequency of the user modality, and the environmental variable stability and device conversion rate of the context modality. The modal sensory factor set is aggregated and normalized in time series to form a modal sensory complex for perception; a weighted scoring function is used to evaluate each perception track in a preset time window. The modal sense composite within the model is scored to obtain the modal attention score; An attention allocation mechanism is introduced. Based on the modal energy complex of each perception track, the modal attention score is normalized and its attention weight in the current multi-track dynamic perception track set is calculated. The attention weight is tuned considering the cross-dependency probability between modalities and mapped to the modal parameter weight factor for attention allocation. Based on the attention allocation results, a corresponding modal parameter weight factor is assigned to each perception track, a modal attention score threshold is preset, feature extraction and enhancement processing is performed on the modalities with a modal attention score greater than the preset modal attention score threshold, and dimensionality reduction operation is performed on the modalities with a modal attention score less than the preset modal attention score threshold; After completing the tuning and allocation of each perception track, the modal perception energy complexes of the multi-track dynamic perception track are weightedly fused based on the modal parameter weight factors, and the fused modal perception energy complexes are encapsulated as the initial index tensor in a unified format.

4. The network intrusion detection method based on multimodal deep learning according to claim 3 is characterized in that: The method for establishing a cross-modal dependency relationship includes: Based on the initial index tensor, a bidirectional response tensor is constructed between the modality and the perception track. The bidirectional response tensor includes the track-to-modality response matrix and the modality-to-track response matrix. The track feature vector corresponding to each perception track in each time window is obtained. The track feature vector is generated by feature aggregation of the different modal sensory energy composite sub-features carried by the track. By using cosine similarity to measure the matching degree between the orbital feature vector and the modal sensory complex, the orbital response weight to the modality is formed; based on the sensory complex of each sensing orbit to different modalities within the time window, the orbital to modal response matrix is obtained; By normalizing the eigenvalues of each dimension in the modal energy complex, multiplying the normalized eigenvalues with the corresponding modal parameter weight factors and summing them, a feature contribution vector is obtained. The modal attention score and feature contribution vector of the modal in each time window are combined to obtain the mode-to-orbit response matrix. The track-to-modal response matrix and the modal-to-track response matrix are fused into a bidirectional response tensor to characterize the interactive response relationship between the modality and the perception track. Based on the bidirectional response tensor, the similarity between the modal sensory energy composite sub-eigenvectors is calculated through the cosine similarity function to generate the inter-modal dependency strength matrix and construct a cross-modal dependency tensor to reflect the dependency strength between different modalities within the time window.

5. The network intrusion detection method based on multimodal deep learning according to claim 4 is characterized in that: The method for constructing the modal linkage node includes: Based on the dependency strength in the cross-modal dependency tensor, the semantic labels of each modality are extracted. The semantic labels include the semantic categories of each modality, contextual variables, and the relationship between modalities. A set of semantic triggering rules is preset to determine whether, within any time window, if a state of the semantic label of modality A appears, it will trigger a semantic response in modality B. Construct a semantic trigger matrix. The semantic trigger matrix is a two-dimensional numerical matrix. Each row in the matrix represents a modality as a trigger source, and each column represents a modality as a response target. If the semantic response of modality B is triggered, a value representing the trigger strength is filled in at position A or B in the semantic trigger matrix, and 0 is filled in at the untriggered position. After obtaining the semantic trigger matrix, a cross-modal residual connection mechanism is introduced to quantify the degree of coupling between the modalities. The similarity between the sensory-energy composite sub-feature vectors of any two modalities is calculated using the cosine similarity function to obtain the modal coupling strength score, and a modal coupling strength score threshold is preset. The modal coupling features with a modal coupling strength score greater than a preset modal coupling strength score threshold are classified as high-order modal coupling features; the modal coupling features with a modal coupling strength score less than or equal to the preset modal coupling strength score threshold are classified as low-order modal coupling features; a parallel residual connection structure is adopted to retain the low-order modal coupling features in the parallel residual path and output the modal linkage node.

6. The network intrusion detection method based on multimodal deep learning according to claim 5, characterized in that: The method for obtaining the abnormal behavior graph structure includes: A multimodal behavior aggregation graph is constructed with modal linkage nodes as vertices. The multimodal behavior aggregation graph is a directed graph structure. The edges of the graph represent the semantic triggering relationship between modalities, and the edge weights represent the triggering strength. The actual number of edges in the subgraph of the multimodal behavior aggregation graph is divided by the preset maximum number of edges to obtain the subgraph node density. A subgraph node density threshold is preset, and subgraphs with a density greater than the preset subgraph node density threshold are defined as highly clustered subgraphs; highly clustered subgraphs and frequent path patterns are extracted. Frequent path patterns include the time series of modal triggering and its occurrence frequency in different time windows; The modal linkage structure of the current time window is compared with the historical highly clustered subgraphs and frequent path patterns to calculate the structural deviation. When the structural deviation is greater than the preset structural deviation threshold, it is marked as an abnormal behavior candidate. A weighted combination function is constructed based on the temporal sequence of the structural deviation and the abnormal behavior to output the abnormality confidence score. Based on the abnormal behavior candidates and abnormal confidence scores, an abnormal behavior graph structure marked with abnormal confidence is generated; the abnormal behavior graph structure contains abnormal path sequence, involved modal identification, time window information, triggering edge attributes and abnormal confidence value.

7. The network intrusion detection method based on multimodal deep learning according to claim 6, characterized in that: The method for obtaining the attack subpath includes: Timestamp the modal nodes and trigger edges in the abnormal behavior graph structure to construct a behavior graph with time-series attributes. Backtrack the graph structure from the end point of the abnormal path in the behavior graph, limit the maximum time span and maximum path depth, and extract the backtracking path that meets the semantic trigger logic and time sequence. The semantic trigger logic includes the existence of the trigger edge and the trigger strength of the edge greater than the preset trigger strength threshold. The backtracking path is structurally and semantically matched with the preset attack behavior pattern library to identify candidate attack sub-paths; the candidate attack sub-paths are judged for temporal consistency, causal logical coherence, and abnormal confidence aggregation, and the attack sub-paths that meet the conditions are extracted.

8. The network intrusion detection method based on multimodal deep learning according to claim 7, characterized in that: The method for obtaining the differential structure of the attack graph includes: Based on the extracted attack subpath, the system continues to backtrack in the abnormal behavior graph structure, limiting the maximum backtracking depth and time span threshold, and recursively traces all reachable abnormal confidence nodes and trigger relationship edges to form an abnormal behavior causal chain graph. In the abnormal behavior causal chain graph, the system identifies the minimum connected subgraph structure that affects any node in the attack subpath as the minimum traceable trigger unit. Define the constraints that the minimum traceable trigger unit must satisfy. These constraints include: the minimum connected subgraph consisting of the nodes and edges contained in the minimum traceable trigger unit must be able to reach the preset target attack node in time sequence; the trigger strength of each edge must be greater than the preset trigger strength threshold; the anomaly confidence score must be greater than the preset anomaly confidence score threshold; and deleting any node or edge in the unit will make the attack subpath unreconstructable. The structural differences between the abnormal behavior causal chain graph and the preset normal behavior template graph are compared to generate the attack graph differential structure; the attack graph differential structure includes a node difference set, an edge difference set, a trigger strength difference set, and an anomaly confidence score set.

9. The network intrusion detection method based on multimodal deep learning according to claim 8, characterized in that: The method for detecting network intrusion behavior includes: The node difference set and edge difference set in the differential structure of the attack graph are marked with modal source traceability to identify the modal source corresponding to each node or edge, and the trigger link information in the attack sub-path is retained; The modal parameter structure of each traceably marked node is extracted and compared with the modal parameter structure of the corresponding node in the preset normal behavior template diagram to determine whether there is modal parameter structure variation. When the modal parameter structure difference is greater than the preset modal parameter structure difference threshold, it is marked as modal parameter structure variation. Combining the timestamps, modal identifications, trigger strengths, and anomaly confidence scores of the nodes in the differential structure of the attack graph, a corresponding multimodal time-series perception vector group is constructed. The multimodal time-series perception vector group is input into a pre-built deep joint discriminant model to detect network intrusion behavior and output the identification results of the network intrusion behavior.

10. A network intrusion detection system based on multimodal deep learning, used to implement the network intrusion detection method based on multimodal deep learning according to any one of claims 1 to 9, characterized in that: include: The multi-track dynamic construction module collects and modally attributes multi-source data from different channels, constructs a temporally continuous multi-track dynamic perception track set, integrates the modal attention allocation mechanism, dynamically allocates modal parameters based on the modal energy complex of each modality in different perception tracks, and outputs the initial index tensor after multi-track modal fusion; The modality-track cross module receives the initial index tensor and establishes cross-modality dependencies by constructing a bidirectional response tensor between the modality and the track; A semantic trigger matrix is constructed based on the semantic labels in the cross-modal dependency relationship. A cross-modal residual connection mechanism is adopted to preserve low-order modal coupling features through parallel residual paths and output modal linkage nodes. The linkage graph mining module constructs a multimodal behavior aggregation graph based on the modal linkage nodes output by the semantic trigger matrix, extracts highly clustered subgraphs and frequent path patterns appearing in the multimodal behavior aggregation graph, and generates an abnormal behavior graph structure annotated with anomaly confidence. The attack trace backtracking module, based on the abnormal behavior graph structure, extracts attack subpaths with time constraints and behavioral logic causal chains through graph structure backtracking and pattern inversion analysis. It uses the abnormal behavior causal chain backtracking mechanism to track the smallest traceable trigger unit of abnormal behavior and generates the differential structure of the attack graph by combining it with the preset normal behavior template graph. The intrusion behavior detection module performs modal source traceability marking on the differential structure of the attack graph, detects the variation of the modal parameter structure, constructs a multimodal time series perception vector group, and inputs it into the pre-built deep joint discriminator for network intrusion behavior detection; when modal drift or track structure abnormality is detected, the intrusion response mechanism is triggered and the detection results are pushed to the network security response terminal.

Citation Information

Patent Citations

  • Multi-modal emotion recognition method based on situational attention neural network

    CN112348075A

  • Unsupervised cross-modal retrieval method based on attention mechanism enhancement

    CN113971209A

  • Intrusion detection and response method and system of satellite internet target range

    CN119155101A

  • Network security method and system based on block chain

    CN119210889A

  • Pumped storage power station construction anomaly detection method and system based on unmanned aerial vehicle image analysis

    CN119888507A

Cited By

  • Computer network intrusion detection system and method based on abnormal behavior analysis

    CN120896779A

  • Head-up display multi-mode abnormity test method, system, device and medium

    CN121051700A