An Intrusion Detection Model and Method Based on Historical Data and GCN

Through an intrusion detection model based on historical data and GCN, the multi-view space-time module is used to fusion time and spatial information, and the problem of insufficient detection capabilities for long sequences in the existing technology is solved, and accurate intrusion detection of the industrial Internet is realized.

CN119420517BActive Publication Date: 2025-07-18INNER MONGOLIA UNIV OF TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411474996.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-22
Publication Date
2025-07-18
Estimated Expiration
2044-10-22

AI Technical Summary

Technical Problem

Existing intrusion detection methods cannot effectively utilize spatial dependencies formed by different data forms, resulting in insufficient processing capabilities for long sequences and low detection accuracy.

Method used

Using an intrusion detection model based on historical data and GCN, multi-view timing data is input through the input module, the embedding module generates embedded results, and information is fused through multiple multi-view space-time modules, and finally the prediction results are generated in the output module, and the interaction between time and space views is used for detection.

Benefits of technology

Accurate detection of long sequence logs and traffic data is realized, improving the accuracy and effectiveness of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119420517B_ABST
    Figure CN119420517B_ABST
Patent Text Reader

Abstract

The present application relates to the field of intrusion detection technology for industrial Internet, and discloses an intrusion detection model and method based on historical data and GCN. By feeding multi-view time-series data into an embedding module and outputting an embedding result, and feeding the embedding result and multiple image data into a multi-view spatio-temporal module, where the number of multi-view spatio-temporal modules is arbitrary, and the input of each multi-view spatio-temporal module is the output of the previous multi-view spatio-temporal module and multiple image data, which can prevent the over-smoothing of the graph neural network. After being processed by the multi-view spatio-temporal module, multiple prediction information is obtained. The multiple prediction information is connected together and input into an output module to output a prediction result. By mining time and space information from multiple views, the spatial dependence relationships formed by different data can be cross-utilized, and the interaction between time and space views can be captured, so as to achieve accurate detection of long-sequence log and traffic data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of intrusion detection in industrial Internet, and particularly relates to an intrusion detection model and method based on historical data and GCN. Background Art

[0002] The application of digital technology enables the industrial Internet to achieve comprehensive perception and highly intelligent operation, strengthening the flexible coordination and interconnection among the source, network, load, and storage links. However, it also brings network security risks to the industrial Internet, impacting the existing technical architecture and security protection system.

[0003] Currently, there are still deficiencies in the network risk perception technology for industrial Internet. Existing intrusion detection methods mainly include rule-based, sequence-based, and graph matching detection methods. The rule-based detection method mainly summarizes the rules according to malicious behaviors, such as a process with low permissions accessing a process with high permissions, downloading files containing insecure information, etc., and tags the log data, which can effectively identify attack behaviors. The sequence-based detection method divides the log data into a series of log behavior data through sessions, static time windows, or dynamic time windows, and then identifies attack behaviors through unsupervised or supervised learning methods. The graph matching detection method constructs malicious behaviors into a graph representation, and then compares the graph nodes and structures with the traceability graph of the log data. If it exceeds a certain threshold, it is an attack behavior. Although the above methods can all detect intrusion methods, the existing detection methods cannot utilize the spatial dependence relationships formed by different data forms. Therefore, although the short-term prediction effect is good, the processing ability for long sequences is insufficient, and the detection accuracy is relatively low. Summary of the Invention

[0004] This application provides an intrusion detection model based on historical data and GCN, aiming to solve the technical problem that the detection methods in the existing technology cannot utilize the spatial dependence relationships formed by different data forms, resulting in insufficient processing ability for long sequences and relatively low detection accuracy.

[0005] This application provides an intrusion detection model based on historical data and GCN, including:

[0006] An input module for inputting multi-view time-series data, where the multi-view time-series data includes historical data at multiple granularities and multiple image data;

[0007] An embedding module for obtaining the multi-view time-series data, generating multiple embedding results according to the multi-view time-series data, and feeding the multiple embedding results and multiple image data to a multi-view spatio-temporal module;

[0008] Multiple multi-view spatio-temporal modules, including a first multi-view spatio-temporal module, a second multi-view spatio-temporal module, and a third multi-view spatio-temporal module, where:

[0009] The first multi-view spatio-temporal module is used to output corresponding multiple temporal information and multiple spatial dependency information according to multiple embedding results and multiple image data, fuse the multiple temporal information and multiple spatial dependency information to obtain first fusion information, and feed the first fusion information, the first prediction information, and the multiple image data to the second multi-view spatio-temporal module;

[0010] The second multi-view spatio-temporal module is used to generate second prediction information according to the first prediction information and the multiple image data, and feed the second prediction information and the multiple image data to the third multi-view spatio-temporal module;

[0011] The third multi-view spatio-temporal module is used to generate corresponding third prediction information according to the second prediction information and the multiple image data;

[0012] The output module is used to generate a prediction result according to the first prediction information, the second prediction information, and the third prediction information generated by the multiple multi-view spatio-temporal modules.

[0013] Preferably, the historical data of multiple granularities includes recent historical data, daily historical data, and weekly historical data. The historical data includes session data, time window time series data divided according to a time window, and neighborhood time series data divided according to a neighborhood. The multiple image data includes multiple groups of images, and each group of images includes an interaction event graph, a time window graph, and a neighbor information graph.

[0014] Preferably, the embedding module includes a first embedding block, a second embedding block, and a third embedding block, where:

[0015] The first embedding block is used to obtain recent historical data and generate a recent embedding result according to the recent historical data and feed it to the multi-view spatio-temporal module;

[0016] The second embedding block is used to obtain daily historical data and generate a daily embedding result according to the daily historical data and feed it to the multi-view spatio-temporal module;

[0017] The third embedding block is used to obtain weekly historical data and generate a weekly embedding result according to the weekly historical data and feed it to the multi-view spatio-temporal module.

[0018] Preferably, each multi-view spatio-temporal module includes a multi-view time learning unit and a first view intelligent fusion unit, where:

[0019] The multi-view time learning unit is used to generate multiple recent temporal information, multiple daily temporal information, and multiple weekly temporal information based on multiple recent embedding results, multiple daily embedding results, and multiple weekly embedding results respectively;

[0020] The first view intelligent fusion unit is used to generate multiple optimal fusion ratios based on multiple recent temporal information, multiple daily temporal information, and multiple weekly temporal information to integrate multiple recent temporal information, multiple daily temporal information, and multiple weekly temporal information, and obtain multiple temporal information.

[0021] Preferably, each multi-view spatio-temporal module includes a multi-view space learning unit and a second view intelligent fusion unit, where:

[0022] The multi-view space learning unit is used to mine image relationship information according to the interaction event graph, time window graph, and neighbor information graph in each group of images, generate corresponding three mining results, and output the three mining results to the next multi-view fusion module;

[0023] The second view intelligent fusion unit is used to generate corresponding attention coefficients according to multiple temporal information, and perform visual fusion operations on the mining results generated by the multi-view learning unit according to the attention coefficients to obtain fusion information.

[0024] Preferably, the second view intelligent fusion unit is used to map multiple temporal information to a key subspace as multiple nodes, and calculate the attention scores of each node in each layer in the key subspace, where the calculation formula is:

[0025]

[0026] where i represents the number of nodes in the key subspace, where i = 1, 2, 3... represents the attention score of the i-th node in the l-th layer in the key subspace, represents the mining result of the i-th node in the l-th layer of the input in the key subspace, W represents the trainable weight matrix, and u represents the query vector;

[0027] It is also used to calculate the attention coefficients of each node in each layer in the key subspace according to the attention scores, where the calculation formula is:

[0028]

[0029] where i represents the number of nodes in the key subspace, where i = 1, 2, 3... represents the attention coefficient of the i-th node in the l-th layer, SoftMax l represents the softmax function, which is used to normalize the attention weights of all nodes in the l-th layer into a probability distribution;

[0030] It is also used to perform a visual fusion operation on the mining results according to the attention coefficient to obtain fusion information, where the formula corresponding to the visual fusion operation is:

[0031]

[0032] where i and P both represent the numbers of nodes in the key subspace, where i = 1, 2, 3... P; represents the fusion information of the i-th node in the k-th layer, represents the attention coefficient of the i-th node in the l-th layer, represents the mining result of the i-th node in the l-th layer input in the key subspace.

[0033] Preferably, the first embedding block, the second embedding block, and the third embedding block all include a first fully connected layer, an activation function, batch normalization, and a second fully connected layer.

[0034] Preferably, the multi-view learning unit includes two graph convolutional network layers connected by a ReLU activation function and batch normalization. Each graph convolutional network layer is used to obtain neighbor information, perform edge transfer on the neighbor information, multiply the neighbor information by the edge weight and aggregate it to obtain the aggregated neighbor information, and perform a linear transformation on the aggregated neighbor information and the node information corresponding to the interaction event graph, the time window graph, and the neighbor information graph respectively and add the linear transformation results to obtain the corresponding three mining results.

[0035] This application also provides an intrusion detection method based on historical data and GCN, including:

[0036] Input multi-view time-series data, where the multi-view time-series data includes historical data of multiple granularities and multiple image data;

[0037] Obtain the multi-view time-series data and generate multiple embedding results according to the multi-view time-series data;

[0038] Output corresponding multiple temporal information and multiple spatial dependence relationship information according to the multiple embedding results and the multiple image data, and fuse the multiple temporal information and the multiple spatial dependence relationship information to obtain the first fusion information;

[0039] Generate second prediction information according to the first prediction information and the multiple image data;

[0040] Generate corresponding third prediction information according to the second prediction information and the multiple image data;

[0041] Generate a prediction result according to the first prediction information, the second prediction information, and the third prediction information.

[0042] Preferably, the historical data of multiple granularities includes recent historical data, daily historical data, and weekly historical data. The historical data includes session data, time window time series data divided according to a time window, and neighborhood time series data divided according to a neighborhood. The multiple image data includes multiple groups of images, and each group of images includes an interaction event graph, a time window graph, and a neighbor information graph.

[0043] Preferably, the step of generating multiple embedding results according to the multi-view time series data includes:

[0044] Obtain recent historical data and generate a recent embedding result according to the recent historical data;

[0045] Obtain daily historical data and generate a daily embedding result according to the daily historical data;

[0046] Obtain weekly historical data and generate a weekly embedding result according to the weekly historical data.

[0047] The beneficial effects of this application are as follows: By feeding the multi-view time series data into the embedding module, the embedding module outputs the embedding result. The embedding result and multiple image data are fed into the multi-view spatio-temporal module. Any number of multi-view spatio-temporal modules can be set according to requirements. The input of each multi-view spatio-temporal module is the output of the previous multi-view spatio-temporal module and multiple image data. This can prevent the over-smoothing of the graph neural network. After being processed by multiple multi-view spatio-temporal modules, multiple prediction information is obtained. The multiple prediction information is connected together and input into the output module to output the prediction result. By setting multiple multi-view spatio-temporal modules, time and space information can be mined from multiple views, the spatial dependence relationships formed by different data can be cross-utilized, and the interaction between the time and space views can be captured. In this way, accurate detection of long-sequence log and traffic data can be achieved. Description of the Drawings

[0048] Figure 1 It is a schematic structural diagram of an intrusion detection model based on historical data and GCN according to an embodiment of this application.

[0049] Figure 2 It is a schematic flow diagram of an intrusion detection method based on historical data and GCN according to an embodiment of this application.

[0050] The realization, functional characteristics, and advantages of the purpose of this application will be further described with reference to the embodiments and the accompanying drawings. Detailed Embodiments

[0051] It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.

[0052] Such as Figure 1 、Figure 2 As shown, this application provides an intrusion detection model based on historical data and GCN, including:

[0053] An input module for inputting multi-view time-series data, where the multi-view time-series data includes historical data at multiple granularities and multiple image data;

[0054] An embedding module for obtaining the multi-view time-series data, generating multiple embedding results according to the multi-view time-series data, and feeding the multiple embedding results and multiple image data to a multi-view spatio-temporal module;

[0055] Multiple multi-view spatio-temporal modules, including a first multi-view spatio-temporal module, a second multi-view spatio-temporal module, and a third multi-view spatio-temporal module, where:

[0056] The first multi-view spatio-temporal module is used to output corresponding multiple temporal information and multiple spatial dependency information according to the multiple embedding results and multiple image data, fuse the multiple temporal information and multiple spatial dependency information to obtain first fusion information, and feed the first fusion information as first prediction information and multiple image data to the second multi-view spatio-temporal module;

[0057] The second multi-view spatio-temporal module is used to generate second prediction information according to the first prediction information and multiple image data, and feed the second prediction information and multiple image data to the third multi-view spatio-temporal module;

[0058] The third multi-view spatio-temporal module is used to generate corresponding third prediction information according to the second prediction information and multiple image data;

[0059] An output module for generating a prediction result according to the first prediction information, second prediction information, and third prediction information generated by the multiple multi-view spatio-temporal modules.

[0060] As described above, the present application considers multi-view time-series data from both temporal and spatial perspectives. Since multiple temporal dependencies include two aspects, trend and periodicity, which respectively reflect the recent trend and periodic influence, in order to describe multiple temporal dependencies from different perspectives, the present application considers historical data with different granularities, where the granularity refers to the time interval, and also considers multi-view spatial data. By feeding the multi-view time-series data into the embedding module, the embedding result is output by the embedding module, and the embedding result and multiple image data are fed into the multi-view spatio-temporal module. Any number of multi-view spatio-temporal modules can be set according to requirements. The input of each multi-view spatio-temporal module is the output of the previous multi-view spatio-temporal module and multiple image data, which can prevent the over-smoothing of the graph neural network. After being processed by multiple multi-view spatio-temporal modules, multiple prediction information is obtained, and the multiple prediction information is concatenated and input into the output module to output the prediction result. The present application proposes a multi-spatio-temporal graph convolution framework, which fully utilizes the multi-view spatio-temporal dependencies and their interactions. By setting multiple multi-view spatio-temporal modules, temporal and spatial information can be mined from multiple views, the spatial dependencies formed by different data can be cross-utilized, and the interactions between temporal and spatial views can be captured, so as to achieve accurate detection of long-sequence log and traffic data.

[0061] In one embodiment, the historical data with multiple granularities includes recent historical data, daily historical data, and weekly historical data. The historical data includes session data, time-window time-series data divided according to time windows, and neighborhood time-series data divided according to neighborhoods. The multiple image data includes multiple groups of images, and each group of images includes an interaction event graph, a time-window graph, and a neighbor information graph.

[0062] As described above, we use image data {G1, G2,..., Gn} to represent a complex data network. For each graph G = {V, E, A}, where V is the vertex set used to represent data, E is the edge set, and A is the adjacency matrix, and Aij represents the edge weight between data transmissions vi and vj. In addition, a path distance graph Gr = {V, Er, Ar} is constructed to characterize the time delay in data transmission in the data information, which contains time data. A common data graph is constructed, and the vertices are connected by various operations, and different operations have intrinsic meanings. In order to more comprehensively describe the data distribution and capture the interactions between operations, we further define a common data graph GT = {V, ET, AT}, Gt = {V, Et, At} to describe the relationships between operations in the common data. The interaction event graph, the time-window graph, and the neighbor information graph can be embodied in the form of a path distance graph or a common data graph.

[0063] In one embodiment, the embedding module includes a first embedding block, a second embedding block, and a third embedding block, where:

[0064] The first embedding block is used to obtain the most recent historical data, generate the most recent embedding result according to the most recent historical data, and feed it to the multi-view spatio-temporal module;

[0065] The second embedding block is used to obtain the daily historical data, generate the daily embedding result according to the daily historical data, and feed it to the multi-view spatio-temporal module;

[0066] The third embedding block is used to obtain the weekly historical data, generate the weekly embedding result according to the weekly historical data, and feed it to the multi-view spatio-temporal module.

[0067] As described above, the most recent historical data is the data within the most recent 10 minutes. The embedding module forms new feature vectors by extracting features from the multi-view time-series data, namely, the most recent embedding result, the daily embedding result, and the weekly embedding result.

[0068] In one embodiment, each multi-view spatio-temporal module includes a multi-view time learning unit and a first view intelligent fusion unit, where:

[0069] The multi-view time learning unit is used to generate multiple most recent temporal information, multiple daily temporal information, and multiple weekly temporal information according to multiple most recent embedding results, multiple daily embedding results, and multiple weekly embedding results respectively;

[0070] The first view intelligent fusion unit is used to generate multiple optimal fusion ratios according to multiple most recent temporal information, multiple daily temporal information, and multiple weekly temporal information to integrate multiple most recent temporal information, multiple daily temporal information, and multiple weekly temporal information, and obtain multiple temporal information.

[0071] As described above, the present application considers historical data with three different granularities. Therefore, the multi-view time learning unit consists of three temporal learning operations, and each operation processes historical observations embedded from a single view. The multi-view time learning unit produces three outputs, representing different temporal information extracted from different perspectives. The first view intelligent fusion unit will receive different temporal information and determine the optimal fusion ratio to integrate the temporal information and generate an input for each subsequent image data.

[0072] In one embodiment, each multi-view spatio-temporal module includes a multi-view space learning unit and a second view intelligent fusion unit, where:

[0073] The multi-view learning unit is used to mine image relationship information according to the interaction event graph, time window graph, and neighbor information graph in each group of images, generate corresponding three mining results, and output the three mining results to the next multi-view fusion module;

[0074] The second-view intelligent fusion unit is used to generate corresponding attention coefficients according to multiple temporal information, and perform visual fusion operations on the mining results generated by the multi-view learning unit based on the attention coefficients to obtain fusion information.

[0075] Specifically, the second-view intelligent fusion unit is used to map multiple temporal information as multiple nodes to the key subspace, and calculate the attention scores of each node in each layer of the key subspace. The calculation formula is:

[0076]

[0077] where i represents the number of nodes in the key subspace, where i = 1, 2, 3... represents the attention score of the i-th node in the l-th layer of the key subspace, represents the mining result of the i-th node in the l-th input layer of the key subspace, W represents the trainable weight matrix, and u represents the query vector;

[0078] It is also used to calculate the attention coefficients of each node in each layer of the key subspace according to the attention scores. The calculation formula is:

[0079]

[0080] where i represents the number of nodes in the key subspace, where i = 1, 2, 3... represents the attention coefficient of the i-th node in the l-th layer, and SoftMax l represents the softmax function, which is used to normalize the attention weights of all nodes in the l-th layer into a probability distribution;

[0081] It is also used to perform visual fusion operations on the mining results according to the attention coefficients to obtain fusion information. The formula corresponding to the visual fusion operation is:

[0082]

[0083] where both i and P represent the numbers of nodes in the key subspace, where i = 1, 2, 3... P; represents the fusion information of the i-th node in the k-th layer, represents the attention coefficient of the i-th node in the l-th layer, represents the mining result of the i-th node in the l-th input layer of the key subspace.

[0084] As described above, multi-view spatial learning is used to mine relationship information from different graphs, which represent different spatial dependency relationships. The input of the multi-view spatial learning unit in this application is three different graphs. Therefore, the output of the multi-view spatial learning unit is three results of passing messages through different graphs. The output of the multi-view spatial learning unit will be fed into the next multi-view fusion module to generate an integrated result for the subsequent process. The second-view intelligent fusion unit connects the multi-view temporal learning unit and the multi-view spatial learning unit. Its role is to determine weights, integrate the output of the upstream multi-view temporal learning unit, and generate an integrated result for the downstream multi-view spatial learning unit. This application determines the ideal weights based on the visual attention mechanism to integrate information from different spatial or temporal views. Since the importance of information from each upstream view may be different for each downstream view, the multi-view fusion block assigns different attention coefficients to each downstream block. For example, assume that the upstream multi-view learning block has P views and the downstream multi-view learning block has Q views. Then Q independent attention-based visual fusion operations are established. Each operation processes all the previous P views and fuses the results of one of the Q subsequent views. The input of each multi-view fusion operation can be formulated as X = {X(1), X(2), …, X(P)}, X(l) ∈ R (l = 1, 2, …, P), and F is the feature dimension of each node. To integrate multi-view temporal information, first, we use a shared linear transformation W ∈ R F×F ' to project the features of each node into the key subspace. Then, the attention scores of each layer are calculated by the similarity with the query vector u, where u ∈ R 1×F ' is an adaptive context embedding, randomly initialized and learned. During the entire training process, after the SoftMax operation, the attention scores are applied to different views (values) to obtain a weighted sum. The mechanism of the fusion operation of each view for node i is expressed as follows:

[0085]

[0086] where, represents the attention score of the i-th node in the l-th layer, represents the multi-view fusion operation of the i-th node in the input l-th layer, W represents the trainable weight matrix, u represents the query vector, represents the attention coefficient of the i-th node in the l-th layer, where the attention coefficient represents the importance of the i-th node to other nodes in the current l-th layer; SoftMax l , represents the softmax function, which is used to normalize the attention weights of all nodes in the l-th layer into a probability distribution. Represents the feature vector of the $i$-th node in the $k$-th layer, that is, the fused information. The multi-view fusion consists of $Q$ view-by-view operations, generating a total of $Q$ outputs for downstream learning, where represents the fused information of the $k$-th output for $1\leq k\leq Q$. This embodiment is based on view attention-based fusion, adaptively identifying the importance of each upstream view, fusing multi-view information, and generating a comprehensive result for the downstream view.

[0087] In one embodiment, the first embedding block, the second embedding block, and the third embedding block are all composed of a multi-layer neural network, including a first fully connected layer, an activation function, batch normalization, and a second fully connected layer.

[0088] The output module includes a fully connected layer, an activation function, batch normalization, a dropout layer, and a fully connected layer. By setting the dropout layer, overfitting can be prevented.

[0089] In one embodiment, the multi-view learning unit includes two graph convolutional networks (GCNs) connected by a ReLU activation function and batch normalization. Each graph convolutional network is used to obtain neighbor information, perform edge passing on the neighbor information, multiply the neighbor information by edge weights and aggregate it to obtain the aggregated neighbor information, and perform a linear transformation on the aggregated neighbor information and the node information corresponding to the interaction event graph, the time window graph, and the neighbor information graph respectively, and add the linear transformation results to obtain the corresponding three mining results.

[0090] As described above, we choose an extended temporal convolutional network (TCN) to process each time series data, and in each time learning, we have two temporal convolutional networks connected by a ReLU activation function and a BN layer. In addition, both the extended operation and the multi-granularity data maintain a large receptive field, which helps capture long-term temporal dependencies.

[0091] This application also provides an intrusion detection method based on historical data and GCN, including:

[0092] S1. Input multi-view time series data, where the multi-view time series data includes historical data of multiple granularities and multiple image data;

[0093] S2. Obtain the multi-view time series data and generate multiple embedding results according to the multi-view time series data;

[0094] S3. Output corresponding multiple temporal information and multiple spatial dependency information according to the multiple embedding results and the multiple image data, and fuse the multiple temporal information and the multiple spatial dependency information to obtain the first fusion information;

[0095] S4. Generate second prediction information according to the first prediction information and the multiple image data;

[0096] S5. Generate corresponding third prediction information based on the second prediction information and multiple pieces of image data;

[0097] S6. Generate a prediction result based on the first prediction information, the second prediction information, and the third prediction information.

[0098] In one embodiment, the historical data of multiple granularities includes recent historical data, daily historical data, and weekly historical data. The historical data includes session data, time-window sequential data divided according to a time window, and neighborhood sequential data divided according to a neighborhood. The multiple pieces of image data include multiple groups of images, and each group of images includes an interaction event graph, a time-window graph, and a neighbor information graph.

[0099] In one embodiment, step S2 of generating multiple embedding results according to the multi-view sequential data includes:

[0100] S21. Obtain recent historical data and generate a recent embedding result according to the recent historical data;

[0101] S22. Obtain daily historical data and generate a daily embedding result according to the daily historical data;

[0102] S23. Obtain weekly historical data and generate a weekly embedding result according to the weekly historical data.

[0103] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium provided in the present application and used in the embodiments can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0104] It should be noted that in this text, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, such that a process, apparatus, article or method comprising a series of elements not only includes those elements but also other elements not expressly listed, or elements inherent to such process, apparatus, article or method. Without further limitation, an element defined by the phrase "comprising an..." does not exclude the presence of additional identical elements in the process, apparatus, article or method comprising such element.

[0105] The above are only the preferred embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall similarly be included in the patent protection scope of the present application.

Claims

1. An intrusion detection system based on historical data and GCN, characterized in that, Including: An input module for inputting multi-view time-series data, where the multi-view time-series data includes historical data at multiple granularities and multiple image data; An embedding module for obtaining the multi-view time-series data, generating multiple embedding results according to the multi-view time-series data, and feeding the multiple embedding results and multiple image data to a multi-view spatio-temporal module; Multiple multi-view spatio-temporal modules, including a first multi-view spatio-temporal module, a second multi-view spatio-temporal module, and a third multi-view spatio-temporal module, where: The first multi-view spatio-temporal module is used to output corresponding multiple temporal information and multiple spatial dependency information according to the multiple embedding results and multiple image data, fuse the multiple temporal information and multiple spatial dependency information to obtain first fusion information, and feed the first fusion information as first prediction information and multiple image data to the second multi-view spatio-temporal module; The second multi-view spatio-temporal module is used to generate second prediction information according to the first prediction information and multiple image data, and feed the second prediction information and multiple image data to the third multi-view spatio-temporal module; The third multi-view spatio-temporal module is used to generate corresponding third prediction information according to the second prediction information and multiple image data; An output module for generating a prediction result according to the first prediction information, second prediction information, and third prediction information generated by the multiple multi-view spatio-temporal modules; The historical data at multiple granularities includes recent historical data, daily historical data, and weekly historical data. The historical data includes session data, time-window time-series data divided according to a time window, and neighborhood time-series data divided according to a neighborhood. The multiple image data includes multiple groups of images, and each group of images includes an interaction event graph, a time-window graph, and a neighbor information graph; The embedding module includes a first embedding block, a second embedding block, and a third embedding block, where: The first embedding block is used to obtain recent historical data, generate a recent embedding result according to the recent historical data, and feed it to the multi-view spatio-temporal module; The second embedding block is used to obtain daily historical data, generate a daily embedding result according to the daily historical data, and feed it to the multi-view spatio-temporal module; The third embedding block is used to obtain weekly historical data, generate a weekly embedding result according to the weekly historical data, and feed it to the multi-view spatio-temporal module; Each multi-view spatio-temporal module includes a multi-view space learning unit and a second-view intelligent fusion unit, where: The multi-view space learning unit is used to mine image relationship information according to the interaction event graph, time-window graph, and neighbor information graph in each group of images, generate corresponding three mining results, and output the three mining results to the next multi-view fusion module; The second-view intelligent fusion unit is used to generate corresponding attention coefficients according to the multiple temporal information, and perform a visual fusion operation on the mining results generated by the multi-view learning unit according to the attention coefficients to obtain fusion information; The second-view intelligent fusion unit is used to map the multiple temporal information as multiple nodes to a key subspace, and calculate the attention scores of each node in each layer in the key subspace. The calculation formula is: Among them, i represents the number of nodes in the key subspace, where i = 1, 2, 3… represents the attention score of the i-th node in the l-th layer of the key subspace, represents the mining result of the i-th node in the l-th layer of the input in the key subspace, W represents the trainable weight matrix, and u represents the query vector; It is also used to calculate the attention coefficient of each node in each layer of the key subspace according to the attention score, where the calculation formula is: Among them, i represents the number of nodes in the key subspace, where i = 1, 2, 3… represents the attention coefficient of the i-th node in the l-th layer, SoftMax l represents the softmax function, which is used to normalize the attention weights of all nodes in the l-th layer into a probability distribution; It is also used to perform a visual fusion operation on the mining results according to the attention coefficient to obtain fusion information, where the formula corresponding to the visual fusion operation is: Among them, both i and P represent the numbers of nodes in the key subspace, where i = 1, 2, 3... P; represents the fusion information of the i-th node in the K-th layer, represents the attention coefficient of the i-th node in the l-th layer, represents the mining result of the i-th node in the l-th layer of the input in the key subspace.

2. The intrusion detection system based on historical data and GCN according to claim 1, wherein Each multi-view spatio-temporal module includes a multi-view time learning unit and a first-view intelligent fusion unit, where: The multi-view time learning unit is used to generate multiple recent temporal information, multiple daily temporal information, and multiple weekly temporal information according to multiple recent embedding results, multiple daily embedding results, and multiple weekly embedding results respectively; The first-view intelligent fusion unit is used to generate multiple optimal fusion ratios according to multiple recent temporal information, multiple daily temporal information, and multiple weekly temporal information to integrate multiple recent temporal information, multiple daily temporal information, and multiple weekly temporal information, and obtain multiple temporal information.

3. The intrusion detection system based on historical data and GCN according to claim 1, characterized in that, The first embedding block, the second embedding block, and the third embedding block all include a first fully connected layer, an activation function, batch normalization, and a second fully connected layer.

4. The intrusion detection system based on historical data and GCN according to claim 2, wherein The multi-view learning unit includes two graph convolutional network layers connected by a ReLU activation function and batch normalization. Each graph convolutional network layer is used to obtain neighbor information, perform edge transfer on the neighbor information, multiply the neighbor information by the edge weight and aggregate it to obtain the aggregated neighbor information, and perform a linear transformation on the aggregated neighbor information and the node information corresponding to the interaction event graph, the time window graph, and the neighbor information graph respectively, and add the linear transformation results to obtain the corresponding three mining results.

5. An intrusion detection method based on historical data and GCN, characterized in that, It includes: Input multi-view time series data, where the multi-view time series data includes historical data of multiple granularities and multiple image data; Obtain the multi-view time series data and generate multiple embedding results according to the multi-view time series data; Output corresponding multiple temporal information and multiple spatial dependency information according to multiple embedding results and multiple image data, fuse the multiple temporal information and multiple spatial dependency information to obtain first fusion information, and use the first fusion information as the first prediction information; Generate second prediction information according to the first prediction information and multiple image data; Generate corresponding third prediction information according to the second prediction information and multiple image data; Generate a prediction result according to the first prediction information, the second prediction information, and the third prediction information; Among them, the historical data of multiple granularities includes recent historical data, daily historical data, and weekly historical data. The historical data includes session data, time window time series data divided according to a time window, and neighborhood time series data divided according to a neighborhood. The multiple image data includes multiple groups of images, and each group of images includes an interaction event graph, a time window graph, and a neighbor information graph; Among them, the step of obtaining the multi-view time series data and generating multiple embedding results according to the multi-view time series data includes: Obtain recent historical data and generate a recent embedding result according to the recent historical data; Obtain daily historical data and generate a daily embedding result according to the daily historical data; Obtain weekly historical data and generate a weekly embedding result according to the weekly historical data; Among them, the step of outputting corresponding multiple temporal information and multiple spatial dependency information according to multiple embedding results and multiple image data, and fusing the multiple temporal information and multiple spatial dependency information to obtain the first fusion information includes: Mining image relationship information based on the interaction event graph, time window graph, and neighbor information graph in each group of images, and generating corresponding three mining results; Generating corresponding attention coefficients according to the multiple temporal information, and performing a visual fusion operation on the mining results according to the attention coefficients to obtain fusion information; Mapping the multiple temporal information to the key subspace as multiple nodes, and calculating the attention scores of each node in each layer of the key subspace, where the calculation formula is: Among them, i represents the number of nodes in the key subspace, where i = 1, 2, 3… represents the attention score of the i-th node in the l-th layer of the key subspace, represents the mining result of the i-th node in the l-th layer of the input in the key subspace, W represents the trainable weight matrix, and u represents the query vector; Calculating the attention coefficients of each node in each layer of the key subspace according to the attention scores, where the calculation formula is: Among them, i represents the number of nodes in the key subspace, where i = 1, 2, 3..., represents the attention coefficient of the i-th node in the l-th layer, SoftMax l represents the softmax function, which is used to normalize the attention weights of all nodes in the l-th layer into a probability distribution; Performing a visual fusion operation on the mining results according to the attention coefficients to obtain fusion information, where the formula corresponding to the visual fusion operation is: Among them, both i and P represent the numbers of nodes in the key subspace, where i = 1, 2, 3... P; represents the fusion information of the i-th node in the K-th layer, represents the attention coefficient of the i-th node in the l-th layer, represents the mining result of the i-th node in the l-th layer of the input in the key subspace.

6. The intrusion detection method based on historical data and GCN according to claim 5, wherein The historical data of multiple granularities includes recent historical data, daily historical data, and weekly historical data. The historical data includes session data, time window time series data divided according to time windows, and neighborhood time series data divided according to neighborhoods. The multiple image data includes multiple groups of images, and each group of images includes an interaction event graph, a time window graph, and a neighbor information graph.

Citation Information

Patent Citations

  • Intrusion detection method and system based on historical data and GCN

    CN118627058A