Real-time intrusion detection method and system based on representation learning

By mapping network traffic packets into picture files and identifying them using convolutional layer technology, combined with deep learning characterization technology and data enhancement strategies, the problems of single data sources, poor reliability and insufficient real-time performance in the existing intrusion detection technology are solved, and efficient and accurate real-time intrusion detection are achieved.

CN120090817APending Publication Date: 2025-06-03BEIHANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411971927.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

Existing intrusion detection technologies face the problems of single data source, poor data set reliability, insufficient real-time detection performance and difficulty in real-time updates, which leads to the inability to effectively deal with complex and rapidly changing cyber attacks.

Method used

A real-time intrusion detection method based on characterization learning is adopted to map traffic packets from the application layer and the network layer into picture files and use convolutional layer technology to identify them to build a standardized and unified data set. This method combines deep learning characterization techniques and data augmentation strategies to optimize intrusion detection models to improve real-time and accuracy.

Benefits of technology

It effectively solves the problems of singularity and unreliability of traditional data sets, improves the real-time and accuracy of intrusion detection, reduces the risks of false positives and overfitting, and maintains the advancedness and adaptability of the model by regularly updating the data set.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120090817A_ABST
    Figure CN120090817A_ABST
Patent Text Reader

Abstract

The invention discloses a real-time intrusion detection method and system based on representation learning, and the method comprises the steps: mapping the data of an application layer and a transmission layer in a network into a picture file of a three-layer structure, constructing a standardized and unified data set through a deep learning technology, especially a convolutional neural network, and achieving the real-time intrusion detection of the network. The problem that a traditional data set is single and unreliable is solved. According to the method, real-time and accurate attack behavior identification is carried out on heterogeneous network flow data through a deep learning model, and a data set is periodically updated according to a detection result so as to optimize model performance. Experiments prove that the method can effectively improve the accuracy and robustness of intrusion detection and reduce the false alarm rate, and is suitable for various network attack scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a real-time intrusion detection method based on representation learning, and also relates to a corresponding real-time intrusion detection system, belonging to the field of network security technology. Background Art

[0002] Intrusion detection technology is a network security mechanism that monitors data such as network node traffic and system logs and uses analysis methods to determine whether there are attack behaviors, thereby providing system security services. With the complexity and novelty of network attack means, the main task of intrusion detection technology is to identify different types of network attacks and malware as early as possible.

[0003] Intrusion detection technology is divided into two parts: data and algorithms. Algorithms are the key to detecting attack behaviors, while data sets ensure the generality and accuracy of the technology. In terms of algorithms, to ensure real-time performance, machine learning algorithms are mostly used. However, traditional machine learning algorithms require a large amount of labeled data and feature selection when the data scale expands, with high costs and time consumption. Deep learning can learn features from raw data, so it is increasingly being used for intrusion detection.

[0004] However, intrusion detection technology faces the following main problems: (1) Single information source and level: Commonly used data sets such as CSE-CIC-IDS2018, CIC-IDS-2017, and KDD Cup’99 are all based on the transport layer, restricting the generality and practicality of intrusion detection models. (2) Poor reliability of data sets: The current data sets are unreliable, do not simulate the real network environment, and it is difficult to obtain transport layer data, affecting the features and information of the data sets. (3) Poor real-time detection performance: The complexity of deep learning intrusion detection models leads to a sacrifice in detection time, affecting the application of intrusion detection technology. (4) Difficult to update in real time: Network attack behaviors are constantly evolving, and existing algorithms and data sets are difficult to update in a timely manner, resulting in the inability to effectively respond to new attacks.

[0005] In summary, while existing intrusion detection technology provides network security protection, it also faces challenges such as a single data source, poor reliability of data sets, insufficient real-time detection performance, and difficulty in real-time updating. Summary of the Invention

[0006] The primary technical problem to be solved by the present invention is to provide a real-time intrusion detection method based on representation learning.

[0007] Another technical problem to be solved by the present invention is to provide a real-time intrusion detection system based on representation learning.

[0008] To achieve the above technical objectives, the present invention adopts the following technical solutions:

[0009] According to a first aspect of an embodiment of the present invention, a real-time intrusion detection method based on representation learning is provided, comprising the following steps:

[0010] Searching for traffic packets of an application layer and / or a network layer in a network, and mapping the traffic packets of the application layer and / or the network layer into image files;

[0011] Inputting the image file into the input layer of the intrusion detection model, and training the corresponding intrusion detection model for the application layer and / or the network layer;

[0012] Use convolutional layers and batch normalization to build a standard, unified dataset;

[0013] It uses multiple residual blocks, pooling layers and fully connected layers connected in sequence, uses cross entropy loss function or ternary loss function for classification, and outputs intrusion detection results.

[0014] Preferably, for the traffic packet of the application layer, the header, the request body and / or the requested URL of the HTTP request in the HTTP request log data are extracted.

[0015] Preferably, preprocessing is performed before mapping to remove the parts of the information in the URL and request body that are consistent with the information in the header, so as to retain the fields that may contain attack behavior characteristics, and splice them according to the corresponding characteristics to map them into image files.

[0016] Preferably, the real-time intrusion detection method further comprises the following sub-steps:

[0017] 1) Cleaning HTTP request log files: Process the original HTTP request log files and remove useless or redundant information;

[0018] 2) Classify HTTP requests: classify according to the type of request;

[0019] 3) Extract key information from HTTP requests: URL, header, and request body;

[0020] 4) Convert the extracted key information into a CSV file: Convert the requested URL, header, and request body into formatted CSV file data according to the request type;

[0021] 5) Assign an ID and label each identified attack behavior to facilitate identification and classification.

[0022] Preferably, for the traffic packets of the network layer, the captured packet data packets are analyzed to extract key information of the network layer, including the source IP address, the destination IP address, the header length and / or the checksum as characteristic data.

[0023] Preferably, during mapping, the traffic packets at the network layer are directly spliced ​​in a fixed order according to their characteristics to be mapped into an image file.

[0024] Preferably, during mapping, the feature text constructed by the feature character string is mapped to an image file; wherein the feature data is concatenated into a complete character string, and the character string is encoded and mapped into a space of a preset size, and the mapped value is used as the pixel point of the image file.

[0025] Preferably, during mapping, the first layer is the positive order after feature data mapping, the second layer is the reverse order after feature data mapping, and the third layer is the frequency of pixel values ​​of the first two layers, and finally the three layers are merged into an image file.

[0026] Preferably, when converting feature data into an image file, the order of the first two layers is exchanged for data types with less data, on the basis of ensuring that the third layer does not change.

[0027] According to a second aspect of an embodiment of the present invention, a real-time intrusion detection system based on representation learning is provided, comprising a processor and a memory, wherein the memory is coupled to the processor and is used to store a computer program, and when the computer program is executed by the processor, the processor implements the real-time intrusion detection method based on representation learning as described above.

[0028] Compared with the prior art, the present invention successfully constructs a standardized and unified data set by converting the application layer and transport layer data into a three-layer structured image file and using convolutional layer technology to identify these images, effectively solving the problems of singleness and unreliability of traditional data sets. In addition, the present invention also provides an intrusion detection model based on deep learning. The model takes real-time requirements into consideration during design, can quickly respond to intrusion behaviors in the network, and reduces the risks of false alarms and overfitting through deep learning characterization technology and optimized data enhancement strategies. In order to maintain the advancement and adaptability of the intrusion detection model, the data set will be updated regularly according to the detection results of the model, thereby achieving continuous improvement and optimization of the intrusion detection model. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 A schematic diagram of the structure of an intrusion detection model provided by the first embodiment of the present invention;

[0030] Figure 2 This is a schematic diagram of traffic characteristic distribution of network attack behavior;

[0031] Figure 3 This is a schematic diagram of merging three layers into a picture file in the first embodiment of the present invention;

[0032] Figure 4This is a schematic diagram comparing the effects before and after data augmentation in the second embodiment of the present invention;

[0033] Figure 5 This is a schematic diagram comparing the effects before and after using the triplet loss function in the second embodiment of the present invention;

[0034] Figure 6 This is a comparison chart of the effects of the real-time intrusion detection methods provided by the third embodiment and the first embodiment;

[0035] Figure 7 This is a schematic diagram of the structure of a real-time intrusion detection system based on representation learning in the fourth embodiment of the present invention. Specific Embodiments

[0036] The technical content of the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0037] The technical concept of the present invention is as follows: By cleaning, labeling, processing, and transforming a real dataset, the application layer data and transport layer data are mapped into picture files with a three-layer structure. Then, convolutional layers are used to identify these picture files, constructing a standardized and unified dataset, effectively solving the problems of singularity and unreliability existing in traditional datasets. At the same time, the present invention also proposes an intrusion detection model based on deep learning representation technology. This model can identify attack behaviors in real-time and accurately for heterogeneous network traffic data, and can regularly update the dataset according to the detection results of the intrusion detection model to achieve continuous improvement and optimization of the model performance.

[0038] First Embodiment

[0039] As Figure 1 shown, the first embodiment of the present invention provides a real-time intrusion detection method based on representation learning, which at least includes the following steps:

[0040] Step 1: Search for traffic packets in the application layer of the network and map the traffic packets in the application layer into picture files.

[0041] First, search for traffic packets in the application layer of the network, clean, label, process the HTTP requests in the HTTP request log data, and transform them into CSV file data.

[0042] Specifically, for traffic packets at the application layer, the headers, request bodies (Body), and request URLs of HTTP requests in the HTTP request log data are extracted. Among them, the headers of HTTP requests contain various request headers (Headers) of HTTP requests, such as User-Agent, Referrer, etc. The request body contains the data sent in the request for POST or PUT requests. The request URL is the address of the request, that is, the path of the page or resource accessed by the user. For common SQL and XSS attack behaviors, their characteristics often lie in the specific content of the URL and the body. SQL injection attack (SQL Injection) refers to an attacker attempting to manipulate the backend database by inputting SQL statements. Cross-site scripting attack (XSS) refers to an attacker attempting to inject malicious scripts to affect other users.

[0043] In various embodiments of the present invention, the picture file is preferably a picture file with a three-layer structure, specifically, it can be an RGB picture, a BGR picture, an ARGB picture, etc., which will not be elaborated one by one here.

[0044] In one embodiment of the present invention, the processing process of the HTTP request log data specifically includes:

[0045] 1) Cleaning the HTTP request log file: Processing the original HTTP request log file to remove useless or redundant information.

[0046] 2) Classifying HTTP requests: Classifying according to the type of request (such as GET, POST, PUT).

[0047] 3) Extracting key information in HTTP requests: URL, Referer (URL of the source page), Header, and Body.

[0048] 4) Converting the extracted key information into a CSV file: Converting the request URL, Refer, Header, and Body into formatted csv file data according to the type of request (GET, POST, PUT).

[0049] 5) Assigning an ID and labeling (Label) to each identified attack behavior for easy identification and classification.

[0050] Secondly, the traffic packets of the application layer in the network are mapped to image files. The intrusion detection model provided by the embodiment of the present invention is an intrusion detection model based on a convolutional neural network (CNN). Therefore, the above-mentioned CSV file data (including ID and corresponding Label) needs to be converted into an image file before it can be input into the intrusion detection model. This is because CNN can show high accuracy in image recognition. Specifically, CNN can automatically extract effective feature information from the original image, avoiding the cumbersome process of manually designing features; moreover, in CNN, the convolution kernel can slide on the entire image, realizing parameter sharing, thus reducing the complexity of the intrusion detection model and improving the computational efficiency. In addition, CNN captures local information of the image through convolution operations, which is in line with the characteristics of image processing, because the pixels in the image usually show local correlation. CNN can learn hierarchical features from simple to complex (hierarchical feature learning) through multi-layer convolution and pooling operations, thereby improving the accuracy of recognition. CNN has translation invariance, that is, it is insensitive to changes in the position of the target object in the image, which makes the intrusion detection model more robust during recognition. At the same time, the CNN deep learning intrusion detection model can be trained end-to-end in an end-to-end manner to learn higher-level abstract features, thereby improving the accuracy and efficiency of image recognition.

[0051] To this end, the present invention takes into account the information of forward semantics and backward semantics, and adopts the following steps to process the feature data and convert it into an RGB three-channel image.

[0052] 1. In data preprocessing, for the traffic packets at the application layer, first remove the parts of the URL, Refer, and Body that are consistent with the information in the Header to retain the fields that may contain attack behavior characteristics, and then splice them according to their characteristics to obtain the spliced ​​string. Splicing means extracting the key information of different parts of each traffic packet at the application layer, organizing them in a unified field order and splicing them into a string for subsequent analysis.

[0053] During the data annotation process, statistics show that the traffic characteristics of most attack behaviors are stored in the URL, Body, and Header. Eliminating the parts of the URL, Refer, and Body that are consistent with the information in the Header can improve accuracy and reduce false alarm rates.

[0054] Regarding the characteristics of the traffic packets at the network layer, since they are all numerical types, they can be directly spliced ​​in a fixed order without additional processing.

[0055] 2. Analyze the length and number of the concatenated feature strings, and use this as a basis to design the size of the image file.

[0056] As Figure 2 shown, the length of more than 86% of the feature data is below 256; only about 14% of the data is above 256. Therefore, considering the training and verification speed of the intrusion detection model and the length effectiveness of the data, in the embodiments of the present invention, the picture size is preferably set to 16*16.

[0057] 3. Map the feature text to a picture file.

[0058] After the foregoing feature data is spliced into a complete string, it is encoded and mapped into a space of a preset size (such as 16*16) here, and the value after mapping is used as the pixel point of the picture file.

[0059] More preferably, in order to make full use of the front-back relationship of the features and utilize as much feature data information as possible, the embodiments of the present invention adopt a three-layer picture merging method for picture synthesis. As Figure 3 shown, the first layer is the forward arrangement after the feature data is mapped, the second layer is the reverse arrangement after the feature data is mapped, and the third layer is the frequency of the pixel point values of the first two layers. Finally, the three layers are merged into a picture file.

[0060] The forward and reverse arrangements after the mapping of the feature data are used as one layer respectively and merged with the third layer. This method of preprocessing data is to improve the intrusion detection model's understanding and recognition ability of the feature data, because the data intercepted from the network is time series data. Using this preprocessing method can provide the following advantages: 1) The sequence information of the feature data can be retained, which is crucial for capturing the time dependence or sequential relationship in the feature data; 2) Increase the data dimension: Convert the one-dimensional feature data into a three-dimensional picture file, increasing the data dimension, enabling the intrusion detection model to learn features from multiple angles and improving the generalization ability of the intrusion detection model for unseen data; 3) The third layer of the picture represents the frequency of the pixel point values of the first two layers, which provides additional information - frequency information about the distribution of the feature data for the intrusion detection model, facilitating the intrusion detection model to understand the patterns and anomalies of the data; 4) It can reduce the over-dependence of the intrusion detection model on specific features, thereby reducing the risk of overfitting; 5) Converting the data into images can utilize existing image processing technologies and hardware acceleration to improve the calculation efficiency.

[0061] Step 2: Search for the traffic packets of the network layer in the network and map the traffic packets of the network layer to a picture file.

[0062] The network layer focuses on the routing and transmission process of traffic packets. For example, in a DDoS attack, source IP forgery and destination IP fixation are usually involved; for another example, forged data packets may involve anomalies in fields such as header length and checksum. Therefore, for the packet data at the network layer, analyze the captured packet data packets and extract the key information at the network layer, such as source IP address, destination IP address, header length, checksum, etc. Convert the key information extracted from the captured packet data packets into a CSV format file. And, assign an ID to each identified attack behavior and label it.

[0063] To this end, the embodiments of the present invention fully consider the information of forward semantics and backward semantics, and adopt the following steps to process the feature data and convert it into an RGB three-channel picture.

[0064] 1. In data preprocessing, for the features of traffic packets at the network layer, since they are all numerical types, they can be directly spliced in a fixed order without additional processing.

[0065] 2. Analyze the length and quantity of the spliced feature strings, and design the size of the picture file based on this.

[0066] 3. Map the feature text to a picture file.

[0067] Take the feature data as a complete string, and encode and map it into a space of a preset size (such as 16*16), and use the mapped value as the pixel point of the picture file.

[0068] More preferably, to make full use of the front-back relationship of the features and utilize as much feature data information as possible, the embodiments of the present invention adopt a three-layer picture merging method for picture synthesis. The specific operations are as Figure 3 shown. As described above, arrange the forward and reverse orders of the mapped feature data respectively as one layer, and merge them with the third layer to form a picture file with a three-layer structure.

[0069] In an embodiment of the present invention, for the network layer and the application layer, Hugging Face and Kaggle open-source real datasets are used for training. The datasets provided by Hugging Face and Kaggle include data such as text and audio, and are used to map the features into the range of 0 to 256. For the network layer data, the public dataset CIC-IDS2017 is used for training. This dataset contains a total of 75 features and is used to map into the range of 0 to 256. It should be noted that the dataset can be flexibly selected according to the classification data type. This is only an example here and does not constitute a limitation to the present invention. In practice, other datasets can also be selected for training.

[0070] Step 3: Input the image file into the input layer of the intrusion detection model, and train the corresponding intrusion detection model for data at different levels.

[0071] Specifically, the operation of this step depends on the available data types:

[0072] First of all, if there is only application layer data, then only the image file mapped from the application layer traffic packet needs to be input into the input layer of the intrusion detection model, and the intrusion detection model is trained specifically for the application layer data. This is done to enable the intrusion detection model to identify and distinguish normal traffic and attack traffic in the application layer.

[0073] Secondly, if there is only network layer data, then the image file mapped from the network layer traffic packet is input into the input layer of the intrusion detection model, and the intrusion detection model is trained specifically for the network layer data. This helps the intrusion detection model to identify abnormal behaviors in the network layer, such as DDoS attacks or IP address fraud, etc.

[0074] Finally, if there is both application layer and network layer data, we input the image files mapped from these two types of traffic packets into the input layer of the intrusion detection model respectively, and train the corresponding intrusion detection models for different traffic levels. This method can provide more comprehensive security protection because it can detect potential threats in both the network layer and the application layer simultaneously.

[0075] In summary, the core of Step 3 is to input the mapped image file into the intrusion detection model according to different data levels (application layer or network layer) and available data types (single level or dual level), and carry out corresponding model training to achieve real-time and accurate monitoring of network traffic.

[0076] When conducting network intrusion analysis on network layer and application layer data, the present invention adopts a unique method to convert text data into image files. Specifically, not only is the text data mapped into an image in the normal order (in ascending order), but it is also mapped in the reverse order (in descending order), and the frequency information of pixel point values is added to generate and merge three-layer images. This design enables the intrusion detection model to capture richer semantic information, enhances the model's ability to completely express features, and thus significantly improves the detection accuracy and robustness.

[0077] Step 4: Use the convolutional layer to construct a standard and unified data set through batch normalization.

[0078] In one embodiment of the present invention, the first few layers of the intrusion detection model are convolutional layers for extracting low-level features. After each convolutional layer, there are batch normalization and ReLU activation functions to improve the training stability and the performance of the intrusion detection model. Among them, by using batch normalization, such as Dropout and L1 / L2 regularization, the intrusion detection model can be prevented from over-relying on the training data and the generalization ability of the intrusion detection model can be improved.

[0079] As described above, the present invention uses convolutional layers for network intrusion feature analysis, drawing on the advantages of high real-time performance and high accuracy of convolutional layers in image recognition, thereby significantly reducing the false alarm rate while achieving high real-time performance.

[0080] Step Five: Process using multiple residual blocks, pooling layers, and fully connected layers connected in sequence, and use the cross-entropy loss function for training to output the intrusion detection result.

[0081] To verify the effect of the embodiments of the present invention, the inventors designed the following experiment. In the experiment, the classification effect of the intrusion detection model is mainly evaluated based on three indicators: precision, recall, and F1-score. At the same time, the classification effect of each category is evaluated through a confusion matrix. In the confusion matrix, the columns represent the predicted categories, and the total number of columns represents the number of data predicted as this category; the rows represent the true categories of the data, and the total number of rows represents the actual number of data instances in this category. The value in each cell of the matrix represents the correspondence between the actual category and the predicted category, that is, the number of actual data predicted as a specific category. Its structure is shown in Table 1:

[0082] Table 1

[0083]

[0084] The definitions of the remaining evaluation indicators are as follows:

[0085]

[0086]

[0087] The experiment was carried out on a server equipped with an RTX 2080Ti GPU, a Core i9 9800K processor, and 32GB of memory. With the support of these high-performance hardware, the performance of the intrusion detection model was evaluated. The specific results and analysis are shown in Table 2, which details the performance of the model in the experiment.

[0088] Table 2

[0089] Intrusion detection model Accuracy (Acc) False alarm rate Detection time The present invention 99.27% 0.81% 0.42ms

[0090] In summary, the present invention proposes a real-time intrusion detection model based on representation learning technology, effectively solving the problems of singularity, unreliability, and insufficient real-time performance existing in traditional data sets. Through experimental verification, this model not only improves the real-time performance and accuracy of intrusion detection, but also reduces the false alarm rate. During the process of constructing the image, the present invention particularly considers the forward and backward semantic information of the data set, and extracts key features and semantic information from the pictures through deep learning representation technology, thereby significantly improving the detection accuracy.

[0091] Second Embodiment

[0092] During the experiment, the inventors found that the precision and recall rates of the real-time intrusion detection method provided in the first embodiment were not satisfactory in some cases. Part of the reason for this problem is that the sample sizes of some types of data (such as label 3 and label 2) are small, only accounting for about 2000 out of a total of approximately 150,000 data; another part of the reason is that the data features of some types (such as label 3) overlap to a certain extent with the data features of other types (such as label 0). To solve these problems, in the second embodiment of the present invention, the inventors introduced data augmentation techniques and a new loss function, aiming to improve the performance of the model through these improvement measures.

[0093] The data augmentation technique is used to expand the data set without changing the original data features. In the second embodiment of the present invention, two methods, offline augmentation and online augmentation, are adopted. Offline augmentation is to exchange the orders of the first two layers while keeping the third layer unchanged for data types with small sample sizes (such as label 3) during the process of converting feature data into picture files, thereby generating new data and enhancing the data set. Online augmentation is to perform real-time random transformations on the data during model training, including rotating, scaling, and flipping the pictures. The data before and after these augmentations are both used for model training to increase the data diversity and reduce the risk of overfitting. As Figure 4 shown, the effect of the augmentation operation on label 3 is particularly obvious, not only making the data more stable, but also increasing the precision by 3% - 5%.

[0094] In the scenario of network intrusion detection, different attack types may have similar feature patterns, and these similarities may cause the intrusion detection model to misclassify one attack behavior as another, thereby reducing the detection accuracy and recall rate. In an embodiment of the present invention, in order to improve the discrimination ability of the intrusion detection model for different attack behaviors, in addition to using the cross-entropy loss function, a triplet loss function is also introduced. This design aims to separate the features of different attack types as much as possible in the feature space during the training process to reduce the risk that the model misclassifies one attack behavior as another with a similar feature pattern, thereby improving the detection accuracy and recall rate. The triplet loss function is particularly suitable for deep learning models, especially in image and text embedding training. It calculates the loss by considering three groups of samples: anchor samples, positive samples, and negative samples, ensuring that similar samples are close in the embedding space and different samples are far apart. By specifically adding the triplet loss to label 3 and label 2, the present invention enhances the discrimination of the model for the feature of these two types of data, increases their distance in the feature space, and thus improves the precision rate.

[0095]

[0096] In the second embodiment of the present invention, by combining the use of offline data augmentation technology and two loss functions - the cross-entropy loss function and the triplet loss function, the experimental results (as Figure 5 shown) indicate that for a specific label 3 and the overall model, after adopting the triplet loss function, its accuracy has been steadily and significantly improved. This design effectively enhances the model's ability to distinguish various attack behaviors by increasing the distance between the features of different attack behaviors in the feature space, reduces the misjudgment caused by feature overlap, and thus improves the overall performance of the intrusion detection model.

[0097] Third Embodiment

[0098] In the third embodiment of the present invention, a simplified real-time intrusion detection method is proposed. This method adjusts the first embodiment and allows step two to be omitted. The following are the adjusted method steps:

[0099] First, the method involves searching for application layer traffic packets in the network and mapping these traffic packets into picture files. Then, these picture files are input into the input layer of the intrusion detection model, and corresponding intrusion detection models are trained for different traffic levels. Next, a standard and unified data set is constructed using convolutional layers and batch normalization techniques. Finally, through the processing of multiple residual blocks, pooling layers, and fully connected layers connected in sequence, classification is performed using the cross-entropy loss function to output the final intrusion detection result. These steps have been described in detail in the first embodiment and will not be elaborated here.

[0100] It should be emphasized that although omitting step two in the first embodiment can simplify the process, this will lead to a corresponding decline in the detection effect. To compare the effects of the two embodiments, the inventor conducted experimental verification, and the results are as Figure 6 shown. In the experimental verification, the data processing method is to match the requested network layer and application layer as a new piece of data. If there is an attack behavior in either the network layer or the application layer of this new piece of data, this data is identified as an attack behavior. 20,000 pieces of data were extracted from the test set of the intrusion detection model for combination and then verified.

[0101] The experimental process includes inputting the network layer and application layer of the data into the corresponding network layer intrusion detection model and application layer intrusion detection model for prediction respectively. For the fused model, both layers of data are input into the model for prediction at the same time. Finally, the prediction results are compared with the previously labeled data to obtain the verification results. This process confirms that omitting step two will affect the model performance, so in practical applications, it is necessary to balance the relationship between process simplification and detection effect.

[0102] Figure 6 Shows the performance comparison of three different methods - only network layer analysis, only application layer analysis, and the fusion of network layer and application layer analysis (i.e., the first embodiment) - in terms of precision, recall, and F1-score (the F1-score is the harmonic mean of precision and recall). The results show that the fusion method performs the best in these three key metrics. This indicates that when detecting network attacks, the fusion method can not only more accurately identify real attacks but also identify more real attacks.

[0103] Further analysis shows that the fused intrusion detection model (the first embodiment) has a better effect compared to the intrusion detection model of only the application layer (the third embodiment). This is because the fusion model can simultaneously identify attack data in both the network layer and the application layer, while the separate application layer model cannot detect attacks in the network layer. Therefore, the fusion model provides more comprehensive security protection.

[0104] Nevertheless, in certain cases, the F1 score of the third embodiment reached 0.5903, and this result is also acceptable in some application scenarios. This indicates that, although some steps in the first embodiment are omitted, the real-time intrusion detection method of the third embodiment still has a certain degree of practicality.

[0105] Fourth Embodiment

[0106] Based on the above real-time intrusion detection method based on representation learning, the fourth embodiment of the present invention further provides a real-time intrusion detection system based on representation learning. As Figure 7 shown, the real-time intrusion detection system includes one or more processors and a memory. Among them, the memory is coupled to the processor and is used to store computer programs. When the computer programs are executed by the processor, the processor implements the real-time intrusion detection method based on representation learning as in the above embodiments.

[0107] Among them, the processor is used to control the overall operation of the real-time intrusion detection system to complete all or part of the steps of the above real-time intrusion detection method based on representation learning. The processor can be a central processing unit (CPU), a graphics processing unit (GPU), a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a digital signal processing (DSP) chip, etc. The memory is used to store various types of data to support the operation of the real-time intrusion detection system. These data can include, for example, instructions for any application program or method operating on the real-time intrusion detection system, as well as application program related data. The memory can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read only memory (EEPROM), erasable programmable read only memory (EPROM), programmable read only memory (PROM), read only memory (ROM), magnetic memory, flash memory, etc.

[0108] In another exemplary embodiment, the present invention also provides a computer-readable storage medium including program instructions. When the program instructions are executed by a processor, the steps of the real-time intrusion detection method based on representation learning in any of the above embodiments are implemented. For example, the computer-readable storage medium can be the above memory including program instructions. The above program instructions can be executed by the processor of the system to complete the above real-time intrusion detection method based on representation learning and achieve the same technical effects as the above method.

[0109] It should be noted that the above multiple embodiments are only examples. The technical solutions of each embodiment can be combined, and the order of each step can be changed, all within the protection scope of the present invention.

[0110] The above has provided a detailed description of the real-time intrusion detection method and system based on representation learning according to the present invention. For those of ordinary skill in the art, any obvious changes made to it without departing from the essence of the present invention will constitute an infringement of the patent right of the present invention and will bear corresponding legal responsibilities.

Claims

1. A real-time intrusion detection method based on representation learning, characterized in that The following steps are involved: Searching for traffic packets of an application layer and / or a network layer in a network, and mapping the traffic packets of the application layer and / or the network layer into image files; Inputting the image file into the input layer of the intrusion detection model, and training the corresponding intrusion detection model for the application layer and / or the network layer; Use convolutional layers and batch normalization to build a standard, unified dataset; It uses multiple residual blocks, pooling layers and fully connected layers connected in sequence, uses cross entropy loss function or ternary loss function for classification, and outputs intrusion detection results.

2. The real-time intrusion detection method based on representation learning as claimed in claim 1, characterized in that: For the traffic packet of the application layer, the header, request body and / or requested URL of the HTTP request in the HTTP request log data are extracted.

3. The real-time intrusion detection method based on representation learning as claimed in claim 2, characterized in that: Before mapping, preprocessing is performed to remove the parts of the URL and request body information that are consistent with the information in the header to retain the fields that may contain attack behavior characteristics, and then splice them according to their characteristics to map them into image files.

4. The real-time intrusion detection method based on representation learning as described in claim 2 is characterized in that The following sub-steps are included: 1) Cleaning HTTP request log files: Process the original HTTP request log files and remove useless or redundant information; 2) Classify HTTP requests: classify according to the type of request; 3) Extract key information from HTTP requests: URL, header, and request body; 4) Convert the extracted key information into a CSV file: Convert the requested URL, header, and request body into formatted CSV file data according to the request type; 5) Assign an ID and label each identified attack behavior to facilitate identification and classification.

5. The real-time intrusion detection method based on representation learning as claimed in claim 1, characterized in that: For the traffic packets of the network layer, the captured packet data packets are analyzed to extract key information of the network layer, including the source IP address, the destination IP address, the header length and / or the checksum as characteristic data.

6. The real-time intrusion detection method based on representation learning as claimed in claim 5, characterized in that: During mapping, the traffic packets at the network layer are directly spliced ​​in a fixed order according to their characteristics to be mapped into image files.

7. The real-time intrusion detection method based on representation learning as claimed in claim 6, characterized in that: During mapping, the feature text constructed by the feature string is mapped to a picture file; The characteristic data are concatenated into a complete character string, and the character string is encoded and mapped into a space of a preset size, and the mapped value is used as the pixel point of the image file.

8. The real-time intrusion detection method based on representation learning as claimed in claim 7, characterized in that: During mapping, the first layer is the positive order arrangement after feature data mapping, the second layer is the reverse order arrangement after feature data mapping, and the third layer is the frequency of pixel values ​​of the first two layers. Finally, the three layers are merged into an image file.

9. The real-time intrusion detection method based on representation learning as claimed in claim 8, characterized in that: When converting feature data into image files, the order of the first two layers is swapped for data types with less data, while ensuring that the third layer does not change.

10. A real-time intrusion detection system based on representation learning, characterized in that It includes a processor and a memory, wherein the memory is coupled to the processor and is used to store a computer program. When the computer program is executed by the processor, the processor implements the real-time intrusion detection method based on representation learning as described in any one of claims 1 to 9.