Container abnormal behavior detection method based on hardware performance counter
Through the detection method based on hardware performance counter and a two-way long and short-term memory network integrating the attention mechanism, the problem of unreliable container abnormal behavior detection and large performance overhead in the prior art is solved, and the detection effect of high accuracy and anti-interference is achieved.
Patent Information
- Application Number
- CN202510064627.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-05-13
AI Technical Summary
The existing software-based container abnormal behavior detection methods have problems such as insufficient data acquisition and high performance overhead, making it difficult to effectively detect malicious behavior in containers.
Using a detection method based on hardware performance counters, the hardware performance counter timing data during container runtime is collected through the Perf tool, and abnormal behavior detection is performed using a two-way long and short-term memory network with a converged attention mechanism.
It improves the accuracy and anti-interference ability of container abnormal behavior detection, reduces the performance overhead of detection methods, and enhances the ability to detect malicious behavior in containers.
Smart Images

Figure CN119989346A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer technology, and more specifically relates to designing a container abnormal behavior detection method based on hardware performance counters in a cloud native environment. Background Art
[0002] With the in-depth development of cloud computing, cloud-native technology has become an important force in promoting the transformation of modern application architecture. As an important part of cloud-native technology, containerized infrastructure, container orchestration platform and cloud-native applications have been widely used in various technical fields. However, this trend also brings new security challenges. Container technology essentially simplifies the deployment and management of applications by packaging applications and their dependencies into lightweight, independent execution environments. However, the widespread use of containers also exposes them to a variety of new security threats, such as image poisoning and container attacks. These threats directly or indirectly affect the security of the entire cloud environment, and then threaten the core data and business security of enterprises and users. Therefore, it is of great practical significance and urgency to study and solve container security issues. Researchers have proposed many methods to detect whether containers have abnormal behavior, which are mainly software-based detection methods.
[0003] Software-based detection methods use agent-based methods such as hooks to obtain software features of containers, such as system API calls and memory. These methods can be further divided into in-band methods and out-of-band methods based on whether the data is obtained directly from inside the container. In-band methods obtain high-level semantic information by installing agents inside the system, including process lists, system calls, and user system files. However, this method has obvious limitations. Since the internal agent and system programs have the same permissions, malware may manipulate kernel data, which seriously reduces the reliability of obtaining data. Out-of-band methods can obtain status information inside the container outside the container, but this method requires crossing the semantic gap and incurs a large performance overhead.
[0004] In recent years, hardware-based abnormal behavior detection methods have been widely used in the field of security detection, such as using hardware performance counters (Hardware Performance Counters, HPCs) to detect malicious behaviors in the Internet of Things and whether virtual machines are under malicious attacks. These methods collect HPCs from computer systems without the need for semantic reconstruction like traditional virtual machine introspection methods, which greatly improves detection efficiency. In addition, in terms of combating malware, the use of HPCs can effectively detect and discover complex malware that uses escape techniques. The present invention proposes a container abnormal behavior detection method based on hardware performance counters, which has the advantages of low overhead, strong anti-interference ability and high accuracy compared to traditional software-based detection methods. Summary of the invention
[0005] In view of the above research status and existing problems, the present invention proposes a container abnormal behavior detection method based on hardware performance counters to improve the accuracy and anti-interference ability of container abnormal behavior detection.
[0006] To achieve the above object, the present invention provides the following technical solutions:
[0007] A method for detecting abnormal behavior of a container based on hardware performance counters comprises the following steps:
[0008] 1) Build a Docker container platform, create a Linux container on it, mount the benign software and malware samples stored in the host machine into the container, run the samples in the container, and use the Perf tool provided by the host Linux system outside the container to collect the hardware performance counter timing data when the container is running;
[0009] 2) Use Python script to preprocess the collected data, remove interference items, retain valid values, and save the processed data into a csv file so that it can be used as input for the deep learning model;
[0010] 3) Construct a bidirectional long short-term memory network integrating attention mechanism, use the preprocessed data as the input of the model for training, and obtain the abnormal behavior detection model.
[0011] In a further optimization of the technical solution, the method of obtaining the hardware performance counter timing data in step 1) is:
[0012] 1.1) Install Ubuntu system on bare metal as host, build Docker container platform on the host, and use the container platform to create Linux container;
[0013] 1.2) docker run-d --name $ContainerName ${ElfFPath_Local}: ${ElfFPath_Container} $ImageName command to mount the benign software and malware samples stored in the host into the Linux container;
[0014] 1.3) Write a shell script in the host machine to realize the automatic running of samples, which are divided into 5 batches. Each batch collects 4 hardware event feature data and saves them in .txt files. This is to prevent the influence of time division multiplexing on data collection and make the collected data more accurate.
[0015] 1.4) The specific command for collecting data is: perfstat -e [hardware event] -I [collection time interval] -G [container name] -o [data storage path]
[0016] The parameter after -e indicates the hardware events to be collected. The present invention collects a total of 20 hardware events, namely: branch-instructions, branch-misses, bus-cycles, cache-misses, cache-references, cpu-cycles, instructions, ref-cycles, L1-dcache-load-misses, L1-dcache-loads, L1-dcache-stores, L1-icache-load-misses, LLC-loads, LLC-stor es, branch-load-misses, dTLB-load-misses, dTLB-loads, dTLB-store-misses, dTLB-stores, iTLB-load-misses; the parameter after -I specifies the collection time interval, which is 100ms in the present invention; the parameter after -G specifies the hardware events of the container to be collected, which is generally in the form of docker / CONTAINER_ID, which utilizes the cgroups module provided by the Linux kernel; the parameter after -o specifies the path where the collected data is to be saved, and the present invention saves the data of each sample into a .txt file.
[0017] The sleep 10s command is used in the shell script to control the total collection time. The present invention collects 100 sets of hardware event timing data within 10 seconds. The sleep 10s command is followed by kill-2 to terminate the perf process, completing the collection of one sample.
[0018] In a further optimization of the technical solution, the method of preprocessing the time series data in step 2) is:
[0019] 2.1) Each sample corresponds to 5 .txt files. Each file is processed separately, and then the 20 features are merged together;
[0020] 2.2) For each .txt file, read it line by line, skip invalid lines starting with the '#' character, and display the valid lines as<not counted> The value of is 0, and the commas are removed from the numerical values, and they are processed in chronological order;
[0021] 2.3) After merging the 5 batches of data corresponding to each sample, merge the time series data of all samples into one csv file.
[0022] The technical solution is further optimized, and the step 3) is as follows:
[0023] 3.1) Build a bidirectional long short-term memory network with an attention mechanism, read the csv file obtained by step 2) preprocessing, where the sample id is 1-1000, each sample corresponds to 100 sets of time series data, the benign software label is 0, and the malicious software label is 1;
[0024] 3.2) The bidirectional long short-term memory network with integrated attention mechanism captures the global time dependency through forward and backward processing. The attention mechanism assigns weights to each time step, calculates the average attention weight of each feature over all time steps, selects the 12 features with the largest contribution, highlights the information of key time steps and important features, ignores the noise of minor time steps and minor features, greatly reduces the gradient attenuation problem of long time series processing, and can better remember previous information.
[0025] The technical solution is further optimized. The bidirectional long short-term memory network structure of the fusion attention mechanism is as follows: the first layer is the Bi-LSTM layer, the input dimension is (32, 100, 20), where 32 is the batch_size, 100 is the time step of each sample, 20 is the feature dimension, the output dimension is (32, 100, 128), and the number of hidden layer neurons is 64. Because it is a bidirectional LSTM, the output dimension is twice the number of hidden layer neurons. The Sigmoid activation function is used internally in the Bi-LSTM to calculate the values of the input gate, forget gate, and output gate. The output range of the Sigmoid function is between 0 and 1, indicating the degree of "open" or "close" of each gate, so that The Tanh activation function is used to calculate the state update. The output range of the Tanh function is between -1 and 1, which can help the LSTM unit maintain information balance. The second layer is the Attention layer, with an input dimension of (32, 100, 128) and an output dimension of (32, 128). The Attention mechanism calculates the attention weight for the hidden state of each time step and generates a context vector. The third layer is the fully connected layer, with an input dimension of (32, 128) and an output dimension of (32, 1). The fourth layer is the Sigmoid layer, with an input dimension of (32, 1) and an output dimension of (32, 1), which is used to convert the output into a probability value to determine whether the container has abnormal behavior.
[0026] Different from the prior art, the above technical solution has the following beneficial effects:
[0027] 1) This detection method obtains feature data in an out-of-band manner and does not need to obtain high-level semantic information of the container through semantic reconstruction. This not only avoids the development required for reverse engineering, but also reduces the technical difficulty and improves the versatility of the detection method.
[0028] 2) This detection method collects the timing data of hardware performance counters as the input of the long short-term memory network that integrates the attention mechanism, which can capture the global time dependency. The attention mechanism assigns weights to each time step, highlights the information of key time steps, and improves the accuracy and anti-interference ability of abnormal behavior detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 The overall flow chart of the abnormal behavior detection method of containers based on hardware performance counters;
[0030] Figure 2 This is the architecture diagram of the long short-term memory network that integrates the attention mechanism. DETAILED DESCRIPTION
[0031] In order to explain the technical content, structural features, achieved objectives and effects of the technical solution in detail, the following is a detailed description in conjunction with specific embodiments and accompanying drawings.
[0032] See also Figure 1 The figure is an overall flow chart of a method for detecting abnormal behavior of a container based on a hardware performance counter. The present invention provides a method for detecting abnormal behavior of a container based on a hardware performance counter, comprising the following steps:
[0033] 1) Build a Docker container platform, create a Linux container on it, mount the benign software and malware samples stored in the host into the container, run the samples in the container, and use the Perf tool provided by the host Linux system outside the container to collect the hardware performance counter timing data when the container is running.
[0034] In this embodiment, step 1) installs Ubuntu system on the bare metal as the host, builds Docker container platform on the host, and uses the container platform to create a Linux container; uses docker run-d--name$ContainerName${ElfFPath_Local}:${ElfFPath_Container}$ImageName command to mount benign software and malware samples stored in the host into the Linux container; writes a shell script in the host to realize automatic sample operation, which is divided into 5 batches, and collects 4 hardware event feature data in each batch and saves them in a .txt file. This is to prevent the influence of time division multiplexing on data collection and make the collected data more accurate; the specific command for collecting data is: perfstat-e[hardware event]-I[collection time interval]-G[container name]-o[data storage path], the parameter after -e indicates the hardware event to be collected, and the present invention collects a total of 20 hardware events, namely: branch-instructions, branch-misses, bus-cycles, cache-misses, cache-references, cpu-cycle s, instructions, ref-cycles, L1-dcache-load-misses, L1-dcache-loads, L1-dcache-stores, L1-icache-load-misses, LLC-loads, LLC-stores, branch-load-misses, dTLB-load-misses, dTLB-loads, dTLB-store-misses, dTLB-stores, iTLB-load-misses. The parameter after -I specifies the time interval for collection, which is 100ms in the present invention. The parameter after -G specifies the hardware events of the container to be collected, which is generally in the form of docker / CONTAINER_ID, which utilizes the cgroups module provided by the Linux kernel. The parameter after -o specifies the path where the collected data is to be saved. The present invention saves the data of each sample in a .txt file. Use sleep in the shell script The 10s command controls the total collection time. The present invention collects a total of 100 sets of hardware event timing data within 10s. The sleep10s command is followed by kill-2 to terminate the perf process to complete the collection of one sample. In order to ensure that each software runs in the same system environment, a "clean" system snapshot is saved before the experiment, that is, the container is in a state where no third-party software is running. Before running the next malware, the system environment is restored to this state.
[0035] 2) Use Python script to preprocess the collected data, remove interference items, retain valid values, and save the processed data into a csv file so that it can be used as input for the deep learning model.
[0036] In this embodiment, in step 2), each sample corresponds to 5 .txt files, each file is processed separately, and then the 20 features are merged together; for each .txt file, read line by line, skip invalid lines starting with the '#' character, and display them in valid lines as<not counted> The value of is 0, and the commas are removed from the numerical values, and they are processed in chronological order. After merging the 5 batches of data corresponding to each sample, the time series data of all samples are merged into one csv file.
[0037] 3) Construct a bidirectional long short-term memory network integrating attention mechanism, use the preprocessed data as the input of the model for training, and obtain the abnormal behavior detection model.
[0038] In this embodiment, step 3) constructs a bidirectional long short-term memory network integrating the attention mechanism, and reads the csv file obtained by preprocessing in step 2), where the sample id is 1-1000, each sample corresponds to 100 sets of time series data, the benign software label is 0, and the malware label is 1; the structure of the bidirectional long short-term memory network integrating the attention mechanism is as shown in the attached figure. Figure 2As shown: the first layer is the Bi-LSTM layer, the input dimension is (32, 100, 20), where 32 is the batch_size, 100 is the time step of each sample, 20 is the feature dimension, the output dimension is (32, 100, 128), the number of hidden layer neurons is 64, because it is a bidirectional LSTM, the output dimension is twice the number of hidden layer neurons, the Bi-LSTM uses the Sigmoid activation function to calculate the values of the input gate, forget gate and output gate, the output range of the Sigmoid function is between 0 and 1, indicating the degree of "open" or "close" of each gate, and the Tanh activation function is used to calculate the state update, The output range of the Tanh function is between -1 and 1, which can help the LSTM unit maintain information balance; the second layer is the Attention layer, with an input dimension of (32, 100, 128) and an output dimension of (32, 128). The Attention mechanism calculates the attention weight for the hidden state of each time step and generates a context vector; the third layer is the fully connected layer, with an input dimension of (32, 128) and an output dimension of (32, 1); the fourth layer is the Sigmoid layer, with an input dimension of (32, 1) and an output dimension of (32, 1), which is used to convert the output into a probability value to determine whether the container has abnormal behavior. The bidirectional long short-term memory network with the fusion attention mechanism captures the global time dependency through forward and backward processing. The attention mechanism assigns weights to each time step, calculates the average attention weight of each feature over all time steps, selects the 12 features with the largest contribution, highlights the information of key time steps and important features, ignores the noise of secondary time steps and secondary features, greatly reduces the gradient decay problem of long time series processing, and can better remember previous information. In this embodiment, the malware dataset comes from the virusshare and virustotal websites, with the attribute of Linux ELF64bitMSB, and contains 500 types of Trojans, virus software, etc., which can cause abnormal behaviors in the container. The normal software dataset comes from the system commands that come with the container, some daily commonly used software installed, CPU and memory benchmark programs, etc., totaling 500.
[0039] In this embodiment, the evaluation indicators include accuracy, precision, recall and F1-score, and the data set is divided according to the ratio of training set: validation set: test set equal to 7:1:2.
[0040] The embodiments in this specification are described in a progressive manner. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the methods.
[0041] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or terminal device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or terminal device. In the absence of further restrictions, the elements defined by the sentence "include..." or "comprise..." do not exclude the existence of other elements in the process, method, article or terminal device including the elements. In addition, in this article, "greater than", "less than", "exceed" and the like are understood to exclude the number itself; "above", "below", "within" and the like are understood to include the number itself.
[0042] Although the above embodiments have been described, once those skilled in the art know the basic creative concepts, they can make additional changes and modifications to these embodiments. Therefore, the above description is only an embodiment of the present invention and does not limit the patent protection scope of the present invention. Any equivalent structure or equivalent process transformation made by using the contents of the specification and drawings of the present invention, or directly or indirectly used in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A method for detecting abnormal behavior of containers based on hardware performance counters, characterized in that: The following steps are included: 1) Build a Docker container platform, deploy Linux containers on it, mount benign software and malware samples stored in the host machine into the Linux container, run the samples in the Linux container, and use the Perf tool provided by the host Linux system outside the Linux container to collect hardware performance counter timing data when the Linux container is running; 2) Use Python script to preprocess the collected data, remove interference items, retain valid values, and save the processed data into a csv file so that it can be used as input for the deep learning model; 3) Construct a bidirectional long short-term memory network integrating the attention mechanism, use the preprocessed data as the input of the deep learning model for training, and obtain an abnormal behavior detection model.
2. The method for detecting abnormal behavior of a container based on hardware performance counters according to claim 1, characterized in that: The method of obtaining the hardware performance counter timing data in step 1) is: 1.1) Install Ubuntu system on bare metal as host, build Docker container platform on the host, and use Docker container platform to deploy Linux container; 1.2) Mount the benign software and malware samples stored in the host machine into the Linux container; 1.3) Write a shell script in the host machine to realize the automatic running of samples. There are N batches in total. Each batch collects X hardware event feature data and saves them in a txt file. This is to prevent the influence of time division multiplexing on data collection and make the collected data more accurate.
3. The method for detecting abnormal behavior of a container based on hardware performance counters according to claim 2, characterized in that: The method of preprocessing the time series data in step 2) is as follows: 2.1) Each sample corresponds to N txt files. The X hardware event feature data in each txt file are processed separately, and then the N*X features are merged together; 2.2) For each txt file, read it line by line, skip invalid lines starting with the '#' character, and display the valid lines as <notcounted> The value of is 0, and the commas are removed from the numerical values, and they are processed in chronological order;< / notcounted> 2.3) After merging the N batches of data corresponding to each sample, merge the time series data of all samples into a csv file.
4. The method for detecting abnormal behavior of a container based on hardware performance counters according to claim 3, characterized in that: The step 3) is as follows: 3.1) Build a bidirectional long short-term memory network with an attention mechanism, read the csv file obtained by step 2) preprocessing, where the sample id is 1-1000, each sample corresponds to 100 sets of time series data, the benign software label is 0, and the malicious software label is 1; 3.2) The bidirectional long short-term memory network with integrated attention mechanism captures the global time dependency through forward and backward processing. The attention mechanism assigns weights to each time step, calculates the average attention weight of each feature over all time steps, selects the 12 features with the largest contribution, highlights the information of key time steps and important features, ignores the noise of minor time steps and minor features, greatly reduces the gradient attenuation problem of long time series processing, and can better remember previous information.
5. The method for detecting abnormal behavior of a container based on hardware performance counters according to claim 4, characterized in that: The bidirectional long short-term memory network structure of the fusion attention mechanism is as follows: the first layer is the Bi-LSTM layer. The Sigmoid activation function is used inside the Bi-LSTM to calculate the values of the input gate, the forget gate and the output gate. The output range of the Sigmoid function is between 0 and 1, indicating the degree of "open" or "close" of each gate. The Tanh activation function is used to calculate the state update. The output range of the Tanh function is between -1 and 1, which can help the LSTM unit maintain the balance of information. The second layer is the Attention layer. The Attention mechanism calculates the attention weight for the hidden state of each time step and generates a context vector. The third layer is the fully connected layer. The fourth layer is the Sigmoid layer, which is used to convert the output into a probability value to determine whether the container has abnormal behavior.