Container anomaly detection method based on improved graph mask auto-encoder network

By improving the graph mask autoencoder network to construct the system call trajectory graph and feature matrix, and combining it with a deep learning model, the problems of low accuracy and network redundancy in container anomaly detection in the existing technology are solved, and efficient and accurate container anomaly detection is achieved.

CN120994511APending Publication Date: 2025-11-21JIANGSU YITONG HIGH TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510975363.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-15
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing container anomaly detection methods do not fully consider system call features, resulting in low detection accuracy. They also have high feature dimensionality and redundant network scale, failing to meet the requirements for high precision and high reliability.

Method used

An improved graph mask autoencoder network is adopted. By constructing a system call trajectory graph and feature matrix, and combining thread features, context features and parameter features, it uses adaptive weights to integrate GCN, GAT and GraphSAGE networks for encoding, and combines deep learning models for anomaly detection.

Benefits of technology

It improves the accuracy of container anomaly detection, reduces false alarm rate, reduces model computation cost and network redundancy, and ensures system security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SMS_1
    Figure SMS_1
  • Figure SMS_2
    Figure SMS_2
  • Figure SMS_3
    Figure SMS_3
Patent Text Reader

Abstract

The invention discloses a container anomaly detection method based on an improved graph mask auto-encoder network. The method comprises the following steps: firstly, acquiring each system call name and call number in a container mirror image with relatively high security, and generating system call statistical characteristics; then, acquiring a system call sequence used when the container application runs, and thread features, context features and parameter features called by each system; then, constructing a trajectory diagram by using the system call sequence, and constructing a feature matrix by combining each feature of the system call; designing an improved graph mask auto-encoder model, inputting an adjacent matrix and an unmasked feature matrix of the trajectory graph into the model, and obtaining a reconstructed graph and features; and finally, constructing a deep learning container anomaly detection model, inputting the reconstructed trajectory diagram and the feature matrix, and outputting a container anomaly judgment result. According to the method, system calling statistical features are creatively fused, an improved graph mask auto-encoder model is designed, and a deep learning model is deeply constructed from two dimensions of statistical feature fusion and feature dimension reduction, so that on one hand, the calculation cost of the model is greatly saved, and the redundancy of a network scale is reduced; and on the other hand, the system detection accuracy is effectively improved, the false alarm rate is reduced, and the safety of containers in the system is ensured in multiple aspects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of anomaly detection, and specifically relates to a container anomaly detection method based on an improved graph mask autoencoder network. Background Technology

[0002] The rapid development of IoT technology has brought to the forefront issues of IoT data security. Specifically in cloud computing, container technology is gaining widespread adoption as technology advances. Compared to traditional virtual machine technology, container technology is more lightweight and offers advantages such as rapid deployment, lower resource consumption, and higher efficiency. Therefore, the application of container technology in specific environments like this system still holds significant research value.

[0003] Currently, container anomaly detection technology mainly employs three methods: rule-based learning, machine learning, and deep learning. While rule-based learning was once widely used for anomaly detection, its reliance on pre-defined rule sets makes it difficult to handle large-scale data and complex application environments, leading to its gradual obsolescence in modern data processing. In contrast, machine learning and deep learning methods, by automatically learning patterns from large amounts of historical data, can more flexibly identify anomalous behavior, significantly improving detection accuracy and efficiency. However, these general-purpose machine learning and deep learning models exhibit significant limitations when applied to specific domains, such as systems. Systems possess unique operational characteristics and security requirements, and their time-series system call data contains specific risk characteristics. Relying solely on general sequence features for analysis often overlooks these critical factors, resulting in inaccurate detection results that fail to meet the system's demands for high precision and reliability. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this invention provides a container anomaly detection method based on an improved graph mask autoencoder network. This method solves the problems of low anomaly detection accuracy caused by insufficient consideration of system call features in existing methods, as well as increased feature coupling and network redundancy due to high feature dimensionality.

[0005] To solve the above-mentioned technical problems, the present invention adopts the following technical solution.

[0006] This invention discloses a container anomaly detection method based on an improved graph mask autoencoder network, comprising the following steps:

[0007] Step 1: Select container images that are frequently used and have high security performance in common scenarios, and generate a collection of container images;

[0008] Step 2: Perform static analysis on each container image to extract system call data and generate system call statistical features;

[0009] Step 3: Dynamically capture the system call sequence during container runtime using the strace tool, and synchronously record the thread characteristics, context characteristics, and parameter characteristics of each system call;

[0010] Step 4: Using a single system call as a node and the relationship between different system calls as edges, construct a system call trajectory graph based on the system call sequence, and integrate the statistical features from Step 2 with the thread features, context features, and parameter features from Step 3 to construct a feature vector;

[0011] Step 5: Input the unmasked feature matrix composed of the adjacency matrix and eigenvectors of the trajectory graph into the improved graph mask autoencoder model, and output the reconstructed trajectory graph and feature matrix;

[0012] The improved graph mask autoencoder model uses adaptive weight integration of three networks, GCN, GAT and GraphSAGE, as the encoder and matrix inner product as the decoder.

[0013] Step 6: Input the reconstructed trajectory map into the deep learning model in the form of an adjacency matrix along with the reconstructed feature matrix, train the model, and finally output the anomaly detection result.

[0014] In a preferred embodiment of the present invention, the statistical feature is the total number of times the same system call appears in different system call sets.

[0015] In a preferred embodiment of the present invention, the static analysis method is as follows:

[0016] Extract the target binary file of the security image, disassemble it, and obtain all the function calls involved in the code; establish the mapping relationship between the function calls and system calls, and obtain all relevant system call data.

[0017] In a preferred embodiment of the present invention, the thread feature is the system call thread ID; the context feature is the association between the system call and the preceding / following calls in the sequence; and the parameter feature is the system call input parameter value.

[0018] In a preferred embodiment of the present invention, the expression for the system call trajectory graph is:

[0019] G = (V, E)

[0020] In the formula, V = {v1, v2, ..., v} i ,...,v j ,...,v n} represents the n nodes of the system call trajectory graph, with each node representing a system call; E = {e 11 ,e 12 ,...,e1n ,e 21 ,...,e 2n ,...,e ij ,...,e nn} represents the edges between two system calls in the system call trajectory graph. An edge is a connection between two different system calls.

[0021] In a preferred embodiment of the present invention, the feature vector is a four-dimensional feature vector, and its expression is:

[0022] X i =[a i ,b i ,c i ,d i ]

[0023] In the formula, a i This represents the statistical characteristics of the current system calls during the runtime of the i-th container; b i Indicates the thread characteristics of the i-th system call; c i Indicates the context characteristics of the i-th system call; d i This represents the parameter characteristics of the i-th system call.

[0024] In a preferred embodiment of the present invention, the training loss function of the improved graph mask autoencoder model is:

[0025]

[0026] Among them, A ij represents the elements of the original adjacency matrix, and W is the adaptive weight matrix.

[0027] In a preferred embodiment of the present invention, the deep learning model includes three sequentially connected graph convolutional layers, an average pooling layer, and a fully connected layer, wherein the three graph convolutional layers employ the ReLU activation function.

[0028] The beneficial effects of this invention are that, compared with the prior art, it provides a container anomaly detection method based on an improved graph mask autoencoder network. By combining thread features, context features, parameter features, and statistical features of system calls to construct a feature matrix, the unmasked feature matrix and the adjacency matrix generated from the dynamically captured system call trajectory graph are input into a pre-designed improved graph mask autoencoder model. The obtained reconstructed graph and features are then input into a pre-constructed deep learning container anomaly detection model, achieving intelligent anomaly detection of containers. Starting from the two dimensions of fusing statistical features and feature dimensionality reduction, a deep learning model is constructed in depth. On the one hand, this significantly saves the computational cost of the model and reduces the redundancy of the network scale; on the other hand, it effectively improves the accuracy of system detection and reduces the false positive rate, thereby ensuring the security of containers in the system from multiple aspects. Attached Figure Description

[0029] Figure 1 This is a flowchart of the container anomaly detection method based on an improved graph mask autoencoder network in this invention.

[0030] Figure 2 This is an example diagram of obtaining the called function in an embodiment of the present invention.

[0031] Figure 3 This is an example diagram of system call representation in an embodiment of the present invention.

[0032] Figure 4 This is an overall structural diagram of the deep learning anomaly detection model in this embodiment of the invention. Detailed Implementation

[0033] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.

[0034] The embodiments described in this application are merely some, not all, embodiments of the present invention. Based on the spirit of the present invention, other embodiments obtained by those skilled in the art without inventive effort are all within the protection scope of the present invention.

[0035] See Figure 1 As shown, the container anomaly detection method based on an improved graph mask autoencoder network disclosed in this invention includes the following steps:

[0036] Step 1: Select container images that are frequently used in common scenarios and have high security performance, and generate a collection of container images.

[0037] Specifically, we searched for and pulled commonly used container images in general scenarios. Based on their usage frequency and security performance, we selected container images with high security that are frequently used in the system to form a set of secure images for future research.

[0038] Step 2: Perform static analysis on each container image to obtain system calls and generate system call statistics.

[0039] See Figure 2In the started container, the "find" command is used to locate the target binary file within the security image. Then, the IDA Pro tool is used to disassemble the binary file, and the "Functions" function in "Open subviews" under the "View" directory is used to directly obtain all the call functions involved in the file. The mapping relationship between the call functions and system calls is established through the system call entry function, and then all system call data related to the call functions is obtained by referring to the mapping relationship.

[0040] See Figure 3 In Linux systems, the system call table in the "unistd_64.h" system file is viewed using the command "cat / usr / include / asm / unistd_64.h". This table is used to obtain a one-to-one mapping between system calls and system call numbers. Finally, the system call number corresponding to each system call is found based on this mapping.

[0041] A system call may appear in multiple sets. The number of times it appears is used as a statistical feature. One appearance is "1", two appearances are "2", and so on. The other two parts record the system call name and the corresponding system call number, respectively. For example, if the system call read appears 5 times, it is recorded as "read 0 5".

[0042] Step 3: Use the strace tool to dynamically capture the system call sequence during container runtime, and synchronously record the thread characteristics, context characteristics, and parameter characteristics of each system call.

[0043] Specifically, install the strace tool in the running container environment and use the "strace ls" command to trace all processes within the container. When it returns information such as the system call thread, context, parameters, return value, and execution time, record all relevant system calls. For example, if the returned information is "openat(AT_FDCWD," / etc / ld.so.cache",O_RDONLY|O_CLOEXEC)=4", record openat. Next, refer to the system call table in the "unistd_64.h" system file to convert the system calls into system call numbers, thereby obtaining the system call sequence. For example, record "execve brkopenat fstat mmap close" as "59 12 257 5 9 3".

[0044] Furthermore, the thread, context, and parameter information of the system call are recorded as features. For example, when the returned information is "725 21:09:22.295305661 0apache2 16680>openat flags=0", "16680" is recorded as the thread feature, the relationship between "openat" and the preceding and following system calls in the "execve brk openat fstat mmap close" sequence is recorded as the context feature, and "flags=0" is recorded as the parameter feature.

[0045] Step 4: Using a single system call as a node and the relationship between different system calls as edges, construct a system call trajectory graph based on the system call sequence, and map the statistical features generated in Step 2 to dynamic nodes through the system call number, and fuse them with the thread features, context features and parameter features in Step 3 to construct a feature vector.

[0046] Specifically, the system call trajectory graph can be abstractly represented as follows:

[0047] G = (V, E)

[0048] In the formula, V = {v1, v2, ..., v} i ,...,v j ,...,v n} represents the n nodes of the system call trajectory graph, with each node representing a system call; E = {e 11 ,e 12 ,...,e 1n ,e 21 ,...,e 2n ,...,e ij ,...,e nn} represents the edges between two system calls in the system call trajectory graph, where one edge connects two different system calls.

[0049] Specifically, if a system call appears repeatedly in a sequence, that system call is counted as only one node, and the resulting trajectory graph is an undirected graph. The connection between two different system calls is only recorded once. For example, if the system call sequence is "11 1 1 3 259 259 3 1 3", then after converting it into a system call trajectory graph, the... Figure 3 The nodes are {1, 3, 259}; the two edges are {<1, 3>, <3, 259>}.

[0050] The construction of the feature vector further includes:

[0051] Constructing a four-dimensional feature vector, abstractly represented as:

[0052] X i =[ai ,b i ,c i ,d i ]

[0053] In the formula a i This represents the statistical characteristics mapped by system call number; b i Indicates the thread characteristics of the i-th system call; c i Indicates the context characteristics of the i-th system call; d i This represents the parameter characteristics of the i-th system call.

[0054] The feature matching mechanism is as follows: static statistical features are associated with dynamically captured nodes through system call numbers.

[0055] Step 5: Input the unmasked feature matrix composed of the adjacency matrix and feature vectors of the trajectory graph into the improved graph mask autoencoder model, and output the reconstructed graph and features.

[0056] The improved graph mask autoencoder model uses adaptive weight ensembles GCN, GAT, and GraphSAGE as the encoder and matrix inner product as the decoder.

[0057] The encoder takes the adjacency matrix A and the unmasked node feature matrix X as input, and obtains node features in the hidden layers through GCN, GAT, and GraphSAGE network layers, respectively. Then, it integrates the node features of each network through adaptive weights to form new node features, and finally obtains the node embedding matrix Z through a fully connected layer. The adaptive weights W are learned through training, and the decoder is used to reconstruct the original graph.

[0058] Furthermore, the original features of the reconstructed features and the masked nodes are applied with a typical reconstruction loss, such as cross-entropy loss, and an L-regularization is added to update the adaptive weights W:

[0059]

[0060] If 30% of the node features are randomly occluded as input to the model, the weights W can be optimized by reconstructing the loss.

[0061] First, a pre-trained image mask autoencoder is trained, and then the detection model is trained after the parameters are fixed.

[0062] Step 6: Input the reconstructed trajectory map into the constructed deep learning model in the form of an adjacency matrix along with the reconstructed feature matrix, train the model, and finally output the anomaly detection result.

[0063] The deep learning model is built based on a graph convolutional neural network, such as Figure 4As shown, the system consists of three graph convolutional layers, one average pooling layer, and one fully connected layer. The fully connected layer connects to a softmax classifier to output anomaly probabilities. The three graph convolutional layers are connected by a ReLU activation function. Then, the average value of all output node vectors from each reconstructed image is passed to the average pooling layer, and the average feature is then passed to the fully connected layer. Finally, the overall feature output from the fully connected layer is passed to the softmax function for anomaly detection result prediction.

[0064] The deep learning model is trained using the Adam optimizer, and the Loss cross-entropy loss function is selected as the objective function. The loss function is specifically expressed as follows:

[0065]

[0066] In the formula, n is the number of training samples, and y is the true label of the sample. This is the predicted result.

[0067] Furthermore, the Batch-size was set to 8, the Epoch to 300, and the learning rate for the first 100 rounds was set to 0.001, while the learning rate for the last 200 rounds was set to 0.0002 to achieve the best detection results.

[0068] The beneficial effects of this invention are that, compared with the prior art, it provides a container anomaly detection method based on an improved graph mask autoencoder network. By combining thread features, context features, parameter features, and statistical features of system calls to construct a feature matrix, the unmasked feature matrix and the adjacency matrix generated from the dynamically captured system call trajectory graph are input into a pre-designed improved graph mask autoencoder model. The obtained reconstructed graph and features are then input into a pre-constructed deep learning container anomaly detection model, achieving intelligent anomaly detection of containers. Starting from the two dimensions of fusing statistical features and feature dimensionality reduction, a deep learning model is constructed in depth. On the one hand, this significantly saves the computational cost of the model and reduces the redundancy of the network scale; on the other hand, it effectively improves the accuracy of system detection and reduces the false positive rate, thereby ensuring the security of containers in the system from multiple aspects.

[0069] The above description is merely an embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent structural or procedural transformations made based on the content of the present invention's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of the present invention.

Claims

1. A method for container anomaly detection based on improved graph mask autoencoder network, characterized in that, The method comprises the following steps: Step 1: selecting a container image with high frequency and high security performance in a general scenario, and generating a container image set; Step 2: performing static analysis on each container image respectively, extracting system call data and generating system call statistical features; Step 3: dynamically capturing the system call sequence of the container runtime through the strace tool, and synchronously recording the thread features, context features and parameter features of each system call; Step 4: taking a single system call as a node and the relationship between different system calls as an edge, constructing a system call trajectory graph according to the system call sequence, and fusing the statistical features of step 2 and the thread features, context features and parameter features of step 3 to construct a feature vector; Step 5: inputting the unmasked feature matrix composed of the adjacency matrix of the trajectory graph and the feature vector into the improved graph mask autoencoder model, and outputting the reconstructed trajectory graph and feature matrix; The improved graph mask autoencoder model integrates GCN, GAT and GraphSAGE three networks as an encoder through adaptive weights, and takes matrix inner product as a decoder; Step 6: inputting the reconstructed trajectory graph in the form of adjacency matrix together with the reconstructed feature matrix into a deep learning model, and training, and finally outputting an abnormality detection result.

2. The container anomaly detection method based on improved graph mask autoencoder network according to claim 1, wherein, The statistical features are the total number of times of the same system call appearing in different system call sets.

3. The container anomaly detection method based on improved graph mask autoencoder network according to claim 2, characterized in that, The method of static analysis is: Extract the target binary file of the security image, disassemble it, and obtain all the calling functions involved in the code; establish the mapping relationship between the calling functions and the system calls, and obtain all the related system call data.

4. The container anomaly detection method based on improved graph mask autoencoder network according to claim 1, wherein, The thread features are system call thread IDs; the context features are the association relationship of the system call with the previous / next call in the sequence; and the parameter features are the parameter value passed in by the system call.

5. The container anomaly detection method based on improved graph mask autoencoder network according to claim 1, wherein, The expression of the system call trajectory graph is: G=(V,E) In the formula, V = {v1, v2, ..., v} i ,...,v j ,...,v n } represents the n nodes of the system call trajectory graph, with each node representing a system call; E = {e 11 ,e 12 ,...,e 1n ,e 21 ,...,e 2n ,...,e ij ,...,e nn } represents the edges between two system calls in the system call trajectory graph. An edge is a connection between two different system calls.

6. The container anomaly detection method based on improved graph mask autoencoder network according to claim 1, wherein, The feature vector is a four-dimensional feature vector, and its expression is: X i = [a i ,b i ,c i ,d i ] In the formula, a i represents the statistical characteristics of the current system call when the i-th container is running; b i thread characteristics of the i-th system call; c i represents a context feature of the i-th system call; d i represents the parameter feature of the i-th system call.

7. The container anomaly detection method based on improved graph mask autoencoder network according to claim 1, wherein, The training loss function of the improved graph mask autoencoder model is: where A ij is the original adjacency matrix element, and W is the adaptive weight matrix.

8. The container anomaly detection method based on improved graph mask autoencoder network according to claim 1, wherein, The deep learning model comprises three graph convolution layers, an average pooling layer and a fully connected layer connected in sequence, wherein the three graph convolution layers adopt ReLU activation function.