Cloud computing platform with container cloud real-time anomaly detection
Patent Information
- Application Number
- CN202311187216.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-14
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2043-09-14
AI Technical Summary
此外,在特征选取方面,多数方法往往只考虑采用系统调用类型作为主要特征,忽略了参数与具体行为,导致信息的丢失,并且无法良好地适用于已有的较成熟日志异常检测方法中
[0017]根据本发明所涉及的具有容器云实时异常检测的云计算平台,因为通过规则检测过滤模块过滤得到原始系统调用事件,再将原始系统调用事件依次通过事件处理模块、候选搜索模块、层次聚类模块、规则模板挖掘模块、条件决策模块和规则适配模块进行处理,得到更新模板并实现无监督的规则更新,使规则集中的规则数据能够适用于各类系统调用且关注具体参数,解决了已有自动化LSM规则生成方法中控制粒度过粗的问题。所以,本发明的具有容器云实时异常检测的云计算平台能够为容器云提供安全监控与低误报的实时异常检测,提高云计算平台的安全性。
Smart Images

Figure CN117097550B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of cloud computing, and more specifically to a cloud computing platform with real-time anomaly detection for container clouds. Background Technology
[0002] Container technology is one of the mainstream technologies in the cloud computing field today, and the container cloud ecosystem, primarily based on Kubernetes and Docker, has been widely used in major cloud computing platforms. Compared to virtual machines, containers share the same kernel as the host machine, thus offering significant advantages such as faster startup speeds, higher portability, and stronger scalability. However, the lower resource isolation and shared kernel characteristics introduce new security risks to containers and cloud platforms, such as exploiting kernel vulnerabilities to allow containers to escape, thereby affecting the confidentiality, integrity, and availability of the host machine and the cloud platform. Therefore, security monitoring of cloud containers and real-time anomaly detection are crucial for ensuring the security and observability of container cloud platforms.
[0003] System call-based rule-based detection is a common method for container anomaly detection and defense, using Linux Security Modules (LSMs) such as Seccomp and AppArmor. A major problem with this approach is its often complex configuration process. The sheer number and rapid iteration of containers on cloud platforms make it virtually impossible to manually write rule files for every container. While some research has proposed automated methods for generating Seccomp and AppArmor policies, these methods, due to limitations in the analyzed data or the corresponding LSM implementations themselves, cannot achieve sufficiently granular anomaly detection rules. This leads to misjudgments of normal behavior or the bypassing of rules by malicious actors. Furthermore, these configuration files are typically static, preventing them from dynamically adapting to environmental changes, such as image updates.
[0004] On the other hand, numerous studies have utilized unsupervised machine learning or neural networks to detect anomalies in container system call behavior. However, based on publicly available datasets and real-world observations, a single container application can generate millions of system calls per second during normal operation. Under high load and the scale of the entire cloud platform, such a volume of data places enormous pressure on machine learning analysis and consumes significant resources. Furthermore, even models with sufficiently high accuracy will experience a deluge of false alarms when faced with such a large amount of data, leading to an overwhelming number of alerts and posing a serious challenge to their use in real-world production environments. Filtering out useless system call data before inputting it into a machine learning model is a crucial step. Moreover, in terms of feature selection, most methods tend to only consider system call type as the primary feature, ignoring parameters and specific behaviors, resulting in information loss and making them unsuitable for use with existing, more mature log anomaly detection methods.
[0005] In summary, existing technologies cannot achieve real-time anomaly detection in container clouds with low false alarms under unsupervised and low-cost conditions. Summary of the Invention
[0006] This invention is made to solve the above problems, and its purpose is to provide a cloud computing platform with real-time anomaly detection for container clouds.
[0007] This invention provides a cloud computing platform with real-time anomaly detection for container clouds, characterized by comprising: multiple container cloud application nodes and an analysis service node. Each container cloud application node includes a system call capture module, a rule detection and filtering module, and an event output module. The system call capture module collects system call event data from containers within the container cloud application nodes. The rule detection and filtering module stores rule data and uses it to detect and filter system call event data, obtaining unfiltered system call events as raw system call events. The event output module outputs the raw system call events to the analysis service node. The analysis service node includes an event processing module, a system call log storage module, an activity template storage module, a candidate search module, a hierarchical clustering module, a rule template mining module, a conditional decision module, a rule adaptation module, a rule base module, and a data parsing module. The event processing module receives raw system call events sent by each container cloud application node and extracts information fields from the raw system call events to obtain system call logs. The system call log storage module stores system call logs. The system consists of several modules: an activity template storage module (containing multiple templates), a candidate search module (selecting m candidate templates corresponding to system call logs from all templates as a candidate template set), a hierarchical clustering module (selecting the candidate template with the highest similarity from the candidate template set as the optimal matching template), a rule template mining module (obtaining updated templates based on the optimal matching template and system call logs), a conditional decision module (storing preset decision conditions to update templates in the activity template storage module and labeling corresponding system call logs in the system call log storage module with anomaly judgment tags), a rule adaptation module (periodicly converting updated templates in the activity template storage module into corresponding rule sets and sending rule update messages to the rule detection and filtering module), a rule base module (storing rule sets), a data parsing module (using the variable content and corresponding differences and distances between system call logs and the optimal matching template as parsed data for log anomaly analysis), and a rule detection and filtering module (extracting the corresponding rule set from the rule base module based on the rule update message and updating the rule data accordingly).
[0008] The cloud computing platform with real-time anomaly detection for container clouds provided by this invention may also have the following features: In the event processing module, information field extraction includes field identification, subject identification, and word segmentation. Field identification is used to extract system call parameter information and metadata related to containers and container cloud application nodes from the original system call event as field data. Subject identification is used to identify the field data based on container ID, image ID, and process-related information, and set corresponding identifiers in the system call log. Word segmentation is used to identify special content in the field data using regular expressions and record the specific content, then mask it with an identifier name. The field data is segmented according to basic delimiters and dynamic segmentation rules to obtain a token sequence as the data for the corresponding field in the system call log.
[0009] The cloud computing platform with real-time anomaly detection for container clouds provided by this invention may also have the following features: In the candidate search module, candidate templates are obtained as follows: when the length of the token sequence is less than or equal to 3, templates with length differences within a preset range are selected from all templates as candidate templates; when the length of the token sequence is greater than 3, the coarse-grained similarity between the token sequence and each template is calculated to obtain a similarity score (D,Q), and all templates with similarity scores (D,Q) greater than a threshold are selected as candidate templates. The formula for calculating the similarity score (D,Q) is as follows: In the formula, D is the template, and Q is the set of morphemes q that are included in the query morpheme set. i For each token sequence, f(q) i D) is the morpheme q i The number of times it appears in template D, |D| is the number of words in template D, avgdl is the average document length, k1 and b are adjustable parameters, N is the total number of templates, n(q i ) contains the morpheme q i The number of templates.
[0010] The cloud computing platform with real-time anomaly detection for container clouds provided by this invention may also have the following feature: when m=0, the data parsing module uses the template with the highest similarity score (D, Q) as the optimal matching template.
[0011] The cloud computing platform with real-time anomaly detection for container clouds provided by this invention may also have the following features: the template is in the form of a tree, the non-leaf nodes of the tree are patterns that match all token sequences belonging to the template, the leaf nodes of the tree are records of each token sequence, and the degree and maximum height of the tree are fixed values. In the hierarchical clustering module, each non-leaf node of each candidate template is aligned with the token sequence using the longest common subsequence (LCS), and then similarity is calculated. The formula for this similarity calculation is as follows: In the formula, |s1| and |s2| are the lengths of the non-leaf node and the corresponding string of the token sequence after alignment, respectively; m is the number of matched characters; t is the number of transposes; D is the template; Q is the input token sequence; Length(LCS(D, Q)) is the LCS length of the template D and the token sequence Q; |Diffs| is the number of differences; sim j This represents the similarity between differences.
[0012] The cloud computing platform with real-time anomaly detection for container clouds provided by this invention may also have the following features: In the rule template mining module, the token sequence of the system call log is merged with the record of the optimal matching template to obtain a merged template. When the merged template is different from the optimal matching template, the merged template is used as the update template. When m=0, the template is constructed according to the corresponding token sequence to obtain the update template.
[0013] The cloud computing platform with real-time anomaly detection for container clouds provided by this invention may also have the following feature: wherein the decision conditions in the conditional decision module are:
[0014] All entities of the same type in the entire cluster generate system calls similar to the original system call event, and the overall frequency and the number of containers reach the preset value; the time interval between the original system call events is small; the context information of the original system call events is similar; and when additional anomaly detection analysis is performed on the original system call event, no anomaly is detected.
[0015] The cloud computing platform with real-time anomaly detection for container cloud provided by this invention may also have the following feature: the rule detection and filtering module updates the rule data by hot-updating or restarting the container and configuring the rule set as new rule data according to update requirements.
[0016] The role and effect of invention
[0017] The cloud computing platform with real-time anomaly detection for container clouds according to this invention obtains original system call events through a rule detection and filtering module. These events are then processed sequentially through an event processing module, a candidate search module, a hierarchical clustering module, a rule template mining module, a conditional decision module, and a rule adaptation module to obtain updated templates and achieve unsupervised rule updates. This ensures that the rule data in the rule set is applicable to various system calls and focuses on specific parameters, solving the problem of excessively coarse control granularity in existing automated LSM rule generation methods. Therefore, the cloud computing platform with real-time anomaly detection for container clouds of this invention can provide security monitoring and low false positive real-time anomaly detection for container clouds, improving the security of the cloud computing platform. Attached Figure Description
[0018] Figure 1 This is a schematic diagram of the framework of the cloud computing platform in an embodiment of the present invention;
[0019] Figure 2 This is a schematic diagram of the framework of a container cloud application node in an embodiment of the present invention;
[0020] Figure 3 This is a schematic diagram of the framework for analyzing service nodes in an embodiment of the present invention;
[0021] Figure 4 A schematic diagram of the process for generating an update template in an embodiment of the present invention;
[0022] Figure 5 This is a schematic diagram illustrating an example of the template update process in an embodiment of the present invention. Detailed Implementation
[0023] To make the technical means, creative features, objectives and effects of the present invention easy to understand, the following embodiments, in conjunction with the accompanying drawings, specifically illustrate the cloud computing platform with real-time anomaly detection for container clouds of the present invention.
[0024] Figure 1 This is a schematic diagram of the framework of the cloud computing platform in an embodiment of the present invention.
[0025] like Figure 1 As shown, the cloud computing platform 100 includes multiple container cloud application nodes 10 and an analysis service node 20. The analysis service node 20 communicates with each container cloud application node 10 through a switch 30.
[0026] Figure 2 This is a schematic diagram of the framework of a container cloud application node in an embodiment of the present invention.
[0027] like Figure 2As shown, the container cloud application node 10 includes a system call capture module 101, a rule detection and filtering module 102, an event output module 103, and a container cloud application control module 104 that controls the above modules.
[0028] The system call capture module 101 is used to collect system call event data of containers in container cloud application nodes.
[0029] The rule detection and filtering module 102 stores rule data and is used to detect and filter system call event data according to the rule data, so as to obtain the unfiltered system call events as the original system call events.
[0030] The event output module 103 is used to output the raw system call events to the analysis service node 20.
[0031] Figure 3 This is a schematic diagram of the framework for analyzing service nodes in an embodiment of the present invention.
[0032] like Figure 3 As shown, the analysis service node 20 includes an event processing module 201, a system call log storage module 202, an activity template storage module 203, a candidate search module 204, a hierarchical clustering module 205, a rule template mining module 206, a conditional decision module 207, a rule adaptation module 208, a rule base module 209, a data parsing module 210, and an analysis service control module 211 that controls the above modules.
[0033] The event processing module 201 is used to receive the raw system call events sent by each container cloud application node 10, and extract the information fields of the raw system call events to obtain the system call log. In this embodiment, the system call log is semi-structured data containing each information field of the system call.
[0034] Information field extraction includes field recognition, subject recognition, and word segmentation.
[0035] Field identification is used to extract system call parameter information and metadata related to containers and container cloud application nodes 10 from the original system call events as field data. In this embodiment, the field data serves as the basis for subject identification and template mining.
[0036] The subject identification is used to identify field data based on container ID, image ID, and process-related information, and set the corresponding identifier in the system call log.
[0037] In this embodiment, the primary function of the image ID is to distinguish the service to which a container belongs within the container cloud application node cluster, thereby generating different templates and rule sets for different services. This field can also be replaced by other tags, such as Pod and Deployment tags. The container ID identifies the source of an event and is used to decide whether to use the template corresponding to a certain event as a new rule based on distribution and other factors. This field can also be further composed of the container cloud application node name, pod ID, etc. For process information, identification is performed through process name and parameter clustering, as well as process chain identification, to accurately obtain the calling entity.
[0038] Tokenization is used to identify and record special content in field data using regular expressions, and then mask it with an identifier. The field data is segmented according to basic delimiters and dynamic segmentation rules to obtain a token sequence, which serves as the data for the corresponding field in the system call log. In this embodiment, the identifier is named "". <url>The basic separators include "()-=\s+ / :;'[".
[0039] The system call log storage module 202 is used to store system call logs. In this embodiment, all system call logs are persistently stored to provide a basis for tracing the source of anomalies.
[0040] The activity template storage module 203 stores multiple templates.
[0041] The template is in the form of a tree. The non-leaf nodes of the tree are patterns that match all token sequences belonging to the template, and the leaf nodes of the tree are records of each token sequence. The degree and maximum height of the tree are fixed values.
[0042] In this embodiment, the basic fields covered by the template include template_info, syscall_type, image_info, process_info, event_action, event_results, and event_data.
[0043] The fields in template_info contain basic information about the template itself, including id, update time, etc.
[0044] The content of the syscall_type field is the system call type information, including the type name and direction, i.e., enter or return. This field is an enumeration value and is not involved in template mining.
[0045] The content of the image_info field is image-related information, including image name and version.
[0046] The fields in process_info contain process-related information, including process name, process path, process parameters, and information about the parent process.
[0047] The event_action field includes three subfields: action_type, action_opt, and action_subjects. The content of these fields contains event operation information, including operation type (e.g., tcp, file), operation options (e.g., read-only or read-write mode when opening a file), and a list of operation targets (e.g., the specific file path or IP address to be opened). Lists are used for operations involving multiple objects, such as rename. The type and options are enumerated values and are not used in template mining.
[0048] The event_results field contains a list of event results, including return values, data size, etc.
[0049] The content of the event_data field is event-related data, such as the content written when the Write function is called.
[0050] Candidate search module 204 is used to select m candidate templates corresponding to system call logs from all templates as candidate template set C = {D1, ..., D2}. m }, where D m For the m-th candidate template, when m=0, the data parsing module 210 will take the template with the highest similarity score (D, Q) as the optimal template.
[0051] The candidate templates are obtained in the following way:
[0052] When the length of the token sequence is less than or equal to 3, select the template whose length difference is within the preset range from all templates as the candidate template;
[0053] When the length of the token sequence is greater than 3, the coarse-grained similarity between the token sequence and each template is calculated to obtain the similarity score score(D,Q). All templates with similarity scores score(D,Q) greater than the threshold are selected as candidate templates.
[0054] The similarity score (score(D, Q)) is calculated using the following formula:
[0055]
[0056]
[0057] In the formula, D is the template, and Q is the set of morphemes q that are included in the query morpheme set. i For each token sequence, f(q) i D) is the morpheme q i The number of times it appears in template D, |D| is the number of words in template D, avgdl is the average document length, k1 and b are adjustable parameters, N is the total number of templates, n(q i ) contains the morpheme q i The number of templates.
[0058] In this embodiment, the assumption that templates of similar length generally hold true in practice. Therefore, when searching for candidate templates, only templates with a certain length difference are included. Using the above search method also avoids the problems in Drain, such as its inability to handle variable-length records and records with different prefixes, making the candidate search space more accurate and its applicability wider.
[0059] The hierarchical clustering module 205 is used to select the candidate template with the highest similarity from the candidate template set as the optimal matching template through similarity calculation.
[0060] In this embodiment, the hierarchical clustering module 205 aligns each non-leaf node of each candidate template with the token sequence using the longest common subsequence (LCS). During alignment, for the wildcard <*> in the template, any number of tokens can be matched until the token after the wildcard is the same as the input token, but this match is not counted in the LCS length.
[0061] After alignment, similarity is calculated using the following formula:
[0062]
[0063]
[0064] In the formula, |s1| and |s2| are the lengths of the non-leaf node and the corresponding string of the token sequence after alignment, respectively; m is the number of matched characters; t is the number of transposes; D is the template; Q is the input token sequence; Length(LCS(D, Q)) is the LCS length of the template D and the token sequence Q; |Diffs| is the number of differences; sim j This represents the similarity between differences.
[0065] In this embodiment, the content of the token sequence during similarity calculation is the original content that is not covered by the identifier name.
[0066] The rule template mining module 206 is used to obtain updated templates based on the optimal matching template and system call logs.
[0067] In the rule template mining module 206, the token sequence of the system call log is merged with the record of the optimal matching template to obtain the merged template. When the merged template is different from the optimal matching template, the merged template is used as the update template. When m=0, the template is constructed based on the corresponding token sequence to obtain the update template.
[0068] Figure 4 A schematic diagram of the process for generating an update template in an embodiment of the present invention.
[0069] like Figure 4 As shown, generating the update template in this embodiment includes the following steps:
[0070] Step S1: Extract the differences between the token sequence and the leaf nodes (records) of the optimal matching template to determine if there are any differences. If yes, proceed to step S2; otherwise, the optimal matching template will not be updated.
[0071] Step S2: Extract the difference between the token sequence and the best-matching record as a variable identifier, and then merge them to obtain the merge pattern. In this embodiment, when the difference is an enumerated value, the variable identifier is... <e>When the difference is a range value, the variable is identified as... <r>When the difference is a free variable value, the variable is identified as <*>, and the free variable value is a string variable.
[0072] Step S3: Determine whether the degree or maximum height of the optimal matching template has reached the maximum value. If yes, proceed to step S5; otherwise, proceed to step S4.
[0073] Step S4: Use the merge pattern as the content of the node where the record is located to generate a new child node, and use the token sequence and the record as leaf nodes under the child node respectively to obtain the merged template, and proceed to step S8.
[0074] Step S5: Determine whether the degree of the optimal matching template has reached the maximum value. If yes, proceed to step S6; otherwise, proceed to step S7.
[0075] Step S6: Delete all child nodes of the node containing the record, make the node a leaf node, and use the merge pattern as the content of the node to obtain the merged template, then proceed to step S8.
[0076] Step S7: Merge the merge pattern with the parent node of the node containing the record, and then use the token sequence as the corresponding leaf node to obtain the merged template, and proceed to step S8.
[0077] Step S8: Use the merged template as the update template.
[0078] In this embodiment, when the number of inputs with the same segmentation pattern that match the differences within the token after each merging reaches a certain number, the segmentation pattern between the differences is recorded in the dynamic segmentation rules of the event handling module 201.
[0079] Figure 4 This is a schematic diagram illustrating an example of the template update process in an embodiment of the present invention.
[0080] like Figure 4 As shown, the token sequence is "curl http: / / baz.ns3.svc.cluster.local:80". (a) is the optimal matching template corresponding to this token sequence, and (b) is the updated template after merging the token sequence and the optimal matching template. This token sequence is the best match for the leaf node "curl http: / / foo.ns1.svc.cluster.local:80" of the optimal matching template, resulting in the merged pattern "curl http: / <*>.<*>.svc.cluster.local:80". At this time, the degree and maximum height of the optimal matching template have reached their maximum values. Therefore, the content of the original leaf node is replaced with the content of the merged pattern, and this node is taken as a child node. "curl http: / / foo.ns1.svc.cluster.local:80" and "curlhttp: / / baz.ns3.svc.cluster.local:80" are taken as leaf nodes under this child node, respectively, thus obtaining the updated template.
[0081] The conditional decision module 207 stores preset decision conditions, which are used to update the template in the activity template storage module 203 according to the decision conditions and the update template, and to mark the corresponding system call log in the system call log storage module 202 with anomaly judgment tags.
[0082] The decision conditions are as follows: all entities of the same type in the entire cluster generate system calls similar to the original system call event, and the overall frequency and the number of containers reach the preset value; the time interval between the original system call events is small; the context information of the original system call events is similar; and when there is additional anomaly detection analysis, no anomalies are detected after anomaly detection analysis of the original system call event.
[0083] In this embodiment, the rule set is not updated immediately by making decisions, thereby avoiding pollution of the rule set caused by abnormal noise in the template itself.
[0084] The rule adaptation module 208 is used to periodically convert the updated templates in the activity template storage module 203 into corresponding rule sets and send rule update messages to the rule detection and filtering module 102.
[0085] The rule base module 209 is used to store rule sets.
[0086] The rule detection and filtering module 102 extracts the corresponding rule set from the rule base module 209 according to the rule update message, and updates the rule data according to the rule set.
[0087] The rule detection and filtering module 102 updates the rule data according to the update requirements by hot-updating or restarting the container and configuring the rule set as new rule data.
[0088] The data parsing module 210 is used to take the variable content and corresponding differences and distances of the system call log and the optimal matching template as parsed data for log anomaly analysis. In this embodiment, the variable content and differences are the content extracted from the differences in the rule template mining module 206. The structure of the parsed data is as follows:
[0089] {template id,
[0090] Variable list: [content of variable 1, content of variable 2, ..., content of variable n],
[0091] Difference list: [{Difference 1 position, Difference 1 distance, Difference 1 content}, ..., {Difference n position, Difference n distance, Difference n content}]}.
[0092] In this embodiment, the cloud computing platform with real-time anomaly detection for container clouds of the present invention (i.e., the method itself) is compared with the SwissLog method and Drain method, which are currently the best performing log parsing algorithms, to verify the correctness and accuracy of the templates obtained by each method. The total loss is used as an evaluation index for correctness and accuracy. The formula for calculating the total loss is as follows:
[0093] Loss=LengthLoss(T)+QualityLoss(T),
[0094] LengthLoss(T)=(2og|T|) θ ,
[0095]
[0096] AvgTokenLost(t)=AvgMatchLength(t)-|t|,
[0097] In the formula, Loss is the total loss of template set T, ti is the template in template set T, AvgTokenLost(t) is the average matching length of all original logs in the log set of evaluating template t and template t, LengthLoss(T) reflects the total number of templates generated in template set T, θ is a hyperparameter set to 1.5, and QualityLoss(T) evaluates whether the generated templates are overgeneralized.
[0098] In addition, extra regular expressions were set for the URL and file path in the Drain method, with the depth and threshold set to 4 and 0.5 respectively. An extra delimiter was set for the SwissLog method, while the default configuration was used for the rest. Both methods were used in online mode, with system call logs entered line by line. The total number of templates generated and the total loss for the three methods in different scenarios are shown in the table below:
[0099]
[0100] The first column in the table above represents different scenarios. The second, third, and fourth columns represent the total number of templates generated by the proposed method, the Drain method, and the SwissLog method in different scenarios, respectively. The fifth, sixth, and seventh columns represent the total loss calculation results of the proposed method, the Drain method, and the SwissLog method in different scenarios, respectively. For example, the cell in the second column of the third row indicates that in the CVE-2017-7529 scenario, the proposed method generated a total of 8 templates. It can be seen that the proposed method, compared with the Drain method and the SwissLog method, has its advantages and disadvantages in terms of the total number of generated templates. However, in most scenarios, the proposed method has the lowest or near-optimal total loss compared to the other two methods. That is, when the total number of log templates mined by the proposed method is close, the quality of the templates is higher, meaning that they match the original logs more accurately and are more stable.
[0101] In this embodiment, the rule set generated by the cloud computing platform with real-time anomaly detection for container cloud of the present invention for Falco is used as the rule set of the present invention. It is compared with the default rules of Falco and manually written rules. The corresponding rules that correctly generate anomaly alarms in different scenarios are marked as "v", and otherwise marked as "cannot be detected". The results of anomaly detection for each rule in different scenarios are shown in the table below:
[0102] CVE-2017-7529 Unable to detect Unable to detect Unable to detect CVE-2017-12635 Unable to detect v v CVE-2018-3760 Unable to detect v v CVE-2019-5418 Unable to detect v v CVE-2020-9484 v v v CVE-2020-13942 Unable to detect v v CVE-2020-23839 v v v EPS_CWE-434 v v v PHP_CWE-434 v v v
[0103] The first column in the table above represents different scenarios. The second, third, and fourth columns show the anomaly detection results of Falco's default rules, the rules of this invention, and manually written rules in the corresponding scenarios, respectively. For example, the cell in the third row and third column indicates that the anomaly detection result of the rules of this invention in the CVE-2017-12635 scenario is "v", meaning that an anomaly alert was correctly generated. It is evident that the rules of this invention are superior to Falco's default rules and have the same anomaly detection results as manually written rules obtained by specifically tailoring Falco rules. Therefore, the rules of this invention have a better anomaly detection effect.
[0104] The role and effect of the embodiments
[0105] According to the cloud computing platform with real-time anomaly detection for container clouds involved in this embodiment, the original system call events are obtained by filtering through the rule detection and filtering module. The original system call events are then processed sequentially through the event processing module, the candidate search module, the hierarchical clustering module, the rule template mining module, the condition decision module, and the rule adaptation module to obtain the updated template and realize unsupervised rule updates. This makes the rule data in the rule set applicable to various system calls and focuses on specific parameters, thus solving the problem of excessively coarse control granularity in existing automated LSM rule generation methods.
[0106] By setting decision conditions to not immediately update the rule set, the pollution of the rule set caused by abnormal noise in the template itself is avoided. It can also mark the system call log with abnormal judgment tags, thereby providing corresponding analytical basis for subsequent anomaly analysis.
[0107] In summary, this method can provide security monitoring and real-time anomaly detection with low false alarms for container clouds, thereby improving the security of cloud computing platforms.
[0108] The above embodiments are preferred embodiments of the present invention and are not intended to limit the scope of protection of the present invention.< / r> < / e> < / url>
Claims
1. A cloud computing platform with real-time anomaly detection for container clouds, characterized in that, include: Multiple container cloud application nodes and one analytics service node, The container cloud application node includes a system call capture module, a rule detection and filtering module, and an event output module. The system call capture module is used to collect system call event data of containers in the container cloud application node. The rule detection and filtering module stores rule data and is used to detect and filter the system call event data according to the rule data, obtaining the unfiltered system call events as the original system call events. The event output module is used to output the original system call event to the analysis service node. The analysis service node includes an event processing module, a system call log storage module, an activity template storage module, a candidate search module, a hierarchical clustering module, a rule template mining module, a conditional decision-making module, a rule adaptation module, a rule base module, and a data parsing module. The event processing module is used to receive the original system call events sent by each of the container cloud application nodes, and to extract information fields from the original system call events to obtain system call logs. The system call log storage module is used to store the system call logs. The activity template storage module stores multiple templates. The candidate search module is used to select m candidate templates corresponding to the system call log from all the templates as a candidate template set. The hierarchical clustering module is used to select the candidate template with the highest similarity from the candidate template set as the optimal matching template through similarity calculation. The rule template mining module is used to obtain an updated template based on the optimal matching template and the system call log. The conditional decision module stores preset decision conditions, which are used to update the template in the activity template storage module according to the decision conditions and the update template, and to mark the corresponding system call log in the system call log storage module with an anomaly judgment tag. The rule adaptation module is used to periodically convert the updated templates in the activity template storage module into corresponding rule sets, and send rule update messages to the rule detection and filtering module. The rule base module is used to store the rule set. The data parsing module is used to take the variable content and corresponding differences and distances between the system call log and the optimal matching template as parsed data for log anomaly analysis. The rule detection and filtering module extracts the corresponding rule set from the rule base module according to the rule update message, and updates the rule data according to the rule set.
2. The cloud computing platform with real-time anomaly detection for container clouds according to claim 1, characterized in that: in, In the event processing module, the information field extraction includes field recognition, subject recognition, and word segmentation. The field identification is used to extract system call parameter information and metadata related to the container and the container cloud application node from the original system call event as field data. The entity identification is used to identify the field data based on the container ID, image ID, and process-related information, and to set the corresponding identifier in the system call log. The word segmentation process is used to identify and record special content in the field data using regular expressions, and then mask it with an identifier name; the field data is segmented according to the basic delimiter and dynamic segmentation rules to obtain a token sequence as the data of the corresponding field in the system call log.
3. The cloud computing platform with real-time anomaly detection for container clouds according to claim 2, characterized in that: in, In the candidate search module, the candidate template is obtained in the following manner: When the length of the token sequence is less than or equal to 3, the template whose length difference is within a preset range is selected from all the templates as the candidate template; When the length of the token sequence is greater than 3, a coarse-grained similarity score (score(D,Q)) is calculated between the token sequence and each of the templates. Templates with similarity scores (score(D,Q)) greater than a threshold are selected as candidate templates. The similarity score (score(D,Q)) is calculated using the following formula: In the formula, D is the template, and Q is the set of morphemes q contained in the query morpheme set. i For each token sequence, f(q) i D) is the morpheme q i The number of times it appears in template D, |D| is the number of words in template D, avgdl is the average document length, k1 and b are adjustable parameters, N is the total number of templates, n(q i ) contains the morpheme q i The number of templates.
4. The cloud computing platform with real-time anomaly detection for container clouds according to claim 3, characterized in that: in, When m=0, the data parsing module uses the template with the highest similarity score (D,Q) as the optimal matching template.
5. The cloud computing platform with real-time anomaly detection for container clouds according to claim 2, characterized in that: in, The template is in the form of a tree, where the non-leaf nodes of the tree are patterns that match all token sequences belonging to the template, and the leaf nodes of the tree are records of each token sequence. The degree and maximum height of the tree are fixed values. In the hierarchical clustering module, each non-leaf node of each candidate template is aligned with the token sequence using the longest common subsequence (LCS) before similarity calculation is performed. The formula for this similarity calculation is as follows: In the formula, |s1| and |s2| are the lengths of the non-leaf node and the corresponding string of the token sequence after alignment, respectively; m is the number of matched characters; t is the number of transposes; D is the template; Q is the input token sequence; Length(LCS(D,Q)) is the LCS length of the template D and the token sequence Q; |Diffs| is the number of differences; sim j This represents the similarity between differences.
6. The cloud computing platform with real-time anomaly detection for container clouds according to claim 5, characterized in that: in, In the rule template mining module, the token sequence of the system call log is merged with the record of the optimal matching template to obtain a merged template. If the merged template is different from the optimal matching template, the merged template is used as the update template. When m=0, the update template is obtained by constructing a template based on the corresponding token sequence.
7. The cloud computing platform with real-time anomaly detection for container clouds according to claim 1, characterized in that: in, The decision conditions in the conditional decision module are: All entities of the same type in the entire cluster generate system calls similar to the original system call event, and the overall frequency and the number of containers reach preset values; and The time interval between the original system call events is relatively short; and The context information of the original system call events is similar; and When additional anomaly detection analysis is available, no anomalies are detected after performing the anomaly detection analysis on the original system call event.
8. The cloud computing platform with real-time anomaly detection for container clouds according to claim 1, characterized in that: in, The rule detection and filtering module updates the rule data according to update requirements by hot-updating or restarting the container and configuring the rule set as new rule data.