An object-driven Linux kernel crash classification method and device

The object-driven Linux kernel crash classification method uses backward taint analysis and PageRank to accurately group crashes by underlying causes, improving classification accuracy and efficiency.

CN116150760BActive Publication Date: 2025-07-15Chinese People's Liberation Army Cyberspace Force Information Engineering University
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211607700.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-14
Publication Date
2025-07-15
Estimated Expiration
2042-12-14

AI Technical Summary

Technical Problem

Existing Linux kernel crash classification methods are inaccurate and lack depth, leading to inefficient bug analysis and incomplete patching due to grouping errors based on similar crash functions rather than underlying causes, and existing methods struggle with PoC-less crashes.

Method used

An object-driven Linux kernel crash classification method using backward taint analysis to extract kernel objects, construct a weighted bipartite graph model, and apply PageRank algorithm to determine crash similarity based on object call relationships.

Benefits of technology

Enhances crash classification accuracy by grouping crashes by underlying causes, reducing false positives and enabling effective classification of PoC-less crashes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116150760B_ABST
    Figure CN116150760B_ABST
Patent Text Reader

Abstract

The present invention provides an object-driven Linux kernel crash classification method and apparatus. The method includes: Step 1: According to the kernel error report and the corresponding kernel source code, use backward taint analysis to extract all kernel objects in the kernel error report; Step 2: Analyze the call relationships between all kernel objects to construct a kernel object call dependency graph; Step 3: According to the kernel object call dependency graph, use the pagerank algorithm to calculate the called degree of each kernel object; Step 4: Use the called degree of the kernel object as the weight information of the weighted bipartite graph, and abstract the reference relationship between the crash and the kernel object into a weighted bipartite graph; Step 5: Calculate the similarity between any two crashes in the weighted bipartite graph, and classify the crashes with a similarity greater than the threshold into the same type.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of network security technology, and in particular, to an object-driven Linux kernel crash classification method and device. Background Art

[0002] Fuzz testing technology is currently the most effective vulnerability mining technology, especially the mining of Linux kernel vulnerabilities has seen an explosive growth in recent years. According to statistics, since the release of the syzkaller project, 4,980 kernel bugs have been discovered, exceeding the number of vulnerabilities discovered in the previous 20 years. Continuous fuzz testing will continuously trigger the same error and generate many duplicate crashes. Currently, the method for crash classification is: if the crashes share the same crashing function and crashing type, then group them under the same error title. As Figure 1 shown, crashes 1-N share the same error title B.

[0003] However, this heuristic-based crash classification method is not accurate. As Figure 1 shown, error titles A, B, and C belong to the same root cause, but the existing method cannot group them into the same crash group. According to statistics, from September 2017 to November 2020, Syzbot reported 2,526 error titles and more than 3.24 million crash reports. According to their patches, they can be divided into 1,686 different crash groups. Among the 2,526 error titles, 1,191 (47.1%) belong to the same crash group as one or more other titles. This will result in different error reports with the same root cause being treated as new bugs and assigned to different security analysts. On the one hand, multiple groups of analysts processing the same bug in parallel and lacking communication with each other will lead to low repair efficiency. On the other hand, limited understanding of the error behavior of the same bug may lead to incomplete patches.

[0004] In addition to the above crash classification method integrated into the syzkaller project, Mu et al.'s latest research (Mu D, Wu Y, Chen Y, et al. An In-depth Analysis of Duplicated Linux Kernel Bug Reports[C] / / Network and Distributed Systems Security (NDSS) Symposium 2022. 2022) selected a part of the bugs for empirical research in response to the problem of a large number of duplicate crashes, obtained the reasons for the diverse error behaviors of the Linux kernel, and proposed a corresponding crash classification method. However, this classification method that only conducts similarity comparison by extracting system calls in the PoC is difficult to analyze in combination with the program execution context, difficult to classify crashes without PoC, and there are also false positive problems in classifying crashes with PoC provided.

[0005] Existing Linux kernel crash classification methods have a relatively coarse granularity and there are significant problems in accuracy, making it difficult to complete the preliminary exploitability assessment and repair work of vulnerabilities. Therefore, proposing an accurate and efficient Linux kernel crash classification method is the primary problem to be solved currently. Summary of the Invention

[0006] Aiming at the problems of relatively coarse granularity and low accuracy existing in the existing Linux kernel crash classification methods, the present invention provides an object-driven Linux kernel crash classification method and device, which deeply analyzes the program execution context through backward taint analysis technology, extracts kernel objects related to the root cause, and constructs a bipartite graph model of crashes and kernel objects to achieve classification.

[0007] On the one hand, the present invention provides an object-driven Linux kernel crash classification method, including:

[0008] Step 1: According to the kernel error report and the corresponding kernel source code, use backward taint analysis to extract all kernel objects in the kernel error report;

[0009] Step 2: Analyze the call relationships between all kernel objects to construct a kernel object call dependency graph;

[0010] Step 3: According to the kernel object call dependency graph, use the pagerank algorithm to calculate the called degree of each kernel object;

[0011] Step 4: Use the called degree of the kernel object as the weight information of the weighted bipartite graph, and abstract the reference relationship between the crash and the kernel object into a weighted bipartite graph;

[0012] Step 5: Calculate the similarity between any two crashes in the weighted bipartite graph, and classify the crashes with similarity greater than the threshold into the same type.

[0013] Further, in Step 1, use backward taint analysis to extract all kernel objects in the kernel error report, specifically including:

[0014] Step 1.1: Use the macro-defined error checking mechanism and the KASAN error checking mechanism in the kernel source code to determine the crash point;

[0015] Step 1.2: Use the function call trace in the kernel error report to construct a control flow graph, then use the crash point as the source point, and propagate taint backward on the control flow graph until the set taint analysis stop condition is met; among them, the objects on the propagation path are the kernel objects to be extracted; the set taint analysis stop condition is: the taint propagates to a variable that has been tainted; or, the taint propagates to any one of the entrances of a system call, the function entrance of an interrupt handler, and the function entrance of a startup work queue scheduler.

[0016] Further, Step 4 specifically includes: Determine the crash set C = {c1, c2,..., c m}, the kernel object set O = {o1, o2,..., o n}; Use the reference relationship between crashes and kernel objects as edges to obtain the edge set E; Use the called degree of kernel objects as the weights of the edges to obtain the weight matrix W; Finally, obtain the weighted bipartite graph G = {C, O, E, W}.

[0017] Further, in Step 5, calculate the similarity between two crashes according to formula (1):

[0018]

[0019] Among them, I(c p ) represents the set of nodes (i.e., the in-neighbor set) in the weighted bipartite graph that all point to the crash node c p , I(c q ) represents the set of all nodes in the weighted bipartite graph that point to the crash node c q , s(I i (c p )I j (c q )) represents the similarity between the i-th node in I(c p ) and the j-th node in I(c q ), |I(c p )| and |I(cq ) represent the number of elements in I(c p ) and I(c q ) respectively, and D is a damping coefficient.

[0020] Furthermore, it also includes: for nodes c p and c q , calculate the similarity evidence between the two according to formula (2):

[0021]

[0022] where E(c p ) represents the set of adjacent crash nodes of the crash node c p , and E(c q ) represents the set of adjacent crash nodes of the crash node c q ;

[0023] Correspondingly, adjust formula (1), and calculate the similarity between two crashes according to formula (3):

[0024] s evidence (c p ,c q ) = evidence(c p ,c q ) · s(c p ,c q )(3).

[0025] On the other hand, the present invention provides an object-driven Linux kernel crash classification device, including:

[0026] A taint analysis module, which is used to extract all kernel objects in the kernel error report by using backward taint analysis according to the kernel error report and the corresponding kernel source code;

[0027] A call dependency graph construction module, which is used to analyze the call relationships between all kernel objects to construct a kernel object call dependency graph; and calculate the called degree of each kernel object according to the kernel object call dependency graph by using the pagerank algorithm;

[0028] A bipartite graph construction module, which is used to abstract the reference relationship between crashes and kernel objects into a weighted bipartite graph with the called degree of kernel objects as the weight information of the weighted bipartite graph;

[0029] A classification module, which is used to calculate the similarity between any two crashes in the weighted bipartite graph, and classify the crashes with similarity greater than the threshold into the same type.

[0030] Advantages of the present invention:

[0031] Through the backward taint technique, the present invention extracts kernel objects related to the root cause of kernel errors, and uses the extracted kernel objects to approximately replace the root cause to guide crash classification, thus making up for the shortcoming of inaccurate classification in the previous classification methods due to the lack of in-depth analysis of the program execution context. For crashes caused by the same root cause, grouping can be achieved, and there will be no false positive situations. Moreover, the present invention only requires error reports as input, and can also classify crashes without a PoC, greatly enhancing the accuracy and effectiveness of classification, and solving the problem of the coarser granularity and lower accuracy of traditional classification methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 In the classification method in the prior art, the relationship among crash groups, error titles, and error reports;

[0033] Figure 2 A schematic flowchart of an object-driven Linux kernel crash classification method provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0034] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0035] Embodiment 1

[0036] As Figure 2 shown, an embodiment of the present invention provides an object-driven Linux kernel crash classification method, which is characterized by including:

[0037] S101: According to the kernel error report and the corresponding kernel source code, use backward taint analysis to extract all kernel objects in the kernel error report;

[0038] Specifically, the kernel source code needs to be compiled before subsequent static analysis. To prevent compilation optimization from affecting subsequent static analysis, in this embodiment, the Clang compiler provided by the LLVM compiler framework is used to compile the Linux kernel source code into IR (Intermediate Representation) in the form of Bitcode (bytecode). In addition, when compiling the kernel source code, the KASAN option is enabled to facilitate the determination of the crash point. At the same time, the compiled Linux kernel version is consistent with the kernel version corresponding to the kernel error report.

[0039] As an implementable manner, it specifically includes the following sub-steps:

[0040] S1011: Use the macro definition error checking mechanism and the KASAN error checking mechanism in the kernel source code to determine the crash point;

[0041] S1012: Use the function call trace in the kernel error report to construct a control flow graph, and then use the crash point as the source point to propagate taint backward on the control flow graph until the set taint analysis stop condition is met; among them, the objects on the propagation path are the kernel objects to be extracted; the set taint analysis stop condition is: the taint propagates to a variable that has been tainted; or, the taint propagates to any one of the entrances of system calls, the functions of interrupt handlers, and the functions of starting the work queue scheduler; this is because when the taint propagates to any one of the above entrances, it means that the taint has propagated to the starting position of this section of the kernel program.

[0042] Specifically, pointer analysis is performed during the process of taint propagation to propagate the taint to the aliases of variables.

[0043] S102: Analyze the call relationships between all kernel objects to construct a kernel object call dependency graph;

[0044] S103: According to the kernel object call dependency graph, use the pagerank algorithm to calculate the called degree of each kernel object;

[0045] Specifically, by analyzing kernel errors, the inventor found that the correlation degrees of different kernel objects with crashes are also different. Specifically, the more a kernel object is called, the lower its correlation with crashes, and the less a kernel object is called, the higher its correlation with crashes. Based on this phenomenon, in this embodiment, by first constructing a kernel object call dependency graph and then calculating the called degree of each kernel object, it can be used as a subsequent measurement index for correlation.

[0046] S104: Use the called degree of the kernel object as the weight information of the weighted bipartite graph, and abstract the reference relationship between the crash and the kernel object into a weighted bipartite graph;

[0047] Specifically, first determine the crash set C = {c1, c2,..., c m}, and the kernel object set O = {o1, o2,..., o n}; then use the reference relationship between the crash and the kernel object as the edge to obtain the edge set E; then use the called degree of the kernel object as the weight of the edge to obtain the weight matrix W; finally, obtain the weighted bipartite graph G = {C, O, E, W}.

[0048] S105: Calculate the similarity between any two crashes in the weighted bipartite graph, and classify the crashes with similarity greater than the threshold into the same type.

[0049] Specifically, the SimRank algorithm is a model for measuring the similarity between any two objects based on the topological structure information of the graph. The core idea of this algorithm is: if two objects are referenced by objects similar to them (i.e., they have similar in-neighbor edge structures), then these two objects are also similar. Therefore, in this embodiment, this algorithm can be used to compare the similarity of kernel objects to determine the similarity of crashes, thereby achieving classification.

[0050] As an implementable manner, calculate the similarity between two crashes according to formula (1):

[0051]

[0052] where I(c p ) represents the set of nodes (i.e., the in-neighbor set) in the weighted bipartite graph that all point to the crash node c p , I(c q ) represents the set of nodes in the weighted bipartite graph that all point to the crash node c q , s(I i (c p )I j (c q )) represents the similarity between the i-th node in I(c p ) and the j-th node in I(c q ), |I(c p )| and |I(c q )| respectively represent the number of elements in I(c p ) and I(c q ), and D is a damping factor, usually taking 0.6 - 0.8.

[0053] For example, I(c p) and I(c q ) are {1, 2, 3, 4} and {3, 4, 5} respectively, then |I(c p )| and |I(c q )| are 4 and 3 respectively; the similarity of s(3, 3) is 1, s(4, 4) is 1, and the rest are 0. The elements in the two sets are compared in sequence, and the calculation results are accumulated to obtain a similarity of 2. Substituting into the expression is

[0054] As can be seen from the above formula (1), the naive simrank algorithm does not consider the weight of the edge, so it cannot achieve classification well. In addition, nodes with simpler relationships (fewer connected edges) are more likely to be similar to other nodes, and nodes with more complex relationships (more connected edges) are more difficult to be similar to other nodes. The naive simrank algorithm also lacks consideration of the size of the common adjacent node set of the nodes. Therefore, in order to better achieve classification, this embodiment also provides another implementable manner. When calculating the similarity, node similarity evidence and edge weights are added to measure the similarity. Specifically:

[0055] For nodes c p and c q , calculate the similarity evidence of the two according to formula (2):

[0056]

[0057] Among them, E(c p ) represents the set of adjacent crash nodes of crash node c p , and E(c q ) represents the set of adjacent crash nodes of crash node c q ;

[0058] Correspondingly, adjust formula (1), and calculate the similarity between two crashes according to formula (3):

[0059] s evidence (c p , c q ) = evidence(c p , c q ) · s(c p , c q ) (3).

[0060] In order to classify crashes with similarity greater than the threshold into the same type, the union-find set algorithm is adopted in this embodiment. Specifically: First, initialize a similarity set, define the similarity comparison as operation ⊕. If c p and c q are similar, then The operation satisfies transitivity. If then Thus, the similarity sets can be further merged.

[0061] In the embodiment of the present invention, the backward taint technique is used to extract kernel objects related to the root cause of the kernel error. The extracted kernel objects are approximately used to replace the root cause to guide the crash classification, making up for the shortcoming of the previous classification method that the program execution context is not deeply analyzed, resulting in inaccurate classification. For crashes caused by the same root cause, grouping can be achieved, and there will be no false positive situation. Moreover, only the error report is required as the input, and crashes without PoC can also be classified, greatly enhancing the accuracy and effectiveness of the classification.

[0062] Embodiment 2

[0063] Corresponding to the above method, the embodiment of the present invention further provides an object-driven Linux kernel crash classification device, including a taint analysis module, a call dependency graph construction module, a bipartite graph construction module, and a classification module;

[0064] Among them, the taint analysis module is used to extract all kernel objects in the kernel error report using backward taint analysis according to the kernel error report and the corresponding kernel source code; the call dependency graph construction module is used to analyze the call relationships between all kernel objects to construct a kernel object call dependency graph; and according to the kernel object call dependency graph, use the pagerank algorithm to calculate the called degree of each kernel object; the bipartite graph construction module is used to use the called degree of the kernel object as the weight information of the weighted bipartite graph, and abstract the reference relationship between the crash and the kernel object into a weighted bipartite graph; the classification module is used to calculate the similarity between any two crashes in the weighted bipartite graph, and classify the crashes with a similarity greater than the threshold into the same type.

[0065] It should be noted that the device provided in the embodiment of the present invention is to implement the above method embodiment, and its functions can be specifically referred to the above method embodiment, which will not be elaborated here.

[0066] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An object-driven Linux kernel crash classification method, characterized in that, Including: Step 1: According to the kernel error report and the corresponding kernel source code, use backward taint analysis to extract all kernel objects in the kernel error report; Step 2: Analyze the call relationships among all kernel objects to construct a kernel object call dependency graph; Step 3: According to the kernel object call dependency graph, use the pagerank algorithm to calculate the called degree of each kernel object; Step 4: Take the called degree of the kernel object as the weight information of the weighted bipartite graph, and abstract the reference relationship between the crash and the kernel object into a weighted bipartite graph; Step 5: Calculate the similarity between any two crashes in the weighted bipartite graph, and classify the crashes with similarity greater than the threshold into the same type.

2. The object-driven Linux kernel crash classification method according to claim 1, wherein In Step 1, using backward taint analysis to extract all kernel objects in the kernel error report specifically includes: Step 1.1: Use the macro-defined error checking mechanism and the KASAN error checking mechanism in the kernel source code to determine the crash point; Step 1.2: Use the function call trace in the kernel error report to construct a control flow graph, then take the crash point as the source point, and propagate taint backward on the control flow graph based on the source point until the set taint analysis stop condition is met; among them, the objects on the propagation path are the kernel objects to be extracted; the set taint analysis stop condition is: the taint propagates to a variable that has been tainted; or, the taint propagates to any one of the entrances of system calls, the functions of interrupt handlers, and the functions of starting the work queue scheduler.

3. An object-driven Linux kernel crash classification method according to claim 1, characterized in that, Step 4 specifically includes: Determine the crash set C = {c1, c2,..., c m}, the kernel object set O = {o1, o2,..., o n}; take the reference relationship between crashes and kernel objects as edges to obtain the edge set E; take the called degree of kernel objects as the edge weights to obtain the weight matrix W; finally obtain the weighted bipartite graph G = {C, O, E, W}.

4. An object-driven Linux kernel crash classification method according to claim 3, characterized in that In Step 5, calculate the similarity between two crashes according to formula (1): Among them, I(c p ) represents the set of all nodes in the weighted bipartite graph that point to the crash node c p , that is, the in-neighbor set. I(c q ) represents the set of all nodes in the weighted bipartite graph that point to the crash node c q . s(I i (c p )I j (c q )) represents the similarity between the i-th node in I(c p ) and the j-th node in I(c q ). |I(c p )| and |I(c q )| respectively represent the number of elements in I(c p ) and I(c q ). D is a damping factor.

5. An object-driven Linux kernel crash classification method according to claim 4, characterized in that It also includes: for node c p and c q , calculate the similarity evidence between the two according to formula (2): Among them, E(c p ) represents the set of adjacent crash nodes of crash node c p , and E(c q ) represents the set of adjacent crash nodes of crash node c q ; Correspondingly, adjust formula (1) and calculate the similarity between two crashes according to formula (3): s evidence (c p ,c q ) = evidence(c p ,c q )·s(c p ,c q )(3).

6. An object-driven Linux kernel crash classification device, characterized in that, Including: A taint analysis module, which is used to extract all kernel objects in the kernel error report according to the kernel error report and the corresponding kernel source code by using backward taint analysis; A call dependency graph construction module, which is used to analyze the call relationships among all kernel objects to construct a kernel object call dependency graph; and according to the kernel object call dependency graph, use the pagerank algorithm to calculate the called degree of each kernel object; A bipartite graph construction module, which is used to abstract the reference relationship between the crash and the kernel object into a weighted bipartite graph by taking the called degree of the kernel object as the weight information of the weighted bipartite graph; A classification module, which is used to calculate the similarity between any two crashes in the weighted bipartite graph and classify the crashes with similarity greater than the threshold into the same type.

Citation Information

Patent Citations

  • Kernel fuzzy test case generation method based on system call dependency graph

    CN112559367A

  • Semantics-aware android malware classification

    US20160057159A1