Performing flow-sensitive SAST on source code representation from which lines identified by flow-insensitive SAST have been removed

By performing flow-insensitive SAST on a simplified source code representation and then flow-sensitive SAST on the pruned version, the execution time and accuracy of SAST are improved, addressing the inefficiencies and false positives of traditional methods.

WO2025151123A1PCT designated stage expired Publication Date: 2025-07-17MICRO FOCUS LLC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/US2024/011320
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-12
Publication Date
2025-07-17

AI Technical Summary

Technical Problem

Existing static application security testing (SAST) methods face challenges with cubic time complexity in execution time due to the need for flow-sensitive heap models, making them impractical for large and complex software applications, and often result in numerous false positives when using flow-insensitive models.

Method used

Perform flow-insensitive SAST on a simplified representation of source code to identify and remove lines not contributing to vulnerabilities, followed by flow-sensitive SAST on the pruned representation to enhance precision and reduce execution time.

Benefits of technology

This approach significantly reduces SAST execution time from cubic to linear, while ensuring accurate identification of security vulnerabilities by eliminating false positives and negatives, thereby making SAST more feasible for complex applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024011320_17072025_PF_FP_ABST
    Figure US2024011320_17072025_PF_FP_ABST
Patent Text Reader

Abstract

A simplified, flow-insensitive graph, such as a Steensgaard graph, of a representation of source code of a program is constructed. Static application security testing (SAST) analysis on the representation is performed using the simplified heap graph to identify potential security vulnerabilities in the source code. The SAST analysis is flow-insensitive due to usage of the simplified heap graph. Any lines of the representation that do not contribute to the potential security vulnerabilities are removed. A non-simplified heap graph of the representation is constructed from which any lines that do not contribute to the potential security vulnerabilities have been removed. The SAST analysis is performed on the representation using a flow-sensitive heap model derived from the non-simplified graph to identify actual security vulnerabilities in the source code. The SAST analysis is flow-sensitive due to usage of the flow-sensitive heap model derived from the non-simplified heap graph.
Need to check novelty before this filing date? Find Prior Art

Description

PERFORMING FLOW-SENSITIVE SAST ON SOURCE CODE REPRESENTATION FROM WHICH LINES IDENTIFIED BY FLOW-INSENSITIVE SAST HAVE BEEN REMOVEDBACKGROUND

[0001] Computing devices like desktops, laptops, and other types of computers, as well as mobile computing devices like smartphones, among other types of computing devices, run software, which can be referred to as applications, to perform intended functionality. An application may be a so-called native application that runs on a computing device directly, or may be a web application or “app” at least partially run on a remote computing device accessible over a network, such as via a web browser running on a local computing device. An application can be tested, or analyzed, in a variety of different ways to ensure that the application correctly performs its intended functionality as well as to ensure that the application does not have any security vulnerabilities.BRIEF DESCRIPTION OF THE DRAWINGS

[0002] FIG. 1 A is a diagram of an example representation of source code of a program.

[0003] FIG. 1 B is a diagram depicting an example non-simplified heap graph of the representation of FIG. 1A.

[0004] FIG. 1 C is a diagram depicting example simplification of the graph of FIG. 1 B.

[0005] FIG. 1 D is a diagram depicting further example simplification of the graph of FIG. 1 B, resulting in an example simplified heap graph of the representation of FIG. 1A on which flow-insensitive static analysis security testing (SAST) is performed to identify security vulnerabilities in the representation.

[0006] FIG. 1 E is a diagram of the example representation of FIG. 1 A after pruning to remove lines that do not contribute to security vulnerabilities identified in the flow-insensitive SAST performed on the graph of FIG. 1 D.

[0007] FIG. 1 F is a diagram of an example non-simplified heap graph of the pruned representation of FIG. 1 E on which flow-sensitive SAST is performed to more precisely identify security vulnerabilities in the representation of FIG. 1A.The security vulnerabilities identified by the flow-sensitive SAST performed on the non-simplified heap graph of FIG. 1 F are a subset of the security vulnerabilities identified by the flow-insensitive SAST performed on the simplified heap graph of FIG. 1 D.

[0008] FIG. 1 G is a diagram of an example flow-sensitive heap model (e.g., alias graph) of the pruned representation of FIG. 1 E that is produced during performance of the flow-sensitive SAST on the non-simplified heap graph of FIG. 1 F.

[0009] FIG. 2 is a diagram of an example computing device storing program code for performing flow-sensitive SAST on a representation of source code of a program from which lines identified by flow-insensitive SAST have been removed to identify security vulnerabilities.

[0010] FIG. 3A is a diagram of another example representation of source code of a program.

[0011] FIG. 3B is a diagram of an example non-simplified heap graph of the representation of FIG. 3A.

[0012] FIG. 3C is a diagram depicting an example simplified heap graph of the representation of FIG. 3A and on which flow-insensitive SAST is performed to identify security vulnerabilities in the representation of FIG. 3A.

[0013] FIG. 3D is a diagram of the example representation of FIG. 3A after pruning to remove lines that do not contribute to security vulnerabilities identified in the flow-insensitive SAST performed on the graph of FIG. 3C.

[0014] FIG. 3E is a diagram of an example non-simplified heap graph of the representation of FIG. 3D on which flow-sensitive SAST is performed to more precisely identify security vulnerabilities in the representation of FIG. 3A. The security vulnerabilities identified by the flow-sensitive SAST performed on the nonsimplified heap graph of FIG. 3E are a subset of the security vulnerabilities identified by the flow-insensitive SAST performed on the simplified heap graph of FIG. 3C.

[0015] FIG. 3F is a diagram of an example flow-sensitive heap model (e.g., alias graph) of the pruned representation of FIG. 3D that is produced during performance of the flow-sensitive SAST on the non-simplified heap graph of FIG. 3E.

[0016] FIG. 4 is a diagram of another example representation of source code of a program that has the non-simplified, dataflow graph of FIG. 3B and the simplified, flow-insensitive graph of FIG. 3C.DETAILED DESCRIPTION

[0017] As noted in the background, an application can be tested to ensure that it performs its intended functionality as well as to ensure that it does not have any security vulnerabilities. One type of application testing that is performed particularly to identify security vulnerabilities is known as static application security testing (SAST). SAST involves analyzing the source code of an application to determine whether, upon generation of executable code from the source code, subsequent execution of the application will have security vulnerabilities. SAST isstatic in that the application is not actually executed (i.e. , executable code for the application is not generated from the source code and / or is not executed) to identify security vulnerabilities. In other words, SAST utilizes only the source code of an application and does not consider the application when it is actually running.

[0018] Other, non-SAST techniques include, among others, dynamic application security testing (DAST) and interactive application security testing (IAST). DAST identifies security vulnerabilities within an application as the application is running (i.e., during execution of the executable code for the application), such as in a production environment in which the application is being used by end users. Unlike SAST, DAST utilizes only the executable code of the application and considers the application when it is actually running. IAST identifies security vulnerabilities within an application during automated or human- assisted testing of the application while the application is running, and can identify the source code responsible for identified security vulnerabilities. Unlike SAST and like DAST, IAST utilizes the executable code of the application and considers the application when it is actually running, but unlike DAST can reference the source code of the application.

[0019] SAST can involve generating a representation of the source code of a program (e.g., a semantic representation of all possible behaviors, or paths, through the source code), and then analyzing that representation to identify security vulnerabilities in the source code, instead of (or in addition to) directly analyzing the source code itself. For instance, the pending US patent application filed on October 31 , 2023, and assigned application no. 18 / 498,961 , which is hereby incorporated by reference, describes examples of such representations, including a generalized lower level representation that is not specific to anyprogramming language. Furthermore, static dataflow analyses may be performed as part of the SAST by applying a superlattice (i.e. , a lattice product) of lattices corresponding to different individual static analyses against such a representation of source code, as described in the pending US patent application filed on August 28, 2023, and assigned application no. 18 / 239,011 , which is also hereby incorporated by reference.

[0020] A difficulty with SAST is that its execution time can increase at least cubically with source code length when performing SAST on source code (e.g., on a representation of the source code) that is sensitive to the direction of the flow of information through the program. That is, if the SAST considers the order in which variables and fields of objects of the source code have values loaded therefrom and stored therein in the source code, then execution time can increase at least cubically with source code length. That the length of time to perform SAST of source code increases at least cubically with source code length can be referred to in shorthand by stating that SAST performance occurs in at least cubic time with source code length.

[0021] Performing dataflow analysis at the level of sophistication needed for SAST requires a model of the program’s heap that describes when two fields or variables bearing distinct names in the program source code may point to the same object. When such a correspondence occurs, information that the dataflow analysis learns about one of them has an effect on what the dataflow analysis learns about the other. Such a heap model may be produced by different algorithms, such as those known as “alias analysis” or “pointer analysis,” for instance, depending on the form of the source code.

[0022] SAST with high precision, in that there is a low frequency of false positives among identified security vulnerabilities, requires a heap model that is “flow-sensitive.” This means that the model consists of statements of the form “the possible values of y include the possible values of x,” which can be succinctly stated as “x flows to y.” Such a unidirectional statement allows dataflow analysis to learn facts about y from facts about x, but not vice versa.

[0023] In contrast to a flow-sensitive heap model, it is possible to construct a “flow-insensitive” heap model. Such a model consists of statements of the form “the possible values of x and the possible values of y coincide,” which can be succinctly stated as “x and y alias.” This is a bidirectional statement that allows a dataflow analysis to both learn facts about y from facts about x and learn facts about x from facts about y. Such bidirectionality makes a flow-insensitive heap model weaker than a flow-sensitive heap model, however. A dataflow analysis using a flow-insensitive heap model is insufficiently precise for SAST, leading to far too many false positives for the identified security vulnerabilities to be practically used in the software development process to drive remedial action.

[0024] Producing a flow-sensitive heap model using standard algorithms, however, can require an execution time that increases at least cubically with source code length. By comparison, producing a flow-insensitive heap model using standard algorithms requires only an execution time that linearly increases with the length of the source code. As such, performing SAST to precisely identify security vulnerabilities can be difficult.

[0025] For instance, even with the availability of large amounts of computing processing capability, performing SAST to precisely identify security vulnerabilities can take hours, days, or longer for programs that are complex andmay hundreds of thousands of lines of source code or more, even without taking into account the supporting libraries that the source code references. For the most complex programs having millions of lines of source code, if not more, performing such SAST may be intractable. That is, even with large amounts of computing processing capability is available now or that may become available in the foreseeable future, performing such SAST may require months, years, or longer, meaning that for all practical purposes it cannot be performed.

[0026] Techniques described herein ameliorate these issues, by pruning the source code of a program to remove lines that definitively do not contribute to security vulnerabilities in the source code. In particular, flow-insensitive SAST is performed on a representation of source code (e.g., such as one or more of the representations described in the 18 / 498,961 application reference above). The flow-insensitive SAST may be performed using the same or similar static dataflow analyses (e.g., those described in the 18 / 239,011 application referenced above) used when later performing flow-sensitive SAST.

[0027] However, the analyses are performed in such a way that the SAST is flow-insensitive in that it does not take into account the direction of information flow in the source code. That is, the SAST is flow-insensitive in that it does not take into account the order in which fields of objects and variables of the source code have values loaded therefrom and stored therein in the source code, together with the fact that the value loaded from a field or variable is the value most recently stored in that field or variable.

[0028] The flow-insensitive SAST identifies security vulnerabilities that include all the security vulnerabilities that flow-sensitive SAST would identify, but which may also include security vulnerabilities that do not actually occur in thesource code. That is, the flow-insensitive SAST may identify false positives, but does not identify false negatives. Therefore, lines of source code that do not contribute to any security vulnerability identified by the flow-insensitive SAST are known to not contribute to any security vulnerability that would be identified by flow-sensitive SAST.

[0029] These lines that do not contribute to any security vulnerability identified by the flow-insensitive SAST are removed, and flow-sensitive SAST is then performed on the resulting pruned representation of the source code (i.e., from which the lines in question have been removed). The flow-sensitive SAST more precisely identifies security vulnerabilities, and specifically identifies just a subset of the vulnerabilities identified by the flow-insensitive SAST. The vulnerabilities identified by the flow-insensitive SAST but not by the flow-sensitive SAST are necessarily false positives (i.e., non-actual vulnerabilities).

[0030] Reducing the size of the source code representation in this manner makes the flow-sensitive SAST more practicable to perform in terms of computation time. As noted above, flow-sensitive SAST is performed in at least cubic time. Therefore, reducing the size of the representation of source code by removing, say, 10%, 20%, or even more of the lines can significantly reduce execution time. The reduction in execution time can be sufficiently significant that flow-sensitive SAST may be able to be performed where previously it was not able to be performed in any practicable manner.

[0031] Performance of the flow-insensitive SAST to identify lines to be removed from the source code representation prior to performance of flowsensitive SAST occurs in linear time. That is, performing flow-insensitive SAST increases in execution time only linearly with source code size (i.e., the number oflines of source code). As such, performing flow-insensitive SAST is likely to be practicable when flow-sensitive SAST is not. Even when flow-sensitive SAST is practicable, performing flow-insensitive SAST first to prune the source code representation on which flow-sensitive SAST is then performed can result in significantly reduced execution time. The time cost to first perform flow-insensitive SAST is more than made up for by the time savings that result from just having to perform flow-insensitive SAST on a pruned version of the source code representation instead of the representation in its entirety.

[0032] The flow-insensitive SAST can be performed using the same techniques as the flow-sensitive SAST, such as by using the dataflow analyses described in the 18 / 239,011 application referenced above. The techniques can be considered as being flow-insensitive when they are performed on a simplified, flow-insensitive graph of the source code representation. By comparison, the techniques can be considered as being flow-sensitive when they are performed on a non-simplified, flow-sensitive graph of the source code representation.

[0033] A simplified heap graph of the representation of source code can be produced using a standard algorithm in linear time (e.g., Steensgaard’s points-to analysis) and can be used as a heap model for a dataflow analysis. The size of the heap model in this case is also linear in the program size and thus causes the execution time of a dataflow analysis to increase linearly with program size.Altogether, a SAST which involves the construction of the simplified heap graph as a heap model and a subsequent dataflow analysis on this model can be performed in linear time with respect to the program size.

[0034] The SAST in this case is flow-insensitive because the simplified heap graph on which the dataflow analysis is performed does not reflect the orderin which fields of objects and variables of the source code have values loaded therefrom and stored therein in the source code. After pruning of the source code representation to remove lines that do not contribute to any security vulnerability identified via the flow-insensitive SAST, a non-simplified heap graph of the pruned representation is then used to perform SAST in at least cubic time.

[0035] The SAST in this case is flow-sensitive because it builds a model of the heap that does reflect the order in which fields of objects and variables have values loaded therefrom and stored therein. Standard algorithms for building a flow-sensitive heap model have an execution time which is at least cubic in the program size (e.g., Andersen’s points-to analysis). The flow-sensitive heap model itself has a size which is quadratic in the program size, so that a dataflow analysis performed on such a heap model executes in time which is quadratic in the program size.

[0036] Altogether, then, a SAST which involves the construction of a flowsensitive heap model and a subsequent dataflow analysis on this model can be performed in at least cubic time with respect to the program size. The cubic time from the heap model construction dominates the quadratic time from the dataflow analysis.

[0037] FIG. 1A shows an example representation 100 of source code of a program that includes six lines. The first line sets the pointer p (e.g., the variable p) to the address at which the variable x (which may be a pointer) is stored, and the second line sets the pointer r (e.g., the variable r) to the address at which the pointer p is stored. The third line sets the pointer q (e.g., the variable q) to the address at which the variable y (which may be a pointer) is stored, and the fourth line sets the pointer s (e.g., the variable s) to the address at which the pointer q isstored. The fifth line sets the pointer r to the value of the pointer s. The sixth line sets the variable a to the value of the variable that the pointer r is currently pointing to. The lines of the source code representation 100 are executed in their order of appearance, such that the first line is executed before the second line, the second line is executed before the third line, and so on.

[0038] FIG. 1 B shows an example non-simplified heap graph 150 of the source code representation 100. The graph 150 may be constructed directly from the source code representation 100 (as opposed to from another graph of the representation 100), and does correspond to the representation 100. In either or both of these respects the graph 150 can therefore be considered non-simplified.

[0039] The graph 150 includes nodes 152A, 152B, 152C, 152D, 152E, 152F, and 152G respectively corresponding to the variables p, q, r, s, x, y, and a, and that are collectively referred to as the nodes 152. The graph 150 further includes edges 154A, 154B, 154C, 154D, 154E, and 154F, collectively referred to as the edges 154, which interconnect the nodes 152. Specifically, each edge 154 connects a corresponding pair of a source node 152 and a target node 152, such that the edge 154 is outgoing from the source node 152 and incoming to the target node 152. Each edge 154 indicates that the variable of the source node 152, if dereferenced, may yield the value of the variable of the target node 152 during some execution path through the program.

[0040] The example program source code representation 100 is simple in that it exhibits only one execution path, where each line of the program is executed in sequence. In general, a program’s source code may exhibit multiple possible execution paths, for example, using branching or looping control flow. Inthat case, the relationship expressed by an edge 154 might not occur in every path through the program at execution, but does occur in at least one path.

[0041] The edge 154A specifically indicates that dereferencing the variable r of the node 152A may yield the value of p of the node 152C. This is a result of the second line of the source code representation 100, “r := &p”. The edge 154B indicates that dereferencing r of the node 152A may also yield the value of q of the node 152D. This is a result of the fourth line, “s := &q” followed by the fifth line, “r := s”. The edge 154C indicates that dereferencing s of the node 152B may yield the value of q of the node 152D, per the fourth line, “s := &q”.

[0042] The edge 154D indicates that dereferencing p of the node 152C may yield the value of x of the node 152E, per the first line, “p := &x”, and the edge 154E indicates that dereferencing q of the node 152D may yield the value of y of the node 152F, per the third line, “q := &y”. Finally, the edge 154F indicates that dereferencing r of the node 152A may yield the value of the variable a of the node 152G, per the sixth line, “a := *r”.

[0043] The graph 150 reflects the heap of the source code representation 100 of the program. The graph 150 is non-simplified in that every node 152 represents only one variable in the source code representation 100 of the program. For example, the node 152C represents variable p and no other variable. The graph 150 moreover can be constructed directly from the representation 100.

[0044] The non-simplified heap graph 150, however, may be iteratively simplified over one or more stages to yield a simplified version of the graph. In each stage, edges 154 having the same source node 152 but different target nodes 152 are replaced with a single merged edge 154 from the common sourcenode 152 to a single merged target node 152, where the single merged target node 152 replaces the multiple original target nodes 152. (More specifically, edges 154 having the same source node 152 but different target nodes 152 are merged when they are labeled with the same field name; in the example the edges 154 are not labeled for explanatory clarity, but a further, more complex example is presented below in which edges are labeled.)

[0045] Therefore, whereas in the non-simplified heap graph 150 no node 152 can correspond to more than one variable, in a simplified version a node 152 can correspond to more than one variable. Once there are no more source nodes 152 that have more than one outgoing edge 154, the graph 150 has been maximally simplified. In this most-simplified version of the graph 150, each node 152 is limited to having one outgoing edge 154, which is not the case in the nonsimplified heap graph 150 prior to simplification.

[0046] FIG. 1 C shows a first stage of simplification of the graph 150 of FIG. 1 B, which as simplified is referenced as the partially simplified graph 150’. In the graph 150’, the edges 154A,154B, and 154F outgoing from the same source node 152A in FIG. 1 B have been replaced by a single edge 154AB, and the target nodes 152C,152D, and 152G of those edges 154A, 154B, and 154F have been replaced by a single merged node 152CDG.

[0047] Because the node 152B is connected to the node 152D in the graph 150 via the edge 154C, it is now connected to the node 152CDG in the graph 150’ via the same edge 154C. Similarly, because the nodes 152C and 152D are respectively connected to the nodes 152E and 152F in the graph 150 via the edges 154D and 154E, the node 152CDG is connected to both nodes 152E and 152F in the graph 150’ via these edges 154D and 154E. Because the newlymerged node 152CDG in the partially simplified graph 150’ has two outgoing edges 154D and 154E, the nodes 152E and 152F now have to be replaced by a merged mode.

[0048] FIG. 1 D accordingly shows a second stage of simplification of the graph 150 of FIG. 1 B, which is a further simplification of the partially simplified graph 150’ of FIG. 1 C, and which as simplified is referenced as the completely simplified graph 150” of the source code representation 100. In the graph 150”, the edges 154D and 154E outgoing from the same source node 152CDG in FIG. 1 C have been replaced by a single edge 154DE, and the target nodes 152E and 152F of those edges 154D and 154E have been replaced by a single merged node 152EF. The newly merged node 152EF does not have more than one outgoing edge 154 (and indeed has no outgoing edges 154), so no further simplification of the graph 150” has to be performed, which is why the graph 150” is considered as being completely simplified.

[0049] The simplified heap graph 150” of the source code representation 100 reflects a superset of the possible ways that variables and object fields may be related on any path through the program during execution. In particular, since the edge 154C connects the source node 152B to the target node 152CDG, this means that dereferencing the variable s of the node 152B may yield the value of variable p of the node 152CDG on some path (as well as the value of q of the node 152C on a different path). However, dereferencing s cannot actually yield p on any path through the source code representation 100. Therefore, the simplified graph 150” represents less precise information about the source code than the non-simplified graph 150.

[0050] Nevertheless, SAST is performed using the simplified heap graph 150”. SAST can be performed by executing the same or similar dataflow analyses that would ordinarily be performed on the source code representation 100 using a flow-sensitive heap model. However, in place of the flow-sensitive heap model, which would ordinarily be constructed from the non-simplified heap graph 150, the simplified heap graph 150” stands in as a flow-insensitive heap model. Whenever two variables are collocated in the same node 152, the dataflow analyses consider bidirectional information flow between the two variables.

[0051] For example, since variables p and q are collocated in node 152CDG, whenever a dataflow analysis learns a fact about p, it learns the same fact about q, and whenever a dataflow analysis learns a fact about q, it learns the same fact about p. That is, the simplified heap graph 150” indicates that p and q may be aliasing. The same bidirected relationship exists between the variables p and a, and between the variables q and a, and between the variables x and y. Because the analyses are performed using the bidirected heap model supplied by the simplified heap graph 150”, the resulting SAST is flow-insensitive.

[0052] Stated another way, flow-insensitive SAST is performed on the representation 100 because the analyses are performed on or using the simplified heap graph 150”. If the analyses were instead performed on or using a standard flow-sensitive heap model constructed from the non-simplified graph 150, flowsensitive SAST would have been performed.

[0053] As a result of the SAST being performed using a flow-insensitive heap model from the simplified graph 150” instead of a flow-sensitive heap model from the non-simplified graph 150, execution time is linear: the execution time of such flow-insensitive SAST of the representation 100 increases linearly withlength or size (e.g., the number of lines) of the source code representation 100. By comparison, if the SAST had been performed using a flow-sensitive heap model from the non-simplified graph 150, execution time would have been at least cubic, with the execution time of such flow-sensitive SAST increasing at least cubically with length or size of the representation 100.

[0054] The performance of flow-insensitive SAST of the representation 100 using the simplified graph 150” still results in identification of security vulnerabilities in the source code representation 100. However, because the security vulnerabilities are identified using the simplified graph 150”, they can include false positives (which are not actual security vulnerabilities) as well as true positives (i.e. , actual security vulnerabilities). For example, suppose a security vulnerability occurs if the value of variable p flows into variable a, and suppose a security vulnerability occurs if the value of variable q flows into variable a. In this case, performing flow-sensitive SAST would identify the actual vulnerability that q flows to a, as evidenced by lines 4, 5, and 6 in the source code representation 100. However, performing a flow-sensitive SAST would correctly not identify any vulnerability pertaining to variable p, since there is no flow from p to a: at the time of execution of line 6, the value of r is no longer the address of p, but the address of q instead.

[0055] However, because flow-insensitive SAST was in fact performed - using the simplified graph 150” - neither security vulnerability can be resolved or identified to its actual level of precision. As to the vulnerability involving variable q flowing to variable a, the graph 150” indicates this relationship by the collocation of q and a in the node 152CDG. However, the same node 152CDG collocatesvariables p and a, causing the potential flow from variable p to variable a to be identified as a vulnerability.

[0056] Therefore, flow-insensitive SAST identifies line 6 in the source code representation 100 where variable a is set by dereferencing variable r as participating in two security vulnerabilities, one originating in line 2 where the address of p is first taken, and one originating in line 4 where the address of q is first taken. That variable q flows to variable a is a security vulnerability (that would have been identified if flow-sensitive SAST were performed), whereas that variable p flows to variable a is a false positive (that would not have been identified if flow-sensitive SAST had been performed).

[0057] The flow-insensitive SAST, because it occurs in linear time as opposed to cubic time, is much faster than flow-sensitive SAST, and the security vulnerabilities that are identified can be used to prune the source code representation 100 to remove lines that definitively do not contribute to any security vulnerability. Because the flow-insensitivity causes false positives but not false negatives, any lines that do not contribute to any identified security vulnerability can be removed without affecting the accuracy of the SAST performed on the resulting pruned version of representation 100 in identifying the security vulnerabilities (which are a subset of the vulnerabilities identified via the flow-insensitive SAST).

[0058] For example, no security vulnerabilities may be identified in the flowinsensitive SAST involving the setting of the variables p or q to the addresses of variables x or y. That is, the flow-insensitive SAST may not identify the edge 154DE as corresponding to any security vulnerability. This means that the lines of the representation 100 corresponding to the edge 154DE can be removed, andindeed the lines corresponding to either variable x or y in the node 152EF insofar as the node 152EF has not been identified as part of any other security vulnerability (e.g., there is no outgoing edge from the node 152EF). In particular, lines 1 and 3 can be removed from the representation 100.

[0059] FIG. 1 E shows the source code representation 100 after lines 1 and 3 have been removed, which is referenced as the pruned source code representation 100’. In actuality, the remaining lines 2, 4, 5, and 6 may be renumbered as lines 1 , 2, 3, and 4, but for clarity the remaining lines are still numbered 2, 4, 5, and 6, and line numbers 1 and 3 depicted in the figure as placeholders with their respective source code removed. Flow-sensitive SAST can then be performed on the pruned source code representation 100’ to more precisely identify security vulnerabilities.

[0060] For instance, such analysis can identify whether the possibility that variable p or variable q (or both) flows to variable a, as captured by node 152CDG of the graph 150”, and being identified as part of two security vulnerabilities in the flow-insensitive SAST, are actually security vulnerabilities. In this way, the security vulnerabilities are more precisely identified by flow-sensitive SAST. The security vulnerabilities identified by the flow-sensitive SAST are specifically a subset of the security vulnerabilities identified by the flow-insensitive SAST, and do not include at least some of the false positives that the flow-insensitive SAST identified.

[0061] Because the flow-sensitive SAST is performed on a version of the source code representation 100 that has fewer lines - i.e. , it is performed on the pruned source code representation 100’ - the flow-sensitive SAST is performed more quickly. As noted, in general flow-sensitive SAST is performed in at leastcubic time. The extra time cost incurred by first performing the flow-insensitive SAST - which occurs more quickly, in linear time - to identify the lines to remove before performing the flow-sensitive SAST is likely to be more than made up for in the cubic-reduction-in-time saving when subsequently performing flow-sensitive SAST. Furthermore, since the flow-insensitivity does not cause false negatives, no lines are removed that would otherwise be identified as contributing to security vulnerabilities during flow-sensitive SAST, so accuracy is not reduced.

[0062] The flow-sensitive SAST of the pruned representation 100’ may be performed by using the same analyses used in the flow-insensitive SAST of the (unpruned) representation 100 in conjunction with the simplified heap graph 150”, but by performing them in conjunction with a non-simplified heap graph of the pruned representation 100’ instead. Therefore, a non-simplified heap graph corresponding to the representation 100’ may be constructed, and then the flow-sensitive SAST performed by using this graph to derive a flow-sensitive heap model.

[0063] FIG. 1 F shows an example non-simplified heap graph 160 of the pruned source code representation 100’. The graph 160 includes the nodes 152A, 152B, 152C, 152D, and 152G of the non-simplified, flow-sensitive graph 150 of the (unpruned) representation 100 of FIG. 1A, but not the nodes 152E and 152F of the graph 150, since the pruned representation 100’ does not have any lines pertaining to the variables x and y. Similarly, the graph 160 includes the edges 154A, 154B, 154C, 154F of the graph 150, but not the edges 154D and 154E, since the pruned representation 100’ does not have any lines pertaining to a relationship between variables p and x or a relationship between variables q and y. When flow-sensitive SAST is performed using the non-simplified heap graph160, a flow-sensitive heap model is derived from the heap graph 160 to take into account the direction of information flow and the ordering of instructions in the pruned source code representation 100”.

[0064] FIG. 1 G shows such an example flow-sensitive heap model 180.Because variable a is obtained by dereferencing variable r in line 6 after variable r is updated to point to q in line 5, the value of q may flow to the value of a.However, the value of p cannot flow to the value of a in line 6, even though r points to p after line 2, again because r is updated to a different value in line 5. Therefore, the flow-sensitive heap model 180 for the source code representation 100 consists of only the edge from variable q to variable a. That is, the model 180 includes the nodes 152D and 152G that respectively correspond to the variables q and a, and an edge 182 from the node 152D to the node 152G.

[0065] When flow-sensitive SAST is performed using the non-simplified heap graph 160 (i.e. , which can include deriving the flow-sensitive heap model 180 from the heap graph 160), one security vulnerability may be identified. This security vulnerability corresponds to the flow edge 182 from the node 152D for the variable q to the node 152G for the variable a in the flow-sensitive heap model 180 constructed from the heap graph 160.

[0066] Because the SAST constructs a flow-sensitive heap model 180 from the graph 160, the SAST itself is flow-sensitive. Since each node 152D and 152G in the heap model 180 corresponds to a single variable, the flow-sensitive SAST reports vulnerabilities with higher precision than the flow-insensitive SAST, in which a heap model is constructed that includes nodes 152 corresponding to multiple variables.

[0067] In the example of FIGs. 1A-1 G that has been described, therefore, a simplified heap graph 150” of the source code representation 100 is constructed in order to perform an initial, flow-insensitive SAST to identify lines that can be removed from the representation 100. In one implementation, the simplified heap graph 150” may be a Steensgaard graph, which may be constructed using Steensgaard’s algorithm.

[0068] FIG. 2 shows an example computing device 200. The computing device 200 is more generally a computing system that can include multiple discrete computing devices. The computing device 200 includes a processor 202 and a memory 204. The memory 204 is more generally a non-transitory computer-readable data storage medium, and stores program code 206 executable by the processor 202 to perform processing, such as a method.

[0069] The processing or method that is performed may be that which has been described by example with reference to FIGs. 1A-1 G. The processing can include receiving a representation 100 of source code for a program (208). The representation 100 may be generated as described in the referenced 18 / 498,961 application.

[0070] The processing can include constructing a simplified, flowinsensitive graph 150” of the representation 100 (including without first generating a non-simplified graph 150 or a partially simplified graph 150’ of the representation 100) (210). The graph 150” may be a Steensgaard graph, and constructed as if it would be used for pointer analysis in Steensgaard’s algorithm.

[0071] The processing includes performing flow-insensitive SAST analysis on the source code representation 100 - e.g., using the simplified graph 150” - to identify security vulnerabilities in the source code (212), albeit imprecisely as hasbeen described. Performing flow-insensitive SAST can include performing dataflow analyses on the graph 150” such as the dataflow analyses as described in the referenced 18 / 239,011 application.

[0072] For instance, the dataflow analyses described in the referenced 18 / 239,011 includes application of a lattice product of lattices corresponding to static analyses specified by a superlattice for the SAST. In one implementation, as is described by way of example in more detail below, the lattices can include a relevance lattice for each value of the representation 100, which is used to identify whether nodes and edges of the graph 150”, and thus variables and fields of the representation 100, are relevant or related to any security vulnerability.

[0073] This information in turn can then be used to identify the source code lines that contribute to any such identified security vulnerability. That is, source code lines that reference any such variables and variables that are relevant or related to any security vulnerability are lines that contribute to any such identified security vulnerability. The processing includes, therefore, removing any lines from the representation 100 that do not contribute to the security vulnerabilities that have been identified (214), resulting in a modified, or pruned, representation 100’ of the source code.

[0074] The processing can include constructing a non-simplified heap graph 160 of the modified source code representation 100’ (216). The graph 160 may be constructed as if it were then going to be simplified to yield a Steensgaard graph used for pointer analysis in Steensgaard’s algorithm. The processing includes performing flow-sensitive SAST analysis on the representation 100’ - e.g., using the non-simplified graph 160 - to more precisely identify security vulnerabilities in the source code (218).

[0075] Performing flow-sensitive SAST can include constructing a flowsensitive heap model 180 from the graph 160, such as by performing Andersen’s algorithm for pointer analysis, and then performing dataflow analyses using this heap model 180, such as the dataflow analyses as described in the referenced 18 / 239,011 application. As noted above, the security vulnerabilities identified by the flow-sensitive SAST are a subset of the security vulnerabilities identified by the previously performed flow-insensitive SAST, and do not include some of the false positives that the flow-insensitive SAST identified. In this way the flowsensitive SAST is more precise than the flow-insensitive SAST.

[0076] Once the vulnerabilities have been identified, a remedial action may be performed (220) with respect to the source code to resolve (or at least lessen the impact of) the more precisely identified security vulnerabilities. For example, the source code may be modified by a developer so that ultimate execution of the program will not result in the security vulnerabilities. As another example, for some types of security vulnerabilities, the source code may be automatically modified to remove the vulnerabilities. Once the remedial action has been performed, the method or processing may again be performed to determine whether the security vulnerabilities have been removed, or whether new vulnerabilities have been introduced.

[0077] FIG. 3A shows another example representation 300 of source code, which may be a generalized lower level representation that is not specific to any programming language as generated by the referenced 18 / 498,961 application. The source code representation 300 includes twelve lines, which reference variables a, b, c, d, e, p, x, and y. The variable p is assigned an object having twofields f and g that reference the variables a, b, c, d, e, x, and y. The variables a, b, c, d, e, x, and y are assigned simple values.

[0078] The first line allocates memory for an object assigned to p. Lines 2- 4 and 9 each set a referenced variable to a denoted integer value. Lines 5-7 each store the right-hand side variable to the field named on the left-hand side. Line 8 loads the field named on the right-hand side to the left-hand side variable. Lines 10 and 11 set their respective variables based on the result of a function, where in line 10 x is true (or 1 ) if d is less than or equal to e and otherwise is false (or 0), and where in line 11 y is the special value ALARM if x is true (or 1 ) and otherwise has no value. Line 12 denotes an alarm condition, indicating that there is a security vulnerability in accordance with the value of y; when y is ALARM there is a security vulnerability, and when y has no value there is not.

[0079] FIG. 3B shows an example non-simplified heap graph 350 of the source code representation 300. The graph 350 includes nodes 352A, 352B, 352C, 352D, 352E, 352P, 352X, and 352Y respectively corresponding to objects a, b, c, d, e, p, x, and y, and which are collectively referred to as the nodes 352. The graph 350 includes edges 354A, 354B, 354C, and 354D, respectively referring to as the edges 354. Each edge 354 identifies a field of the object p of the node 352D, and indicates that the field can be set to the value of a corresponding node 352 on a path through the program at execution. For example, the edges 354A and 354B indicate that the field f of p may have the same value as a and b, respectively, whereas the edge 354D indicates that d may have the same value as the field f, on a path through the program. The edge 354C indicates that the field g of p can be set to the field c on a path through the program.

[0080] Each of the nodes 352E, 352X, and 352Y are not connected to any other node 352 in the graph 350. This is because they are neither stored to any fields or variables nor loaded from any fields or variables in the program. That is, the graph 350 (specifically the edges 354 thereof) reflects just relationships among the fields and variables in the source code representation 300, and not other aspects as to their values in the program. For example, e is set to 1 in line 9 of the representation 300, but 1 is an integer and not another field. The same is true as to a, b, and c in lines 2, 3, and 4, respectively. Similarly, x and y are set in lines 10 and 11 in accordance with functions that evaluate other values, and are not set directly to other fields or variables.

[0081] FIG. 3C shows an example simplified heap graph 350’ of the source code representation 300. The edges 354A, 354B, and 354D outgoing from the node 352P and having the same label of field f in the non-simplified graph 350 are merged into a single edge 354ABD in the simplified graph 350’, and likewise the nodes 352A, 352B, and 352D are merged into a single node 352ABD. Otherwise, the simplified graph 350’ is the same as the graph 350, and includes the edge 354C from the node 352P to the node 352C, and the nodes 352E, 352X, and 352Y that are not connected to any other nodes. In particular, edge 354C and node 352C are not merged into the edge 354ABD and node 352ABD because the edges are labeled with distinct fields, f and g.

[0082] A flow-insensitive analysis is then performed on the representation 300 using the simplified graph 350’. Because the analysis uses the simplified graph 350’, it only has knowledge (and can only take into account) that a, b, and d may be referring to the same value because of their common connection to field f of object p. The analysis does not have knowledge, and cannot take into account,that d is obtained only from b and not from a in the source code representation 300. The analysis further cannot differentiate between the possibility that d is obtained from b and that b is obtained from d.

[0083] Using the lattice-based analysis described in the referenced 18 / 239,011 application, each node 352ABD, 352C, 352E, 352P, 352X, and 352Y is associated with a single lattice element and is evaluated in accordance with the dataflow induced by the representation 300. For example, if the lattice is the lattice of sets of possible integer values, the interpretation of line 2, “a =1”, requires that the lattice element associated to the node 352ABD be a set including the value 1 . Line 3, “b=2” requires that the same lattice element associated with node 352ABD be a set including the value 2. Solving for a fixed point of this analysis can produce the following result: p: {}; abd: { 1 , 2 }; c: { 3 }; e: { 1 }; x: {}; and y: {}.

[0084] A richer lattice may be used via a product of multiple lattices (i.e. , a superlattice of multiple lattices) per the 18 / 239,011 application. For instance, a lattice dimension can be introduced for sets of possible Boolean values and a lattice dimension can be introduced for determining whether an alarm may be firing (and thus whether there is a security vulnerability). In this case, the fixed point solution becomes p: {}; abd: { 1 , 2 }; c: { 3 }; e: { 1 }; x: { true, false }; and y: { ALARM }.

[0085] Therefore, the alarm on y is deemed to be potentially firing because x is potentially true, since merged abd may be 1 , which is than or equal to e, which may be 1. In actuality, there is no dataflow path through the program that results in e being 1 , since line 6 overwrites the value stored in line 5. However, the flow-insensitive analysis based on the simplified graph 350’ conflates thevalues of a, b, and d. The flow-insensitive analysis thus overapproximates what can actually occur when executing the program.

[0086] As noted above, the lattices can include a relevance lattice, which is used to identify whether nodes and edges of a graph, and thus variables and fields of the source code representation, are relevant or related to any security vulnerability. That is, an additional lattice dimension can be included to determine which variables and fields in the program contribute to each alarm instruction, and thus which variables and fields accordingly contribute to security vulnerabilities.

[0087] For example, when such a relevance lattice is used, the lattice element for this lattice that is associated with y includes the fact that knowing the value of y is relevant to determining whether the instruction “alarm y” is firing (and thus whether a security vulnerability has occurred). This information can be expressed as “relevant(y)”, and is due to the presence of the instruction “alarm y” in line 12. Line 11 further induces that x is relevant to the alarm instruction, line 10 induces that each of d and e are relevant to the alarm instruction, and so on.

[0088] In general, for each load or store instruction, “u.f = v” or “v = u.f”, if v (whether loaded or stored) is relevant to an alarm, then so is u. That is, in order to understand where v may flow from or flow to, since v may be stored at the field f of u, it is necessary to also understand where u may flow from or flow to.

[0089] The fixed point of the analysis including the sets of integers, Boolean values, facts regarding the alarm, and the relevancy facts is as follows: p: < {}, {}, {}, { relevant(y) } >; abd: < {1 , 2}, {}, {}, { relevant(y) } >; c: < {3}, {}, {}, {} >; e: < {1}, {}, {}, { relevant(y) } >; x: < {}, { true, false }, {}, { relevant(y) } >; y: < {}, {}, { ALARM }, { relevant(y) } >. The analysis converges to a solution in which the node 352C is not deemed relevant to instruction “alarm y”, and thus is not part ofa security vulnerability. Therefore, all instructions (i.e. , all lines) involving c of node 352C can be safely ignored and thus removed from the representation 300 during pruning.

[0090] FIG. 3D accordingly shows the source code representation 300 after lines relating to c have been removed, and is referenced as the pruned source code representation 300’. Lines 4 and 7 of the representation 300 have been removed in the pruned representation 300’. Flow-sensitive analysis of the pruned source code representation 300’ can then be performed to identify security vulnerabilities within the source code. For instance, such flow-sensitive analysis of the pruned representation 300’ can be performed via a non-simplified heap graph corresponding to the representation 300’.

[0091] FIG. 3E shows in this respect an example non-simplified heap graph 350” of the pruned source code representation 300’. Like the non-simplified graph 350 of the (unpruned) representation 300, the graph 350” includes nodes 352A, 352B, 352D, 352E, 352P, 352X, and 352Y respectively corresponding to fields a, b, d, e, p, x, and y. Unlike the graph 350, however, the graph 350” does not include node 352C, since lines representing field c have been removed from the pruned representation 300’.

[0092] FIG. 3F shows an example flow-sensitive heap model 360 which may be produced as part of performing SAST via derivation from the nonsimplified heap graph 350”. The heap model graph includes just a single edge 354DB from the node 352D to node 352B. This is because in any dataflow path through the pruned representation 300’, d is obtained from b. That is, even though field f of p is set to the a in line 5, this is immediately overwritten in line 6 by b in line 5. Therefore, the subsequent setting of d to field f of p in line 8 meansthat d is obtained from b in any dataflow path through the program. This fact may be obtained, and the flow-sensitive heap model 360 may thus be constructed from the source code representation 300 and the non-simplified heap graph 350”, using flow-sensitive alias analysis techniques as part of SAST.

[0093] The dataflow analysis that is performed using the flow-sensitive heap model 360 is therefore flow-sensitive as well. As compared to if SAST were performed on the non-simplified heap graph 350 of the (unpruned) representation 300, performing SAST on the non-simplified graph 350” of the pruned representation 300’ (i.e. , using the heap model 360) is faster. This is at least because the graph 350” has fewer edges 354 and nodes 352 as compared to the graph 350. That is, the flow-sensitive heap model 360 is easier to construct (and is more quickly constructed) using the graph 350” than if such a model were constructed using the graph 350. Moreover, as compared to a flow-insensitive analysis performed on the simplified graph 350’, performing flow-sensitive analysis on the non-simplified graph 350” using the flow-sensitive heap model 360 may result in fewer false positives.

[0094] For instance, in the graph 350’, d may be evaluated as 1 by flowinsensitive dataflow analysis, and thus considered a security vulnerability that fires the alarm. This is because a, b, and d are part of the same merged node 352ABD in the graph 350’, and therefore d can be construed as taking on the value of a that is set to 1 in line 2, as a result of both a and d having a relationship with the field f of p per the merged edge 352ABD. By comparison, d cannot be evaluated as 1 by a flow-sensitive dataflow analysis using the flow-sensitive heap model 360, because there is no edge from 352A to 352D in the heap model 360.Rather, d can only be evaluated as 2 since the heap model 360 has only oneedge directed into 352D, namely the edge 354DB from 352B, and b is evaluated as 2.

[0095] FIG. 4 shows another example source code representation 400. The source code representation 400 is identical to the representation 300 of FIG. 3A, except that in line 9 the field e is set to 0 in the representation 400 instead of 1 as in the representation 300. Because this difference does not concern a loading or storing relationship between fields, the representation 400 has the same non-simplified heap graph 350 of FIG. 3B and the same simplified heap graph 350’ of FIG. 3C.

[0096] The fixed point solution of the simplified graph 350’ as to the source code representation 400, including the sets of integers, Boolean values, facts regarding the alarm, and the relevant facts is as follows: p: < {}, {}, {},{ relevant(y) } >; abd: < {1 , 2}, {}, {}, { relevant(y) } >; c: < {3}, {}, {}, {} >; e: < {0}, {}, {}, { relevant(y) } >; x: < {}, { false }, {}, { relevant(y) } >; y: < {}, {}, {}, { relevant(y) } >. As compared to the fixed point solution as to the source code representation 300, in the solution as to the representation 400, the set of values that the lattice for e can have is 0 and relevant(y), as opposed to 1 and relevant(y). Further, the set of values for x in the solution as to the representation 400 includes just false and relevant(y), as opposed to true, false, and relevant(y) in the solution as to the representation 300. More significantly, the set of values for y in the solution as to the representation 400 includes just relevant(y), as opposed to ALARM and relevant(y).

[0097] This means that alarm y is not deemed as potentially firing, such that no security vulnerabilities are identified in the flow-insensitive analysis of the source code representation 400 using the simplified graph 350’. That is, evenwith the over approximation that results from the graph 350’, there is no way for x to be true. This is because the d has to have a value of 1 or 2 per its relationship with a and b as a result of being in the same merged node 352ABD, and it does not matter which of the two integers d actually is, since both are not equal to or less than 0 (i.e. , the value of e).

[0098] This means, therefore, that there are no lines in the source code representation 400 that contribute to even a security vulnerability that is a false positive. As a result, all the lines are removed from the source code representation 400 during pruning. No subsequent flow-sensitive analysis has to be performed on the representation 400 after pruning, because there are no lines left. The example of FIG. 4 thus shows how in some cases, the flow-insensitive analysis may itself be sufficient to identify that there are no security vulnerabilities, when the initial flow-insensitive analysis itself indicates that there are not any.

[0099] The examples of FIGs. 3A-3F and 4 have been described in relation to the dataflow analysis technique that is performed according to the 18 / 239,011 patent application. A source code representation may contain multiple alarms, some of which may be deemed to potentially fire in the initial flow-insensitive analysis, and some which may not be. In general, an instruction (i.e., a line) is not removed from the representation during pruning if it uses one or more fields or variables whose associated lattice element includes relevant(z) for a variable z for which the associated lattice element includes an alarm, indicating that there is a security vulnerability. All other instructions (i.e., lines) can be removed during pruning.

[0100] Techniques have been described herein for initially performing flow-insensitive SAST on a representation of source code for aprogram to identify security vulnerabilities in the program in an imprecise manner such that the security vulnerabilities may include false positive. Lines of the source code representation that do not contribute to any such imprecisely identified vulnerabilities are removed. Flow-sensitive SAST is then performed on the resulting pruned source code representation to more precisely identify security vulnerabilities in the program (and source code lines that contribute to them).

[0101] The security vulnerabilities more precisely identified by flowsensitive SAST are a subset of the security vulnerabilities that were (less precisely) identified by the initially performed flow-insensitive SAST. Specifically, the security vulnerabilities identified by the more precise flow-sensitive SAST do not include some security vulnerabilities identified by the less precise flowinsensitive SAST that are false positives.

[0102] As a result of this process, SAST is more quickly performed.Flow-sensitive SAST is performed in at least cubic time, whereas flow-insensitive SAST is performed in linear time. Therefore, removing lines from the source code representation prior to performing flow-sensitive SAST can significantly reduce how much time it takes to perform this analysis, even taking into account the time that it takes to perform the initial flow-insensitive SAST to identify which lines are to be removed.

Claims

We claim:1 . A non-transitory computer-readable data storage medium storing program code executable by a processor to perform processing comprising: performing flow-insensitive static application security testing (SAST) analysis on a representation of source code of a program to identify security vulnerabilities in the source code; removing any lines of the representation that do not contribute to the security vulnerabilities identified by the flow-insensitive SAST analysis; and performing flow-sensitive SAST analysis on the representation from which any lines that do not contribute to the security vulnerabilities identified by the flowinsensitive SAST analysis have been removed, to more precisely identify the security vulnerabilities in the source code.

2. The non-transitory computer-readable data storage medium of claim 1 , wherein the security vulnerabilities identified by the flow-sensitive SAST analysis are a subset of the security vulnerabilities performed by the flow-insensitive SAST analysis.

3. The non-transitory computer-readable data storage medium of claim 1 , wherein the processing further comprises: performing a remedial action regarding the source code to resolve the security vulnerabilities that have been identified.

4. The non-transitory computer-readable data storage medium of claim 1 , wherein the processing further comprises:constructing a simplified heap graph of the representation, wherein performing the flow-insensitive SAST analysis on the representation comprises performing SAST analysis using the simplified heap graph, such that the SAST analysis is flow-insensitive due to usage of the simplified heap graph.

5. The non-transitory computer-readable data storage medium of claim 4, wherein constructing the simplified heap graph of the representation comprises constructing a Steensgaard graph of the representation.

6. The non-transitory computer-readable data storage medium of claim 4, wherein the simplified heap graph of the representation comprises: a plurality of nodes; and a plurality of edges, each node directly connecting a corresponding pair of the nodes including a source node and a target node such that the edge is outgoing from the source node and is incoming to the target node, wherein each node has no more than one outgoing edge with a same label.

7. The non-transitory computer-readable data storage medium of claim 6, wherein one or more of the nodes each correspond to more than one field of the representation of the source code, such that the simplified heap graph is simplified and induces flow-insensitive analysis due to the one or more of the nodes that each correspond to more than one field.

8. The non-transitory computer-readable data storage medium of claim 4, wherein the processing further comprises:constructing a non-simplified heap graph of the representation from which any lines that do not contribute to the security vulnerabilities identified by the flowinsensitive SAST analysis have been removed, wherein performing the flow-sensitive SAST analysis on the representation from which any lines that do not contribute to the security vulnerabilities identified by the flow-insensitive SAST analysis have been removed comprises performing the SAST analysis using the non-simplified heap graph, such that the SAST analysis is flow-sensitive due to usage of a flow-sensitive heap model derived from the non-simplified heap graph.

9. The non-transitory computer-readable data storage medium of claim 8, wherein the non-simplified heap graph of the representation comprises: a plurality of nodes; and a plurality of edges, each node directly connecting a corresponding pair of the nodes including a source node and a target node such that the edge is outgoing from the source node and is incoming to the target node, wherein each node is not limited to having no more than one outgoing edge with a same label.

10. The non-transitory computer-readable data storage medium of claim 9, wherein every node corresponds to one value of the representation of the source code, such that the non-simplified heap graph is non-simplified.11 . The non-transitory computer-readable data storage medium of claim 8, wherein performing the SAST analysis comprises: executing generalized dataflow analysis executable code on therepresentation of the source code using a lattice product of lattices corresponding to static analyses specified by a superlattice for SAST.

12. The non-transitory computer-readable data storage medium of claim 11 , wherein the lattices include a relevance lattice having a lattice element for each of a plurality of values of the representation, wherein the lattice is used to identify whether variables and fields of the source code are relevant to the security vulnerabilities that are identified.

13. The non-transitory computer-readable data storage medium of claim 1 , wherein the representation of the source code comprises a generalized lower level representation that is not specific to any programming language.

14. The non-transitory computer-readable data storage medium of claim 1 , wherein the security vulnerabilities that are identified by performing the flowinsensitive SAST analysis comprise: actual security vulnerabilities that are true positives; and non-actual security vulnerabilities that are false positives due to the flowinsensitive SAST analysis being flow-insensitive.

15. A computing device comprising: a processor; and a memory storing instructions executable by the processor to: construct a simplified heap graph of a representation of source code of a program; perform static application security testing (SAST) analysis on therepresentation using the simplified heap graph to identify security vulnerabilities in the source code, wherein the SAST analysis is flow-insensitive due to usage of a flow-insensitive heap model derived from the simplified heap graph; remove any lines of the representation that do not contribute to the security vulnerabilities identified by the flow-insensitive SAST analysis; construct a non-simplified heap graph of the representation from which any lines that do not contribute to the security vulnerabilities identified by the flow-insensitive SAST have been removed; and perform the SAST analysis on the representation using the nonsimplified graph to more precisely identify the security vulnerabilities in the source code, wherein the SAST analysis is flow-sensitive due to usage of a flow-sensitive heap model derived from the non-simplified heap graph.

16. The computing device of claim 15, wherein the program code is executable by the processor to further: perform a remedial action regarding the source code to resolve the security vulnerabilities that have been identified.

17. The computing device of claim 15, wherein the simplified heap graph of the representation comprises a Steensgaard graph of the representation.

18. The computing device of claim 15, wherein the simplified heap graph of the representation comprises: a plurality of nodes; and a plurality of edges, each node directly connecting a corresponding pair of the nodes including a source node and a target node such that the edge isoutgoing from the source node and is incoming to the target node, wherein each node has no more than one outgoing edge with a same label.

19. The non-transitory computer-readable data storage medium of claim 6, wherein one or more of the nodes each correspond to more than one field of the representation of the source code, such that the simplified heap graph is simplified.

20. A method comprising: constructing, by a processor, a simplified Steensgaard graph of a representation of source code of a program; performing, by the processor, static application security testing (SAST) analysis on the representation using the simplified Steensgaard graph to identify security vulnerabilities in the source code, wherein the SAST analysis is flowinsensitive due to usage of a flow-insensitive heap model derived from the simplified Steensgaard graph; removing, by the processor, any lines of the representation that do not contribute to the security vulnerabilities identified by the flow-insensitive SAST; constructing, by the processor, a non-simplified heap graph of the representation from which any lines that do not contribute to the security vulnerabilities identified by the flow-insensitive SAST have been removed; and performing, by the processor, the SAST analysis on the representation using the non-simplified graph to more precisely identify the security vulnerabilities in the source code, wherein the SAST analysis is flow-sensitive due to usage of a flow-sensitive heap model derived from the non-simplified heap graph.

Citation Information

Patent Citations

  • Machine-checkable code-annotations for static application security testing

    US20170169228A1

  • Source code clustering for automatically identifying false positives generated through static application security testing

    US20230169177A1

  • Position analysis of source code vulnerabilities

    US9792443B1