Code analysis chart construction method and related product

By converting the source code into intermediate code and building a code analysis diagram, the problem of insufficient support for multiple programming languages ​​in the prior art is solved, and a wider scope of application and higher compatibility are achieved.

CN119938050APending Publication Date: 2025-05-06TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311470801.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-06
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing code attribute graph (CPG) method is mainly applicable to C language and cannot support other programming languages, resulting in a small scope of application and poor compatibility, especially the analysis needs of object-oriented languages.

Method used

By identifying the file type in the source code directory, converting the file into an intermediate code that matches the type, and constructing a code analysis diagram based on the intermediate code. Intermediate code has nothing to do with the characteristics of the programming system, can serve different programming systems, and is suitable for multiple programming languages.

Benefits of technology

It expands the scope of application of code analysis diagrams, improves the compatibility of program static analysis, and can effectively analyze programs in many different programming languages, including object-oriented languages.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938050A_ABST
    Figure CN119938050A_ABST
Patent Text Reader

Abstract

The invention discloses a code analysis chart construction method and a related product. The method comprises the following steps: identifying the type of a file under a source code directory; based on the type of the file, converting the file into an intermediate code matched with the type of the file; nodes for constructing the code analysis graph are determined based on code composition elements of the intermediate code, edges between the nodes are determined based on element relationships between the code composition elements, and the code analysis graph is constructed based on the nodes and the edges between the nodes. Since the intermediate code is irrelevant to the characteristics of the programming system and can serve different programming systems, the file is converted into the intermediate code before composition, and the code analysis chart can be constructed by means of the intermediate code without considering the programming language used by the file under the source code directory. Therefore, the constructed code analysis chart can be suitable for static analysis of programs of various different programming languages, so that the application range is expanded, and the compatibility of static analysis of the programs is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a method for constructing a code analysis graph and related products. Background Art

[0002] Static analysis generally refers to a method of analyzing program code without running the code. Code Property Graph (CPG) is a commonly used static analysis method that can be used to discover various vulnerabilities in source code.

[0003] However, in actual applications, CPG can generally only process C language (a computer programming language) and cannot support other programming languages. Therefore, it has a small scope of application and poor compatibility. Summary of the invention

[0004] The embodiments of the present application provide a method for constructing a code analysis graph and related products, so that the constructed code analysis graph can be applied to static analysis of programs in a variety of different programming languages, thereby expanding the scope of application and improving the compatibility of program static analysis.

[0005] The embodiments of the present application disclose the following technical solutions:

[0006] In a first aspect, an embodiment of the present application provides a method for constructing a code analysis graph, comprising:

[0007] Identify the types of files in the source code directory;

[0008] Based on the type of the file, converting the file into an intermediate code matching the type of the file;

[0009] Nodes for constructing a code analysis graph are determined based on the code constituent elements of the intermediate code, edges between the nodes are determined based on element relationships between the code constituent elements, and the code analysis graph is constructed based on the nodes and the edges between the nodes.

[0010] In a second aspect, an embodiment of the present application provides a device for constructing a code analysis graph, including:

[0011] Type identification module, used to identify the types of files in the source code directory;

[0012] A file conversion module, used for converting the file into an intermediate code matching the type of the file based on the type of the file;

[0013] A graph construction module determines nodes for constructing a code analysis graph based on code constituent elements of the intermediate code, determines edges between the nodes based on element relationships between the code constituent elements, and constructs the code analysis graph based on the nodes and the edges between the nodes.

[0014] In a third aspect, an embodiment of the present application provides a device for constructing a code analysis graph, the device comprising a processor and a memory:

[0015] The memory is used to store a computer program and transmit the computer program to the processor;

[0016] The processor is used to execute the steps of the method for constructing a code analysis graph provided in the first aspect according to the instructions in the computer program.

[0017] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which is used to store a computer program. When the computer program is executed by a computer device, it implements the steps of the method for constructing a code analysis graph provided in the first aspect.

[0018] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program or instructions, which, when executed by a computer device, implements the steps of the method for constructing a code analysis graph provided in the first aspect.

[0019] It can be seen from the above technical solutions that the embodiments of the present application have the following advantages:

[0020] In an embodiment of the present application, the type of the file under the source code directory can be first identified to determine the type of the file, and then based on the type of the file, the file can be converted into an intermediate code that matches the type of the file; then, the nodes used to construct the code analysis graph can be determined based on the code constituent elements of the intermediate code, the edges between the nodes can be determined based on the element relationship between the code constituent elements, and the code analysis graph can be constructed based on the edges between the nodes and the nodes. Since the intermediate code is independent of the characteristics of the programming system, it can serve different programming systems, that is, the code constituent elements of the intermediate code can match the nodes and edges used to construct the static analysis graph. Therefore, before composing the graph, the file is first converted into the intermediate code, and the code analysis graph can be constructed with the help of the intermediate code, without considering the programming language used by the file under the source code directory. In this way, the constructed code analysis graph can be applied to static analysis of programs in a variety of different programming languages, thereby expanding the scope of application and improving the compatibility of program static analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 A flowchart of a method for constructing a code analysis graph provided in an embodiment of the present application;

[0022] Figure 2 A schematic diagram of an abstract syntax tree provided in an embodiment of the present application;

[0023] Figure 3a A schematic diagram of a code provided in an embodiment of the present application;

[0024] Figure 3b A schematic diagram of a first graph structure provided in an embodiment of the present application;

[0025] Figure 3c A schematic diagram of a second graph structure provided in an embodiment of the present application;

[0026] Figure 3d A schematic diagram of a code analysis diagram provided in an embodiment of the present application;

[0027] Figure 4a A schematic diagram of a third diagram structure provided in an embodiment of the present application;

[0028] Figure 4b A schematic diagram of a target map provided in an embodiment of the present application;

[0029] Figure 5 A schematic diagram of the structure of a device for constructing a code analysis graph provided in an embodiment of the present application;

[0030] Figure 6 A schematic diagram of the structure of a server provided in an embodiment of the present application;

[0031] Figure 7 A schematic diagram of the structure of a terminal device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0032] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0033] As mentioned above, in the actual application of static analysis scenarios, CPG can generally only process C language (a computer programming language) and cannot support other programming languages. Therefore, it has a small scope of application and poor compatibility. Specifically, in order to analyze a procedural language such as C language, the called methods of function nodes in CPG belong to fixed class methods. For object-oriented languages, the called methods of function nodes need to change with the changes in context. Therefore, the currently constructed CPG cannot meet the analysis needs of object-oriented languages.

[0034] Based on the above problems, an embodiment of the present application provides a method for constructing a code analysis graph, which includes: first, the type of the file under the source code directory can be identified, the type of the file can be determined, and then based on the type of the file, the file can be converted into an intermediate code that matches the type of the file; then, the nodes used to construct the code analysis graph can be determined based on the code constituent elements of the intermediate code, the edges between the nodes can be determined based on the element relationship between the code constituent elements, and the code analysis graph can be constructed based on the edges between the nodes and the nodes. Since the intermediate code is independent of the characteristics of the programming system, it can serve different programming systems, that is, the code constituent elements of the intermediate code can match the nodes and edges used to construct the static analysis graph. Therefore, before composing the graph, the file is first converted into the intermediate code, and the code analysis graph can be constructed with the help of the intermediate code, without considering the programming language used by the file under the source code directory. In this way, the constructed code analysis graph can be applied to static analysis of programs in a variety of different programming languages, thereby expanding the scope of application and improving the compatibility of program static analysis.

[0035] In addition, in an embodiment of the present application, another graph structure (hereinafter referred to as the third graph structure) can be introduced based on the above-constructed code analysis graph. The graph structure can be constructed based on the inheritance relationship, attribution relationship or implementation relationship between the elements of the object-oriented language. In this way, the graph structure can be used to implement the analysis of the object-oriented language, solve the problem that CPG cannot meet the analysis requirements of the object-oriented language, further improve the scope of application of the above-mentioned code analysis graph, and thus further improve the compatibility of the static analysis of the program. For the specific implementation method, please refer to the introduction made below.

[0036] Next, several terms that may be involved in the embodiments of the present application are explained.

[0037] Reverse engineering is a process of technological imitation, which involves reverse analysis and research of a target product, thereby deducing and deriving the product's processing flow, organizational structure, functional performance specifications and other design elements, in order to produce products with similar functions but not completely the same.

[0038] The abstract syntax tree is an abstract representation of the syntax structure of the source code. It represents the syntax structure of the programming language in a tree form, and each node in the tree represents a data structure in the source code.

[0039] A control flow graph is a graphical data structure that represents the execution order of data in a program and the conditional branches that need to be met. In a control flow graph, nodes represent code statements, and nodes are connected by directed edges to represent the execution order and branches.

[0040] A data flow graph is a graphical data structure that represents the flow of data in a program. A data flow graph can be used to capture the relationship between the definition and use of variables in a program. In a data flow graph, nodes represent instructions or statements in a program, and edges represent the flow of data from one instruction or statement to another.

[0041] The code property graph is a graph structure that integrates the abstract syntax tree, control flow graph, and data flow graph, which facilitates the analysis of multiple data types. For example, during the graph traversal process, vulnerability scanning templates can be elegantly organized to analyze buffer overflows, integer overflows, format string vulnerabilities, memory leaks, and other issues.

[0042] Three-address code is a machine instruction that can use three operands to write efficient and flexible machine instructions. Three-address code instructions are commonly used in assembly language and embedded systems.

[0043] It should be noted that the embodiments of the present application do not limit the execution subject for executing the technical solution of the present application. For example, the method for constructing the code analysis diagram of the embodiment of the present application can be applied to a terminal device or a server, or, the terminal device and the server can be collaboratively processed. As an example, the terminal device includes but is not limited to a mobile phone, a desktop computer, a tablet computer, a laptop computer, a PDA, an intelligent voice interaction device, a smart home appliance, a vehicle-mounted terminal, an aircraft, etc. The server can be a stand-alone server, a cluster server, or a cloud server.

[0044] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0045] Figure 1 A flowchart of a method for constructing a code analysis graph provided in an embodiment of the present application. Figure 1 As shown, the method for constructing a code analysis graph provided in the embodiment of the present application may include:

[0046] S101: Identify the types of files in the source code directory.

[0047] The source code directory may refer to a directory of files related to the source code. In an embodiment of the present application, the files under the source code directory may include source code files and binary files. Among them, the source code file may be a file including source code, for example, a file of C language type, a file of C++ (a computer programming language) type, or a file of Python (a computer programming language) type. The binary file may be a file generated after compiling or assembling the source code, for example, a portable executable file (Portable Executable, PE), a Mach object (a file format) type file, or an executable and linkable format (Executable and Linkable Format, ELF) file. Correspondingly, in an embodiment of the present application, for the above-mentioned step S101, there are different ways to identify the types of the two files, source code files and binary files. The following can be divided into two cases to exemplify the process of identifying the file type.

[0048] As an example, for a source code file, the type of the source code file can be determined based on the suffix information of the source code file. For example, in a C language type file, the suffix information of the file can be ".c"; in a C++ type file, the suffix information of the file can be ".cpp" or ".C"; in a Java (a computer programming language) type file, the suffix information of the file can be ".java".

[0049] As another example, for a binary file, the type of the binary file can be determined based on the file header information of the binary file. The file header information may refer to a byte sequence in the header of the binary file, and the first 4 bytes of the byte sequence, that is, the magic number of the binary file, are used to identify the format of the binary file. For example, in a PE type file, the file header information of the file may include the byte "0x5A" or "0x4D"; in a Mach object type file, the file header information of the file may include the byte "0xFE", "0xED", "0xFA", "0xCE", "0x0C" or "0x00"; in an ELF type file, the file header information of the file may include the byte "0x7F", "0x45", "0x4C" or "0x46".

[0050] Thus, the purpose of step S101 to identify the types of files in the source code directory is to adaptively select a method that matches the type of the file according to the type of the file to convert the file into a corresponding intermediate code.

[0051] S102: Based on the type of the file, convert the file into an intermediate code matching the type of the file.

[0052] Intermediate code generally refers to an equivalent internal representation code of a syntax-oriented source code. Since the intermediate code is independent of the characteristics of the programming system and can serve different programming systems, it is not necessary to consider the programming language used by the files in the source code directory when building a code analysis diagram based on the intermediate code. Generally speaking, the form of the intermediate code can include reverse Polish notation, ternary form, quadruple form and tree representation. Correspondingly, in an embodiment of the present application, the intermediate code corresponding to the source code file can be an abstract syntax tree; the intermediate code corresponding to the binary file can be a three-address code.

[0053] S103: Determine nodes for constructing a code analysis graph based on the code constituent elements of the intermediate code, determine edges between nodes based on element relationships between the code constituent elements, and construct a code analysis graph based on the edges between the nodes.

[0054] The code constituent elements of the intermediate code are, for example, various elements such as classes, objects, interfaces, functions, structures, and enumeration bodies. In an embodiment of the present application, the method for obtaining the code constituent elements of the intermediate code can be reflected in first traversing the intermediate code, and then identifying the code structure obtained by the traversal, so as to obtain the corresponding code constituent elements. In addition, in an embodiment of the present application, the obtained code constituent elements can also be recorded. For example, the code constituent elements can be recorded in the form of the following Table 1. In Table 1, in addition to recording the code constituent elements themselves, the relevant information of the code constituent elements can also be recorded.

[0055] Table 1

[0056]

[0057] As mentioned above, when building a code analysis graph based on the intermediate code, it is not necessary to consider the programming language used by the files in the source code directory. In this way, the nodes are subsequently constructed based on the code components of the intermediate code, and the edges between the nodes are constructed based on the element relationships between the code components. The resulting code analysis graph can be applied to files in various programming languages, thereby expanding the scope of application and improving the compatibility of program static analysis.

[0058] In an embodiment of the present application, for the above step S102, the file can be converted into an intermediate code that matches its type based on the type of the source code file and the type of the binary file. In this way, the method that matches the type of the file can be adaptively selected in combination with the file type to convert the file into the corresponding intermediate code, which is convenient for the subsequent construction of a code analysis diagram with a wider range of applications and improves the compatibility of the static analysis of the program. For ease of understanding, the following is an exemplary description of the process of obtaining the intermediate code, that is, step S102, in combination with the type of the source code file and the type of the binary file, according to different situations.

[0059] As an example, for the source code file, step S102 may specifically include step 1A-step 1B (it should be noted that step 1A-step 1B are not shown in the figure):

[0060] Step 1A: Based on the type of the source code file, determine the parsing program corresponding to the source code file.

[0061] The type of the source code file is determined based on the suffix information of the source code file. The specific determination method can refer to the relevant content of the aforementioned step S101, which will not be repeated here.

[0062] In actual applications, different types of source code files correspond to different parsing programs. For example, the parsing programs corresponding to C language type and C++ type files can both be clang (a compiler); the parsing program corresponding to Python type files can be PyAST (a compiler); the parsing program corresponding to Java type files can be Eclipse JDT (a compiler).

[0063] Step 1B: Convert the source code file through a parsing program to obtain an abstract syntax tree corresponding to the source code file as the intermediate code of the source code file.

[0064] Here, the parsing program can be specifically used to perform lexical analysis, syntax analysis and semantic analysis on the source code file, so as to convert the source code file into an abstract syntax tree. For ease of understanding, the embodiment of the present application can also provide an example of converting data into an abstract syntax tree. Figure 2 A schematic diagram of an abstract syntax tree provided in an embodiment of the present application, specifically taking an operation expression "x = a*b+c" as an example, then in Figure 2In the example, the root node is the operator "=", the first-level child nodes adjacent to the root node may include the character "x" and the character "+", the second-level child nodes adjacent to the child node "+" may include the operator "*" and the character "c", and the third-level child nodes adjacent to the child node "*" may include the character "a" and the character "b". In this way, through the root node and each child node, as well as the edges used to represent the adjacent relationship, the grammatical abstract graph corresponding to the operation expression "x=a*b+c" can be obtained.

[0065] In this way, the source code file is converted by a parsing program corresponding to the type of the source code file, and an abstract syntax tree corresponding to the source code file can be obtained.

[0066] As another example, for a binary file, step S102 may specifically include step 2A-step 2B (it should be noted that step 2A-step 2B are not shown in the figure):

[0067] Step 2A: Based on the type of the binary file, determine the reverse analysis program corresponding to the binary file.

[0068] The type of the binary file is determined based on the file header information of the binary file. The specific determination method can refer to the relevant content of the aforementioned step S101, which will not be repeated here.

[0069] In actual applications, different types of binary files correspond to different reverse analysis programs. For example, the reverse analysis programs corresponding to ELF type files, PE type files, and Mach object type files can all be Ghidra (a reverse analysis tool); the reverse analysis program corresponding to Java bytecode type files can be Soot (a reverse analysis tool); the reverse analysis program corresponding to Low Level Virtual Machine (LLVM) bytecode type files can be the LLVM tool suite.

[0070] Step 2B: Convert the binary file through a reverse analysis program to obtain the three-address code corresponding to the binary file as the intermediate code of the binary file.

[0071] Here, the reverse analysis program can be specifically used to reverse analyze the binary file to understand the working principle, algorithm, data structure and logic information inside the compiled program, so as to convert the binary file into a three-address code. For ease of understanding, the embodiment of the present application can also provide an example of converting data into a three-address code. Still taking an operation expression "x=a*b+c" as an example, the following formulas (2-1), (2-2) and (2-3) are the three-address codes corresponding to the operation expression:

[0072] T1=a*b (2-1)

[0073] T2=T1+c (2-2)

[0074] x=T2 (2-3)

[0075] Among them, T1 and T2 can be represented as temporary variables generated by the reverse analysis program. In this way, by converting the binary code file through the reverse analysis program corresponding to the type of the binary file, the three-address code corresponding to the binary code file can be obtained.

[0076] Further, for the above step S103, the code composition elements and the element relationship between the code composition elements can be determined based on the intermediate code, so as to obtain the nodes and edges used to construct the code analysis graph, and then the code analysis graph is constructed by combining the nodes and edges. Based on this, in the embodiment of the present application, the construction process of the above code analysis graph, that is, step S103, can specifically include steps 31-34 (it should be noted that steps 31-34 are not shown in the figure):

[0077] Step 31: extract multiple code component elements and related information of the multiple code component elements from the intermediate code.

[0078] As mentioned above, in the embodiment of the present application, the obtained code component elements and related information can be recorded, as shown in the above Table 1. Therefore, multiple code component elements and corresponding related information can be directly extracted from Table 1 here.

[0079] Step 32: Based on the multiple code component elements and the related information, determine the control flow relationship and the data flow relationship between the multiple code component elements.

[0080] Among them, control flow relationship and data flow relationship are both subordinate to element relationship. For ease of understanding, the meaning of control flow relationship and data flow relationship are respectively explained below with examples. For control flow relationship, taking conditional statement element as an example, it can be determined that there is a control flow relationship between conditional statement element and the corresponding "yes" branch element and "no" branch element; for data flow relationship, taking function element as an example, it can be determined that there is a data flow relationship between function and the actual parameter element passed by the corresponding caller.

[0081] Step 33: Create multiple nodes one-to-one based on multiple code component elements; and create edges between the multiple nodes based on the control flow relationship and the data flow relationship.

[0082] For example, a Property node can be created for a property element in a code component element, an assignment node can be created for an assignment operation element, and a Call node can be created for a function element. In addition, the relevant information of the above-mentioned multiple code component elements can be used as node attribute information for representing the corresponding code component elements, and stored in the code analysis graph after obtaining the graph.

[0083] Further, the present embodiment may not specifically limit the implementation process of creating edges between multiple nodes in step 33. For ease of understanding, a possible implementation is described below.

[0084] As a possible implementation, the process of creating edges between multiple nodes may include: creating a first type of edge between multiple nodes based on a control flow relationship to obtain a first graph structure; and creating a second type of edge between multiple nodes based on a data flow relationship to obtain a second graph structure. Correspondingly, still taking the control flow relationship corresponding to the above-mentioned conditional statement element as an example, after creating node 1 for representing the conditional statement element, node 2 for representing the "yes" branch element, and node 3 for representing the "no" branch element, node 1 can be connected to node 2 and node 3 respectively, thereby determining the edge between node 1, node 2 and node 3, that is, the edge corresponding to the control flow relationship between the three. Still taking the data flow relationship corresponding to the above-mentioned function element as an example, after creating node 4 for representing the function element and node 5 for representing the actual parameter element passed by the caller corresponding to the function element, node 4 and node 5 can be connected, thereby determining the edge between node 4 and node 5, that is, the edge corresponding to the data flow relationship between the two.

[0085] For ease of understanding, the present application embodiment may also provide a schematic diagram of the first graph structure and the second graph structure. Specifically, Figure 3a A schematic diagram of a code provided in an embodiment of the present application; Figure 3b A schematic diagram of a first graph structure provided in an embodiment of the present application; Figure 3c A schematic diagram of a second graph structure provided in an embodiment of the present application. Figure 3a The code shown is taken as an example, combined with Figure 3b As shown, the first graph structure includes multiple frame lines and edges between the frame lines; wherein the frame lines are nodes, and the edges between the frame lines are control flow relationships between the nodes. Figure 3c As shown, the second graph structure also includes multiple frame lines and edges between the multiple frame lines; wherein the frame lines are nodes, and the edges between the multiple frame lines are data flow relationships between the nodes.

[0086] Step 34: Construct a code analysis graph based on the intermediate code, the multiple nodes, and the edges between the multiple nodes.

[0087] In an embodiment of the present application, the construction process of the code analysis graph, that is, step S34, may include: fusing the intermediate code, the first graph structure and the second graph structure to obtain a code analysis graph. As mentioned earlier, for the source code files under the source code directory, the intermediate code can be in the form of an abstract syntax tree. Therefore, the code analysis graph can be obtained by fusing the abstract syntax tree, the first graph structure and the second graph structure. For the binary files under the source code directory, the intermediate code can be in the form of a three-address code, and the three-address code is actually a linear representation of the abstract syntax tree. Therefore, the three-address code needs to be converted into the corresponding graph structure first, and then merged with the first graph structure and the second graph structure, so that the code analysis graph can be obtained.

[0088] For ease of understanding, the embodiment of the present application can also illustrate the construction of the code analysis diagram in conjunction with the accompanying drawings. Figure 3d A schematic diagram of a code analysis diagram provided in an embodiment of the present application. Figure 3d Specifically, different lines are used to distinguish the edges in the abstract syntax tree, the first graph structure, and the second graph structure. In addition, when performing graph fusion, duplicate nodes and edges can be removed to further optimize the constructed code analysis graph.

[0089] In addition, as mentioned above, the embodiment of the present application can also introduce a third graph structure based on the above-constructed code analysis graph, so as to realize the analysis of object-oriented languages, solve the problem that CPG cannot meet the analysis requirements of object-oriented languages, further improve the scope of application of the above-mentioned code analysis graph, and further improve the compatibility of program static analysis. Based on this, in the embodiment of the present application, the process of introducing the third graph structure based on the code analysis graph can specifically include steps 41-43 (it should be noted that steps 41-43 are not shown in the figure):

[0090] Step 41: traverse the intermediate code, the first graph structure, the second graph structure and the code analysis graph to determine the first elements in the code component elements and the first relationship between the first elements.

[0091] The first element is an element of an object-oriented language, such as an object, a class, an interface, an attribute, and a function. Correspondingly, the first relationship between the first elements may include an inheritance relationship, an attribution relationship, or an implementation relationship. In practical applications, if a class element and a function element are used as first elements, the first relationship between the two is an attribution relationship; if a class element and an interface element are used as first elements, the first relationship between the two is an implementation relationship; if different class elements are used as first elements, the first relationship between them is an implementation relationship.

[0092] Step 42: Construct a third graph structure with the first element as the first node and the first relationship as the edge between the first nodes.

[0093] For ease of understanding, the embodiment of the present application may illustrate the construction of the third structural diagram in conjunction with the accompanying drawings. Figure 4a A schematic diagram of a third structure provided in an embodiment of the present application. Figure 4a As shown, the third graph structure includes multiple frames and edges between the multiple frames. Specifically, the multiple frames are the first nodes; "Class" represents the nodes of the class element, including the node "Class1" and the node "Class2"; "Method" represents the nodes of the function element, including the node "Method1", the node "Method2" and the node "Method3"; "Interface" represents the node of the interface element, Figure 4a The node in the figure is "Interface1". The edges between multiple frames are the first relationships between the first nodes; among them, "extend" indicates an inheritance relationship, "interface" indicates an implementation relationship, and "has" indicates an ownership relationship.

[0094] Step 43: Fuse the third graph structure and the code analysis graph to obtain a target graph.

[0095] Here, the target graph can be used to analyze object-oriented languages. This is because the first node in the third graph structure is created with elements in the object-oriented language, and the edges between the first nodes are also created with the relationship between the elements in the object-oriented language. Therefore, the third graph structure can realize the analysis of the object-oriented language. Based on this, the target graph obtained by fusing the third graph structure and the code analysis graph can also be used to analyze object-oriented languages, solving the problem that CPG cannot meet the analysis requirements of object-oriented languages, thereby further improving the scope of application of the above-mentioned code analysis graph and improving the compatibility of program static analysis.

[0096] In addition, in actual applications, when there is a Call node for representing a function element in the above code analysis graph, the Call node will be difficult to determine which class definition function is called due to the lack of type information, which will cause errors in the subsequent static analysis based on the code analysis graph. Based on this, in an embodiment of the present application, the third graph structure and the code analysis graph can also be fused according to the third graph structure to obtain the target graph. The method for obtaining the target graph, that is, step 43, is described in detail below.

[0097] For ease of understanding, the implementation of step 43 is described below with reference to the accompanying drawings. Figure 4bA schematic diagram of a target graph provided for an embodiment of the present application. The above step 43 may specifically include: traversing the code analysis graph based on the third graph structure to determine the type information of the second node among the multiple nodes; the second node is a node used to represent a function element in the code composition element; based on the type information of the second node, the third graph structure and the code analysis graph are merged to obtain the target graph. Here, the second node is the above-mentioned Call node. Combined Figure 4b As shown, the code analysis graph is traversed based on each node in the third graph structure, and based on the graph data obtained by the traversal, the nodes related to the second node can be determined from the third graph structure, for example, there is an association relationship between the node "Method2'" and the Call node, that is, the second node. Therefore, the type information of the second node can be determined based on the "Method2'" node, and the third graph structure and the code analysis graph are connected based on the type information to complete the fusion, thereby obtaining the target graph.

[0098] In addition, in the embodiment of the present application, after obtaining the code analysis graph, the code analysis graph can be used to analyze the files in the source code directory, thereby realizing the static analysis of the program. Regarding the analysis process of the files in the source code directory, the embodiment of the present application may not be specifically limited. For ease of understanding, it is described below in combination with a variety of possible implementation methods.

[0099] As a possible implementation, the analysis process of the files under the source code directory may include: storing the code analysis graph in the graph database; receiving an operation command for the graph database; the operation command is used to instruct the analysis of the file; traversing the code analysis graph based on the operation command; analyzing the file based on the graph data in the code analysis graph obtained by traversal. Here, file analysis is relatively simple and fast with the help of graph database technology, which can reduce the cost of transformation. In practical applications, the operation command for the graph database can be reflected in detecting whether there are vulnerabilities or other potential problems in the file based on the code analysis graph. Accordingly, when the code analysis graph is traversed through the interface based on the operation command to implement code analysis, when it is determined that there are vulnerabilities or other potential problems in the file, an alarm can be issued, thereby realizing vulnerability identification of the source code. In addition, when the code analysis graph is traversed based on the operation command, the operation command can be a batch query statement in the graph database, so that batch query operations can be performed, reducing the number of I / O operations of the graph database, thereby reducing the waiting time of I / O operations.

[0100] As another possible implementation, the analysis process of the files under the source code directory may include: extracting the target code based on the code analysis graph; the target code is used to analyze part or all of the code in the file; encapsulating the target code into a code analysis toolkit; and calling the code analysis toolkit when receiving an analysis instruction for part or all of the code in the file. Here, the code analysis toolkit can be implemented in the form of a software development kit (SDK). In this way, by extracting the target code and encapsulating it into an SDK, it is convenient to directly call it later to realize the corresponding analysis function. For example, if you want to obtain all functions directly from the files under the source code directory, you can determine the code corresponding to all functions as the target code and encapsulate it into an SDK. In this way, when the SDK is called later, all functions under the source code directory can be returned.

[0101] In addition, after the target code is packaged as an SDK, it can be further packaged as a domain-specific language (DSL). For example, the DSL can be used to determine the number of functions in a file under the source code directory. If it exceeds the preset number, an alarm will be issued, thereby realizing the detection of the number of functions in the file. In this way, the corresponding key identifier in the DSL can be directly called to complete the analysis of part or all of the code in the file without the need for complex programming language compilation or configuration of the development environment, which can greatly reduce the difficulty and time cost of development.

[0102] Furthermore, in actual applications, there may be multiple files in the source code directory, and each of the multiple files corresponds to a code analysis graph. In the static analysis process, the same code analysis graph may be analyzed multiple times. Therefore, in an embodiment of the present application, the current analysis result of the code analysis graph can be stored, and when the code analysis graph is jumped to for analysis next time, the analysis result can be directly obtained, thereby reducing the time-consuming repeated calculations and improving the analysis performance.

[0103] Based on the method for constructing a code analysis graph provided in the above embodiment, the embodiment of the present application can also provide a device for constructing a code analysis graph. The device for constructing a code analysis graph is described below in conjunction with the embodiments and drawings.

[0104] Figure 5 A schematic diagram of a device for constructing a code analysis graph provided in an embodiment of the present application. Figure 5 As shown, the device 500 for constructing a code analysis graph provided in an embodiment of the present application includes:

[0105] A type identification module 501 is used to identify the type of files in the source code directory;

[0106] A file conversion module 502, configured to convert the file into an intermediate code matching the type of the file based on the type of the file;

[0107] The graph construction module 503 determines nodes for constructing a code analysis graph based on the code constituent elements of the intermediate code, determines edges between the nodes based on element relationships between the code constituent elements, and constructs the code analysis graph based on the edges between the nodes.

[0108] Optionally, the files under the source code directory include source code files under the source code directory; the file conversion module 502 includes:

[0109] A first program determination module, configured to determine a parsing program corresponding to the source code file based on the type of the source code file; the type of the source code file is determined based on suffix information of the source code file;

[0110] The first conversion module is used to convert the source code file through the parsing program to obtain an abstract syntax tree corresponding to the source code file as the intermediate code of the source code file.

[0111] Optionally, the files under the source code directory include binary files under the source code directory; the file conversion module 502 includes:

[0112] A second program determination module, configured to determine a reverse analysis program corresponding to the binary file based on the type of the binary file; the type of the binary file is determined based on file header information of the binary file;

[0113] The second conversion module is used to convert the binary file through the reverse analysis program to obtain the three-address code corresponding to the binary file as the intermediate code of the binary file.

[0114] Optionally, the graph construction module 503 includes:

[0115] A first extraction module, used to extract a plurality of code components and related information of the plurality of code components from the intermediate code;

[0116] A relationship determination module, configured to determine a control flow relationship and a data flow relationship between the plurality of code component elements based on the plurality of code component elements and the relevant information; the control flow relationship and the data flow relationship both belong to the element relationship;

[0117] A first construction module is used to create a plurality of nodes based on the plurality of code component elements one-to-one; and to create edges between the plurality of nodes based on the control flow relationship and the data flow relationship;

[0118] The second construction module is used to construct the code analysis graph based on the intermediate code, the multiple nodes, and the edges between the multiple nodes.

[0119] Optionally, the first building block includes:

[0120] A first construction submodule is configured to create a first type of edge between the plurality of nodes based on the control flow relationship to obtain a first graph structure; and to create a second type of edge between the plurality of nodes based on the data flow relationship to obtain a second graph structure;

[0121] The second building block comprises:

[0122] The second sub-construction module is used to merge the intermediate code, the first graph structure and the second graph structure to obtain the code analysis graph.

[0123] Optionally, the device 500 for constructing the code analysis graph further includes:

[0124] A first traversal module, used for traversing the intermediate code, the first graph structure, the second graph structure and the code analysis graph, and determining a first element in the code component elements and a first relationship between the first elements; the first element is an element of an object-oriented language; the first relationship includes an inheritance relationship, an attribution relationship or an implementation relationship;

[0125] A third construction module, configured to construct a third graph structure using the first elements as first nodes and the first relationships as edges between the first nodes;

[0126] The second fusion module is used to fuse the third graph structure and the code analysis graph to obtain a target graph; the target graph is used to analyze the object-oriented language.

[0127] Optionally, the second fusion module includes:

[0128] A second traversal module, used to traverse the code analysis graph based on the third graph structure, and determine type information of a second node among the multiple nodes; the second node is a node used to represent a function element in the code component element;

[0129] The second fusion submodule is used to fuse the third graph structure and the code analysis graph based on the type information of the second node to obtain the target graph.

[0130] Optionally, the device 500 for constructing the code analysis graph further includes:

[0131] A storage module, used for storing the code analysis graph in a graph database;

[0132] A communication module, used for receiving an operation command for the graph database; the operation command is used for instructing to analyze the file;

[0133] A third traversal module, used for traversing the code analysis graph based on the operation command;

[0134] The analysis module is used to analyze the file based on the graph data in the code analysis graph obtained by traversal.

[0135] Optionally, the device 500 for constructing the code analysis graph further includes:

[0136] A second extraction module is used to extract target code based on the code analysis graph; the target code is used to analyze part or all of the code in the file;

[0137] A code packaging module, used for packaging the target code into a code analysis toolkit;

[0138] The calling module is used to call the code analysis toolkit when receiving an analysis instruction for part or all of the code in the file.

[0139] The structures of the control devices for implementing the method of constructing the above code analysis diagram are introduced in the following in the server form and the terminal device form respectively.

[0140] Figure 6 : This is a schematic diagram of a server structure provided in an embodiment of the present application. The server 900 may have relatively large differences due to different configurations or performances, and may include one or more central processing units (CPU) 922 (for example, one or more processors) and a memory 932, and one or more storage media 930 (for example, one or more mass storage devices) storing application programs 942 or data 944. Among them, the memory 932 and the storage medium 930 can be short-term storage or permanent storage. The program stored in the storage medium 930 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Furthermore, the central processing unit 922 can be configured to communicate with the storage medium 930 and execute a series of instruction operations in the storage medium 930 on the server 900.

[0141] The server 900 may also include one or more power supplies 926, one or more wired or wireless network interfaces 950, one or more input and output interfaces 958, and / or one or more operating systems 941, such as Windows Server 2000.TM , Mac OS X TM , Unix TM ,Linux TM , FreeBSD TM etc.

[0142] The steps performed by the server in the above embodiment can be based on the Figure 6 The server structure shown.

[0143] The CPU 922 is used to execute the following steps:

[0144] Identify the types of files in the source code directory;

[0145] Based on the type of the file, converting the file into an intermediate code matching the type of the file;

[0146] Nodes for constructing a code analysis graph are determined based on the code constituent elements of the intermediate code, edges between the nodes are determined based on element relationships between the code constituent elements, and the code analysis graph is constructed based on the nodes and the edges between the nodes.

[0147] The present application embodiment also provides another device for constructing a code analysis graph, and the acquisition device may be a terminal device. Figure 7 For the sake of convenience, only the parts related to the embodiments of the present application are shown. For specific technical details not disclosed, please refer to the method part of the embodiments of the present application. The terminal can be any terminal device including a mobile phone, a tablet computer, a personal digital assistant (English full name: Personal Digital Assistant, English abbreviation: PDA), a sales terminal (English full name: Point of Sales, English abbreviation: POS), a car computer, etc., taking the terminal as a mobile phone as an example:

[0148] Figure 7 FIG. 1 is a block diagram showing a partial structure of a mobile phone related to a terminal provided in an embodiment of the present application. Figure 7 The mobile phone includes: a radio frequency (RF) circuit 1010, a memory 1020, an input unit 1030, a display unit 1040, a sensor 1050, an audio circuit 1060, a wireless fidelity (WiFi) module 1070, a processor 1080, and a power supply 1090. Those skilled in the art can understand that Figure 7 The mobile phone structure shown in the figure does not constitute a limitation on the mobile phone, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.

[0149] Combine the following Figure 7 A detailed introduction to the various components of the mobile phone:

[0150] The RF circuit 1010 can be used for receiving and sending signals during information transmission or calls. In particular, after receiving the downlink information of the base station, it is sent to the processor 1080 for processing; in addition, the designed uplink data is sent to the base station. Generally, the RF circuit 1010 includes but is not limited to an antenna, at least one amplifier, a transceiver, a coupler, a low noise amplifier (full name: LowNoiseAmplifier, English abbreviation: LNA), a duplexer, etc. In addition, the RF circuit 1010 can also communicate with the network and other devices through wireless communication. The above-mentioned wireless communications may use any communication standard or protocol, including but not limited to Global System of Mobile communications (Global System of Mobile communication, English abbreviation: GSM), General Packet Radio Service (General Packet Radio Service, GPRS), Code Division Multiple Access (Code Division Multiple Access, English abbreviation: CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (Long Term Evolution, English abbreviation: LTE), e-mail, Short Messaging Service (SMS), etc.

[0151] The memory 1020 can be used to store software programs and modules. The processor 1080 executes various functional applications and data processing of the mobile phone by running the software programs and modules stored in the memory 1020. The memory 1020 can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area can store data created according to the use of the mobile phone (such as audio data, a phone book, etc.), etc. In addition, the memory 1020 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage devices.

[0152] The input unit 1030 can be used to receive input digital or character information, and to generate key signal input related to the user settings and function control of the mobile phone. Specifically, the input unit 1030 may include a touch panel 1031 and other input devices 1032. The touch panel 1031, also known as a touch screen, can collect the user's touch operation on or near it (such as the user's operation on the touch panel 1031 or near the touch panel 1031 using any suitable object or accessory such as a finger, stylus, etc.), and drive the corresponding connection device according to a pre-set program. Optionally, the touch panel 1031 may include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the user's touch orientation, detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into contact coordinates, and then sends it to the processor 1080, and can receive and execute commands sent by the processor 1080. In addition, the touch panel 1031 can be implemented in various types such as resistive, capacitive, infrared, and surface acoustic waves. In addition to the touch panel 1031, the input unit 1030 may further include other input devices 1032. Specifically, the other input devices 1032 may include but are not limited to one or more of a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, a joystick, and the like.

[0153] The display unit 1040 can be used to display information input by the user or information provided to the user and various menus of the mobile phone. The display unit 1040 may include a display panel 1041. Optionally, the display panel 1041 may be configured in the form of a liquid crystal display (full name in English: Liquid Crystal Display, English abbreviation: LCD), an organic light-emitting diode (full name in English: Organic Light-Emitting Diode, English abbreviation: OLED), etc. Further, the touch panel 1031 may cover the display panel 1041. When the touch panel 1031 detects a touch operation on or near it, it is transmitted to the processor 1080 to determine the type of touch event. Subsequently, the processor 1080 provides a corresponding visual output on the display panel 1041 according to the type of touch event. Although in Figure 7 In the embodiment, the touch panel 1031 and the display panel 1041 are used as two independent components to realize the input and output functions of the mobile phone, but in some embodiments, the touch panel 1031 and the display panel 1041 can be integrated to realize the input and output functions of the mobile phone.

[0154] The mobile phone may also include at least one sensor 1050, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor, wherein the ambient light sensor may adjust the brightness of the display panel 1041 according to the brightness of the ambient light, and the proximity sensor may turn off the display panel 1041 and / or the backlight when the mobile phone is moved to the ear. As a type of motion sensor, the accelerometer sensor can detect the magnitude of acceleration in all directions (generally three axes), and can detect the magnitude and direction of gravity when stationary. It can be used for applications that identify the posture of the mobile phone (such as horizontal and vertical screen switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc.; as for other sensors that can be configured in the mobile phone, such as gyroscopes, barometers, hygrometers, thermometers, infrared sensors, etc., they will not be repeated here.

[0155] The audio circuit 1060, the speaker 1061, and the microphone 1062 can provide an audio interface between the user and the mobile phone. The audio circuit 1060 can transmit the received audio data to the speaker 1061 after converting the received audio data into an electrical signal, which is converted into a sound signal for output; on the other hand, the microphone 1062 converts the collected sound signal into an electrical signal, which is received by the audio circuit 1060 and converted into audio data, and then the audio data is output to the processor 1080 for processing, and then sent to another mobile phone through the RF circuit 1010, or the audio data is output to the memory 1020 for further processing.

[0156] WiFi is a short-range wireless transmission technology. The mobile phone can help users send and receive emails, browse web pages and access streaming media through the WiFi module 1070. It provides users with wireless broadband Internet access. Figure 7 A WiFi module 1070 is shown, but it is understandable that it is not an essential component of the mobile phone and can be omitted as needed without changing the essence of the invention.

[0157] The processor 1080 is the control center of the mobile phone. It uses various interfaces and lines to connect various parts of the entire mobile phone. By running or executing software programs and / or modules stored in the memory 1020, and calling data stored in the memory 1020, it executes various functions of the mobile phone and processes data, thereby collecting overall data and information of the mobile phone. Optionally, the processor 1080 may include one or more processing units; preferably, the processor 1080 may integrate an application processor and a modem processor, wherein the application processor mainly processes the operating system, user interface, and application programs, and the modem processor mainly processes wireless communications. It is understandable that the above-mentioned modem processor may not be integrated into the processor 1080.

[0158] The mobile phone also includes a power supply 1090 (such as a battery) for supplying power to various components. Preferably, the power supply can be logically connected to the processor 1080 through a power management system, so that the power management system can manage functions such as charging, discharging, and power consumption.

[0159] Although not shown, the mobile phone may also include a camera, a Bluetooth module, etc., which will not be described in detail here.

[0160] In the embodiment of the present application, the processor 1080 included in the terminal also has the following functions:

[0161] Identify the types of files in the source code directory;

[0162] Based on the type of the file, converting the file into an intermediate code matching the type of the file;

[0163] Nodes for constructing a code analysis graph are determined based on the code constituent elements of the intermediate code, edges between the nodes are determined based on element relationships between the code constituent elements, and the code analysis graph is constructed based on the nodes and the edges between the nodes.

[0164] An embodiment of the present application also provides a computer-readable storage medium for storing a computer program. When the computer program is executed on a computer device, the computer device is used to execute any one of the methods for constructing a code analysis graph described in the aforementioned embodiments.

[0165] An embodiment of the present application also provides a computer program product including a computer program, which, when executed on a computer device, enables the computer device to execute any one of the implementations of the code analysis graph construction method described in the aforementioned embodiments.

[0166] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0167] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0168] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0169] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0170] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (full name in English: Read-Only Memory, English abbreviation: ROM), random access memory (full name in English: Random Access Memory, English abbreviation: RAM), disk or optical disk and other media that can store program codes.

[0171] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for constructing a code analysis graph, characterized in that: include: Identify the types of files in the source code directory; Based on the type of the file, converting the file into an intermediate code matching the type of the file; Nodes for constructing a code analysis graph are determined based on the code constituent elements of the intermediate code, edges between the nodes are determined based on element relationships between the code constituent elements, and the code analysis graph is constructed based on the nodes and the edges between the nodes.

2. The method for constructing a code analysis graph according to claim 1, characterized in that: The files under the source code directory include source code files under the source code directory; and based on the type of the file, converting the file into an intermediate code matching the type of the file includes: Based on the type of the source code file, determining a parsing program corresponding to the source code file; the type of the source code file is determined based on suffix information of the source code file; The source code file is converted by the parsing program to obtain an abstract syntax tree corresponding to the source code file as the intermediate code of the source code file.

3. The method for constructing a code analysis graph according to claim 1, characterized in that: The files under the source code directory include binary files under the source code directory; The step of converting the file into an intermediate code matching the type of the file based on the type of the file comprises: Based on the type of the binary file, determining a reverse analysis program corresponding to the binary file; the type of the binary file is determined based on file header information of the binary file; The binary file is converted by the reverse analysis program to obtain the three-address code corresponding to the binary file as the intermediate code of the binary file.

4. The method for constructing a code analysis graph according to claim 1, characterized in that: The determining of nodes for constructing a code analysis graph based on the code constituent elements of the intermediate code, determining edges between the nodes based on element relationships between the code constituent elements, and constructing the code analysis graph based on the nodes and the edges between the nodes, comprises: Extracting a plurality of code component elements and related information of the plurality of code component elements from the intermediate code; Based on the multiple code component elements and the related information, determine a control flow relationship and a data flow relationship between the multiple code component elements; the control flow relationship and the data flow relationship are both subordinate to the element relationship; Creating a plurality of nodes one-to-one based on the plurality of code constituent elements; and creating edges between the plurality of nodes based on the control flow relationship and the data flow relationship; The code analysis graph is constructed based on the intermediate code, the multiple nodes, and the edges between the multiple nodes.

5. The method for constructing a code analysis graph according to claim 4, characterized in that: The step of creating edges between the plurality of nodes based on the control flow relationship and the data flow relationship comprises: Creating a first type of edge between the plurality of nodes based on the control flow relationship to obtain a first graph structure; and creating a second type of edge between the plurality of nodes based on the data flow relationship to obtain a second graph structure; The step of constructing the code analysis graph based on the intermediate code, the plurality of nodes, and the edges between the nodes includes: The intermediate code, the first graph structure and the second graph structure are merged to obtain the code analysis graph.

6. The method for constructing a code analysis graph according to claim 5, characterized in that: The method further comprises: Traversing the intermediate code, the first graph structure, the second graph structure and the code analysis graph, determining a first element in the code component elements and a first relationship between the first elements; the first element is an element of an object-oriented language; the first relationship includes an inheritance relationship, an attribution relationship or an implementation relationship; Constructing a third graph structure using the first elements as first nodes and the first relationships as edges between the first nodes; The third graph structure and the code analysis graph are merged to obtain a target graph; the target graph is used to analyze the object-oriented language.

7. The method for constructing a code analysis graph according to claim 6, characterized in that: The step of fusing the third graph structure and the code analysis graph to obtain a target graph includes: Traversing the code analysis graph based on the third graph structure, determining type information of a second node among the plurality of nodes; the second node is a node used to represent a function element in the code component element; The third graph structure and the code analysis graph are merged based on the type information of the second node to obtain the target graph.

8. The method for constructing a code analysis graph according to any one of claims 1 to 7, characterized in that: The method further comprises: Storing the code analysis graph in a graph database; receiving an operation command for the graph database; the operation command is used to instruct to analyze the file; Traversing the code analysis graph based on the operation command; The file is analyzed based on the graph data in the code analysis graph obtained through traversal.

9. The method for constructing a code analysis graph according to any one of claims 1 to 7, characterized in that: The method further comprises: Extracting target code based on the code analysis graph; the target code is used to analyze part or all of the code in the file; Encapsulating the target code into a code analysis toolkit; When an analysis instruction for part or all of the code in the file is received, the code analysis toolkit is called.

10. A device for constructing a code analysis graph, characterized in that: include: Type identification module, used to identify the types of files in the source code directory; A file conversion module, used for converting the file into an intermediate code matching the type of the file based on the type of the file; A graph construction module determines nodes for constructing a code analysis graph based on code constituent elements of the intermediate code, determines edges between the nodes based on element relationships between the code constituent elements, and constructs the code analysis graph based on the nodes and the edges between the nodes.

11. A device for constructing a code analysis graph, characterized in that: The device comprises a processor and a memory: The memory is used to store a computer program and transmit the computer program to the processor; The processor is used to execute the method for constructing a code analysis graph according to any one of claims 1 to 9 according to the instructions in the computer program.

12. A computer-readable storage medium, characterized in that: The computer-readable storage medium is used to store a computer program, and when the computer program is executed by a device for constructing a code analysis graph, the steps of the method for constructing a code analysis graph according to any one of claims 1 to 9 are implemented.

13. A computer program product, characterized in that It comprises a computer program, which, when executed by a device for constructing a code analysis graph, implements the steps of the method for constructing a code analysis graph as described in any one of claims 1 to 9.