Data flow diagram generation method and device, equipment, storage medium and program product
By constructing and replacing the control flow graph and generating the data flow graph in combination with the abstract syntax tree, the problem of low accuracy in data flow graph generation in the prior art is solved, and higher accuracy is achieved.
Patent Information
- Application Number
- CN202311454223.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-02
- Publication Date
- 2025-05-06
AI Technical Summary
In the prior art, data flow analysis of the code is performed to generate data flow diagrams, which has low accuracy and is greatly affected by the compilation environment and compiled scripts.
By obtaining the abstract syntax tree of the object code, an initial control flow chart is constructed, and the extended control flow chart of the target subcode segment is obtained, the basic blocks in the initial control flow chart are replaced to obtain the reference control flow chart, and the target data flow chart is generated by combining the abstract syntax tree and the reference control flow chart.
The accuracy of the generation of data flow graphs is effectively improved, so that the resulting target data flow graph can accurately reflect the data flow in the target code.
Smart Images

Figure CN119938048A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a method, device, equipment, storage medium and program product for generating a data flow graph. Background Art
[0002] Data Flow Diagram (DFD) graphically expresses the logical functions of the system, the logical flow of data within the system, and the logical transformation process from the perspective of data transmission and processing. It is the main expression tool of structured system analysis methods and a graphical method for representing software models.
[0003] In the related art, for the generation of a data flow graph, data flow analysis is usually performed directly on the code to obtain the data flow graph. Since data flow analysis of the code often relies on the accuracy of the analysis tool and is highly affected by the compilation environment and compilation scripts, the accuracy of the data flow graph generation is low. Summary of the invention
[0004] The embodiments of the present application provide a method, device, electronic device, computer-readable storage medium and computer program product for generating a data flow graph, which can effectively improve the generation accuracy of the data flow graph.
[0005] The technical solution of the embodiment of the present application is implemented as follows:
[0006] The present application provides a method for generating a data flow graph, including:
[0007] Obtaining an abstract syntax tree of the target code, the abstract syntax tree comprising a plurality of nodes, each of the nodes corresponding to a different set of sub-code segments in the target code, the sub-code segment set comprising a target sub-code segment having a target logic;
[0008] For each of the nodes, based on the sub-code segment set corresponding to the node, construct an initial control flow graph of the node, wherein the initial control flow graph includes initial basic blocks corresponding to each sub-code segment in the sub-code segment set one by one;
[0009] Obtaining an extended control flow graph of each of the target sub-code segments, and replacing an initial basic block corresponding to each of the target sub-code segments in each of the initial control flow graphs with a corresponding extended control flow graph, to obtain a reference control flow graph corresponding to each of the initial control flow graphs;
[0010] The target data flow graph of the target code is generated by combining the abstract syntax tree and the reference control flow graph.
[0011] The present application provides a data flow graph generation device, including:
[0012] An acquisition module, used for acquiring an abstract syntax tree of a target code, wherein the abstract syntax tree includes a plurality of nodes, each of the nodes corresponds to a different set of sub-code segments in the target code, and the sub-code segment set includes a target sub-code segment having a target logic;
[0013] A construction module, configured to construct, for each of the nodes, an initial control flow graph of the node based on a set of sub-code segments corresponding to the node, wherein the initial control flow graph includes initial basic blocks corresponding one-to-one to each sub-code segment in the set of sub-code segments;
[0014] A replacement module is used to obtain an extended control flow graph of each target sub-code segment, and replace the initial basic block corresponding to each target sub-code segment in each initial control flow graph with the corresponding extended control flow graph to obtain a reference control flow graph corresponding to each initial control flow graph;
[0015] A generation module is used to combine the abstract syntax tree and the reference control flow graph to generate a target data flow graph of the target code.
[0016] In the above scheme, the sub-code segment set includes multiple sub-code segments, and the sub-code segment includes multiple code statements. The above construction module is also used to select the first sub-code segment from the multiple sub-code segments, create the first initial basic block corresponding to the first sub-code segment, and determine the first initial basic block as the first initial control flow graph; wherein the starting code statement of the first sub-code segment is the same as the starting code statement of the sub-code segment set; traverse i to perform the following processing: select the i-th sub-code segment from the multiple sub-code segments, and create the i-th initial basic block corresponding to the i-th sub-code segment; connect the i-th initial basic block with the i-1th initial control flow graph to obtain the i-th initial control flow graph; wherein the starting code statement of the i-th sub-code segment is logically associated with the ending code statement of the i-1th sub-code segment, 2≤i≤N, N is used to indicate the number of the sub-code segments in the sub-code segment set; determine the N-th initial control flow graph as the initial control flow graph of the node.
[0017] In the above scheme, the above construction module is also used to obtain the starting code statement of the sub-code segment set, and perform the following processing for each sub-code segment: obtain the starting code statement of the sub-code segment, and compare the starting code statement of the sub-code segment with the starting code statement of the sub-code segment set to obtain a comparison result corresponding to the sub-code segment; when the comparison result indicates that the starting code statement of the sub-code segment is the same as the starting code statement of the sub-code segment set, determine the sub-code segment as the first sub-code segment.
[0018] In the above scheme, the above construction module is also used to obtain the end code statement of the i-1th sub-code segment, and perform the following processing for each sub-code segment: obtain the start code statement of the sub-code segment, and determine the association information between the start code statement of the sub-code segment and the end code statement of the i-1th sub-code segment; when the association information indicates that there is a logical association between the start code statement of the sub-code segment and the end code statement of the i-1th sub-code segment, determine the sub-code segment as the i-th sub-code segment.
[0019] In the above scheme, the device for generating the above data flow graph also includes: a code segment module, which is used to obtain the logical identification character of the target logic, and perform the following processing for each of the sub-code segment sets: extract code characters for each of the sub-code segments in the sub-code segment set to obtain a code character set corresponding to each of the sub-code segments; for each of the code character sets, compare each code character in the code character set with the logical identification character to obtain a comparison result corresponding to the code character set; when the comparison result indicates that the logical identification character exists in the code character set, determine the sub-code segment corresponding to the code character set as the target sub-code segment.
[0020] In the above scheme, the above replacement module is also used to obtain the mapping relationship between multiple preset logical identifiers and corresponding preset control flow graphs, and perform the following processing for each target sub-code segment: obtain the logical identifier of the target logic in the target sub-code segment, and compare the logical identifier with each of the preset logical identifiers in the mapping relationship to obtain the comparison result corresponding to the logical identifier; when the comparison result indicates that the logical identifier exists in the preset logical identifier, the preset control flow graph corresponding to the logical identifier in the mapping relationship is determined as the extended control flow graph of the target sub-code segment.
[0021] In the above scheme, the above generation module is also used to replace each of the nodes in the abstract syntax tree with the corresponding reference control flow graph to obtain the target control flow graph corresponding to the target code; perform data flow graph conversion on the target control flow graph to obtain the initial data flow graph of the target code; perform redundancy identification on the data flow in the initial data flow graph to obtain the redundant data flow in the initial data flow graph, and delete the redundant data flow in the initial data flow graph to obtain the target data flow graph.
[0022] In the above solution, the above-mentioned generation module is further configured to perform the first data flow graph conversion on the target control flow graph to obtain the first data flow graph of the target code; traverse j and perform the following processing until the initial data flow graph of the target code is obtained: perform the j-th data flow graph conversion on the target control flow graph to obtain the j-th data flow graph of the target code; when the j-th data flow graph is the same as the (j-1)-th data flow graph, determine the j-th data flow graph as the initial data flow graph of the target code; where 1 < j ≤ M, and the M-th data flow graph is the same as the (M-1)-th data flow graph.
[0023] In the above solution, the target control flow graph includes multiple target basic blocks, and the target basic blocks correspond one-to-one with the target code segments in the target code. The above-mentioned generation module is further configured to run the target code and record the running data flow of each target code segment during the running process of the target code; for each target basic block, cache the running data flow of the target code segment corresponding to the target basic block into the target basic block to obtain the first updated basic block corresponding to the target basic block; replace each target basic block in the target control flow graph with the corresponding first updated basic block to obtain the first data flow graph of the target code.
[0024] In the above solution, the initial data flow graph includes multiple data flows, and the data flow is used to indicate the data migration process from the start point to the end point of the data flow; the above-mentioned generation module is further configured to perform the following processing on each data flow in the initial data flow graph respectively: obtain the start level of the start point of the data flow in the target code and the end level of the end point of the data flow in the target code; compare the start level and the end level to obtain a level comparison result; when the level comparison result indicates that the start level is greater than the end level, determine the data flow as the redundant data flow.
[0025] In the above solution, the acquisition module is further configured to acquire the target code with the target logic and perform lexical analysis on the target code to obtain multiple sub-code segment sets of the target code, and different sub-code segment sets correspond to different logical functions; perform syntax analysis on each sub-code segment set to obtain a syntax analysis result indicating the logical association between the sub-code segment sets; combine the syntax analysis result and the multiple sub-code segment sets to construct an abstract syntax tree of the target code.
[0026] An embodiment of the present application provides an electronic device, including:
[0027] A memory for storing computer-executable instructions or computer programs;
[0028] The processor is used to implement the method for generating a data flow graph provided in an embodiment of the present application when executing the computer executable instructions or computer program stored in the memory.
[0029] An embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions for causing a processor to execute and implement the method for generating a data flow graph provided in the embodiment of the present application.
[0030] The embodiment of the present application provides a computer program product, which includes a computer program or a computer executable instruction, and the computer program or the computer executable instruction is stored in a computer-readable storage medium. The processor of the electronic device reads the computer executable instruction from the computer-readable storage medium, and the processor executes the computer executable instruction, so that the electronic device executes the method for generating the data flow graph described in the embodiment of the present application.
[0031] The embodiments of the present application have the following beneficial effects:
[0032] By obtaining the abstract syntax tree of the target code, the initial control flow graph of the node is constructed based on the set of sub-code segments corresponding to each node, and the extended control flow graph of each target sub-code segment is obtained, and the initial basic block corresponding to each target sub-code segment in the initial control flow graph is replaced with the corresponding extended control flow graph, and the reference control flow graph corresponding to each initial control flow graph is obtained. The target flow graph of the target code is generated by combining the abstract syntax tree and the reference control flow graph. By replacing the initial basic block corresponding to the target sub-code segment of the initial control flow graph with the corresponding extended control flow graph, the control flow graph corresponding to each target sub-code segment in the obtained reference control flow graph is more accurate, and then by combining the abstract syntax tree and the reference control flow graph, the accuracy of the target data flow graph of the generated target code can be effectively improved, so that the obtained target data flow graph can accurately reflect the data flow in the target code, and the generation accuracy of the data flow graph is effectively improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 It is a schematic diagram of the architecture of a data flow graph generation system provided in an embodiment of the present application;
[0034] Figure 2 It is a structural schematic diagram of an electronic device for generating a data flow graph provided in an embodiment of the present application;
[0035] Figures 3 to 7 It is a flowchart of a method for generating a data flow graph provided in an embodiment of the present application;
[0036] Figure 8 It is a schematic diagram of the structure of a preset control flow graph T1 provided in an embodiment of the present application;
[0037] Fig. 9 It is a schematic diagram of the structure of a preset control flow graph T2 provided in an embodiment of the present application;
[0038] Fig.10 It is a schematic diagram of the structure of a preset control flow graph T3 provided in an embodiment of the present application;
[0039] Fig.11 It is a schematic diagram of the structure of a preset control flow graph T4 provided in an embodiment of the present application;
[0040] Fig.12 It is a schematic diagram of the structure of a preset control flow graph T5 provided in an embodiment of the present application;
[0041] Fig.13 It is a schematic diagram of the structure of a preset control flow graph T6 provided in an embodiment of the present application;
[0042] Fig.14 It is a schematic diagram of the structure of a preset control flow graph T7 provided in an embodiment of the present application;
[0043] Fig.15 It is a schematic diagram of the structure of a preset control flow graph T8 provided in an embodiment of the present application;
[0044] Fig.16 It is a schematic diagram of the structure of a preset control flow graph T9 provided in an embodiment of the present application;
[0045] Fig.17 is a schematic diagram of the structure of a preset control flow graph T10 provided in an embodiment of the present application;
[0046] Fig.18 It is a flowchart of a method for generating a data flow graph provided in an embodiment of the present application;
[0047] Fig.19 It is a schematic diagram of the principle of the method for generating a data flow graph provided in an embodiment of the present application;
[0048] Fig. 20 It is a schematic diagram of the principle of the method for generating a data flow graph provided in an embodiment of the present application. DETAILED DESCRIPTION
[0049] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings. The described embodiments should not be regarded as limiting the present application. All other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of this application.
[0050] In the following description, reference is made to “some embodiments”, which describe a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0051] In the following description, the terms "first\second\third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described here can be implemented in an order other than that illustrated or described here.
[0052] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0053] Before further describing the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations.
[0054] 1) Abstract Syntax Tree: It is an abstract representation of the syntax structure of the source code. It represents the syntax structure of the programming language in a tree form, and each node on the tree represents a structure in the source code. It is also called an abstract syntax tree. The reason why the syntax is "abstract" is that the syntax here does not represent every detail that appears in the real syntax. For example, nested parentheses are implied in the structure of the tree and are not presented in the form of nodes; and conditional jump statements such as if-condition-then can be represented by nodes with two branches. As an important intermediate representation of the program, the basic formation process of the abstract syntax tree can be roughly divided into three steps, namely, firstly, the source program code is lexically analyzed, the word stream generated by the lexical analysis is grammatically analyzed, and finally the semantic analysis is performed to form the abstract syntax tree of the source code. As a good intermediate representation, the syntax tree contains complete source program information. Using the abstract syntax tree, a variety of source program processing tools can be implemented, such as intelligent editors, source program browsers, etc. In addition, the parser of the syntax tree can also be used as the front end of the program static analysis tool to provide it with an input that is easy to analyze. The construction of the abstract syntax tree first requires the creation of nodes corresponding to the source code, each node representing a category of code. The types of nodes are divided into compound nodes and base nodes. The non-terminal symbols in the code correspond to compound nodes, while the terminal symbols correspond to base nodes. Secondly, according to the grammatical structure of the language to be parsed, a suitable tree structure should be designed to represent different types of program statements. The quality of the design directly affects the efficiency of the subsequent program analysis.
[0055] 2) Control Flow Graph (CFG): An abstract representation of the program execution process. The control flow graph simulates the program execution by showing the flow of data between basic blocks during program execution. The control flow graph is an abstract representation of a process or program. It is an abstract data structure used in the compiler and maintained internally by the compiler. It represents all the paths that will be traversed during the execution of a program. It uses a graph to represent the possible flow of execution of all basic blocks in a process, and can also reflect the real-time execution process of a process. The control flow graph is the basic basis for static code analysis. The biggest feature of the basic blocks in the control flow graph is that the first line of code in the basic block contains all the entries, and the last line of code contains all the exits. This requires that the basis for generating the control flow graph needs to contain enough information.
[0056] 3) Data Flow Diagram (DFD): From the perspective of data transmission and processing, it graphically expresses the logical functions of the system, the logical flow of data within the system, and the logical transformation process. It is the main expression tool of structured system analysis methods and a graphical method for representing software models.
[0057] 4) Basic block: refers to a program - a sequence of statements executed sequentially. A basic block has only one entry and one exit. The entry is the first statement in it, and the exit is the last statement in it. For a basic block, it is only entered from its entry and exited from its exit. Specifically: a basic block has only one entry, which means that there is no other place in the program that can enter this basic block through jump instructions. A basic block has only one exit, which means that only the last instruction of the program can lead to entering other basic blocks for execution. A typical feature of a basic block is that as long as the first instruction in the basic block is executed, all executions in the basic block will be executed only once in sequence.
[0058] 5) Code: It is a source file written by programmers in a language supported by development tools. It is a set of clear rules that represent information in discrete form by characters, symbols or signal code elements. The principles of code design include unique certainty, standardization and universality, extensibility and stability, easy recognition and memory, short and uniform format, and easy modification. Source code is a branch of code. In a sense, source code is equivalent to code. In modern programming languages, source code can appear in the form of books or tapes, but the most commonly used format is text files. The purpose of this typical format is to compile computer programs. The ultimate goal of computer source code is to translate human-readable text into binary instructions that can be executed by computers. This process is called compilation, which is completed by a compiler.
[0059] During the implementation of the embodiments of the present application, the applicant discovered that the related technology has the following problems:
[0060] In the related art, for the generation of a data flow graph, data flow analysis is usually performed directly on the code to obtain the data flow graph. Since data flow analysis of the code often relies on the accuracy of the analysis tool and is highly affected by the compilation environment and compilation scripts, the accuracy of the data flow graph generation is low.
[0061] The embodiments of the present application provide a method, device, electronic device, computer-readable storage medium and computer program product for generating a data flow graph, which can effectively improve the generation accuracy of the data flow graph. The following describes an exemplary application of the data flow graph generation system provided in the embodiments of the present application.
[0062] See also Figure 1 , Figure 1 It is an architectural diagram of a data flow graph generation system 100 provided in an embodiment of the present application. A terminal (terminal 400 is shown as an example) is connected to a server 200 via a network 300. The network 300 may be a wide area network or a local area network, or a combination of the two.
[0063] The terminal 400 is used for the user to use the client 410, and the target data flow diagram is displayed on the graphical interface 410-1 (the graphical interface 410-1 is shown as an example). The terminal 400 and the server 200 are connected to each other via a wired or wireless network.
[0064] In some embodiments, server 200 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content distribution networks (Content Delivery Network, CDN), and big data and artificial intelligence platforms. Terminal 400 can be a smart phone, tablet computer, laptop computer, desktop computer, smart speaker, smart TV, smart watch, car terminal, etc., but is not limited to this. The electronic device provided in the embodiment of the present application can be implemented as a terminal or as a server. The terminal and the server can be directly or indirectly connected by wired or wireless communication, which is not limited in the embodiment of the present application.
[0065] In some embodiments, the server 200 obtains an abstract syntax tree of the target code, and for each node, builds an initial control flow graph of the node based on the set of sub-code segments corresponding to the node, replaces the initial basic blocks corresponding to each target sub-code segment in each initial control flow graph with the corresponding extended control flow graph, and obtains a reference control flow graph corresponding to each initial control flow graph; combines the abstract syntax tree and the reference control flow graph to generate a target data flow graph of the target code, and sends the target data flow graph to the terminal 400.
[0066] In other embodiments, the terminal 400 obtains an abstract syntax tree of the target code, and for each node, constructs an initial control flow graph of the node based on a set of sub-code segments corresponding to the node, replaces the initial basic blocks corresponding to each target sub-code segment in each initial control flow graph with a corresponding extended control flow graph, and obtains a reference control flow graph corresponding to each initial control flow graph; combines the abstract syntax tree and the reference control flow graph to generate a target data flow graph of the target code, and sends the target data flow graph to the server 200.
[0067] In other embodiments, the embodiments of the present application can be implemented with the aid of cloud technology (Cloud Technology). Cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, and network within a wide area network or a local area network to achieve data computing, storage, processing, and sharing.
[0068] Cloud technology is a general term for network technology, information technology, integration technology, management platform technology, and application technology based on the cloud computing business model. It can form a resource pool that can be used on demand and is flexible and convenient. Cloud computing technology will become an important support. The background services of the technical network system require a large amount of computing and storage resources.
[0069] See also Figure 2 , Figure 2 is a schematic diagram of the structure of an electronic device 500 for generating a data flow graph provided in an embodiment of the present application, wherein: Figure 2 The electronic device 500 shown may be Figure 1 The server 200 or the terminal 400 in Figure 2 The electronic device 500 shown includes: at least one processor 430, a memory 450, and at least one network interface 420. The various components in the electronic device 500 are coupled together via a bus system 440. It is understood that the bus system 440 is used to achieve connection and communication between these components. In addition to the data bus, the bus system 440 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, the bus system 440 is not described in detail. Figure 2 Various buses are labeled as bus system 440 .
[0070] The processor 430 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0071] The memory 450 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard drives, optical drives, etc. The memory 450 may optionally include one or more storage devices that are physically remote from the processor 430.
[0072] The memory 450 includes a volatile memory or a non-volatile memory, and may also include both volatile and non-volatile memories. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 450 described in the embodiments of the present application is intended to include any suitable type of memory.
[0073] In some embodiments, memory 450 can store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as exemplarily described below.
[0074] Operating system 451, including system programs for processing various basic system services and performing hardware-related tasks, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks;
[0075] The network communication module 452 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 420. Exemplary network interfaces 420 include: Bluetooth, Wireless Fidelity (WiFi), and Universal Serial Bus (USB).
[0076] In some embodiments, the data flow graph generation device provided in the embodiments of the present application can be implemented in software. Figure 2 The data flow graph generation device 455 stored in the memory 450 is shown, which can be software in the form of a program and a plug-in, etc., including the following software modules: an acquisition module 4551, a construction module 4552, a replacement module 4553, and a generation module 4554. These modules are logical, so they can be arbitrarily combined or further split according to the functions implemented. The functions of each module will be explained below.
[0077] In other embodiments, the data flow graph generation device provided in the embodiments of the present application can be implemented in hardware. As an example, the data flow graph generation device provided in the embodiments of the present application can be a processor in the form of a hardware decoding processor, which is programmed to execute the data flow graph generation method provided in the embodiments of the present application. For example, the processor in the form of a hardware decoding processor can adopt one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs) or other electronic components.
[0078] In some embodiments, the terminal or server can implement the method for generating a data flow graph provided in the embodiment of the present application by running a computer program or a computer executable instruction. For example, the computer program can be a native program (e.g., a dedicated data flow graph generation program) or a software module in an operating system, for example, a data flow graph generation module that can be embedded in any program (such as an instant messaging client, an album program, an electronic map client, a navigation client); for example, it can be a local (Native) application (APP, Application), that is, a program that needs to be installed in the operating system to run. In short, the above-mentioned computer program can be an application, module or plug-in in any form.
[0079] The method for generating a data flow graph provided in the embodiment of the present application will be explained in combination with the exemplary application and implementation of the server or terminal provided in the embodiment of the present application.
[0080] See also Figure 3 , Figure 3 is a flow chart of a method for generating a data flow graph provided in an embodiment of the present application, which will be combined with Figure 3 Steps 101 to 105 are shown for illustration. The method for generating a data flow graph provided in the embodiment of the present application can be implemented by a server or a terminal alone, or by a server and a terminal in collaboration. The following will be described using the server alone as an example.
[0081] In step 101, an abstract syntax tree of a target code is obtained.
[0082] In some embodiments, the abstract syntax tree includes multiple nodes, each of which corresponds to a different sub-code segment set in the target code, and the sub-code segment set includes a target sub-code segment with target logic. The abstract syntax tree is used to describe the logical association between the sub-code segment sets corresponding to each node.
[0083] In some embodiments, the above step 101 can be implemented as follows: obtain a target code with target logic, and perform lexical analysis on the target code to obtain multiple sub-code segment sets of the target code, where different sub-code segment sets correspond to different logical functions; perform syntactic analysis on each sub-code segment set to obtain a syntactic analysis result for indicating the logical relationship between each sub-code segment set; and construct an abstract syntax tree of the target code by combining the syntactic analysis result and multiple sub-code segment sets.
[0084] In some embodiments, the above lexical analysis, as the first stage of the program compilation process, inputs the program source code as a character stream in sequence, scans and analyzes the read character sequence, and finally converts it into a word sequence output, where "word" is the smallest unit that makes up the program. This process does not consider the relationship between words. Lexical analysis provides symbol nodes for the generation of an abstract syntax tree.
[0085] In some embodiments, the above-mentioned syntax analysis is mainly to perform logical analysis on the program, which is used to analyze the basic syntax structure of the target code, such as methods, variables, expressions and other components that make up the program, based on a set of multiple sub-code segments and according to the syntax rules of the analyzed language.
[0086] In some embodiments, the above-mentioned combination of syntax analysis results and multiple sub-code segment sets to construct an abstract syntax tree of the target code can be achieved in the following way: based on the syntax analysis results, the logical associations between different sub-code segment sets are determined, and the abstract syntax tree of the target code is constructed with the logical associations between different sub-code segment sets as edges and each sub-code segment set as a node.
[0087] As an example, see Figure 4 , Figure 4 is a flow chart of a method for generating a data flow graph provided in an embodiment of the present application, Figure 4 The abstract syntax tree of the target code shown includes node 1, node 2, node 3, node 4, node 5 and node 6, wherein node 6 is the syntax tree corresponding to the first line of code content of the target code, node 5 is the syntax tree corresponding to the second line of code content of the target code, and node 3 is the syntax tree corresponding to the third line of code content of the target code.
[0088] In this way, by obtaining the abstract syntax tree of the target code, the abstract representation of the target code can be accurately realized. Since the abstract syntax tree contains the complete source program information of the target code, it lays an accurate data support for the subsequent accurate determination of the target data flow graph of the target code.
[0089] In step 102, for each node, an initial control flow graph of the node is constructed based on a set of sub-code segments corresponding to the node.
[0090] In some embodiments, the above-mentioned sub-code segment set includes multiple sub-code segments, the sub-code segment includes multiple code statements, and the initial control flow graph corresponding to the sub-code segment set includes initial basic blocks that correspond one-to-one to each sub-code segment in the sub-code segment set.
[0091] In some embodiments, a control flow graph refers to an abstract representation of the program execution process. The control flow graph achieves the effect of simulating program execution by showing the flow of data between basic blocks when the program is executed within the process. A control flow graph is an abstract representation of a process or program. It is an abstract data structure used in the compiler, maintained internally by the compiler, and represents all paths that will be traversed during the execution of a program. It uses a graph to represent the possible flow of execution of all basic blocks within a process, and can also reflect the real-time execution process of a process. The control flow graph is the basic basis for static code analysis. The biggest feature of the basic blocks in the control flow graph is that the first line of code in the basic block contains all the entries, and the last line of code contains all the exits, which requires that the basis for generating the control flow graph needs to contain enough information.
[0092] In some embodiments, a basic block refers to a sequence of statements executed sequentially in a program. A basic block has only one entry and one exit, the entry is the first statement in the basic block, and the exit is the last statement in the basic block. For a basic block, execution only enters from its entry and exits from its exit. Specifically: a basic block has only one entry, which means that there is no other place in the program that can enter this basic block through a jump instruction. A basic block has only one exit, which means that only the last instruction of the program can lead to entering other basic blocks for execution. A typical feature of a basic block is that as long as the first instruction in the basic block is executed, all executions in the basic block will be executed only once in sequence.
[0093] In some embodiments, see Figure 5 , Figure 5 is a flow chart of a method for generating a data flow graph provided in an embodiment of the present application, Figure 3 The step 102 shown can be performed for each sub-code segment set separately. Figure 5 Steps 1021 to 1024 are implemented as shown.
[0094] In step 1021, a first sub-code segment is selected from a plurality of sub-code segments.
[0095] In some embodiments, the first sub-code segment is a sub-code segment whose starting code statement is the same as the starting code statement of the sub-code segment set.
[0096] In some embodiments, the above-mentioned step 1021 can be implemented in the following manner: obtain the starting code statement of the sub-code segment set, and perform the following processing for each sub-code segment: obtain the starting code statement of the sub-code segment, and compare the starting code statement of the sub-code segment with the starting code statement of the sub-code segment set to obtain a comparison result corresponding to the sub-code segment; when the comparison result indicates that the starting code statement of the sub-code segment is the same as the starting code statement of the sub-code segment set, determine the sub-code segment as the first sub-code segment.
[0097] In some embodiments, when the comparison result indicates that the start code statement of the sub-code segment is different from the start code statement of the sub-code segment set, the sub-code segment is not determined as the first sub-code segment.
[0098] In some embodiments, the above-mentioned sub-code segment set includes multiple sub-code segments arranged in sequence, and the sub-code segment includes multiple code statements arranged in sequence. The starting code statement of the sub-code segment set is the code statement with the earliest arrangement order in the sub-code segment set, and the starting code statement of the sub-code segment is the code statement with the earliest arrangement order in the sub-code segment.
[0099] In some embodiments, the above-mentioned sub-code segment includes multiple code statements, and the code statements included in the sub-code segment include a start code statement and an end code statement, wherein the start code statement has no preceding code statement in the sub-code segment, the start code statement has a subsequent code statement in the sub-code segment, the end code statement has a preceding code statement in the sub-code segment, and the end code statement has no subsequent code statement in the sub-code segment.
[0100] As an example, sub-code segment set A is {sub-code segment A1, sub-code segment A2, sub-code segment A3}, that is, sub-code segment set A includes sub-code segment A1, sub-code segment A2, and sub-code segment A3, wherein the starting code statement of sub-code segment set A is code statement A11, the starting code statement of sub-code segment A1 is code statement A11, the starting code statement of sub-code segment A2 is code statement A21, and the starting code statement of sub-code segment A3 is code statement A31. The starting code statement A11 of sub-code segment A1 is compared with the starting code statement A11 of sub-code segment set A to obtain a comparison result corresponding to sub-code segment A1; when the comparison result indicates that the starting code statement A11 of sub-code segment A1 is the same as the starting code statement A11 of sub-code segment set A, sub-code segment A1 is determined as the first sub-code segment.
[0101] In step 1022, a first initial basic block corresponding to the first sub-code segment is created, and the first initial basic block is determined as a first initial control flow graph.
[0102] In some embodiments, the start code statement of the first sub-code segment is the same as the start code statement of the sub-code segment set. The first initial control flow graph includes a first initial basic block, and the first initial basic block corresponds to the first sub-code segment one by one.
[0103] In step 1023, traverse i and perform the following processing: select the i-th sub-code segment from multiple sub-code segments, and create the i-th initial basic block corresponding to the i-th sub-code segment; connect the i-th initial basic block with the i-1-th initial control flow graph to obtain the i-th initial control flow graph.
[0104] In some embodiments, a start code statement of the i-th sub-code segment is logically associated with an end code statement of the i-1-th sub-code segment, 2≤i≤N, where N indicates the number of sub-code segments in the sub-code segment set.
[0105] As an example, when N=4, the second sub-code segment is selected from multiple sub-code segments, and a second initial basic block corresponding to the second sub-code segment is created; the second initial basic block is connected to the first initial control flow graph to obtain the second initial control flow graph; the third sub-code segment is selected from multiple sub-code segments, and a third initial basic block corresponding to the third sub-code segment is created; the third initial basic block is connected to the second initial control flow graph to obtain the third initial control flow graph; the fourth sub-code segment is selected from multiple sub-code segments, and a fourth initial basic block corresponding to the fourth sub-code segment is created; the fourth initial basic block is connected to the third initial control flow graph to obtain the fourth initial control flow graph.
[0106] As an example, when N=2, the second sub-code segment is selected from multiple sub-code segments, and a second initial basic block corresponding to the second sub-code segment is created; the second initial basic block is connected to the first initial control flow graph to obtain a second initial control flow graph.
[0107] In some embodiments, the above-mentioned selection of the i-th sub-code segment from multiple sub-code segments can be implemented in the following manner: obtain the end code statement of the i-1-th sub-code segment, and perform the following processing for each sub-code segment: obtain the start code statement of the sub-code segment, and determine the association information between the start code statement of the sub-code segment and the end code statement of the i-1-th sub-code segment; when the association information indicates that there is a logical association between the start code statement of the sub-code segment and the end code statement of the i-1-th sub-code segment, determine the sub-code segment as the i-th sub-code segment.
[0108] In some embodiments, when the association information indicates that there is no logical association between the start code statement of the sub-code segment and the end code statement of the (i-1)th sub-code segment, the sub-code segment is not determined as the (i)th sub-code segment.
[0109] As an example, sub-code segment set A is {sub-code segment A1, sub-code segment A2, sub-code segment A3}, that is, sub-code segment set A includes sub-code segment A1, sub-code segment A2, and sub-code segment A3, wherein the end code statement of sub-code segment A1 is logically associated with the start code statement of sub-code segment A2, and the end code statement of sub-code segment A2 is logically associated with the start code statement of sub-code segment A3.
[0110] Continuing with the above example, when i is equal to 2, the end code statement of the first subcode segment (subcode segment A1) is obtained, and for subcode segment A2, the start code statement of subcode segment A2 is obtained, and the association information between the start code statement of subcode segment A2 and the end code statement of the first subcode segment (subcode segment A1) is determined; when the association information indicates that there is a logical association between the start code statement of subcode segment A2 and the end code statement of the first subcode segment (subcode segment A1), subcode segment A2 is determined as the second subcode segment.
[0111] Continuing with the above example, when i is equal to 2, the end code statement of the first subcode segment (subcode segment A1) is obtained, and for subcode segment A3, the start code statement of subcode segment A3 is obtained, and the association information between the start code statement of subcode segment A3 and the end code statement of the first subcode segment (subcode segment A1) is determined; when the association information indicates that there is a logical association between the start code statement of subcode segment A3 and the end code statement of the first subcode segment (subcode segment A1), subcode segment A3 is not determined as the second subcode segment.
[0112] In step 1024, the Nth initial control flow graph is determined as the initial control flow graph of the node.
[0113] Continuing with the above example, when N=4, the second sub-code segment is selected from multiple sub-code segments, and the second initial basic block corresponding to the second sub-code segment is created; the second initial basic block is connected to the first initial control flow graph to obtain the second initial control flow graph; the third sub-code segment is selected from multiple sub-code segments, and the third initial basic block corresponding to the third sub-code segment is created; the third initial basic block is connected to the second initial control flow graph to obtain the third initial control flow graph; the fourth sub-code segment is selected from multiple sub-code segments, and the fourth initial basic block corresponding to the fourth sub-code segment is created; the fourth initial basic block is connected to the third initial control flow graph to obtain the fourth initial control flow graph, and the fourth initial control flow graph is determined as the initial control flow graph of the node.
[0114] Continuing with the above example, when N=2, the second sub-code segment is selected from multiple sub-code segments, and a second initial basic block corresponding to the second sub-code segment is created; the second initial basic block is connected to the first initial control flow graph to obtain a second initial control flow graph, and the second initial control flow graph is determined as the initial control flow graph of the node.
[0115] In this way, by creating the 1st initial basic block corresponding to the 1st sub-code segment, and determining the 1st initial basic block as the 1st initial control flow graph, and traversing i to perform the following processing: select the ith sub-code segment from multiple sub-code segments, and create the ith initial basic block corresponding to the ith sub-code segment; connect the ith initial basic block with the i-1th initial control flow graph to obtain the ith initial control flow graph, and determine the Nth initial control flow graph as the initial control flow graph of the node, so that the control flow transfer process of each basic block in the initial control flow graph of the node is consistent with the logical transfer process of each sub-code segment in the sub-code segment set corresponding to the node, thereby effectively improving the accuracy of the determined initial control flow graph.
[0116] In step 103, an extended control flow graph of each target sub-code segment is obtained.
[0117] In some embodiments, the target sub-code segment refers to a sub-code segment with target logic. For example, the target sub-code segment may be a sub-code segment with jump logic, the target sub-code segment may be a sub-code segment with judgment logic, or the target sub-code segment may be a sub-code segment with loop logic.
[0118] In some embodiments, see Figure 6 , Figure 6 is a flow chart of a method for generating a data flow graph provided in an embodiment of the present application, Figure 3 Before step 103 shown, it is also possible to execute Figure 6 Steps 106 to 109 are shown to determine the target sub-code segment.
[0119] In step 106, the logic identification character of the target logic is obtained, and the following steps 107 to 109 are respectively performed for each sub-code segment set.
[0120] As an example, the target logic may be jump logic, loop logic, enhanced loop logic, loop condition logic, etc. Different target logics correspond to different logic identification characters. The following describes the logic identification characters corresponding to different target logics.
[0121] Continuing with the above example, for jump logic, the functional functions with jump logic can be "if function" and "switch function", so the logical identification characters of the jump logic can be: "if" and "switch".
[0122] Continuing with the above example, for loop logic, the functional function with loop logic can be a "for function", and the logical identification character of the loop logic can be: "for".
[0123] Continuing with the above example, for enhanced loop logic, the functional function with loop logic can be "while", and the logical identification character of the loop logic can be: "while".
[0124] Continuing with the above example, for loop conditional logic, the functional function with loop conditional logic can be a "do-while function", and the logical identification character of the loop logic can be: "do-while".
[0125] In step 107, code characters are extracted from each sub-code segment in the sub-code segment set to obtain a code character set corresponding to each sub-code segment.
[0126] In some embodiments, a sub-code segment includes multiple code statements, each code statement includes multiple code characters, and by extracting code characters from each sub-code segment in the sub-code segment set, a code character set corresponding to each sub-code segment can be obtained, and the code character set includes all the code characters in the corresponding sub-code segment.
[0127] In step 108, for each code character set, each code character in the code character set is compared with the logical identification character to obtain a comparison result corresponding to the code character set.
[0128] In some embodiments, the comparison result indicates whether the logical identification character is present in the set of code characters.
[0129] As an example, for code character set Q, each code character in code character set Q is compared with the logical identification character to obtain the comparison result corresponding to code character set Q; for code character set W, each code character in code character set W is compared with the logical identification character to obtain the comparison result corresponding to code character set W.
[0130] In step 109, when the comparison result indicates that the logic identification character exists in the code character set, the sub-code segment corresponding to the code character set is determined as the target sub-code segment.
[0131] In some embodiments, when the comparison result indicates that there is no logical identification character in the code character set, the sub-code segment corresponding to the code character set is not determined as the target sub-code segment.
[0132] In some embodiments, when the comparison result indicates that there is a logical identification character in the code character set, it means that the sub-code segment corresponding to the code character set can realize the target logical function indicated by the logical identification character, and the sub-code segment can be determined as a target sub-code segment with the target logical function.
[0133] In this way, by obtaining the logical identification character of the target logic, and for each sub-code segment set, code characters are extracted for each sub-code segment in the sub-code segment set, and code character sets corresponding to each sub-code segment are obtained, each code character in the code character set is compared with the logical identification character to obtain a comparison result corresponding to the code character set. When the comparison result indicates that there is a logical identification character in the code character set, the sub-code segment corresponding to the code character set is determined as the target sub-code segment, so as to accurately determine whether the sub-code segment is a target sub-code segment with a target logical function by comparing the logical identification character of the target logic with the code character of the sub-code segment, thereby facilitating subsequent processing of the target sub-code segment to obtain reference control flow graphs corresponding to the initial control flow graphs, thereby effectively improving the accuracy of the reference control flow graphs determined subsequently.
[0134] In some embodiments, see Figure 7 , Figure 7 is a flow chart of a method for generating a data flow graph provided in an embodiment of the present application, Figure 3 Step 103 shown may be performed by executing Figure 7 Steps 1031 to 1033 are shown to be implemented.
[0135] In step 1031, a mapping relationship between a plurality of preset logic identifiers and corresponding preset control flow graphs is obtained, and the following steps 1032 to 1033 are respectively performed for each target sub-code segment.
[0136] As an example, refer to Table 1 below, which is a schematic table of the mapping relationship between multiple preset logic identifiers and corresponding preset control flow graphs provided in an embodiment of the present application.
[0137] Table 1 Schematic diagram of mapping relationship between multiple preset logic identifiers and corresponding preset control flow graphs
[0138] Preset logical identifier "for" Preset control flow graph T1 Preset logical identifier "continue" Preset control flow graph T2 Preset logical identifier "While" Preset control flow graph T3 Preset logical identifier "Do-while" Preset control flow graph T4 Preset logical identifier "Goto" Preset control flow graph T5 Preset logical identifier "Label" Preset control flow graph T6 Preset logical identifier "try-catch" Preset control flow graph T7 Preset logical identifier "if" Preset control flow graph T8 Preset logical identifier "ifelse" Preset control flow graph T9 Preset logical identifier "switch" Preset control flow graph T10
[0139] The following is a detailed description of each preset control flow graph provided in Table 1 above.
[0140] In some embodiments, for the preset control flow graph T1, see Figure 8 , Figure 8 It is a structural schematic diagram of a preset control flow graph T1 provided in an embodiment of the present application. The preset control flow graph T1 is a control flow graph corresponding to the preset logic identifier "for", which is used to describe the control flow of the for loop statement. The preset control flow graph T1 includes a predecessor basic block, an entry basic block, an initialization basic block, a conditional basic block, a loop basic block, an exit basic block, an incremental basic block, and a subsequent basic block.
[0141] In some embodiments, for the preset control flow graph T2, see Fig. 9 , Fig. 9 It is a structural schematic diagram of a preset control flow graph T2 provided in an embodiment of the present application. The preset control flow graph T2 is a control flow graph corresponding to the preset logic identifier "continue", which is used to describe the control flow of the enhanced loop statement. The preset control flow graph T2 includes a predecessor basic block, an entry basic block, an initialization basic block, a conditional basic block, a loop basic block, an exit basic block and a subsequent basic block.
[0142] In some embodiments, for the preset control flow graph T3, see Fig.10 , Fig.10 It is a structural diagram of a preset control flow graph T3 provided in an embodiment of the present application. The preset control flow graph T3 is a control flow graph corresponding to the preset logic identifier "While", which is used to describe the control flow of the While loop statement. The preset control flow graph T3 includes a predecessor basic block, an entry basic block, a conditional basic block, a loop basic block, an exit basic block and a subsequent basic block.
[0143] In some embodiments, for the preset control flow graph T4, see Fig.11 , Fig.11It is a structural diagram of a preset control flow graph T4 provided in an embodiment of the present application. The preset control flow graph T4 is a control flow graph corresponding to the preset logic identifier "Do-while", which is used to describe the control flow of the Do-while loop statement. The preset control flow graph T4 includes a predecessor basic block, an entry basic block, a conditional basic block, a loop basic block, an exit basic block and a subsequent basic block.
[0144] In some embodiments, for the preset control flow graph T5, see Fig.12 , Fig.12 It is a structural diagram of a preset control flow graph T5 provided in an embodiment of the present application. The preset control flow graph T5 is a control flow graph corresponding to a preset logic identifier "Goto", which is used to describe the control flow of a Goto loop statement. The preset control flow graph T5 includes a Goto statement basic block set, and the Goto statement basic block set includes multiple Goto statement basic blocks.
[0145] In some embodiments, for the preset control flow graph T6, see Fig.13 , Fig.13 It is a structural diagram of a preset control flow graph T6 provided in an embodiment of the present application. The preset control flow graph T6 is a control flow graph corresponding to the preset logic identifier "Label", which is used to describe the control flow of the Label statement, and the preset control flow graph T6 includes a Label statement basic block.
[0146] In some embodiments, for the preset control flow graph T7, see Fig.14 , Fig.14 It is a structural diagram of a preset control flow graph T7 provided in an embodiment of the present application. The preset control flow graph T7 is a control flow graph corresponding to the preset logic identifier "try-catch", which is used to describe the control flow of the try-catch statement. The preset control flow graph T7 includes a try statement basic block, a catch statement basic block, a predecessor basic block, and a subsequent basic block.
[0147] In some embodiments, for the preset control flow graph T8, see Fig.15 , Fig.15 It is a schematic diagram of the structure of the preset control flow graph T8 provided in the embodiment of the present application. The preset control flow graph T8 is a control flow graph corresponding to the preset logic identifier "if", which is used to describe the control flow of the if statement. The preset control flow graph T8 includes a predecessor basic block and a subsequent basic block.
[0148] In some embodiments, for the preset control flow graph T9, see Fig.16 , Fig.16It is a structural diagram of a preset control flow graph T9 provided in an embodiment of the present application. The preset control flow graph T9 is a control flow graph corresponding to the preset logic identifier "ifelse", which is used to describe the control flow of the ifelse statement. The preset control flow graph T9 includes a predecessor basic block, a subsequent basic block, an if statement basic block, and an else statement basic block.
[0149] In some embodiments, for the preset control flow graph T10, see Fig.17 , Fig.17 It is a structural diagram of a preset control flow graph T10 provided in an embodiment of the present application. The preset control flow graph T10 is a control flow graph corresponding to the preset logic identifier "switch", which is used to describe the control flow of the switch statement. The preset control flow graph T10 includes a predecessor basic block, a subsequent basic block, and a case basic block.
[0150] In step 1032, a logic identifier of the target logic in the target sub-code segment is obtained, and the logic identifier is compared with each preset logic identifier in the mapping relationship to obtain a comparison result corresponding to the logic identifier.
[0151] In some embodiments, the comparison result corresponding to the above-mentioned logical identifier is used to indicate whether the logical identifier exists in the preset logical identifier.
[0152] Continuing with the above example, when the logical identifier of the target logic in the target sub-code segment is switch, the logical identifier is compared with each preset logical identifier in the mapping relationship in Table 1 above to obtain the comparison result corresponding to the logical identifier switch. The comparison result corresponding to the logical identifier switch is used to indicate whether the logical identifier switch exists in the preset logical identifier.
[0153] Continuing with the above example, when the logical identifier of the target logic in the target sub-code segment is if, the logical identifier is compared with each preset logical identifier in the mapping relationship in Table 1 above to obtain the comparison result corresponding to the logical identifier if. The comparison result corresponding to the logical identifier if is used to indicate whether the logical identifier if exists in the preset logical identifier.
[0154] In step 1033, when the comparison result indicates that a logical identifier exists in the preset logical identifier, the preset control flow graph corresponding to the logical identifier in the mapping relationship is determined as the extended control flow graph of the target sub-code segment.
[0155] Continuing with the above example, when the comparison result corresponding to the logic identifier switch is used to indicate that the logic identifier switch exists in the preset logic identifier, the preset control flow graph T10 corresponding to the logic identifier switch in the mapping relationship is determined as the extended control flow graph of the target sub-code segment.
[0156] Continuing with the above example, when the comparison result corresponding to the logic identifier if is used to indicate that the logic identifier if exists in the preset logical identifier, the preset control flow graph T8 corresponding to the logic identifier if in the mapping relationship is determined as the extended control flow graph of the target sub-code segment.
[0157] In this way, by comparing the logical identifier of the target logic in the target sub-code segment with each preset logical identifier in the mapping relationship, a comparison result corresponding to the logical identifier is obtained. When the comparison result indicates that the logical identifier exists in the preset logical identifier, the preset control flow graph corresponding to the logical identifier in the mapping relationship is determined as the extended control flow graph of the target sub-code segment, thereby achieving accurate search of the extended control flow graph of the target sub-code segment.
[0158] In step 104, the initial basic blocks corresponding to the target sub-code segments in the initial control flow graphs are replaced with the corresponding extended control flow graphs to obtain reference control flow graphs corresponding to the initial control flow graphs.
[0159] Continuing with the above example, when the target sub-code segment in the initial control flow graph includes a target sub-code segment with a logical identifier switch and a target sub-code segment with a logical identifier if, the initial basic block corresponding to the target sub-code segment with the logical identifier switch in the initial control flow graph is replaced with an extended control flow graph (preset control flow graph T10), and the initial basic block corresponding to the target sub-code segment with the logical identifier if in the initial control flow graph is replaced with an extended control flow graph (preset control flow graph T8), to obtain a reference control flow graph corresponding to the initial control flow graph.
[0160] In this way, the initial basic blocks corresponding to each target sub-code segment in each initial control flow graph are replaced by the corresponding extended control flow graph, and the reference control flow graphs corresponding to each initial control flow graph are obtained, thereby realizing a control flow graph expression that refines the target sub-code segment with relatively complex control flow in the initial control flow graph, so that the determined reference control flow graph can more accurately express the control flow of the target code, thereby effectively improving the accurate expression of the control flow of the target code.
[0161] In step 105, the target data flow graph of the target code is generated by combining the abstract syntax tree and the reference control flow graph.
[0162] In some embodiments, a data flow diagram graphically expresses the logical functions of a system, the logical flow of data within the system, and the logical transformation process from the perspective of data transmission and processing. It is the main expression tool of structured system analysis methods and a graphical method for representing software models.
[0163] In some embodiments, see Fig.18 , Fig.18 It is a schematic flowchart of the method for generating a data flow diagram provided by an embodiment of the present application. Figure 3 The step 105 shown can be implemented by executing Fig.18 the steps 1051 to 1054 shown.
[0164] In step 1051, each node in the abstract syntax tree is respectively replaced with a corresponding reference control flow diagram to obtain a target control flow diagram corresponding to the target code.
[0165] In some embodiments, the nodes in the abstract syntax tree are in one-to-one correspondence with the initial control flow Figure 1 and the initial control flow diagram is in one-to-one correspondence with the reference control flow Figure 1 By respectively replacing each node in the abstract syntax tree with the corresponding reference control flow diagram of each node, the conversion between the abstract syntax tree and the control flow diagram is realized.
[0166] In some embodiments, since the determined reference control flow diagram can more accurately represent the control flow of the target code, and the abstract syntax tree of the target code can accurately implement the abstract representation of the target code, by respectively replacing each node in the abstract syntax tree with the corresponding reference control flow diagram, a target control flow diagram corresponding to the target code is obtained, so that the obtained target control flow diagram can both accurately represent the control flow of the target code and accurately implement the abstract representation of the target code.
[0167] In step 1052, a data flow diagram conversion is performed on the target control flow diagram to obtain an initial data flow diagram of the target code.
[0168] In some embodiments, the above step 1052 can be implemented in the following manner: perform a first data flow diagram conversion on the target control flow diagram to obtain a first data flow diagram of the target code; traverse j and perform the following processing until the initial data flow diagram of the target code is obtained: perform a j-th data flow diagram conversion on the target control flow diagram to obtain a j-th data flow diagram of the target code; when the j-th data flow diagram is the same as the (j - 1)-th data flow diagram, determine the j-th data flow diagram as the initial data flow diagram of the target code.
[0169] In some embodiments, 1 < j ≤ M, and the M-th data flow diagram is the same as the (M - 1)-th data flow diagram.
[0170] As an example, perform a first data flow diagram conversion on the target control flow diagram to obtain a first data flow diagram of the target code; perform a second data flow diagram conversion on the target control flow diagram to obtain a second data flow diagram of the target code; when the second data flow diagram is the same as the first data flow diagram, determine the first data flow diagram as the initial data flow diagram of the target code.
[0171] As an example, the target control flow graph is subjected to the first data flow graph conversion to obtain the first data flow graph of the target code; the target control flow graph is subjected to the second data flow graph conversion to obtain the second data flow graph of the target code; when the second data flow graph is different from the first data flow graph, the target control flow graph is subjected to the third data flow graph conversion to obtain the third data flow graph of the target code, and when the third data flow graph is the same as the second data flow graph, the second data flow graph is determined as the initial data flow graph of the target code.
[0172] In this way, by performing at least two data flow graph conversions on the target control flow graph, when the data flow graphs obtained from two adjacent data flow graph conversions are the same, the data flow graph obtained from the previous data flow graph conversion is determined as the initial data flow graph of the target code, thereby reducing the probability of inaccurate data flow graph caused by unstable data flow during the data flow graph conversion process, thereby effectively improving the accuracy of the determined data flow graph.
[0173] In some embodiments, the target control flow graph includes a plurality of target basic blocks, and the target basic blocks correspond one-to-one to the target code segments in the target code.
[0174] In some embodiments, the above-mentioned first data flow graph conversion of the target control flow graph to obtain the first data flow graph of the target code can be achieved in the following manner: running the target code, and in the process of running the target code, recording the running data flow of each target code segment during the running process; for each target basic block, caching the running data flow of the target code segment corresponding to the target basic block into the target basic block to obtain the first updated basic block corresponding to the target basic block; replacing each target basic block in the target control flow graph with the corresponding first updated basic block, respectively, to obtain the first data flow graph of the target code.
[0175] In some embodiments, during the process of running the target code, each target code segment of the target code will generate an operation data flow during the operation. By recording the generated operation data flow and adding the recorded operation data flow to the cache of the corresponding target basic block, an updated basic block is obtained, and each target basic block in the target control flow graph is replaced with the corresponding updated basic block, thereby realizing the conversion of the control flow graph to the data flow graph.
[0176] In some embodiments, the above-mentioned j-th data flow graph conversion of the target control flow graph to obtain the j-th data flow graph of the target code can be achieved in the following manner: obtaining the j-1th data flow graph, the j-1th data flow graph including the j-1th operation data flow of each target code segment during the j-1th operation; and running the target code for the jth time, and in the process of running the target code for the jth time, recording the j-th operation data flow of each target code segment during the j-1th operation; and comparing each j-th operation data flow with the corresponding j-1th operation data flow respectively to obtain the comparison result of each j-th operation data flow; when the comparison result of the j-th operation data flow indicates that the j-th operation data flow and the corresponding j-1th operation data flow are different, deleting the j-1th operation data flow in the target basic block that caches the j-1th operation data flow in the j-1th data flow graph, and caching the j-1th operation data flow in the target basic block, thereby achieving the j-th data flow graph conversion of the target control flow graph by updating the j-1th data flow graph.
[0177] In other embodiments, the above-mentioned j-th data flow graph conversion of the target control flow graph to obtain the j-th data flow graph of the target code can also be achieved in the following way: running the target code for the j-th time, and during the j-th running of the target code, recording the running data flow of each target code segment during the j-th running process; for each target basic block, caching the j-th running data flow of the target code segment corresponding to the target basic block into the target basic block to obtain the j-th updated basic block corresponding to the target basic block; replacing each target basic block in the target control flow graph with the corresponding j-th updated basic block, to obtain the j-th data flow graph of the target code.
[0178] In this way, by performing at least two data flow graph conversions on the target control flow graph, when the data flow graphs obtained from two adjacent data flow graph conversions are the same, the data flow graph obtained from the previous data flow graph conversion is determined as the initial data flow graph of the target code, thereby reducing the probability of inaccurate data flow graph caused by unstable data flow during the data flow graph conversion process, thereby effectively improving the accuracy of the determined data flow graph.
[0179] In step 1053, redundant data flows in the initial data flow graph are identified to obtain redundant data flows in the initial data flow graph.
[0180] In some embodiments, the initial data flow graph includes multiple data flows, and the data flows are used to indicate the data migration process from the starting point of the data flow to the end point of the data flow. The above-mentioned redundancy identification is a process for identifying redundant data flows.
[0181] In some embodiments, the above-mentioned step 1053 can be implemented as follows: perform the following processing for each data flow in the initial data flow graph: obtain the starting level of the starting point of the data flow in the target code, and the end level of the data flow in the target code; compare the starting level and the end level to obtain a level comparison result; when the level comparison result indicates that the starting level is greater than the end level, determine the data flow as a redundant data flow.
[0182] In some embodiments, when the level comparison result indicates that the start level is less than or equal to the end level, the data stream is not determined as a redundant data stream.
[0183] In some embodiments, the starting point of the above-mentioned data flow is at the starting point level in the target code, which is the code level of the code segment in the target code where the starting point of the data flow is in the target code, and the end point of the above-mentioned data flow is at the end point level in the target code, which is the code level of the code segment in the target code where the end point of the data flow is in the target code.
[0184] In some embodiments, the above-mentioned code hierarchy is used to indicate the number of logical nesting levels at the target position of the target code. Since the scope of a code segment with a lower code hierarchy can cover the scope of a code segment with a higher code hierarchy, and the scope of a code segment with a higher code hierarchy cannot cover the scope of a code segment with a lower code hierarchy, the data flow generated during the code execution can flow from a lower code hierarchy to a higher code hierarchy, but cannot flow from a higher code hierarchy to a lower code hierarchy (since a higher code hierarchy cannot cover the scope of a lower code hierarchy, even if the data flow flows from a higher code hierarchy to a lower code hierarchy, the code segment corresponding to the lower code hierarchy cannot call the corresponding data). Therefore, when the hierarchy comparison result indicates that the starting level is greater than the ending level, the data flow corresponding to the hierarchy comparison result can be determined as a redundant data flow.
[0185] In step 1054, the redundant data flows in the initial data flow graph are deleted to obtain the target data flow graph.
[0186] In this way, by redundancy identification of data flows in the initial data flow graph, redundant data flows in the initial data flow graph are obtained, and the redundant data flows in the initial data flow graph are deleted to obtain the target data flow graph, so that the obtained target data flow graph does not contain redundant data flows, thereby effectively improving the accuracy of the target data flow graph.
[0187] In this way, by obtaining the abstract syntax tree of the target code, the initial control flow graph of the node is constructed based on the set of sub-code segments corresponding to each node, and the extended control flow graph of each target sub-code segment is obtained, and the initial basic block corresponding to each target sub-code segment in the initial control flow graph is replaced with the corresponding extended control flow graph, and the reference control flow graph corresponding to each initial control flow graph is obtained. The target flow graph of the target code is generated by combining the abstract syntax tree and the reference control flow graph. By replacing the initial basic block corresponding to the target sub-code segment of the initial control flow graph with the corresponding extended control flow graph, the control flow graph corresponding to each target sub-code segment in the obtained reference control flow graph is more accurate, and then combined with the abstract syntax tree and the reference control flow graph, the accuracy of the target data flow graph of the generated target code can be effectively improved, so that the obtained target data flow graph can accurately reflect the data flow in the target code, and the accuracy of the data flow graph generation is effectively improved.
[0188] The following is an explanation of an exemplary application of an embodiment of the present application in an actual data flow graph generation application scenario.
[0189] A data flow diagram is a graphical representation of the logical functions of a system, the logical flow of data within the system, and the logical transformation process from the perspective of data transmission and processing. It is the main expression tool of the structured system analysis method and a graphical representation method for representing software models. The process of generating a data flow diagram in the embodiment of the present application will be described in detail below.
[0190] The data flow graph obtained by the data flow graph generation method provided in the embodiment of the present application can be implemented in code defect tools. From the effect of use, the parsing mechanism from the abstract syntax tree to the control flow graph basically covers most of the code scenarios, and can support large-scale (such as game client) code engineering with reliable performance guarantee.
[0191] In some embodiments, for generating a control flow graph based on an abstract syntax tree, the abstract syntax tree is a tree structure, each code has a StatementContext type node as a parent node, and adjacent StatementContext nodes are connected by StatementSeqContext nodes. When the target code is int i=j; intj=3; i=j+1, the corresponding abstract syntax tree is as follows Figure 4 As shown, Figure 4 Shown is the abstract syntax tree of the target code.
[0192] In some embodiments, the control flow graph, as a graph representing program logic and data flow, contains two types of members: basic blocks and edges. A basic block contains a list of codes belonging to the block. The biggest feature of a basic block is that the basic block has one and only one entry, that is, the first code of the block; the basic block has one and only one exit, that is, the last code of the block. There is an edge connection between basic block A and basic block B if and only if there are two situations: there is a conditional or unconditional jump from the end of basic block A to the beginning of basic block B; the end code of basic block A is adjacent to the beginning code of basic block B, and the end of basic block A does not contain an unconditional jump. Convert the tree-like abstract syntax tree into a control flow graph structure as required. During the entire conversion process, by creating a task queue and initializing the task, the current basic block, and the node stack to be processed by the current basic block.
[0193] In some embodiments, the task model in the queue includes two members: the current basic block + the node stack to be processed by the current basic block, where ParserRuleContext is the unified parent class of the Antlr4 syntax node. The basic task processing logic is to obtain the task members in the queue, pop the first node in the stack, and process it according to the type of the node. Under normal circumstances, according to the default processing logic, the child nodes of the node will be obtained, and they will be pressed into the node stack of the current member respectively, and then the current member will be put back into the queue. At the same time, it will be judged whether this node is a StatementContext node of a non-forked code line. If so, it means that the node is the beginning of a new line of code, and the content contained in the node is added to the code list of the current basic block (hereinafter referred to as InstructionList). From the characteristics of the basic blocks and edges, it can be seen that the core point of the conversion is to find out the nodes involved in jumps or forks in the abstract syntax tree and perform special processing, that is, to process the logic outside the ordinary situation. Here, in the embodiment of the present application, the edge pointing to the basic block is called the input edge of the basic block, and the edge pointing outward from the basic block is called the output edge; the basic block to which the edge points is called the destination block, and the basic block from which the edge originates is called the source block. The control flow structure at the processed fork node is called a fork set. Due to the process before and after the process, the fork node will be expanded from one basic block to a fork set with several basic blocks, and the downstream edges of the original basic blocks need to be reorganized and connected to the basic block facing downstream at the bottom of the processed fork set (hereinafter collectively referred to as the output basic block). In order to complete this common action, three common actions are abstracted for each processing logic, namely: collecting the output edges of the current basic block (that is, the downstream edges of the original basic block mentioned above), hereinafter referred to as collecting output edges; configuring follow-up basic blocks. If there are nodes in the node stack of the current basic block that have not been popped out, it means that there are still lines of code to be processed after the node, and then configure follow-up basic blocks for the remaining nodes. If the follow-up basic blocks exist, they can be uniformly used as the processed output basic blocks. If they do not exist, the basic blocks at the bottom of the processed fork set itself need to be used as the output basic blocks. This process is hereinafter referred to as obtaining follow-up blocks; after the processing is completed, reconnect the output edges facing the downstream basic blocks and the original collection, hereinafter referred to as reconnection. The following will explain how to handle fork jump nodes.
[0194] In some embodiments, see Figure 15 to Figure 16 , call the ctx.If() function. If it is not empty, it means it is an if statement. Then call the ctx.Else() function to determine whether it contains an else clause. It is necessary to give two different scenarios based on whether it contains an else clause. Figures 15 to 16The control flow graphs of these two fork sets (and the extended control flow graphs described above). For these two scenarios, first perform the operations of obtaining the follow-up block and collecting the external edges uniformly, call the ctx.statement() function to obtain all the clauses of this node, including if and else (if existing) clauses and traverse them. For each clause node, create a new basic block and a new node stack, push the clause node into the stack, and assemble the new task and add it to the queue. Each newly created basic block also needs to be connected with the current basic block (that is, the predecessor basic block) (IF and IF_ELSE type edges). At the same time, it is also necessary to determine whether the follow-up block has been obtained. If so, connect the new basic block with the follow-up block. If not, add the new basic block to the output basic block set. The difference between the above two scenarios is that if it is the second scenario and there is a follow-up basic block, it is necessary to connect the predecessor basic block and the follow-up basic block (corresponding to Figure 4 The leftmost vertical edge in the (IF_END type edge), finally, perform the reconnection operation uniformly.
[0195] In some embodiments, see Fig.17 , call ctx.Switch() function to determine whether it is a switch statement. The switch statement fork set control flow graph (the extended control flow graph described above) is as follows Fig.17 As shown. Call the ctx.Switch() function to determine whether it is a switch statement. The switch statement also first obtains the follow-up block and collects the output edges. Considering that the content of the switch statement is relatively complex, another subclass is defined to traverse the content of the switch statement in visitor mode to collect complete case clause information, including whether the case clause contains break\return\goto\continue and other jump statement information (note that here we can only collect whether the jump statement is included. Since the case clause information has not been parsed, the specific location of the jump statement cannot be determined and is uniformly processed in the subsequent jump node). After the collection is completed, traverse the collected case clause list, create a new basic block and a new node stack for each case clause, push the case clause content, generate a new task and add it to the queue, and create an edge connection with the current basic block. For a case clause with a break, or the last case clause, determine whether a follow-up block is obtained. If so, connect the new basic block with the follow-up block (SWITCH_BREAK_EDGE type edge). If not, add the new basic block to the output basic block set. For a case clause without a break, it needs to be connected with the next case clause basic block. Finally, perform the reconnection operation uniformly.
[0196] In some embodiments, the loop statement of C++ is slightly simpler than other types of jump statement processing logic, mainly because the loop statements are unified single-input and single-output structures. Because of this, unified entry blocks and exit blocks are pre-defined during processing. They do not contain code themselves, but can be used to connect the upstream and downstream of the fork set externally, and can be used as the unified entry and exit of the fork set internally, which is convenient for simplifying logical processing. There are four main forms of loop statements: ordinary for loops, enhanced for loops, while statements, and do-while statements. Before executing the four specific types of logic, first uniformly execute the common actions of obtaining follow-up blocks and collecting output edges, and define unified entry and exit blocks. Before introducing the four types of loop statement processing logic respectively, it is also necessary to introduce that during the processing process, specific basic blocks need to be marked with continue jump identifiers and break jump identifiers, which are used for subsequent continue / break jump statements to find the target basic blocks for connection.
[0197] In some embodiments, see Figure 8 , the common for loop statement fork set control flow graph is as follows Figure 8 As shown, ctx.For() is called to determine whether it is a for loop statement, and then ctx.forinitstatement() is called to obtain the initialization clause to determine whether it is a normal for loop statement. If so, the ctx.condition() function is called to obtain the conditional clause, the ctx.statement() function is called to obtain the loop content clause, and the ctx.expression() function is called to obtain the post-loop incremental processing clause. The processing flow of this type of statement is to first create an initialization basic block and a conditional basic block for the initialization statement and the conditional statement respectively, and connect the entry block and the initialization basic block (ITER_INIT_EDGE type edge), the initialization basic block and the conditional basic block (ITER_CONDITION_EDGE type edge). The conditional basic block has a fork, which connects the loop basic block (ITER_EDGE type edge) and the exit block (ITER_END_EDGE type edge) created for the loop content clause. The loop basic block is connected to the basic block created for the post-loop increment processing clause (ITER_BACK_EDGE type edge), and the incremental basic block needs to be connected back to the conditional basic block (ITER_CONDITION_EDGE type edge), representing the completion of a loop; on the other hand, the exit basic block represents the unified exit of the entire basic block set when the loop condition is not met. Here, the incremental basic block is configured with the continue flag, and the exit block is set to the break flag.
[0198] In some embodiments, see Fig. 9 , the enhanced loop statement fork set control flow graph is as follows Fig. 9As shown, the enhanced statement control flow graph is similar to the common statement, except that the enhanced loop statement control flow graph does not have an increment basic block. Here, the conditional basic block is configured with a continue flag, and the exit basic block is configured with a break flag.
[0199] In some embodiments, see Fig.10 , the control flow graph of the while statement fork set is as follows Fig.10 As shown in the figure, the control flow graph of the while loop statement is similar to that of the common statement, except that the control flow graph of the while loop statement does not have an initialization basic block. Here, the conditional basic block is configured with a continue flag, and the exit basic block is configured with a break flag.
[0200] In some embodiments, see Fig.11 , the Do-while statement fork set control flow graph is as follows Fig.11 As shown in the figure, this type of loop statement is quite different from the previous three. First, the entry basic block is connected to the loop basic block, then to the conditional basic block, and finally to the exit basic block. When the loop condition is not met, the conditional basic block needs to be connected back to the loop basic block, indicating the data flow when the loop condition is not met. Here, the conditional basic block is configured with a continue flag, and the exit basic block is configured with a break flag.
[0201] In some embodiments, the original output edges of the current basic block are deleted for the Return statement processing logic. In addition, it can be determined whether to add the input and total output blocks on the control flow graph according to the actual situation. If added, the current basic block is connected to the total output block (RETURN type edge).
[0202] In some embodiments, for Break / Continue statements, it is necessary to solve the jump problem mentioned in the switch statement node and loop statement node above. Taking the current basic block as the base point, starting from the input edge set, the upstream basic block is traversed in a breadth-first manner until a basic block with a SWITCH_BREAK_EDGE output edge or an entry basic block of a loop statement is encountered, thereby determining whether the jump statement is in a switch statement or a loop statement. If it is in a loop statement, for a break statement, it is necessary to connect the current basic block to the basic block (BREAK type edge) with a break mark in the corresponding loop statement fork set; for a Continue statement, it is necessary to connect the current basic block to the basic block with a continue mark in the corresponding loop statement fork set. If it is in a switch statement, take out the current most recent basic block with a SWITCH_BREAK_EDGE type output edge, and switch the source block of the edge to the current basic block. The significance of source block switching is that there may be a complex set of basic blocks in the case clause, so it is necessary to switch the SWITCH_BREAK_EDGE edge that is initially configured to the basic block where the real break statement is located.
[0203] In some embodiments, see Figure 12 to Figure 13 For the Goto statement, connect the Goto statement with the label, and create the goto source basic block cache and the label statement basic block cache from the global dimension. For the goto statement, the current basic block will be stored in the first cache, and the label statement basic block will be taken from the second cache with the label name as the identifier. If there is a value, it will be connected (GOTO type edge). Using two caches avoids the situation of missing connections due to the uncertain order of goto and label statements. The label statement introduced below is exactly the opposite of this processing logic.
[0204] In some embodiments, as the target of a goto statement, the operation of the Label node is exactly the opposite of that of the goto statement. First, the label name is obtained, and the current basic block is placed in the goto target basic block cache using the label name as an identifier, indicating that the basic block is a jump target basic block; at the same time, the corresponding basic block is obtained from the goto source basic block cache using the label name as an identifier, and if the obtained value is not empty, the current basic block is linked to the obtained basic block.
[0205] In some embodiments, see Fig.14 , the try-catch statement fork set control flow graph is as follows Fig.14As shown, the try clause content is obtained by calling the ctx.compound() function, and the catch clause list is obtained by calling the ctx.handlerSeq() function. First, a new basic block is created for the try clause content, and connected with the predecessor basic block, and connected with each subsequent basic block respectively (TRY_CATCH_END_EDGE type edge), then traverse to obtain the catch clause list, create new basic blocks respectively, and connect with the try statement basic block (TRY_CATCH_EDGE type edge), and connect with each subsequent basic block respectively (TRY_CATCH_END_EDGE type edge).
[0206] In some embodiments, the processing logic of the Throw node is similar to Break / Continue, and it is necessary to traverse upward to the nearest try statement basic block and modify the connection edge between the original try statement basic block and the catch statement basic block to the connection edge between the current basic block and the catch statement basic block.
[0207] In some embodiments, during the data flow analysis process, a visitor approach can be used to perform single-line code transformation equation calculations, and edge transformation equations can be used to embed domain-sensitive mechanisms in the data flow analysis.
[0208] In some embodiments, the process of calculating the single-line code conversion equation using the visitor method is shown in the following code:
[0209] For(each basic block){
[0210] OUT[B] = nil;
[0211] }
[0212] While(changes to any OUT){
[0213] For(each basic block){
[0214] IN[B]=U papredecessor of B OUT[P];
[0215] OUT[B]=TransferFunction(IN[B])
[0216] }
[0217] }
[0218] The transfer function is the conversion equation for each code instruction in the basic block. For reachability analysis, the conversion equation is usually in the form of:
[0219] OUT[B]=gen B U(IN[B]-kill B )
[0220] The actual meaning of the conversion equation represents the union of the definition of variable v in this line of code and the definition information of all variables except variable v in the predecessor input. IN / OUT is called the intermediate result of each basic block calculation. It can be seen that due to the kill B In the iteration process, a variable needs to be repeatedly overwritten / updated. For this reason, the intermediate result type is defined as a mapping type with variable name as key and bit variable as value.<String,BitSet> ), each bit in the bit variable represents a different code line type according to the rule requirements. This definition makes it convenient to perform overwriting operations when converting equations, thereby achieving the effect of eliminating the variable definition information elsewhere in the code. Each element in the InstructionList in the basic block is of StatementContext type, which contains all the information of a whole line of code, and its form is still the tree structure of the abstract syntax tree generated by Antlr4. Antlr4 itself provides a visitor base class BaseVisitor for traversing the tree structure. The embodiment of the present application inherits and uses this class to obtain the specific content of each code. By overwriting the access of specific types of nodes, the embodiment of the present application can capture the code lines involved in the rules, thereby realizing the conversion equation calculation logic of the embodiment of the present application. Here is an example, such as the null pointer dereference rule, the rule is concerned with the statement that assigns the null pointer to NULL, and the dereference statement (* pointer variable name or pointer variable name->). The pseudo code of the implemented visitor method is as follows:
[0221] Bitset bit variable;
[0222] The 0th bit is marked as empty;
[0223] The first bit identifies the dereference operation;
[0224] It is passed in as input when the class is initialized and is output after traversal and processing;
[0225] Override the assignment access statement (the value can be assigned to null);
[0226] If a statement with a left value of NULL is encountered, the operation resultFact.put (left value variable name) will be performed;
[0227] Override the initialization access statement (can be initialized to NULL);
[0228] If a statement with a right value of NULL is encountered, the operation resultFact.put (left value variable name) will be performed;
[0229] Override->dereference access statement;
[0230] When encountering the -> pointer dereference method, and the 0th bit of the original bitset value of the variable is true, the 1st bit is also assigned the value of true, indicating a null pointer reference prompt;
[0231] Override->prefix access statement;
[0232] When a * prefixed pointer dereference method is encountered, and the 0th bit of the original bitset value of the variable is already true, the 1st bit is also assigned a value of true, indicating a null pointer dereference prompt.
[0233] From the above code, we can see that the 0th bit in the bit variable BitSet represents the code statement that assigns it to NULL, and the 1st bit represents the pointer dereference statement. In iterative calculation, if the 0th bit of the BitSet value of a variable is already true and a dereference code line is encountered, the 1st bit of the BitSet will be assigned to true, and a prompt will be displayed.
[0234] The domain-sensitive mechanism is embedded in the data flow analysis by using the edge conversion equation. In addition, since most rules involve variable scope issues, the intermediate calculation information of the variable cannot be passed on outside the scope of the variable definition. As the variable becomes invalid, the intermediate result information of the variable needs to be cleared. One point that needs to be pointed out here is that in the previous step, in the process of generating the control flow graph from the Antlr4 abstract syntax tree, the embodiment of the present application adds hierarchical attributes to the basic block. The level of the root basic block of the control flow graph is 0. When the {} bracket code (CompoundStatementContext) is encountered during the traversal process, the level of the basic block in the bracket will be increased by 1. Through this mechanism, the generated control flow graph itself carries domain information. In order to achieve the effect of domain sensitivity, the original data flow analysis calculation formula is slightly changed here to achieve a simple domain-sensitive effect for use in data flow analysis. The changed code is as follows:
[0235] For(each basic block){
[0236] OUT[B] = nil;
[0237] }
[0238] While(changes to any OUT){
[0239] For(each basic block){
[0240] IN[B]=U papredecessor of B OUT[P];
[0241] OUT[B]=TransferFunction(IN[B])
[0242] }
[0243] For(each edge E){
[0244] S for source block ofE;
[0245] IN[E]=OUT[S];
[0246] OUT[E]=EdgeTranferFunction(IN[E])
[0247] }
[0248] }
[0249] Among them, the newly added edge conversion equation for each edge does not involve specific code processing, but mainly processes domain-sensitive information. The processing logic is mainly performed by comparing the hierarchical size between the source basic block and the target basic block of the edge. First, the embodiment of the present application defines a cache of valid variables of basic blocks in the global dimension of control flow. The cache structure is as follows: Fig.19 As shown, Fig.19 It is a schematic diagram of the principle of the method for generating a data flow graph provided by an embodiment of the present application. As can be seen from the figure, the cache mainly includes two layers of mapping relationships, the first layer is mapped to the mapping of the basic block object and the level (level 0, level 1, level 2 and level 3), indicating that the basic block contains valid variables defined in which levels, and the second layer is mapped to the mapping relationship of the level to the variable name set, indicating that the level contains the set of valid variables, for example, level 0 contains variables A1, variable B1 and variable C1, level 1 contains variables A2, variable B2 and variable C2, level 2 contains variables A3, variable B3 and variable C3, level 3 contains variables A4, variable B4 and variable C4, and the value of the global cache is updated with the conversion equation of the code line. When encountering a variable definition statement, the basic block object, level information and variable name where the definition statement is located are added to the cache. In the variable conversion equation, it is calculated according to the hierarchical relationship between the source basic block and the target basic block.
[0250] In some embodiments, see Fig. 20 , Fig. 20 The schematic diagram of the principle of the method for generating a data flow graph provided in the embodiment of the present application is shown in FIG. Fig. 20The process of transferring data from the source basic block to the target basic block shown in the figure indicates that the scopes with a higher level than the target basic block have become invalid, and only the variable definitions with a level less than or equal to the target basic block are still valid. Therefore, it is necessary to transfer the value of the target basic block in the cache from the corresponding value of the source basic block with a level less than or equal to the target basic block in the global valid variable cache. Fig. 20 It can be seen that only the level 0 information in the source basic block cache is passed to the target basic block cache.
[0251] In some embodiments, when the source basic block level is larger than the target basic block, it means that some scopes will be invalid, and the variables defined in these scopes will also be invalid, so these variables need to be removed from the intermediate results of data flow analysis. The rule for removal is to find the names of some defined variables in the source basic block cache that have a higher level than the target basic block, find the corresponding information in the intermediate variables, and delete them. In this way, the domain-sensitive effect can be achieved in the process of data flow analysis.
[0252] In this way, the access threshold is significantly lowered. Users no longer need to write complex compilation scripts to access. They can scan after selecting the platform. The scanning fuse mechanism is strong and the scanning process is robust. Different from the compilation process, the failure of a certain file compilation will cause the failure of the entire process. The tool performs failure fuse at the file dimension. The failure of parsing a certain file will not affect the generation of prompts for other files. The generation of the abstract syntax tree does not rely on the time-consuming compilation process, and there is much room for improvement in terms of time consumption. The control flow parsing process and data flow optimization mechanism are concise and clear, and can quickly respond to new personalized needs of users and achieve rapid secondary development.
[0253] It is understandable that in the embodiments of the present application, related data such as target codes are involved. When the embodiments of the present application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of relevant data need to comply with relevant laws, regulations and standards of relevant countries and regions.
[0254] The following further describes the exemplary structure of the data flow graph generation device 455 provided in the embodiment of the present application implemented as a software module. In some embodiments, for example Figure 2As shown, the software modules in the data flow graph generation device 455 stored in the memory 450 may include: an acquisition module, used to acquire an abstract syntax tree of the target code, the abstract syntax tree includes multiple nodes, each of the nodes corresponds to a different sub-code segment set in the target code, and the sub-code segment set includes a target sub-code segment with target logic; a construction module, used to construct an initial control flow graph of each node based on the sub-code segment set corresponding to the node, the initial control flow graph includes an initial basic block corresponding to each sub-code segment in the sub-code segment set; a replacement module, used to acquire an extended control flow graph of each target sub-code segment, and replace the initial basic block corresponding to each target sub-code segment in each initial control flow graph with the corresponding extended control flow graph, so as to obtain a reference control flow graph corresponding to each initial control flow graph; a generation module, used to combine the abstract syntax tree and the reference control flow graph to generate a target data flow graph of the target code.
[0255] In some embodiments, the sub-code segment set includes multiple sub-code segments, and the sub-code segment includes multiple code statements. The construction module is also used to select the first sub-code segment from the multiple sub-code segments, create the first initial basic block corresponding to the first sub-code segment, and determine the first initial basic block as the first initial control flow graph; wherein the starting code statement of the first sub-code segment is the same as the starting code statement of the sub-code segment set; traverse i to perform the following processing: select the i-th sub-code segment from the multiple sub-code segments, and create the i-th initial basic block corresponding to the i-th sub-code segment; connect the i-th initial basic block with the i-1th initial control flow graph to obtain the i-th initial control flow graph; wherein the starting code statement of the i-th sub-code segment is logically associated with the ending code statement of the i-1th sub-code segment, 2≤i≤N, N is used to indicate the number of the sub-code segments in the sub-code segment set; determine the N-th initial control flow graph as the initial control flow graph of the node.
[0256] In some embodiments, the above-mentioned construction module is also used to obtain the starting code statement of the sub-code segment set, and perform the following processing for each sub-code segment: obtain the starting code statement of the sub-code segment, and compare the starting code statement of the sub-code segment with the starting code statement of the sub-code segment set to obtain a comparison result corresponding to the sub-code segment; when the comparison result indicates that the starting code statement of the sub-code segment is the same as the starting code statement of the sub-code segment set, determine the sub-code segment as the first sub-code segment.
[0257] In some embodiments, the above-mentioned construction module is also used to obtain the end code statement of the i-1th sub-code segment, and perform the following processing for each sub-code segment: obtain the start code statement of the sub-code segment, and determine the association information between the start code statement of the sub-code segment and the end code statement of the i-1th sub-code segment; when the association information indicates that there is a logical association between the start code statement of the sub-code segment and the end code statement of the i-1th sub-code segment, determine the sub-code segment as the i-th sub-code segment.
[0258] In some embodiments, the above-mentioned data flow graph generation device also includes: a code segment module, which is used to obtain the logical identification character of the target logic, and perform the following processing for each of the sub-code segment sets: extract code characters for each of the sub-code segments in the sub-code segment set to obtain a code character set corresponding to each of the sub-code segments; for each of the code character sets, compare each code character in the code character set with the logical identification character to obtain a comparison result corresponding to the code character set; when the comparison result indicates that the logical identification character exists in the code character set, determine the sub-code segment corresponding to the code character set as the target sub-code segment.
[0259] In some embodiments, the above-mentioned replacement module is also used to obtain the mapping relationship between multiple preset logical identifiers and corresponding preset control flow graphs, and perform the following processing for each target sub-code segment: obtain the logical identifier of the target logic in the target sub-code segment, and compare the logical identifier with each of the preset logical identifiers in the mapping relationship to obtain the comparison result corresponding to the logical identifier; when the comparison result indicates that the logical identifier exists in the preset logical identifier, the preset control flow graph corresponding to the logical identifier in the mapping relationship is determined as the extended control flow graph of the target sub-code segment.
[0260] In some embodiments, the above-mentioned generation module is also used to replace each of the nodes in the abstract syntax tree with the corresponding reference control flow graph to obtain the target control flow graph corresponding to the target code; perform data flow graph conversion on the target control flow graph to obtain the initial data flow graph of the target code; perform redundancy identification on the data flow in the initial data flow graph to obtain the redundant data flow in the initial data flow graph, and delete the redundant data flow in the initial data flow graph to obtain the target data flow graph.
[0261] In some embodiments, the above-mentioned generation module is further configured to perform the first data flow graph conversion on the target control flow graph to obtain the first data flow graph of the target code; traverse j and perform the following processing until the initial data flow graph of the target code is obtained: perform the jth data flow graph conversion on the target control flow graph to obtain the jth data flow graph of the target code; when the jth data flow graph is the same as the (j - 1)th data flow graph, determine the jth data flow graph as the initial data flow graph of the target code; where 1 < j ≤ M, and the Mth data flow graph is the same as the (M - 1)th data flow graph.
[0262] In some embodiments, the target control flow graph includes a plurality of target basic blocks, and the target basic blocks correspond one-to-one with target code segments in the target code. The above-mentioned generation module is further configured to run the target code and record the running data flow of each target code segment during the running process of the target code; for each target basic block, cache the running data flow of the target code segment corresponding to the target basic block into the target basic block to obtain the first updated basic block corresponding to the target basic block; replace each target basic block in the target control flow graph with the corresponding first updated basic block to obtain the first data flow graph of the target code.
[0263] In some embodiments, the initial data flow graph includes multiple data flows, and the data flow is used to indicate the data migration process from the starting point to the ending point of the data flow. The above-mentioned generation module is further configured to perform the following processing for each data flow in the initial data flow graph respectively: obtain the starting level of the starting point of the data flow in the target code and the ending level of the ending point of the data flow in the target code; compare the starting level and the ending level to obtain a level comparison result; when the level comparison result indicates that the starting level is greater than the ending level, determine the data flow as the redundant data flow.
[0264] In some embodiments, the acquisition module is further configured to acquire the target code with the target logic and perform lexical analysis on the target code to obtain multiple sub-code segment sets of the target code, where different sub-code segment sets correspond to different logical functions; perform syntax analysis on each sub-code segment set to obtain a syntax analysis result indicating the logical association between the sub-code segment sets; combine the syntax analysis result and the multiple sub-code segment sets to construct an abstract syntax tree of the target code.
[0265] The embodiment of the present application provides a computer program product, which includes a computer program or a computer executable instruction, and the computer program or the computer executable instruction is stored in a computer-readable storage medium. The processor of the electronic device reads the computer executable instruction from the computer-readable storage medium, and the processor executes the computer executable instruction, so that the electronic device executes the method for generating the data flow graph described in the embodiment of the present application.
[0266] The embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, the processor will execute the method for generating a data flow graph provided in the embodiment of the present application, for example, Figure 2 A method for generating a data flow graph is shown.
[0267] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface storage, optical disk, or CD-ROM; or it may be various electronic devices including one or any combination of the above memories.
[0268] In some embodiments, computer executable instructions may be in the form of a program, software, software module, script or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine or other unit suitable for use in a computing environment.
[0269] As an example, computer-executable instructions may, but need not, correspond to a file in a file system, may be stored as part of a file storing other programs or data, such as in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files storing one or more modules, subroutines, or code portions).
[0270] As an example, computer executable instructions may be deployed to be executed on one electronic device, or on multiple electronic devices located at one site, or on multiple electronic devices distributed at multiple sites and interconnected by a communication network.
[0271] In summary, the embodiments of the present application have the following beneficial effects:
[0272] (1) By obtaining the abstract syntax tree of the target code, based on the set of sub-code segments corresponding to each node, the initial control flow graph of the node is constructed, and the extended control flow graph of each target sub-code segment is obtained, and the initial basic block corresponding to each target sub-code segment in the initial control flow graph is replaced with the corresponding extended control flow graph, and the reference control flow graph corresponding to each initial control flow graph is obtained. The target flow graph of the target code is generated by combining the abstract syntax tree and the reference control flow graph. By replacing the initial basic block corresponding to the target sub-code segment of the initial control flow graph with the corresponding extended control flow graph, the control flow graph corresponding to each target sub-code segment in the obtained reference control flow graph is more accurate. Then, by combining the abstract syntax tree and the reference control flow graph, the accuracy of the target data flow graph of the generated target code can be effectively improved, so that the obtained target data flow graph can accurately reflect the data flow in the target code, and the accuracy of the data flow graph generation is effectively improved.
[0273] (2) By obtaining the abstract syntax tree of the target code, the abstract representation of the target code can be accurately realized. Since the abstract syntax tree contains the complete source program information of the target code, it lays an accurate data support for the subsequent accurate determination of the target data flow graph of the target code.
[0274] (3) By creating the first initial basic block corresponding to the first sub-code segment, and determining the first initial basic block as the first initial control flow graph, and traversing i to perform the following processing: select the ith sub-code segment from multiple sub-code segments, and create the ith initial basic block corresponding to the ith sub-code segment; connect the ith initial basic block with the i-1th initial control flow graph to obtain the ith initial control flow graph, and determine the Nth initial control flow graph as the initial control flow graph of the node, so that the control flow transfer process of each basic block in the initial control flow graph of the node is consistent with the logical transfer process of each sub-code segment in the sub-code segment set corresponding to the node, thereby effectively improving the accuracy of the determined initial control flow graph.
[0275] (4) By obtaining the logic identification character of the target logic, and for each sub-code segment set, extracting the code characters of each sub-code segment in the sub-code segment set respectively, obtaining the code character set corresponding to each sub-code segment respectively, and comparing each code character in the code character set with the logic identification character respectively, to obtain the comparison result corresponding to the code character set; when the comparison result indicates that the logic identification character exists in the code character set, the sub-code segment corresponding to the code character set is determined as the target sub-code segment, thereby accurately determining whether the sub-code segment is the target sub-code segment with the target logic function by comparing the logic identification character of the target logic with the code character of the sub-code segment, thereby facilitating the subsequent processing of the target sub-code segment to obtain the reference control flow graph corresponding to the initial control flow graph respectively, thereby effectively improving the accuracy of the reference control flow graph determined subsequently.
[0276] (5) By comparing the logical identifier of the target logic in the target sub-code segment with each preset logical identifier in the mapping relationship, a comparison result corresponding to the logical identifier is obtained. When the comparison result indicates that the logical identifier exists in the preset logical identifier, the preset control flow graph corresponding to the logical identifier in the mapping relationship is determined as the extended control flow graph of the target sub-code segment, thereby realizing accurate search of the extended control flow graph of the target sub-code segment.
[0277] (6) The initial basic blocks corresponding to the target sub-code segments in each initial control flow graph are replaced with the corresponding extended control flow graphs to obtain the reference control flow graphs corresponding to the initial control flow graphs, thereby achieving a control flow graph expression that refines the target sub-code segments with relatively complex control flows in the initial control flow graphs, so that the determined reference control flow graph can more accurately express the control flow of the target code, thereby effectively improving the accurate expression of the control flow of the target code.
[0278] (7) Since the determined reference control flow graph can more accurately express the control flow of the target code, the abstract syntax tree of the target code can accurately implement the abstract representation of the target code. By replacing each node in the abstract syntax tree with the corresponding reference control flow graph, the target control flow graph corresponding to the target code is obtained, so that the obtained target control flow graph can accurately express the control flow of the target code and accurately implement the abstract representation of the target code.
[0279] (8) By performing at least two data flow graph conversions on the target control flow graph, when the data flow graphs obtained from two adjacent data flow graph conversions are the same, the data flow graph obtained from the previous data flow graph conversion is determined as the initial data flow graph of the target code.
[0280] (9) By performing at least two data flow graph conversions on the target control flow graph, when the data flow graphs obtained from two adjacent data flow graph conversions are the same, the data flow graph obtained from the previous data flow graph conversion is determined as the initial data flow graph of the target code, thereby reducing the probability of inaccurate data flow graph caused by unstable data flow during the data flow graph conversion process, thereby effectively improving the accuracy of the determined data flow graph.
[0281] (10) During the running of the target code, each target code segment of the target code will generate a running data flow during the running process. By recording the generated running data flow and adding the recorded running data flow to the cache of the corresponding target basic block, an updated basic block is obtained, and each target basic block in the target control flow graph is replaced with the corresponding updated basic block, thereby realizing the conversion from the control flow graph to the data flow graph.
[0282] (11) By redundancy identification of data flows in the initial data flow graph, redundant data flows in the initial data flow graph are obtained, and the redundant data flows in the initial data flow graph are deleted to obtain a target data flow graph, so that the obtained target data flow graph does not contain redundant data flows, thereby effectively improving the accuracy of the target data flow graph.
[0283] The above is only an embodiment of the present application and is not intended to limit the protection scope of the present application. Any modifications, equivalent substitutions and improvements made within the spirit and scope of the present application are included in the protection scope of the present application.
Claims
1. A method for generating a data flow graph, characterized in that: The method comprises: Obtaining an abstract syntax tree of the target code, wherein the abstract syntax tree includes a plurality of nodes, each of the nodes corresponds to a different set of sub-code segments in the target code, and the sub-code segment set includes a target sub-code segment having a target logic; For each of the nodes, based on the sub-code segment set corresponding to the node, construct an initial control flow graph of the node, wherein the initial control flow graph includes initial basic blocks corresponding to each sub-code segment in the sub-code segment set; Obtaining an extended control flow graph of each of the target sub-code segments, and replacing an initial basic block corresponding to each of the target sub-code segments in each of the initial control flow graphs with the corresponding extended control flow graph, to obtain a reference control flow graph corresponding to each of the initial control flow graphs; The target data flow graph of the target code is generated by combining the abstract syntax tree and the reference control flow graph.
2. The method according to claim 1, characterized in that The sub-code segment set includes a plurality of sub-code segments, and the sub-code segment includes a plurality of code statements. The initial control flow graph of the node is constructed based on the sub-code segment set corresponding to the node, including: Selecting a first sub-code segment from the multiple sub-code segments, creating a first initial basic block corresponding to the first sub-code segment, and determining the first initial basic block as a first initial control flow graph; The starting code statement of the first sub-code segment is the same as the starting code statement of the sub-code segment set; Traversing i, the following processing is performed: selecting the i-th sub-code segment from the multiple sub-code segments, and creating the i-th initial basic block corresponding to the i-th sub-code segment; connecting the i-th initial basic block with the i-1-th initial control flow graph to obtain the i-th initial control flow graph; There is a logical association between the start code statement of the i-th sub-code segment and the end code statement of the i-1-th sub-code segment, 2≤i≤N, and N is used to indicate the number of the sub-code segments in the sub-code segment set; The Nth initial control flow graph is determined as the initial control flow graph of the node.
3. The method according to claim 2, characterized in that The selecting the first sub-code segment from the multiple sub-code segments comprises: Obtain the starting code statement of the sub-code segment set, and perform the following processing on each of the sub-code segments: Obtaining a starting code statement of the sub-code segment, and comparing the starting code statement of the sub-code segment with a starting code statement of the sub-code segment set to obtain a comparison result corresponding to the sub-code segment; When the comparison result indicates that the starting code statement of the sub-code segment is the same as the starting code statement of the sub-code segment set, the sub-code segment is determined as the first sub-code segment.
4. The method according to claim 2, characterized in that: The step of selecting the i-th sub-code segment from the plurality of sub-code segments comprises: Get the end code statement of the i-1th sub-code segment, and perform the following processing on each of the sub-code segments: Obtaining a start code statement of the sub-code segment, and determining association information between the start code statement of the sub-code segment and the end code statement of the (i-1)th sub-code segment; When there is a logical association indicated by the associated information between the starting code statement of the sub-code segment and the ending code statement of the (i - 1)-th sub-code segment, determine the sub-code segment as the i-th sub-code segment.
5. The method according to claim 1, characterized in that Before obtaining the extended control flow graphs of the target sub-code segments, the method further includes: Obtain the logical identification characters of the target logic, and perform the following processing for each of the sub-code segment sets respectively: Extract code characters from each of the sub-code segments in the sub-code segment set to obtain code character sets corresponding to the sub-code segments respectively; For each of the code character sets, compare each code character in the code character set with the logical identification characters to obtain the comparison results corresponding to the code character sets; When the comparison results indicate that the logical identification characters exist in the code character set, determine the sub-code segment corresponding to the code character set as the target sub-code segment.
6. The method according to claim 1, characterized in that The obtaining of the extended control flow graphs of the target sub-code segments includes: Obtain the mapping relationships between multiple preset logical identifiers and corresponding preset control flow graphs, and perform the following processing for each of the target sub-code segments respectively: Obtain the logical identifier of the target logic in the target sub-code segment, and compare the logical identifier with each of the preset logical identifiers in the mapping relationship to obtain the comparison results corresponding to the logical identifier; When the comparison results indicate that the logical identifier exists in the preset logical identifiers, determine the preset control flow graph corresponding to the logical identifier in the mapping relationship as the extended control flow graph of the target sub-code segment.
7. The method according to claim 1, characterized in that The generating of the target data flow graph of the target code by combining the abstract syntax tree and the reference control flow graph includes: Replace each of the nodes in the abstract syntax tree with the corresponding reference control flow graph to obtain the target control flow graph corresponding to the target code; Perform data flow graph conversion on the target control flow graph to obtain the initial data flow graph of the target code; Identify redundant data flows in the initial data flow graph, obtain the redundant data flows in the initial data flow graph, and delete the redundant data flows in the initial data flow graph to obtain the target data flow graph.
8. The method according to claim 7, characterized in that The performing of data flow graph conversion on the target control flow graph to obtain the initial data flow graph of the target code includes: Perform the first data flow graph conversion on the target control flow graph to obtain the first data flow graph of the target code; Traverse j and perform the following processing until the initial data flow graph of the target code is obtained: Perform the j-th data flow graph conversion on the target control flow graph to obtain the j-th data flow graph of the target code; When the j-th data flow graph is the same as the (j - 1)-th data flow graph, determine the j-th data flow graph as the initial data flow graph of the target code; where 1 < j ≤ M, and the M-th data flow graph is the same as the (M - 1)-th data flow graph.
9. The method according to claim 8, characterized in that The target control flow graph includes a plurality of target basic blocks, and the target basic blocks correspond to target code segments in the target code one by one. The first data flow graph conversion is performed on the target control flow graph to obtain the first data flow graph of the target code, including: Running the target code, and in the process of running the target code, recording the running data flow of each target code segment during the running process; For each of the target basic blocks, cache the running data flow of the target code segment corresponding to the target basic block into the target basic block to obtain a first updated basic block corresponding to the target basic block; Each of the target basic blocks in the target control flow graph is replaced with the corresponding first updated basic block to obtain the first data flow graph of the target code.
10. The method according to claim 7, characterized in that The initial data flow graph includes a plurality of data flows, wherein the data flows are used to indicate a data migration process from a starting point of the data flow to an end point of the data flow; The redundancy identification of the data flows in the initial data flow graph to obtain the redundant data flows in the initial data flow graph includes: The following processing is performed for each of the data flows in the initial data flow graph: Obtaining a starting point level of the starting point of the data stream in the target code, and an ending point level of the ending point of the data stream in the target code; Compare the starting point level with the end point level to obtain a level comparison result; When the level comparison result indicates that the start level is greater than the end level, the data stream is determined as the redundant data stream.
11. The method according to claim 1, characterized in that: The step of obtaining an abstract syntax tree of the target code comprises: Obtaining a target code having the target logic, and performing lexical analysis on the target code to obtain a plurality of sub-code segment sets of the target code, wherein different sub-code segment sets correspond to different logical functions; Performing grammatical analysis on each of the sub-code segment sets to obtain a grammatical analysis result indicating a logical association between each of the sub-code segment sets; An abstract syntax tree of the target code is constructed by combining the syntax analysis result and the plurality of sub-code segment sets.
12. A device for generating a data flow graph, characterized in that: The device comprises: An acquisition module, used for acquiring an abstract syntax tree of a target code, wherein the abstract syntax tree includes a plurality of nodes, each of the nodes corresponds to a different set of sub-code segments in the target code, and the sub-code segment set includes a target sub-code segment having a target logic; A construction module, configured to construct, for each of the nodes, an initial control flow graph of the node based on a set of sub-code segments corresponding to the node, wherein the initial control flow graph includes initial basic blocks corresponding one-to-one to each sub-code segment in the set of sub-code segments; A replacement module is used to obtain an extended control flow graph of each target sub-code segment, and replace the initial basic block corresponding to each target sub-code segment in each initial control flow graph with the corresponding extended control flow graph to obtain a reference control flow graph corresponding to each initial control flow graph; A generation module is used to combine the abstract syntax tree and the reference control flow graph to generate a target data flow graph of the target code.
13. An electronic device, characterized in that: The electronic device comprises: A memory for storing computer executable instructions or computer programs; A processor, used to implement the method for generating a data flow graph as described in any one of claims 1 to 11 when executing computer executable instructions or computer programs stored in the memory.
14. A computer-readable storage medium storing computer-executable instructions, characterized in that: When the computer executable instructions are executed by a processor, the method for generating a data flow graph as described in any one of claims 1 to 11 is implemented.
15. A computer program product comprising a computer program or computer executable instructions, characterized in that: When the computer program or computer executable instructions are executed by a processor, the method for generating a data flow graph as described in any one of claims 1 to 11 is implemented.
Citation Information
Cited By
Source code hierarchy positioning method, electronic equipment and medium
CN120743345A
Source code level location methods, electronic devices and media
CN120743345B