Method and device for processing syntax tree, electronic equipment and storage medium

By using preconfigured node mapping tables to map and combine the syntax tree structures of different code languages, a unified structure syntax tree is generated, which solves the problem of redundant information of syntax trees and the difference in languages ​​in the existing technology, and improves the efficiency of static analysis.

CN119938049APending Publication Date: 2025-05-06TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311466450.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-03
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing technology generates a lot of redundant information and has a low effective information density, which leads to high cost and difficulty in the analysis and implementation of static code scanning tools. The structure of abstract syntax trees in different languages ​​is different, making static code scanning tools difficult to be compatible with multiple languages.

Method used

By obtaining the source code file, parse the syntax tree, and use the preconfigured node mapping table to map and combine syntax nodes of different code languages ​​to generate a syntax tree with a unified structure.

Benefits of technology

It realizes the streamlined unification of the syntax tree structure of different code languages, reduces the cost and difficulty of subsequent static analysis, and improves the efficiency of static analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938049A_ABST
    Figure CN119938049A_ABST
Patent Text Reader

Abstract

The invention provides a syntax tree processing method and device, electronic equipment, a computer program product and a computer readable storage medium. The method comprises the steps of obtaining a source code file; analyzing the source code file to obtain a first syntax tree; a pre-configured node mapping table is obtained, and the node mapping table comprises to-be-mapped grammar nodes of multiple code languages and self-defined grammar nodes in incidence relation with the to-be-mapped grammar nodes; the first syntax tree is traversed, a second syntax node is inquired from a node mapping table based on the traversed first syntax node of the first syntax tree, and the second syntax node is a self-defined syntax node having an association relationship with the first syntax node in the node mapping table; and combining the second syntax nodes having the association relationship with each first syntax node according to the traversal sequence to obtain a second syntax tree. Through the method, syntax tree structures of different code languages can be simplified and unified, and the efficiency of subsequent static analysis is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to computer technology, and in particular to a method, device, electronic device and storage medium for processing a syntax tree. Background Art

[0002] In the field of computer static code scanning tools, syntax tree processing is an important research direction. The abstract syntax tree is an intermediate structure for converting source code into executable binary files in the computer. Static code scanning tools perform analysis based on the abstract syntax tree. However, the existing technology generates syntax trees with a lot of redundant information and a low density of effective information, resulting in high cost and difficulty in analysis implementation. In addition, there are differences in the abstract syntax tree structures of different languages, which makes it difficult for static code scanning tools to be compatible with multiple languages. Summary of the invention

[0003] The embodiments of the present application provide a syntax tree processing method, device, electronic device, computer program product and computer-readable storage medium, which can simplify and unify the syntax tree structure of different code languages ​​and improve the efficiency of subsequent static analysis.

[0004] The technical solution of the embodiment of the present application is implemented as follows:

[0005] The present application provides a method for processing a syntax tree, the method comprising:

[0006] Obtain a source code file, wherein the source code file is developed based on a first code language, and the first code language is any code language;

[0007] Parsing the source code file to obtain a first syntax tree;

[0008] Obtaining a preconfigured node mapping table, wherein the node mapping table includes syntax nodes to be mapped in multiple code languages ​​and user-defined syntax nodes associated with the syntax nodes to be mapped, the multiple code languages ​​including the first code language;

[0009] Traversing the first syntax tree, and obtaining a second syntax node by querying from the node mapping table based on the first syntax node of the traversed first syntax tree, wherein the second syntax node is the custom syntax node in the node mapping table that has the association relationship with the first syntax node;

[0010] The second syntax nodes that have the association relationship with each of the first syntax nodes are combined in the order of traversal to obtain a second syntax tree.

[0011] The present application provides a syntax tree processing device, including:

[0012] An acquisition module, configured to acquire a source code file, wherein the source code file is developed based on a first code language, and the first code language is any code language;

[0013] Obtaining a preconfigured node mapping table, wherein the node mapping table includes syntax nodes to be mapped in multiple code languages ​​and user-defined syntax nodes associated with the syntax nodes to be mapped, the multiple code languages ​​including the first code language;

[0014] A parsing module, used for parsing the source code file to obtain a first syntax tree;

[0015] A traversal module, used to traverse the first syntax tree, and query from the node mapping table to obtain a second syntax node based on the first syntax node of the traversed first syntax tree, wherein the second syntax node is the custom syntax node in the node mapping table that has the association relationship with the first syntax node;

[0016] The processing module is used to combine the second syntax nodes that have the association relationship with each of the first syntax nodes in the order of traversal to obtain a second syntax tree.

[0017] An embodiment of the present application provides an electronic device, the electronic device comprising:

[0018] A memory for storing computer executable instructions;

[0019] The processor is used to implement the syntax tree processing method provided in the embodiment of the present application when executing the computer executable instructions stored in the memory.

[0020] An embodiment of the present application provides a computer-readable storage medium storing a computer program or computer-executable instructions for implementing a syntax tree processing method provided in an embodiment of the present application when executed by a processor.

[0021] An embodiment of the present application provides a computer program product, including a computer program or computer executable instructions. When the computer program or computer executable instructions are executed by a processor, the method for processing a syntax tree provided in the embodiment of the present application is implemented.

[0022] The embodiments of the present application have the following beneficial effects:

[0023] By establishing a mapping table for the to-be-mapped syntax nodes of multiple code languages ​​and the associated custom syntax nodes, the original nodes (i.e., the first syntax nodes) in the syntax trees of different code languages ​​can all be queried to obtain the corresponding custom syntax nodes by querying the mapping table, and then the queried custom syntax nodes are sequentially combined to form a second syntax tree. In this way, no matter which code language the first syntax tree (i.e., the first code language) is, it can be converted into a code tree with a unified structure (i.e., the second code tree) by the above method, thereby facilitating the subsequent analysis of the syntax tree with a unified structure and improving the efficiency of static analysis of the source code. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 It is a structural diagram of the system architecture for processing syntax trees provided in an embodiment of the present application;

[0025] Figure 2 It is a structural schematic diagram of a syntax tree processing device provided in an embodiment of the present application;

[0026] Figure 3A It is a first flow chart of the method for processing a syntax tree provided in an embodiment of the present application;

[0027] Figure 3B is a schematic diagram of a process for generating a first syntax tree provided in an embodiment of the present application;

[0028] Figure 3C is a schematic diagram of a process for determining a first grammar node provided in an embodiment of the present application;

[0029] Figure 3D It is a schematic diagram of a process for obtaining a pre-configured node mapping table provided in an embodiment of the present application;

[0030] Figure 3E It is a flowchart of obtaining a pre-configured custom syntax node list provided in an embodiment of the present application;

[0031] Figure 3F is a schematic diagram of a process of traversing a first syntax tree provided in an embodiment of the present application;

[0032] Figure 3G A schematic diagram of the process of traversing a syntax tree in pre-order provided in an embodiment of the present application;

[0033] Figure 4 is a schematic diagram of a process for determining a pre-configured attribute type provided in an embodiment of the present application;

[0034] Figure 5A This is a schematic diagram of the first defect scanning process provided by an embodiment of the present application;

[0035] Figure 5BIt is a schematic diagram of the second defect scanning process provided by an embodiment of the present application;

[0036] Fig. 6A This is a schematic diagram of the Java native syntax tree structure provided by an embodiment of the present application;

[0037] Figure 6B It is a schematic diagram of the CPP native syntax tree structure provided in an embodiment of the present application;

[0038] Figure 6C It is a schematic diagram of the structure of a custom syntax tree provided in an embodiment of the present application;

[0039] Fig. 7A It is a second flow chart of the method for processing a syntax tree provided in an embodiment of the present application;

[0040] Figure 7B It is a third flow chart of the method for processing a syntax tree provided in an embodiment of the present application;

[0041] Figure 7C It is a schematic diagram of the first structure of the native syntax tree provided in an embodiment of the present application;

[0042] Fig.7D It is a schematic diagram of the first structure of the native syntax tree provided in an embodiment of the present application;

[0043] Fig. 7E It is a schematic diagram of the custom syntax tree structure provided in an embodiment of the present application. DETAILED DESCRIPTION

[0044] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings. The described embodiments should not be regarded as limiting the present application. All other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of this application.

[0045] In the following description, reference is made to “some embodiments” which describe a subset of all possible embodiments, but it will be appreciated that the “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0046] In the following description, the terms "first\second\third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described here can be implemented in an order other than that illustrated or described here.

[0047] Unless otherwise defined, all technical and scientific terms used in the embodiments of the present application have the same meanings as those commonly understood by those skilled in the art. The terms used in the embodiments of the present application are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.

[0048] Before further describing the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations.

[0049] 1) Another Tool for Language Recognition (ANTLR) is an open source framework that includes recognizers, parsers, and translators that automatically construct custom languages ​​based on grammatical descriptions for Java, C++, C#, and other code languages.

[0050] 2) Abstract Syntax Tree (ST) is an abstract representation of the syntax structure of the source code. It represents the syntax structure of the programming language in a tree form. Each node on the tree represents a structure in the source code.

[0051] 3) Static analysis: It is an analysis method that obtains endogenous variables based on the given exogenous variable values. It is an analysis method that conducts a comprehensive comparative analysis of the source code. Specifically, it analyzes the code content without running the code to discover defects and risks in the code.

[0052] In the prior art syntax tree processing method, ANTLR or similar syntax analyzers are usually used to generate syntax analysis code of the corresponding language based on the syntax file of the specified language. Then the Programming Mistake Detector (PMD) parses the project source code based on the syntax analysis code, and traverses the syntax nodes by using the syntax tree observer or visitor framework of ANTLR or similar syntax analyzers, thereby implementing the scanning rules of the corresponding nodes.

[0053] However, the above method has the following technical problems:

[0054] First, language differences make it impossible to reuse the analysis framework. The grammars of different languages ​​are quite different, so the syntax analysis codes of different languages ​​are also quite different. Furthermore, the framework codes or rule codes based on syntax tree traversal are also very different. The rules finally developed have language differences and cannot be universal.

[0055] Secondly, the complexity of the syntax tree structure makes rule development difficult. Code compilation to generate an abstract syntax tree is a part of the compilation process, and the compilation process itself is to ultimately generate machine code binary data for machine execution. Therefore, the abstract syntax tree itself carries information for subsequent execution, while static analysis itself discovers code defects by analyzing project code. The two have different purposes, and therefore have different requirements for the structure and content of the abstract syntax tree. It is difficult to analyze problems based on the original abstract syntax tree.

[0056] Based on the above analysis of the technical problems in the prior art, the applicant designs a custom abstract syntax tree structure to carry the core information required for static analysis, and then analyzes the analysis codes of different languages ​​and performs mapping conversion to convert the code files of different languages ​​into abstract syntax tree data of a unified structure. This eliminates the language differences in rule development. The embodiment of the present application provides a syntax tree processing method that simplifies and unifies the syntax tree structures of different code languages ​​and improves the efficiency of subsequent static analysis.

[0057] The embodiments of the present application provide a method, device, electronic device, computer-readable storage medium and computer program product for processing a syntax tree, which can simplify and unify the syntax tree structure of different code languages ​​and improve the efficiency of static analysis. The following describes an exemplary application of an electronic device for processing a syntax tree provided by an embodiment of the present application. The device provided by an embodiment of the present application can be implemented as various types of user terminals such as a laptop computer, a tablet computer, a desktop computer, a set-top box, a mobile device (e.g., a mobile phone, a portable music player, a personal digital assistant, a dedicated messaging device, a portable gaming device), a smart phone, a smart speaker, a smart watch, a smart TV, and a vehicle-mounted terminal, and can also be implemented as a server.

[0058] See also Figure 1 , Figure 1 1 is a schematic diagram of the architecture of a syntax tree processing system 100 provided in an embodiment of the present application, for implementing a processing application supporting a syntax tree, for example, Figure 1 The syntax tree processing system 100 involves a server 200, a network 300 and a terminal 400. The terminal 400 is connected to the server 200 via the network 300. The network 300 may be a wide area network or a local area network, or a combination of the two.

[0059] In some embodiments, a code analysis tool can be run in the terminal 400 to obtain a source code file in a first code language from a local development environment, where the first code language can be any code language, and send a processing request to the server 200 through the network 300, carrying the source code file. The server 200 extracts a first syntax tree based on the first code language from the source code file, converts it into a second syntax tree of a unified structure, and sends the second syntax tree to the terminal 400. The terminal 400 presents the second syntax tree in the graphical interface 410 (graphical interface 410-1 is shown as an example). The developer can call the code analysis tool in the terminal 400 to perform static analysis on the second syntax tree to detect whether there are problems or defects in the source code file.

[0060] In other embodiments, a code analysis tool can be run in the terminal 400 to obtain a source code file of a first code language from a local development environment, where the first code language can be any code language. A first syntax tree is extracted from the source code file by the code analysis tool, and a process carrying the first syntax tree is sent to the server 200. The server 200 converts the first syntax tree into a second syntax tree of the same structure and sends it to the terminal. The terminal 400 presents the second syntax tree on the graphical interface 410 (graphical interface 410-1 is shown as an example).

[0061] In some other embodiments, the terminal 400 runs a code analysis tool to obtain a source code file of a first code language from a local development environment, where the first code language can be any code language. The first syntax tree is extracted from the source code file by the code analysis tool, and then each first syntax node in the first syntax tree is converted into a corresponding second syntax node with a mapping relationship by querying a preconfigured node mapping table, and finally the first syntax tree is converted into a second syntax tree with a unified structure.

[0062] In some embodiments, server 200 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDN), as well as big data and artificial intelligence platforms.

[0063] The solution of the embodiment of the present application can be implemented through artificial intelligence. Artificial Intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that the machines have the functions of perception, reasoning and decision-making.

[0064] For example, the language type of source code files in different code languages ​​can be analyzed through a machine learning model. By collecting source code file samples in different code languages ​​and marking the type label of the code language for the source code files of each code language, the collected source code file samples are called by the machine learning model to obtain the predicted type label, and the difference between the predicted type label and the marked type label (for example, the difference between the embedding vectors of the two) is brought into the loss function to calculate the error signal, and the error signal is back-propagated in the machine learning model through the back-propagation algorithm to update the parameters of the machine learning model. Among them, the machine learning model can be a statistical model, such as a hidden Markov model, or an algorithmic model, such as a conditional random field model, a probability model, such as a maximum entropy model, or a linear model, such as a logistic regression model, etc.; the loss function can be a mean square error loss function, a cross entropy loss function, a KL divergence loss function, etc.

[0065] See also Figure 2 , Figure 2 is a schematic diagram of the structure of the terminal 400 provided in an embodiment of the present application, Figure 2 The terminal 400 shown includes: at least one processor 410, a memory 450, at least one network interface 420 and a user interface 430. The various components in the terminal 400 are coupled together via a bus system 440. It is understood that the bus system 440 is used to achieve connection and communication between these components. In addition to the data bus, the bus system 440 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, the bus system 440 is not shown in FIG. Figure 2 Various buses are labeled as bus system 440 .

[0066] The processor 410 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0067] The user interface 430 includes one or more output devices 431 that enable presentation of media content, including one or more speakers and / or one or more visual display screens. The user interface 430 also includes one or more input devices 432, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.

[0068] The memory 450 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard disk drives, optical disk drives, etc. The memory 450 may optionally include one or more storage devices that are physically remote from the processor 410.

[0069] The memory 450 includes a volatile memory or a non-volatile memory, and may also include both volatile and non-volatile memories. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 450 described in the embodiments of the present application is intended to include any suitable type of memory.

[0070] In some embodiments, memory 450 can store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as exemplarily described below.

[0071] Operating system 451, including system programs for processing various basic system services and performing hardware-related tasks, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks;

[0072] A network communication module 452, used to reach other electronic devices via one or more (wired or wireless) network interfaces 420, exemplary network interfaces 420 include: Bluetooth, Wireless Compatibility Certification (WiFi), and Universal Serial Bus (USB), etc.;

[0073] a presentation module 453 for enabling presentation of information via one or more output devices 431 (e.g., display screen, speaker, etc.) associated with the user interface 430 (e.g., a user interface for operating peripherals and displaying content and information);

[0074] The input processing module 454 is used to detect one or more user inputs or interactions from one of the one or more input devices 432 and translate the detected inputs or interactions.

[0075] In some embodiments, the device provided in the embodiments of the present application can be implemented in software. Figure 2 The processing device 455 of the syntax tree stored in the memory 450 is shown, which can be software in the form of a program and a plug-in, etc., including the following software modules: an acquisition module 4551, a parsing module 4552, a traversal module 4553 and a processing module 4554. These modules are logical, so they can be arbitrarily combined or further split according to the functions implemented. The functions of each module will be explained below.

[0076] In some embodiments, the terminal or server can implement the processing method of the syntax tree provided in the embodiment of the present application by running various computer executable instructions or computer programs. For example, computer executable instructions can be commands, machine instructions or software instructions at the microprogram level. The computer program can be a native program or software module in the operating system; it can be a local (Native) application (APPlication, APP), that is, a program that needs to be installed in the operating system to run, such as code analysis software. In short, the above-mentioned computer executable instructions can be instructions in any form, and the above-mentioned computer program can be an application, module or plug-in in any form.

[0077] For example, the syntax tree processing method provided in the embodiment of the present application can be integrated into a code analysis tool for static analysis in the form of a module, plug-in or dedicated program, as the basic part of project code parsing, to provide data for subsequent defect problem matching engines.

[0078] The method for processing the syntax tree provided in the embodiment of the present application will be explained in combination with the exemplary application and implementation of the terminal provided in the embodiment of the present application.

[0079] The following describes the syntax tree processing method provided in the embodiment of the present application. As mentioned above, the electronic device that implements the syntax tree processing method in the embodiment of the present application may be a terminal, a server, or a combination of the two. Therefore, the execution subject of each step will not be repeatedly described below.

[0080] See also Figure 3A , Figure 3A is a first flow chart of the method for processing a syntax tree provided in an embodiment of the present application, which will be combined with Figure 3A The steps shown are explained.

[0081] In step 101, a source code file is obtained, wherein the source code file is developed based on a first code language, and the first code language is any code language.

[0082] In some embodiments, the first code language can be any code language, such as Java, C#, Python2, Python3, Go, JavaScript, TypeScript, C++, Swift, PHP, and Dart.

[0083] In step 102, the source code file is parsed to obtain a first syntax tree.

[0084] In some embodiments, see Figure 3B , Figure 3B is a schematic diagram of a process for generating a first syntax tree provided in an embodiment of the present application, Figure 3A Step 102 shown may be performed by Figure 3B The processing from step 1021 to step 1023 is implemented as described in detail below.

[0085] In step 1021, the source code file is parsed to obtain a syntax analysis code corresponding to the first code language of the source code file.

[0086] In some embodiments, a syntax analyzer may be used to perform syntax analysis on the source code file to obtain a syntax analysis code corresponding to the first code language of the source code file.

[0087] In some embodiments, syntax parsing is to check whether the word sequence output by the lexical analyzer conforms to the grammatical rules of the source code language. The word sequence generated by the lexical analysis of the source code file is read to see whether it satisfies the grammar of the source code language. For example, in the process of syntax parsing the source code file of the C language through the syntax analyzer, the statement "int double=" appears, which is an error that does not conform to the language syntax specification. Syntax errors in the code can be detected and repaired through syntax parsing to obtain the syntax analysis code of the source code file corresponding to the first code language.

[0088] In step 1022, a plurality of first syntax nodes in a first syntax tree are determined based on the syntax analysis code.

[0089] In some embodiments, see Figure 3C , Figure 3C is a schematic diagram of a process for determining a first grammar node provided in an embodiment of the present application, Figure 3B Step 1022 shown may be performed by Figure 3C The processing implementation of steps 10221 to 10223 is described in detail below.

[0090] In step 10221, a preconfigured list of syntax nodes to be mapped is obtained, wherein the list of syntax nodes to be mapped includes multiple syntax nodes to be mapped in multiple code languages.

[0091] In some embodiments, the terminal device can provide developers with a setting function for setting syntax nodes to be mapped through a code analysis tool, obtain a list of syntax nodes to be mapped preconfigured by the developer according to different rules of multiple code languages, and assist the developer to quickly set the list of nodes to be mapped through human-computer interaction.

[0092] As an example, the preconfigured list of syntax nodes to be mapped can be obtained by the following method: first, display the syntax node setting interface to be mapped in the code analysis tool of the terminal. Secondly, in response to the syntax node setting operation to be mapped in the developer setting interface, obtain multiple syntax nodes to be mapped preconfigured by the developer in the setting interface; finally, combine the multiple syntax nodes to be mapped to obtain the syntax node list to be mapped, and display the list of syntax nodes to be mapped in the code analysis tool.

[0093] In step 10222, the syntax analysis code is traversed, and the syntax nodes obtained by the traversal are compared with the syntax nodes to be processed in the syntax nodes to be mapped.

[0094] In some embodiments, each syntax node obtained through traversal is compared with the syntax node to be processed in the syntax nodes to be mapped.

[0095] As an example, when a syntax node with a type definition appears when traversing the syntax analysis code, the type definition is used as the attribute type of the syntax node and compared with the syntax node to be processed in the syntax node to be mapped. If the attribute type of the syntax node to be processed in the syntax node to be mapped is also a type definition, the comparison is consistent.

[0096] As an example, when a syntax node of an execution statement appears while traversing the syntax analysis code, such as a switch statement, the type definition is used as the attribute type of the syntax node and compared with the syntax node to be processed in the syntax node to be mapped. If the attribute type of the syntax node to be processed in the syntax node to be mapped is also an execution statement, the comparison is consistent.

[0097] In step 10223, in response to the comparison being consistent, the grammar node obtained by traversal is used as the first grammar node.

[0098] In some embodiments, the node among the grammar nodes obtained through traversal that is consistent with the grammar node to be mapped is used as the first grammar node.

[0099] Continue to see Figure 3B , and the description continues with step 1022 above.

[0100] In step 1023, the first syntax nodes of the plurality of first syntax trees are combined according to the logical relationship of the syntax analysis code to obtain a first syntax tree.

[0101] For example, see Fig. 7A , Fig. 7A is a second flow chart of the method for processing a syntax tree provided in an embodiment of the present application, Fig. 7A The source code file 701A in the source code file 701A may be a CPP project code file or a Java code file. Fig. 7A The Antlr parser 702A in the source code parses the source code to obtain Fig. 7A The Antlr syntax tree structure 703A in the source code file 701A is a CPP project code, and the corresponding Antlr CPP syntax tree structure in the Antlr syntax tree structure 703A is parsed; when the source code 701A is a Java project code, the corresponding AntlrJava syntax tree structure in the Antlr syntax tree structure 703A is parsed.

[0102] Continue to see Figure 3A , and the description continues with step 102 above.

[0103] In step 103, a preconfigured node mapping table is obtained, wherein the node mapping table includes syntax nodes to be mapped in multiple code languages ​​and user-defined syntax nodes associated with the syntax nodes to be mapped, and the multiple code languages ​​include a first code language.

[0104] In some embodiments, see Figure 3D , Figure 3D is a schematic diagram of a process for obtaining a pre-configured node mapping table provided in an embodiment of the present application. Figure 3A Step 103 shown can be performed by executing Figure 3D The processing from step 1031 to step 1034 is implemented as described in detail below.

[0105] In step 1031, a preconfigured list of syntax nodes to be mapped is obtained, wherein the list of syntax nodes to be mapped includes multiple syntax nodes to be mapped in multiple code languages.

[0106] For the specific implementation of step 1031 in some embodiments, please refer to the content of step 10221 above, which will not be repeated here.

[0107] In step 1032, a preconfigured custom syntax node list is obtained, wherein the custom syntax node list includes a plurality of custom syntax nodes.

[0108] In some embodiments, the terminal device can provide developers with a setting function for configuring custom syntax nodes through a code analysis tool, obtain a custom syntax node list pre-configured by the developer according to different rules of multiple code languages, and assist the developer in quickly setting the custom syntax node list through human-computer interaction. The custom syntax node list includes the names of multiple custom syntax nodes and the attribute values ​​corresponding to each custom syntax node.

[0109] As an example, a node list setting interface can be displayed in the code analysis tool of the terminal, wherein the node list setting interface includes multiple types of configuration entries of the custom syntax nodes to be configured. In response to a trigger operation for any type of configuration entry, the corresponding node list setting interface is displayed, and the node list setting interface can include commonly used configuration items and configuration items previously used by developers, so that developers can quickly configure the corresponding syntax nodes according to the needs of static analysis.

[0110] As an example, the preconfigured custom syntax node list may be shown in the following table. Table 1 shows the names of the custom syntax nodes and the corresponding attribute values ​​of the custom syntax nodes.

[0111]

[0112]

[0113] Table 1

[0114] In some embodiments, see Figure 3E , Figure 3E is a flowchart of obtaining a preconfigured custom syntax node list provided in an embodiment of the present application, Figure 3A Step 1032 shown may be performed by Figure 3E The processing implementation of steps 10321 to 10325 is described in detail below.

[0115] In step 10321, preconfigured types of statements in multiple code languages ​​are obtained, wherein the statements include definition statements and execution statements.

[0116] In some embodiments, the types of definition statements include type definition, variable definition, and function definition. For example, inti=1, the definition type is int, the definition variable name is i, and the initialization value is 1.

[0117] In some embodiments, the types of execution statements include type instantiation, variable assignment, variable reference, function reference, pointer dereference, loop, determination, and jump. For example, the for loop and while loop in the preconfigured custom syntax node list given in the above application embodiment are all execution statements.

[0118] In some embodiments, a type setting function may be provided to developers through a code analysis tool to improve the efficiency of setting types. For example, the code analysis tool running on the terminal displays a node list setting interface, and in response to a type setting operation in the node list setting interface, obtains preconfigured types of statements in multiple code languages ​​input by the developer.

[0119] In step 10322, a preconfigured attribute value of at least one preconfigured attribute type for the custom syntax node is obtained.

[0120] In some embodiments, in response to a property setting operation in a node list setting interface, a preconfigured property value of at least one preconfigured property type for a custom syntax node is obtained.

[0121] In some embodiments, the preconfigured attribute types are determined based on static analysis requirements. Figure 4 , Figure 4 is a schematic diagram of a flow chart for determining a pre-configured attribute type provided in an embodiment of the present application. Figure 3E Before step 10322, you can execute Figure 4 The processing from step 201 to step 204 determines the pre-configured attribute type, which is described in detail below.

[0122] In step 201, a lexical analysis is performed on a source code file to obtain lexical units of the source code file.

[0123] In some embodiments, a lexical analysis is performed on the source code file to decompose the source code into lexical units, such as identifiers, keywords, operators, etc. A lexical unit consists of a lexical unit name and an optional attribute. For example, the lexical unit name is id, and the attribute value is a pointer.

[0124] In step 202, according to the grammatical rules of the first code language, the lexical units of the source code file are grouped into a grammatical structure as a first grammar tree.

[0125] In some embodiments, according to the grammatical rules of the source code language, the lexical units obtained in step 201 are combined into a grammatical structure to form an abstract syntax tree as a first syntax tree. In the process of programming based on the code language, the grammatical rules are used to define the structure and behavior of the program, including rules on variables and data types, expressions and operators, control flow, functions and modules, etc.

[0126] As an example, control flow statements include conditional statements (if-else), loop statements (for, while), and jump statements (break, continue). These statements control the execution path of the program through certain grammatical rules to implement different logic and functions. For example, in a conditional statement, different code blocks are executed depending on the truth or falsity of the condition. The grammatical rules require that the conditional expression return a Boolean value, and the result of the Boolean value determines which code block to execute. Loop statements repeatedly execute a section of code through certain conditions until the condition is not met.

[0127] In step 203, based on the first syntax tree, the source code file is analyzed in multiple dimensions to obtain analysis results of attribute types in each dimension, wherein the multiple dimensions include semantics, data flow, control flow, different execution paths and defects.

[0128] As an example, when multiple dimensions include semantics, a semantic check is performed on the abstract syntax tree to detect syntax errors, type errors, etc.; when multiple dimensions include data flow, the data flow of the source code program corresponding to the abstract syntax tree is analyzed to detect unsafe data access, uninitialized variables, and other problems; when multiple dimensions include control flow, the control flow of the source code program corresponding to the abstract syntax tree is analyzed to detect potential logical errors, dead loops, etc.; when multiple dimensions include different execution paths, symbolic operations are performed on the source code program corresponding to the abstract syntax tree to explore different execution paths of the program to detect potential vulnerabilities and errors; when multiple dimensions include defects, predefined defect scanning rules are used to detect common defects and errors in the source code program corresponding to the abstract syntax tree.

[0129] In step 204, in response to the analysis result indicating that any attribute type is abnormal, the abnormal attribute type is used as at least one preconfigured attribute type of the custom syntax node.

[0130] In some embodiments, according to the analysis results of the attribute types in each dimension obtained in step 203, a corresponding analysis report is generated to identify the location and attribute type of the anomaly.

[0131] By using the abnormal attribute type as at least one preconfigured attribute type of the custom syntax node, the source code file corresponding to the custom syntax tree can be analyzed more specifically during subsequent static analysis, thereby improving the efficiency of static analysis.

[0132] Continue to see Figure 3E , continue with step 10322 above for explanation.

[0133] In step 10323, each statement of the preconfigured type is taken as a custom syntax node, and a node name preconfigured for the custom syntax node is obtained.

[0134] For example, int i=1;, the defined type is int, the defined variable name is i, and the initialization value is 1. Then, in the preconfigured custom syntax node list, a type instantiation or variable definition syntax node VarDefStmt is designed. For this code statement, the core information required for static analysis is to define the type int, define the variable name i, and initialize the value 1. Therefore, when designing the syntax node, the core attribute values ​​are type, variable name, and initialization node.

[0135] In step 10324, the node name and attribute value of the custom syntax node are combined to obtain a custom syntax node.

[0136] For example, in the preconfigured custom syntax node list given in the above text application embodiment, each custom syntax node includes the name of the custom syntax node and the corresponding attribute value of the custom syntax node.

[0137] In step 10325, each custom syntax node obtained by combining is combined into a preconfigured custom syntax node list.

[0138] In some embodiments, when developers perform static analysis on code files in different code languages, they will configure corresponding custom syntax nodes. The configured custom syntax nodes can be continuously accumulated to fill the preconfigured custom syntax node list in an additive manner to be suitable for a wider range of static analysis.

[0139] The custom syntax node list obtained through configuration can increase the information carrying capacity of the syntax nodes in the abstract syntax tree compared to the syntax nodes in the abstract syntax tree constructed by the prior art, thereby improving the readability in the subsequent static analysis process.

[0140] Continue to see Figure 3D , and the description continues with step 1032 above.

[0141] In step 1033, the association relationship configured between the syntax node to be mapped and the custom syntax node is obtained.

[0142] In some embodiments, the pre-configuration of the association relationship between the syntax node to be mapped and the custom syntax node configuration can be achieved by displaying a setting interface on the terminal, wherein the setting interface includes configuration entries for the syntax node to be mapped and the custom syntax node to be configured. In response to a triggering operation for a configuration entry for any set of association relationships, a corresponding setting interface is displayed, and the developer can configure the corresponding association relationship between the syntax node to be mapped and the custom syntax node configuration according to the static analysis requirements.

[0143] In step 1034, a node mapping table representing the association relationship between the syntax nodes to be mapped and the user-defined syntax nodes is constructed.

[0144] In some embodiments, a node mapping table representing the association relationship between the syntax node to be mapped and the custom syntax node can be as shown in Table 2 below, where the list includes the name of the syntax node to be mapped of the Antlr syntax node, the name of the syntax node to be customized, and the attribute value of the corresponding custom syntax node.

[0145]

[0146] Table 2

[0147] Continue to see Figure 3A , and the description continues with step 103 above.

[0148] In step 104, the first syntax tree is traversed, and a second syntax node is obtained by querying from a node mapping table based on the first syntax node of the traversed first syntax tree, wherein the second syntax node is a custom syntax node in the node mapping table that has an association relationship with the first syntax node;

[0149] In some embodiments, the first syntax node is any syntax node in the first syntax tree, rather than a specific node.

[0150] In some embodiments, see Figure 3F , Figure 3F is a schematic diagram of a process of traversing a first syntax tree provided in an embodiment of the present application, Figure 3A The "traversal of the first syntax tree" in step 104 shown can be performed by executing Figure 3F The processing from step 1041 to step 1043 is implemented as described in detail below.

[0151] In step 1041, the root node of the first syntax tree is traversed to obtain the first syntax node corresponding to the root node of the first syntax tree.

[0152] In step 1042, a recursive pre-order traversal is performed on the left subtree of the first syntax tree to obtain a first syntax node corresponding to the left subtree of the first syntax tree.

[0153] In step 1043, a recursive pre-order traversal is performed on the right subtree of the first syntax tree to obtain a first syntax node corresponding to the right subtree of the first syntax tree.

[0154] As an example, see Figure 3G , Figure 3G The following is a flow chart of the pre-order traversal of the syntax tree provided in the embodiment of the present application. Figure 3G As shown, first Figure 3G The root node of the syntax tree in is the syntax node 1, and the syntax node 1 is used as the first syntax node corresponding to the syntax node; then Figure 3G The left subtree of the grammar node 1 in the pre-order traversal is performed, that is, the left subtree of the grammar node 1 is traversed in sequence. Figure 3G Traverse the syntax nodes 2, 4, and 5 in the grammar node, and obtain the first syntax nodes corresponding to the syntax nodes 2, 4, and 5 respectively; finally, Figure 3G The right subtree of the syntax node 1 in the pre-order traversal is performed, that is, the Figure 3G The syntax node 3 and the syntax node 6 in the grammar are traversed to obtain the first syntax node corresponding to the syntax node 3 and the syntax node 6 respectively.

[0155] Continue to see Figure 3A , and the description continues with step 104 above.

[0156] In step 105, the second syntax nodes associated with each first syntax node are combined in a traversal order to obtain a second syntax tree.

[0157] In some embodiments, according to the syntax tree processing method provided in the embodiments of the present application, the number of syntax nodes in the abstract syntax tree obtained by parsing in the syntax analysis can be effectively reduced. The specific implementation is to determine the key attribute types in the statement according to the static analysis requirements, and then convert the ANTLR native syntax tree structure into a single custom syntax node containing key attribute information according to the mapping relationship queried from the node mapping table.

[0158] For example, see Fig. 7A , can be Fig. 7A The AST mapping converter 704A in implements the conversion from the ANTLR native syntax tree structure to the custom syntax tree structure.

[0159] For example, for a simple definition statement int i=1, the corresponding syntax analysis code content is:

[0160]

[0161] In some embodiments, the ANTLR native syntax tree structure is obtained by parsing the code, and then according to the syntax tree processing method provided in the embodiment of the present application, the key attribute type information in the statement is determined as the definition type, variable name, assignment value and other basic information according to the needs of static analysis, and the syntax nodes in the Antlr native syntax tree are traversed in turn, and the ANTLR native syntax tree is converted into a custom syntax tree, that is, the second syntax tree in the embodiment of the present application.

[0162] In some embodiments, the effect of converting the ANTLR native syntax tree into a custom syntax tree is shown in Fig. 6A , Figure 6B and Figure 6C , Fig. 6A is a schematic diagram of the Java native syntax tree structure provided in an embodiment of the present application, Figure 6B is a schematic diagram of the CPP native syntax tree structure provided in the embodiment of the present application, Figure 6C is a schematic diagram of a custom syntax tree structure provided in an embodiment of the present application. Fig. 6A , Figure 6B and Figure 6C It can be seen that the syntax tree processing method proposed in the embodiment of the present application obtains Figure 6C The custom syntax tree in Fig. 6A , Figure 6B The size of the native syntax tree shown in the figure is significantly reduced. In subsequent static analysis, for source code files of the same code language, Fig. 6A or Figure 6B The cost of traversing the native syntax tree shown in is much higher than Figure 6C The read cost of the custom syntax tree shown in , that is, Figure 6C The efficiency of static analysis of custom syntax trees in Fig. 6A Or the efficiency of static analysis of the native syntax tree in 6B.

[0163] The syntax tree processing method of the embodiment of the present application can integrate the syntax features of different code languages, effectively simplify and unify the syntax tree structure of different code languages, eliminate the differences in subsequent static analysis, reduce reading costs, and improve the efficiency of static analysis.

[0164] In some embodiments, see Figure 5A , Figure 5A is a schematic diagram of the first defect scanning process provided by the embodiment of the present application. Figure 3A After step 105, you can execute Figure 5A The processing from step 301 to step 302 is used to scan the second syntax tree, which is described in detail below.

[0165] In step 301, defect scanning rules are obtained.

[0166] In some embodiments, the defect scanning rule includes a scanning object and a scanning target, wherein the scanning object includes a first target defect grammar node; and the scanning target includes a scanning sample defect grammar node.

[0167] In step 302, the second syntax tree is scanned based on the defect scanning rule to obtain a scanning result.

[0168] As an example, when the scan object is a switch statement and the scan target is a default statement, the switch syntax node in the second syntax tree can be intercepted through the custom syntax tree visitor interface, and a flag representing the default syntax node can be defined; then all child nodes of the switch syntax node are traversed, and it is verified whether there is a default syntax node. If the scan finds the default syntax node, the location information of the default syntax node is identified in the scan result, and the abnormal node with the default syntax node is displayed; if the scan does not find the default syntax node, the scan result is that the abnormal node of the default syntax node does not exist in the current syntax tree.

[0169] In some embodiments, see Figure 5B , Figure 5B is a schematic diagram of the second defect scanning process provided by an embodiment of the present application, Figure 5A Step 302 can be performed by executing Figure 5B The processing from step 3021 to step 3023 is implemented as described in detail below.

[0170] In step 3021, the syntax node to be scanned in the second syntax node is intercepted through a preconfigured syntax tree visitor interface.

[0171] In some embodiments, the syntax nodes to be scanned in the second syntax node are intercepted by using a syntax tree observer or visitor interface preconfigured in ANTLR or a similar syntax analyzer.

[0172] In step 3022, the child nodes in the syntax node to be scanned are traversed, and the child nodes are compared with the sample target defect syntax nodes.

[0173] In some embodiments, by traversing the child nodes in the syntax node to be scanned in the syntax tree, the child nodes are compared with the preset sample target defect syntax nodes, and the subsequent step 3023 is executed according to the comparison result.

[0174] In step 3023, in response to the matching, the child node is used as a defective syntax node, and an abnormal scanning result is generated based on the position information of the child node.

[0175] In some embodiments, in response to the comparison being inconsistent, the scanning result is output as an alarm to indicate that there is no abnormality in the second syntax tree.

[0176] For example, see Fig. 7A ,according to Fig. 7A Defect scanning rule 705A, for Fig. 7A The custom syntax tree obtained by the AST mapping converter 704A executes a custom syntax tree scan 706A to obtain a project scan result 707A.

[0177] The following describes an exemplary application of the syntax tree processing method provided in an embodiment of the present application in an actual application scenario.

[0178] For example, the application scenario can be the development scenario of the game source code. The game development software project types cover various types such as web pages, games, and mobile clients, and the code language almost covers all currently popular programming languages. Therefore, in the process of developing static analysis tools, it is necessary to develop code defect rules for different languages, especially different languages ​​for the same type of projects, resulting in the need to implement the same defect rule for different languages ​​separately. The processing method of the syntax tree provided in the embodiment of the present application realizes the generation of a simplified syntax tree structure by parsing the game source code during the development process of the game source code, thereby improving the efficiency of subsequent static analysis and improving the efficiency of static analysis.

[0179] See also Figure 7B , Figure 7B This is a third flow chart of the method for processing a syntax tree provided in an embodiment of the present application, which is described in detail below.

[0180] Taking the source code statement int i=1 as an example, an exemplary application of the syntax tree processing method provided in the embodiment of the present application is described.

[0181] First, the project source code file 701B is parsed by the Antlr parser 702B to obtain AntlrAST 703B. Figure 7C As shown, Figure 7C is a schematic diagram of the first structure of the native syntax tree provided in an embodiment of the present application, Figure 7C The syntax tree contains node information such as function definition syntax node functionDefinition, functionBody node, simpleDeclaration node, etc.

[0182] Then, configure as Figure 7B The custom AST node library 704B in the syntax tree is the custom syntax node list in the syntax tree processing method provided in the embodiment of the present application.

[0183] In the code int i=1;, the type is defined as int, the variable name is defined as i, and the initialization value is 1. Therefore, a type instantiation or variable definition syntax node VarDefStmt is designed in the preconfigured custom syntax node library. The core information required for static analysis of this code statement is to define the type int, define the variable name i, and initialize the value 1. Therefore, when designing the syntax node, the core attribute values ​​are type, variable name, and initialization node. Other syntax nodes in the custom syntax node library are configured in the same way.

[0184] The preconfigured custom syntax node library may be as shown in Table 1 above, and the library includes the names of the custom syntax nodes and the corresponding attribute values ​​of the custom syntax nodes.

[0185] Secondly, according to Figure 7B Custom AST node library 704B in configuration Figure 7B The Antlr-customized AST node mapping table 705B is the preconfigured node mapping table in the syntax tree processing method provided in the embodiment of the present application.

[0186] Take the code int i=1 as an example. The structure of the source code file after parsing in Antlr AST is shown in Fig.7D , Fig.7D is a schematic diagram of the second structure of the native syntax tree provided in the embodiment of the present application, such as Fig.7D As shown, the statement takes the SimpleDeclaration syntax node as the root node, and separates: declSpecifierSeq as the type root node. After deriving 5 child nodes downward, the type name is int; then initDeclaratorList is used as the variable name and initialization statement, and 8 and 23 child nodes are derived downward respectively to finally obtain the variable name i and the initialization value 1.

[0187] From the above, we can see that Figure 7B The custom AST node library 704B in the example is designed with a type instantiation or variable definition syntax node VarDefStmt, the attribute values ​​of which include var, type, and value, where the attribute value var of the syntax node records the variable name, the attribute value type of the syntax node records the type, and the attribute value value of the syntax node records the initialization value. Therefore, the Antlr syntax node SimpleDeclaration can be associated with the custom syntax node VarDefStmt. By analogy, we can finally obtain the syntax type node mapping table shown in Table 2 above, as Figure 7BThe Antlr-custom AST node mapping table 705B in the table includes the names of the syntax nodes to be mapped of the Antlr syntax nodes, the names of the syntax nodes to be customized, and the attribute values ​​of the corresponding customized syntax nodes.

[0188] Finally, through Figure 7B The Antlr AST 703B in the code executes the Antlr AST node traversal 706B, and at the same time queries the mapping relationship between the native syntax tree nodes and the custom syntax nodes in the Antlr-custom AST node mapping table 705B, and converts the syntax tree nodes of the Antlr AST 703B into the custom syntax nodes in the custom AST 708B through the custom AST generator 707B.

[0189] In the process of traversing the syntax nodes of Antlr AST 703B, Figure 7C As shown, first identify the function definition syntax node functionDefinition, and generate the corresponding mapping custom node MethodDeclaration. Extract the core information in the child nodes of the original node functionDefinition, such as function name, return type, parameter type, etc., and fill them in the custom function definition syntax node MethodDeclaration. And continue to traverse down with the functionBody node. After identifying the simpleDeclaration node, analyze the type subnode and variable definition subnode of its child nodes, extract the type name, variable name, and initialization value, and generate the corresponding node mapping custom VarDefStmt node, and set the extracted information as the custom node attribute value. Finally, link the node to the parent node MethodDeclaration to obtain the custom abstract syntax tree structure data. The final generated custom syntax tree is as follows Fig. 7E shown.

[0190] In an embodiment of the present application, a preconfigured node mapping table is used to realize the conversion of the syntax node to be mapped to the custom syntax node based on the preconfigured node mapping table, thereby simplifying the structure of the ANTLR syntax tree, effectively reducing the language differences of the abstract syntax trees generated based on different code languages ​​in the static analysis stage, and improving the efficiency of subsequent static analysis.

[0191] The following further describes an exemplary structure of the syntax tree processing device 455 provided in the embodiment of the present application implemented as a software module. In some embodiments, for example Figure 2 As shown, the software modules stored in the syntax tree processing device 455 of the memory 440 may include:

[0192] The acquisition module 4551 is used to acquire a source code file, wherein the source code file is developed based on a first code language, and the first code language is any code language; and acquire a preconfigured node mapping table, wherein the node mapping table includes syntax nodes to be mapped of multiple code languages ​​and custom syntax nodes associated with the syntax nodes to be mapped, and the multiple code languages ​​include the first code language.

[0193] The parsing module 4552 is used to parse the source code file to obtain a first syntax tree.

[0194] The traversal module 4553 is used to traverse the first syntax tree, and obtain a second syntax node from the node mapping table based on the first syntax node of the traversed first syntax tree, wherein the second syntax node is a custom syntax node in the node mapping table that has an association relationship with the first syntax node.

[0195] The processing module 4554 is used to combine the second syntax nodes associated with each first syntax node in a traversal order to obtain a second syntax tree.

[0196] In some embodiments, the parsing module 4552 is also used to perform syntax analysis on the source code file to obtain a syntax analysis code of the first code language corresponding to the source code file; based on the syntax analysis code, determine multiple first syntax nodes in the first syntax tree; and combine the first syntax nodes of the multiple first syntax trees according to the logical relationship of the syntax analysis code to obtain a first syntax tree.

[0197] In some embodiments, the parsing module 4552 is also used to obtain a preconfigured list of syntax nodes to be mapped, wherein the list of syntax nodes to be mapped includes multiple syntax nodes to be mapped in multiple code languages; traverse the syntax analysis code, and compare the syntax nodes obtained through the traversal with the syntax nodes to be processed in the syntax nodes to be mapped; in response to the comparison being consistent, use the syntax node obtained through the traversal as the first syntax node.

[0198] In some embodiments, the acquisition module 4551 is also used to display a setting interface for syntax nodes to be mapped; in response to a setting operation for syntax nodes to be mapped in the setting interface, obtain a plurality of pre-configured syntax nodes to be mapped; and combine the plurality of syntax nodes to be mapped to obtain a list of syntax nodes to be mapped.

[0199] In some embodiments, the acquisition module 4551 is also used to obtain a preconfigured list of syntax nodes to be mapped, wherein the list of syntax nodes to be mapped includes multiple syntax nodes to be mapped in multiple code languages; obtain a preconfigured list of custom syntax nodes, wherein the custom syntax node list includes multiple custom syntax nodes; obtain the association relationship configured for the syntax nodes to be mapped and the custom syntax nodes; and construct a node mapping table representing the association relationship between the syntax nodes to be mapped and the custom syntax nodes.

[0200] In some embodiments, the acquisition module 4551 is also used to obtain preconfigured types of statements in multiple code languages, wherein statements include definition statements and execution statements; obtain preconfigured attribute values ​​of at least one preconfigured attribute type for a custom syntax node; treat each statement of the preconfigured type as a custom syntax node, and obtain the preconfigured node name for the custom syntax node; combine the node name and attribute value of the custom syntax node to obtain a custom syntax node; and combine each combined custom syntax node into a preconfigured custom syntax node list.

[0201] In some embodiments, the processing module 4554 is further used to perform lexical analysis on the source code file to obtain lexical units of the source code file before obtaining the preconfigured attribute value of at least one preconfigured attribute type for the custom syntax node; according to the grammatical rules of the first code language, the lexical units of the source code file are combined into a grammatical structure as a first syntax tree; based on the first syntax tree, the source code file is analyzed in multiple dimensions to obtain analysis results of the attribute types in each dimension, wherein the multiple dimensions include semantics, data flow, control flow, different execution paths and defects; in response to the analysis result indicating that any attribute type is abnormal, the abnormal attribute type is used as at least one preconfigured attribute type of the custom syntax node.

[0202] In some embodiments, the types of definition statements include type definition, variable definition, and function definition; the types of execution statements include type instantiation, variable assignment, variable reference, function reference, pointer dereference, loop, determination, and jump.

[0203] In some embodiments, the defect scanning rules include a scanning object and a scanning target, wherein the scanning object includes a first target defect syntax node; the scanning target includes a scanning sample defect syntax node; the processing module 4554 is also used to obtain the defect scanning rules; based on the defect scanning rules, the second syntax tree is scanned to obtain the scanning results.

[0204] In some embodiments, the processing module 4554 is also used to intercept the syntax node to be scanned in the second syntax node through a preconfigured syntax tree visitor interface; traverse the child nodes in the syntax node to be scanned, and compare the child nodes with the sample target defective syntax node; in response to the comparison being consistent, treat the child node as a defective syntax node, and generate an abnormal scanning result based on the position information of the child node.

[0205] In some embodiments, the traversal module 4553 is also used to traverse the root node of the first syntax tree to obtain the first syntax node corresponding to the root node of the first syntax tree; to perform a recursive pre-order traversal on the left subtree of the first syntax tree to obtain the first syntax node corresponding to the left subtree of the first syntax tree; and to perform a recursive pre-order traversal on the right subtree of the first syntax tree to obtain the first syntax node corresponding to the right subtree of the first syntax tree.

[0206] The embodiment of the present application provides a computer program product, which includes a computer program or a computer executable instruction, and the computer program or the computer executable instruction is stored in a computer-readable storage medium. The processor of the electronic device reads the computer executable instruction from the computer-readable storage medium, and the processor executes the computer executable instruction, so that the electronic device executes the syntax tree processing method described in the embodiment of the present application.

[0207] The embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions or computer programs are stored. When the computer-executable instructions or computer programs are executed by a processor, the processor will execute the method for processing a syntax tree provided in the embodiment of the present application, for example, Figure 3A The method for processing the syntax tree is shown.

[0208] In some embodiments, the computer-readable storage medium may be a memory such as RAM, ROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or may be various devices including one or any combination of the above memories.

[0209] In some embodiments, computer executable instructions may be in the form of a program, software, software module, script or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine or other unit suitable for use in a computing environment.

[0210] As an example, computer-executable instructions may, but need not, correspond to a file in a file system, may be stored as part of a file that stores other programs or data, such as in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files storing one or more modules, subroutines, or code portions).

[0211] As an example, computer executable instructions may be deployed to be executed on one electronic device, or on multiple electronic devices located at one site, or on multiple electronic devices distributed at multiple sites and interconnected by a communication network.

[0212] In summary, it can be learned from the embodiments of the present application that a first syntax tree is obtained by parsing the source code file, and then a preconfigured node mapping table is obtained, which includes syntax nodes to be mapped of multiple code languages ​​and custom syntax nodes that are associated with the syntax nodes to be mapped; secondly, the first syntax tree is traversed to determine the first syntax node, and a second syntax node that is associated with the first syntax node is obtained from the node mapping table; finally, each second syntax node that is associated with the first syntax node is combined in the order of traversal to obtain a second syntax tree, thereby parsing code files of different code languages ​​into data of a syntax tree structure with a unified structure, which can streamline and unify the syntax tree structures of different code languages, facilitate subsequent static analysis using a unified structure for processing, and improve the efficiency of subsequent static analysis.

[0213] The above is only an embodiment of the present application and is not intended to limit the protection scope of the present application. Any modifications, equivalent substitutions and improvements made within the spirit and scope of the present application are included in the protection scope of the present application.

Claims

1. A method for processing a syntax tree, characterized in that: The method comprises: Obtain a source code file, wherein the source code file is developed based on a first code language, and the first code language is any code language; Parsing the source code file to obtain a first syntax tree; Obtaining a preconfigured node mapping table, wherein the node mapping table includes syntax nodes to be mapped in multiple code languages ​​and user-defined syntax nodes associated with the syntax nodes to be mapped, the multiple code languages ​​including the first code language; Traversing the first syntax tree, and obtaining a second syntax node by querying from the node mapping table based on the first syntax node of the traversed first syntax tree, wherein the second syntax node is the custom syntax node in the node mapping table that has the association relationship with the first syntax node; The second syntax nodes that have the association relationship with each of the first syntax nodes are combined in the order of traversal to obtain a second syntax tree.

2. The method according to claim 1, characterized in that The step of parsing the source code file to obtain a first syntax tree includes: Performing syntax analysis on the source code file to obtain a syntax analysis code corresponding to the first code language of the source code file; Based on the syntax analysis code, determining a plurality of the first syntax nodes in the first syntax tree; The first syntax nodes of the plurality of the first syntax trees are combined according to the logical relationship of the syntax analysis code to obtain a first syntax tree.

3. The method according to claim 2, characterized in that The determining, based on the syntax analysis code, a plurality of the first syntax nodes in the first syntax tree comprises: Acquire a preconfigured list of syntax nodes to be mapped, wherein the list of syntax nodes to be mapped includes a plurality of the syntax nodes to be mapped in the plurality of code languages; Traversing the syntax analysis code, and comparing the syntax nodes obtained by the traversal with the syntax nodes to be processed in the syntax nodes to be mapped; In response to the comparison being consistent, the syntax node obtained by the traversal is used as the first syntax node.

4. The method according to claim 3, characterized in that The obtaining of a preconfigured list of syntax nodes to be mapped includes: Display the syntax node setting interface to be mapped; In response to the to-be-mapped syntax node setting operation in the setting interface, obtaining a plurality of pre-configured to-be-mapped syntax nodes; Combine a plurality of the syntax nodes to be mapped to obtain the syntax node list to be mapped.

5. The method according to any one of claims 1 to 4, characterized in that: The obtaining of the pre-configured node mapping table includes: Acquire a preconfigured list of syntax nodes to be mapped, wherein the list of syntax nodes to be mapped includes a plurality of the syntax nodes to be mapped in the plurality of code languages; Obtain a preconfigured custom syntax node list, wherein the custom syntax node list includes a plurality of the custom syntax nodes; Acquire the association relationship configured for the syntax node to be mapped and the custom syntax node; A node mapping table representing the association relationship between the to-be-mapped syntax node and the custom syntax node is constructed.

6. The method according to claim 5, characterized in that The method of obtaining a pre-configured custom syntax node list includes: Obtaining preconfigured types of statements in the plurality of code languages, wherein the statements include definition statements and execution statements; Obtaining a preconfigured attribute value of at least one preconfigured attribute type for the custom syntax node; Taking each of the statements of the preconfigured type as a custom syntax node, and obtaining a node name preconfigured for the custom syntax node; Combine the node name and the attribute value of the custom syntax node to obtain the custom syntax node; Each of the combined custom syntax nodes is combined into a preconfigured custom syntax node list.

7. The method according to claim 6, characterized in that Before obtaining a preconfigured attribute value of at least one preconfigured attribute type for the custom syntax node, the method further includes: Performing lexical analysis on the source code file to obtain lexical units of the source code file; According to the grammatical rules of the first code language, composing the lexical units of the source code file into a grammatical structure as the first grammar tree; Based on the first syntax tree, the source code file is analyzed in multiple dimensions to obtain analysis results of attribute types in each dimension, wherein the multiple dimensions include semantics, data flow, control flow, different execution paths, and defects; In response to the analysis result indicating that any one of the attribute types is abnormal, the abnormal attribute type is used as at least one preconfigured attribute type of the custom syntax node.

8. The method according to claim 6, characterized in that The types of definition statements include type definition, variable definition and function definition; The types of execution statements include type instantiation, variable assignment, variable reference, function reference, pointer dereference, loop, determination and jump.

9. The method according to any one of claims 1 to 4, characterized in that: After combining the second syntax nodes that have the association relationship with each of the first syntax nodes in the order of traversal to obtain a second syntax tree, the method further includes: Get defect scanning rules; Based on the defect scanning rule, the second syntax tree is scanned to obtain a scanning result.

10. The method according to claim 9, characterized in that The defect scanning rule includes a scanning object and a scanning target, wherein the scanning object includes a first target defect syntax node; the scanning target includes scanning the sample defect syntax node; The step of scanning the second syntax tree based on the defect scanning rule to obtain a scanning result includes: Intercepting the to-be-scanned syntax node in the second syntax node through a preconfigured syntax tree visitor interface; Traversing the child nodes in the to-be-scanned syntax node, and comparing the child nodes with the sample target defect syntax node; In response to the comparison being consistent, the child node is used as a defective grammar node, and an abnormal scanning result is generated based on the position information of the child node.

11. The method according to any one of claims 1 to 4, characterized in that: The traversing the first syntax tree comprises: Traversing the root node of the first syntax tree to obtain a first syntax node corresponding to the root node of the first syntax tree; Recursively traverse the left subtree of the first syntax tree in pre-order order to obtain a first syntax node corresponding to the left subtree of the first syntax tree; Recursively perform a pre-order traversal on the right subtree of the first syntax tree to obtain a first syntax node corresponding to the right subtree of the first syntax tree.

12. A syntax tree processing device, characterized in that: The device comprises: An acquisition module, configured to acquire a source code file, wherein the source code file is developed based on a first code language, and the first code language is any code language; Obtaining a preconfigured node mapping table, wherein the node mapping table includes syntax nodes to be mapped in multiple code languages ​​and user-defined syntax nodes associated with the syntax nodes to be mapped, the multiple code languages ​​including the first code language; A parsing module, used for parsing the source code file to obtain a first syntax tree; A traversal module, used to traverse the first syntax tree, and query from the node mapping table to obtain a second syntax node based on the first syntax node of the traversed first syntax tree, wherein the second syntax node is the custom syntax node in the node mapping table that has the association relationship with the first syntax node; The processing module is used to combine the second syntax nodes that have the association relationship with each of the first syntax nodes in the order of traversal to obtain a second syntax tree.

13. An electronic device, characterized in that: The electronic device comprises: A memory for storing computer executable instructions; A processor, configured to implement the method for processing a syntax tree according to any one of claims 1 to 11 when executing computer executable instructions stored in the memory.

14. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that: When the computer executable instructions or computer program are executed by a processor, the method for processing a syntax tree according to any one of claims 1 to 11 is implemented.

15. The present application embodiment provides a computer program product, including a computer program or a computer executable instruction, characterized in that: When the computer executable instructions or computer program are executed by a processor, the method for processing a syntax tree according to any one of claims 1 to 11 is implemented.