Automatic protocol vulnerability mining method and device based on enhanced data flow diagram

Through the generation of static analysis technology of enhanced data flow diagrams, protocol vulnerabilities are automatically identified, which solves the problems of low efficiency and limited scope in the existing technology, and realizes efficient and automated protocol vulnerability identification and analysis.

CN120281567AActive Publication Date: 2025-07-08TSINGHUA UNIVERSITY +1
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202510741801.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-07-08
Estimated Expiration
2045-06-05

AI Technical Summary

Technical Problem

Existing protocol vulnerability mining solutions rely on expert manual analysis or fuzz testing, are inefficient and difficult to meet the analysis needs of modern complex network protocols, and lack real-timeness, making it impossible to effectively identify vulnerabilities caused by multi-level interactions.

Method used

Through static analysis technology, the protocol interaction semantic information in the source code is extracted, the enhanced data flow diagram is generated, the potential vulnerability risks are automatically identified, and the analysis scope is expanded by using automated methods to handle the communication process of different protocol participants across platforms.

Benefits of technology

It realizes efficient identification of vulnerabilities brought about by protocol interaction, improves vulnerability mining efficiency, shortens analysis time, expands analysis scope, and avoids dependence on expert manual analysis and test cases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120281567A_ABST
    Figure CN120281567A_ABST
Patent Text Reader

Abstract

The invention provides a protocol vulnerability automatic mining method and device based on an enhanced data flow diagram, and the method comprises the steps: determining an analysis range of a source code, and compiling the source code into an intermediate language file based on the analysis range; performing control flow analysis on the intermediate language file to obtain a total control flow diagram; based on the total control flow diagram, performing data flow analysis and improvement on the intermediate language to generate an enhanced data flow diagram; determining an analysis target, and obtaining a calling path of the analysis target based on the enhanced data flow diagram; and performing potential risk judgment on the calling path of the analysis target, printing a path with a potential vulnerability risk, and warning a user. The protocol interaction semantic information in the source code is extracted through the static analysis technology, potential risk judgment is performed on the interaction information, vulnerabilities brought by protocol interaction can be effectively recognized, dependence on manual analysis or test cases in the protocol vulnerability mining process is reduced, the analysis range of the protocol vulnerabilities is expanded, and the protocol vulnerability mining efficiency is improved. And the protocol vulnerability mining capability and efficiency are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of cyberspace security, and particularly to a method and device for automatically mining protocol vulnerabilities based on an enhanced data flow graph. Background Art

[0002] Protocols are the basis of network communication and data exchange and play an important role in computer networks. However, with the continuous development of network technologies and the continuous evolution of attackers' attack means, the security of protocols is facing increasingly severe challenges. Malicious attackers can use protocol vulnerabilities to perform malicious behaviors such as unauthorized access, denial of service, and remote control. These attacks not only cause network system paralysis but also lead to user privacy leakage and property losses, posing a serious threat to cyberspace security. Therefore, mining protocol vulnerabilities has become an important research direction in the field of cyberspace security.

[0003] The existing protocol vulnerability mining solutions mainly have the following defects: 1) Traditional protocol vulnerability mining solutions either rely on experts to manually analyze the interaction logic of protocol communication or rely on fuzz testing to explore the path execution results. For these two solutions, the former's vulnerability mining rate and speed are completely limited by the ability of experts, and potential vulnerabilities are easily missed; the latter is limited by the coverage and effectiveness of test cases, and it is difficult to construct complete test cases.

[0004] 2) Existing protocol vulnerability mining solutions usually analyze one or several specific protocols. However, modern network protocols are usually very complex, and the generation of a protocol vulnerability may involve interactions between multiple levels and multiple parties. Analyzing a single protocol alone obviously cannot meet the requirements.

[0005] 3) The evolution and update of protocols is a continuous process, and the attack means of attackers are also continuously improving. Therefore, the mining and repair of protocol vulnerabilities need to be carried out continuously, and the real-time requirement for vulnerability mining is relatively high. However, the existing protocol vulnerability mining solutions require a high time cost and are easy to give attackers an opportunity. Summary of the Invention

[0006] This application aims to solve at least one of the technical problems in the related technologies to some extent.

[0007] To this end, the first object of this application is to propose a method for automatically mining protocol vulnerabilities based on an enhanced data flow graph. For protocol vulnerabilities caused by semantic loss, the protocol interaction semantic information in the source code is extracted through static analysis technology, and potential risks of the interaction information are judged, so as to effectively identify the vulnerabilities brought by protocol interactions.

[0008] The second object of the present application is to propose a protocol vulnerability automatic mining device based on an enhanced data flow graph.

[0009] The third object of the present application is to propose an electronic device.

[0010] To achieve the above object, an embodiment of the first aspect of the present application proposes a protocol vulnerability automatic mining method based on an enhanced data flow graph, including: Determine the analysis scope of the source code, and compile the source code into an intermediate language file based on the analysis scope; Perform control flow analysis on the intermediate language file to obtain a total control flow graph; Based on the total control flow graph, perform data flow analysis and improvement on the intermediate language to generate an enhanced data flow graph; Determine the analysis target, and obtain the call path of the analysis target based on the enhanced data flow graph; Perform potential risk judgment on the call path of the analysis target, print the path with potential vulnerability risks, and warn the user.

[0011] Optionally, the determining the analysis scope of the source code and compiling the source code into an intermediate language file based on the analysis scope includes: Determine that the analysis scope of the source code is global analysis, include all source codes of the linux kernel in the analysis scope, analyze the linux kernel as a project, and obtain the intermediate language file; Determine that the analysis scope of the source code is local analysis, include the target file in the linux kernel in the analysis scope, separately compile the target file and then link it uniformly to obtain the intermediate language file.

[0012] Optionally, the determining that the analysis scope of the source code is global analysis, including all source codes of the linux kernel in the analysis scope, analyzing the linux kernel as a project, and obtaining the intermediate language file includes: Use the wllvm script under the llvm compiler project to compile the entire linux kernel to obtain a compiled kernel image vmlinux file. During the kernel compilation process, specify the Clang tool as the compiler and llvm as the compiler backend; Use the extract-bc tool to extract the compiled vmlinux file into an LLVM bytecode file, and use the LLVM bytecode file as the intermediate language file.

[0013] Optionally, the analysis scope for determining the source code is local analysis. The target files in the Linux kernel are included in the analysis scope, and after separately compiling the target files and then performing unified linking, the intermediate language file is obtained, including: Using the Clang tool to compile each of the target files into intermediate language sub-files; Using the llvm-link tool to link multiple intermediate language sub-files together, and taking the merged result as the intermediate language file.

[0014] Optionally, performing control flow analysis on the intermediate language file to obtain a total control flow graph, including: Collecting the target statements of the jump statements within the functions of the intermediate language file and the next statements of the jump statements. Both the target statements and the next statements of the jump statements mark the start of a new basic block. Determining the transfer relationship between basic blocks according to the jump instructions and adding control flow edges to obtain a control flow graph within the function; Traversing the control flow graph within the function to collect function call statements, determining the call relationship between functions according to the function call statements, and constructing a function call graph; Combining the control flow graph within the function and the function call graph to obtain the total control flow graph.

[0015] Optionally, based on the total control flow graph, performing data flow analysis and improvement on the intermediate language file to generate an enhanced data flow graph, including: Invoking the data flow graph generation function of the open-source static program analysis framework SVF tool to analyze and transform the intermediate language file, obtaining the basic data pointing relationship graph PAG of the intermediate language file, and taking this basic data pointing relationship graph PAG as the basic data flow graph. During the transformation process, pointer variables or abstract objects are set as nodes, and the relationships between pointer variables are transformed into edges; Traversing each node in the basic data flow graph to obtain the node instructions and the first node positions corresponding to the nodes; If the node instruction corresponding to the first node position is a jump statement, obtaining the second node position where the node instruction is located in the total control flow graph, traversing all subsequent nodes after the second node position, and sequentially adding the node corresponding to the second node position and all its subsequent nodes to the back of the node in the basic data flow graph with the first node position; If the subsequent node of the second node position is a function call statement, adding a control dependency edge between the subsequent node corresponding to the second node position and the node representing the function in the basic data flow graph; After adding, obtaining the enhanced data flow graph.

[0016] Optionally, determining the analysis target and obtaining the call path of the analysis target based on the enhanced data flow graph includes: Determine the key variables and structures related to vulnerability formation, apply the alias analysis method to determine the aliases corresponding to the key variables and the structures, and record the key variables, the structures, and their corresponding aliases in a linked list in memory; Traverse the key variables, the structures, or the aliases in the memory, and trace the call paths of the key variables, the structures, or the aliases in the enhanced data flow graph; Among them, during the tracing process, search for the code segments before the variables are assigned values or after they are used.

[0017] Optionally, performing a potential risk judgment on the call path of the analysis target, printing the paths with potential vulnerability risks, and warning the user includes: Determine potential risk instructions according to the risk measurement indicators, and use the Elist linked list to record all the risk instructions; Traverse the call path, and determine whether each instruction in the call path appears in the Elist linked list; If at least one instruction in the call path appears in the Elist linked list, determine that the call path has potential vulnerability risks, print the call path, and send a vulnerability warning to the user.

[0018] To achieve the above object, the second aspect embodiment of the present application proposes a protocol vulnerability automatic mining device based on an enhanced data flow graph, including: An initialization module, configured to determine the analysis scope of the source code, and compile the source code into an intermediate language file based on the analysis scope; A control flow analysis module, configured to perform control flow analysis on the intermediate language file to obtain a total control flow graph; A data flow analysis module, configured to perform data flow analysis and improvement on the intermediate language based on the total control flow graph to generate an enhanced data flow graph; A path tracing module, configured to determine an analysis target and obtain the call path of the analysis target based on the enhanced data flow graph; A security check module, configured to perform a potential risk judgment on the call path of the analysis target, print the paths with potential vulnerability risks, and warn the user.

[0019] To achieve the above object, the third aspect embodiment of the present application proposes an electronic device, including: a processor, and a memory communicatively connected to the processor; The memory stores computer execution instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of the above first aspects.

[0020] The technical solutions provided by the embodiments of the present application at least bring the following beneficial effects: Regarding the protocol vulnerabilities caused by semantic loss, by extracting the protocol interaction semantic information in the source code through static analysis technology and judging the potential risks of the interaction information, the vulnerabilities brought by protocol interaction can be effectively identified; this solution does not rely on expert manual analysis nor on the completeness of test cases; in addition, this method is not limited to specific analysis of one or several protocols, but adds all protocol participants involved in the protocol communication process to the analysis scope, expanding the analysis scope of the protocol; furthermore, this method uses an automated method to mine protocol vulnerabilities for different analysis targets, can process a large amount of protocol code in a short time and give potential vulnerability risks, thereby improving the efficiency of vulnerability mining.

[0021] The additional aspects and advantages of the present application will be partly given in the following description, partly will become obvious from the following description, or will be understood through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The above and / or additional aspects and advantages of the present application will become obvious and easy to understand from the following description of the embodiments in conjunction with the drawings, where: Figure 1 is a flowchart of an automated protocol vulnerability mining method based on an enhanced data flow graph according to an embodiment of the present application; Figure 2 is a flowchart of a method for performing control flow analysis on an intermediate language file according to an embodiment of the present application; Figure 3 is a block diagram of an automated protocol vulnerability mining device based on an enhanced data flow graph according to an embodiment of the present application; Figure 4 is a block diagram of an electronic device. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0023] The embodiments of the present application will be described in detail below. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions from beginning to end. The embodiments described below by referring to the drawings are exemplary and are intended to explain the present application and should not be construed as limiting the present application.

[0024] The term explanations that appear in the present application are shown in Table 1.

[0025]

[0026] Table 1 The following describes a method and apparatus for automatically mining protocol vulnerabilities based on an enhanced data flow graph according to an embodiment of the present application with reference to the accompanying drawings.

[0027] Figure 1 is a flowchart of a method for automatically mining protocol vulnerabilities based on an enhanced data flow graph shown according to an embodiment of the present application, as Figure 1 shown, the method includes the following steps: Step 101, determine the analysis scope of the source code, and compile the source code into an intermediate language file based on the analysis scope.

[0028] It can be understood that the generation of modern protocol vulnerabilities is very complex, and the generation of a protocol vulnerability may involve interaction behaviors between multiple levels or participants.

[0029] Therefore, the present application includes all protocol participants involved in the protocol communication process in the analysis scope to discover more complex vulnerabilities, such as vulnerabilities caused by data interaction between multiple protocols. In addition, an ideal vulnerability mining method should have cross-platformness, scalability, and high-level abstraction. The present application analyzes the source code after compiling it into an intermediate language. The intermediate language is a common intermediate representation between different programming languages, which means that the same analysis tools and techniques can be used to process code on different platforms and different types of high-level languages, with high scalability.

[0030] The solution disclosed in the present application supports two analysis scope selections, which can meet different memory and speed requirements. And it can be understood that if the analysis scope selected by the user is different, the compilation method is also different. The following takes the mining of protocol vulnerabilities in the linux kernel as an example for explanation.

[0031] As a possible implementation, if the analysis scope of the source code is determined to be global analysis, all source code of the linux kernel is included in the analysis scope, and the linux kernel is analyzed as a project to obtain an intermediate language file.

[0032] As an example, the process of analyzing the linux kernel as a project is implemented in the following way: Use the wllvm script under the llvm compiler project to compile the entire linux kernel to obtain a compiled kernel image vmlinux file. Among them, the Clang tool is specified as the compiler during the kernel compilation process, and llvm is used as the compiler backend; use the extract-bc tool to extract the compiled vmlinux file into an LLVM bytecode file, and use the LLVM bytecode file as the intermediate language file.

[0033] It can be understood that through global analysis, all communication interaction processes between kernel protocol stacks can be captured by analyzing the source code, so as to discover more types of protocol vulnerabilities. However, this option requires more memory and computing resources.

[0034] As another possible implementation, if the analysis scope of the source code is determined to be local analysis, the target files in the Linux kernel are included in the analysis scope, and the target files are separately compiled and then linked together to obtain an intermediate language file.

[0035] Among them, the target files are certain specific files in the Linux kernel selected by the user, and this application does not make specific limitations on this.

[0036] As an example, the process of separately compiling and then linking the target files is implemented in the following way: Use the Clang tool to compile each target file into an intermediate language sub-file; use the llvm-link tool to link multiple intermediate language sub-files together, and use the merged result as the intermediate language file for overall analysis.

[0037] It can be understood that through local analysis, only the communication processes of specific modules or protocol participants under the kernel protocol stack need to be concerned. This option requires less memory and computing resources.

[0038] Step 102, perform control flow analysis on the intermediate language file to obtain the overall control flow graph.

[0039] In the embodiments of this application, based on Step 101, the analysis scope of the source code has been determined and the final intermediate language file has been obtained, which provides an analysis basis for this application. Since the flow and dependency relationships of data are usually affected by the control flow structure, therefore, this application will first construct a control flow graph according to the intermediate language file to represent the control flow transfer relationships in the program and prepare for the subsequent data flow analysis. As Figure 2 shown, Step 102 further includes the following steps: Step 201, collect the target statements of the intra-function jump statements in the intermediate language file and the next statement of the jump statement. The target statement of the jump statement and the next statement both mark the start of a new basic block. Determine the transfer relationship between basic blocks according to the jump instruction and add control flow edges to obtain the intra-function control flow graph.

[0040] It should be noted that the intra-function control flow graph is a graph structure representing the program execution flow within a function. The nodes in the graph are instructions, and the edges represent the control flow.

[0041] As a possible implementation, the present application first collects the target statements of the intra-function jump statements in the intermediate language file and the next statement after the jump statement. These statements all mark the start of a new basic block. Then, based on the jump instructions, the transfer relationships between the basic blocks are determined and control flow edges are added to obtain the intra-function control flow graph.

[0042] It should be noted that each basic block only contains sequential statements. The present application will traverse each basic block internally. When encountering an instruction, it will be separately designated as a node, and the next instruction will be used as the next node. A control flow edge will be connected between the two nodes, and finally the intra-function control flow graph will be obtained.

[0043] As an example, the final form of the intra-function control flow graph is , where is the set of nodes , is the set of directed edges, n is the number of nodes, and each node is an instruction.

[0044] Step 202: Traverse the intra-function control flow graph to collect function call statements, determine the call relationships between functions based on the function call statements, and construct a function call graph.

[0045] Among them, it is necessary to record the file name where the function is located, the function name, the function parameters, and the function return value to form a function signature, so as to record which function calls which function exactly, in order to draw the control flow edges.

[0046] As a possible implementation, a HashMap is used to record the mapping relationship between functions and the called functions.

[0047] Step 203: Combine the intra-function control flow graph and the function call graph to obtain the total control flow graph.

[0048] In the embodiment of the present application, the intra-function control flow graph constructed in step 201 and the function call graph constructed in step 202 are combined to generate a more comprehensive total control flow graph.

[0049] Among them, the total control flow graph includes the function call relationships of the entire intermediate language file and the control flow relationships between each instruction.

[0050] Step 103: Based on the total control flow graph, perform data flow analysis and improvement on the intermediate language to generate an enhanced data flow graph.

[0051] In the embodiments of the present application, based on step 102, for the intermediate language file, the present application has generated a total control flow graph, but only the control flow cannot enable the present application to judge the call path of key data. In order to obtain the call path of key data, the present application also needs to construct a data flow graph.

[0052] As a possible implementation manner, the present application calls the data flow graph generation function of the open-source static program analysis framework SVF tool to analyze and transform the intermediate language file, obtains the basic data pointing relationship graph PAG of the intermediate language file, and uses the basic data pointing relationship graph PAG as the basic data flow graph.

[0053] It should be noted that during the transformation process, pointer variables or abstract objects are set as nodes, and the relationships between pointer variables are transformed into edges. And this basic data flow graph is flow-insensitive and also context-insensitive. In addition, the basic data flow graph generated in the above steps only contains the data dependency relationships between variables, while the inducing reason for the semantic missing vulnerability in the protocol is the inappropriate or incorrect processing of the program under certain conditional statements. Therefore, the data flow graph must also contain some control flow information.

[0054] As an example, the data flow graph needs to contain the basic block content after the br instruction of the intermediate language. In this way, when tracing the call path of a certain variable, if it is used to determine the control flow, the present application can also traverse the basic block contents of different branches caused by this variable, so as to judge which instruction operations the program has performed due to different values of this variable, so as to judge whether the program processing method is appropriate.

[0055] Therefore, as a possible implementation manner, traverse each node in the basic data flow graph, obtain the node instruction corresponding to each node and the first node position a. If the node instruction corresponding to the first node position a is a jump statement, obtain the second node position b where the node instruction is located in the total control flow graph, traverse all subsequent nodes after the second node position b, and sequentially add the node corresponding to the second node position b and all its subsequent nodes behind the node at the first node position a in the basic data flow graph; if the subsequent node of the second node position is a function call statement, add a control dependency edge between the subsequent node corresponding to the second node position and the node representing the function in the basic data flow graph.

[0056] After performing the above operations on all nodes that meet the requirements, an enhanced data flow graph is obtained.

[0057] Step 104, determine the analysis target, and obtain the call path of the analysis target based on the enhanced data flow graph.

[0058] In the embodiments of the present application, based on step 103, there is already a basic graph data structure. Next, the user needs to determine their analysis target, and then trace the call path of the target on the current data flow graph. The purpose of this step is to allow the user to select a part of the analysis program according to the known situation of the vulnerability, and there is no need to analyze all variables or instructions in the program.

[0059] First, the user needs to determine the analysis target, that is, the user needs to determine the starting point of the program analysis, that is, to determine the key variables and structures related to the formation of the vulnerability, and then apply the alias analysis method to determine the aliases corresponding to the key variables and structures. Record the key variables, structures and their corresponding aliases in a linked list in memory, and name this memory Glist.

[0060] It can be understood that there are various alias analysis methods, and the present application does not make specific limitations on this.

[0061] It should be noted that through the above method, the types of mined vulnerabilities are more targeted. In the subsequent analysis process, the user traverses the key variables, structures and aliases in Glist, and traces the call paths of the key variables, structures or aliases on the existing data flow graph.

[0062] Moreover, the present application designs a variable call path tracing algorithm based on depth-first traversal to search for code segments before a variable is assigned or after it is used during the tracing process. For these two cases, specifically, starting from the current node, recursively traverse its neighbor nodes, and add the unvisited neighbor nodes to the queue. If there is a repetition between the current node and the nodes in the queue, it indicates that the current path traversal ends.

[0063] Step 105, perform a potential risk judgment on the call path of the analysis target, print the paths with potential vulnerability risks and warn the user.

[0064] It can be understood that in the above step 104, the key variables and their call paths in the kernel have been obtained. Next, it is necessary to verify the call paths to filter out the paths that may actually cause potential vulnerabilities, and print the paths and warn the user.

[0065] First, the user needs to determine the risk measurement indicators to determine the potential risk instructions, and use the Elist linked list to record all the risk instructions, and the risk instructions are related to the type of vulnerability.

[0066] It can be understood that if a path contains these risk instructions, it is considered that the path has potential vulnerability risks.

[0067] Then, traverse the call path. For each instruction in the call path, determine whether the instruction appears in the Elist linked list. If at least one instruction in the call path appears in the Elist linked list, it is determined that there is a potential vulnerability risk in the call path, and the call path needs to be printed and a vulnerability warning needs to be sent to the user.

[0068] As another possible implementation, if you want to improve the path recognition accuracy, symbolic execution can be used to filter potential vulnerability paths. This application does not provide specific descriptions about this, but only describes a solution to improve the path recognition accuracy.

[0069] In the embodiments of this application, for protocol vulnerabilities caused by semantic loss, semantic information of protocol interactions in the source code is extracted through static analysis technology, and potential risks of the interaction information are judged, which can effectively identify vulnerabilities brought by protocol interactions. This solution neither depends on expert manual analysis nor on the completeness of test cases. In addition, this method is not limited to specific analysis of one or several protocols, but adds all protocol participants involved in the protocol communication process to the analysis scope, expanding the analysis scope of the protocol. In addition, for different analysis targets, this method uses an automated method to discover protocol vulnerabilities, can process a large amount of protocol code in a short time and give potential vulnerability risks, thereby improving the efficiency of vulnerability discovery.

[0070] Figure 3 It is a block diagram of a protocol vulnerability automated mining device 10 shown according to the embodiments of this application, including: An initialization module 100, configured to determine the analysis scope of the source code and compile the source code into an intermediate language file based on the analysis scope; A control flow analysis module 200, configured to perform control flow analysis on the intermediate language file to obtain a total control flow graph; A data flow analysis module 300, configured to perform data flow analysis and improvement on the intermediate language based on the total control flow graph to generate an enhanced data flow graph; A path tracing module 400, configured to determine the analysis target and obtain the call path of the analysis target based on the enhanced data flow graph; A security check module 500, configured to perform potential risk judgment on the call path of the analysis target, print the path with potential vulnerability risk, and warn the user.

[0071] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.

[0072] Figure 4FIG. shows a schematic block diagram of an exemplary electronic device 700 that can be used to implement an embodiment of the present application. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present application described and / or claimed herein.

[0073] As Figure 4 shown, the device 700 includes a computing unit 701 that can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the device 700 can also be stored. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0074] A plurality of components in the device 700 are connected to the I / O interface 705, including: an input unit 706, such as a keyboard, a mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a magnetic disk, an optical disk, etc.; and a communication unit 709, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 709 allows the device 700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0075] The computing unit 701 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 701 executes the various methods and processes described above, such as the voice instruction response method. For example, in some embodiments, the voice instruction response method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, one or more steps of the voice instruction response method described above can be executed. Alternatively, in other embodiments, the computing unit 701 can be configured to execute the voice instruction response method in any other suitable manner (e.g., by means of firmware).

[0076] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0077] The program code for implementing the methods of the present application can be written in any combination of one or more programming languages. These program codes can be provided to the processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program code can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0078] In the context of this application, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0079] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).

[0080] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), the Internet, and a blockchain network.

[0081] A computer system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The relationship between the client and the server is created by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services ("Virtual Private Server", or simply "VPS"). The server may also be a server of a distributed system or a server combined with a blockchain.

[0082] It should be understood that various forms of processes shown above can be used, steps can be reordered, added or deleted. For example, the steps described in this application can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solution of this application can be achieved, and no limitations are imposed herein.

[0083] The above specific embodiments do not constitute a limitation on the protection scope of this application. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of this application shall be included within the protection scope of this application.

Claims

1. An automated protocol vulnerability mining method based on an enhanced data flow graph, characterized in that, Including: Determine the analysis scope of the source code, and compile the source code into an intermediate language file based on the analysis scope; Perform control flow analysis on the intermediate language file to obtain a total control flow graph; Based on the total control flow graph, perform data flow analysis and improvement on the intermediate language to generate an enhanced data flow graph; Determine the analysis target, and obtain the call path of the analysis target based on the enhanced data flow graph; Perform potential risk judgment on the call path of the analysis target, print the path with potential vulnerability risks, and warn the user.

2. The method according to claim 1, wherein The determining the analysis scope of the source code and compiling the source code into an intermediate language file based on the analysis scope includes: Determine that the analysis scope of the source code is global analysis, include all source codes of the linux kernel in the analysis scope, analyze the linux kernel as a project, and obtain the intermediate language file; Determine that the analysis scope of the source code is local analysis, include the target files in the linux kernel in the analysis scope, separately compile the target files and then link them together uniformly to obtain the intermediate language file.

3. The method according to claim 2, wherein The determining that the analysis scope of the source code is global analysis, including all source codes of the linux kernel in the analysis scope, analyzing the linux kernel as a project, and obtaining the intermediate language file, includes: Use the wllvm script under the llvm compiler project to compile the entire linux kernel to obtain a compiled kernel image vmlinux file. Among them, the Clang tool is specified as the compiler during the kernel compilation process, and llvm is used as the compiler backend; Use the extract-bc tool to extract the compiled vmlinux file into an LLVM bytecode file, and use the LLVM bytecode file as the intermediate language file.

4. The method according to claim 2, wherein The determining that the analysis scope of the source code is local analysis, including the target files in the linux kernel in the analysis scope, separately compiling the target files and then linking them together uniformly to obtain the intermediate language file, includes: Use the Clang tool to compile each target file into an intermediate language sub-file; Use the llvm-link tool to link multiple intermediate language sub-files together, and use the merged result as the intermediate language file.

5. The method according to claim 1, wherein The performing control flow analysis on the intermediate language file to obtain a total control flow graph includes: Collect the target statement of the intra-function jump statement and the next statement of the jump statement in the intermediate language file. Both the target statement and the next statement of the jump statement mark the start of a new basic block. Determine the transfer relationship between basic blocks according to the jump instruction and add control flow edges to obtain an intra-function control flow graph; Traverse the intra-function control flow graph to collect function call statements, determine the call relationship between functions according to the function call statements, and construct a function call graph; Composite the intra-function control flow graph and the function call graph to obtain the total control flow graph.

6. The method according to claim 1, wherein The performing data flow analysis and improvement on the intermediate language file based on the total control flow graph to generate an enhanced data flow graph includes: Invoke the data flow graph generation function of the open-source static program analysis framework SVF tool to analyze and transform the intermediate language file, and obtain the basic data pointing relationship graph PAG of the intermediate language file. Use this basic data pointing relationship graph PAG as the basic data flow graph. During the transformation process, pointer variables or abstract objects are set as nodes, and the relationships between pointer variables are transformed into edges; Traverse each node in the basic data flow graph to obtain the node instructions and the first node position corresponding to each node; If the node instruction corresponding to the first node position is a jump statement, obtain the second node position where the node instruction is located in the total control flow graph. Traverse all subsequent nodes after the second node position, and sequentially add the node corresponding to the second node position and all its subsequent nodes to the back of the node in the basic data flow graph where the first node position is located; If the subsequent node of the second node position is a function call statement, add a control dependence edge between the subsequent node corresponding to the second node position and the node representing the function in the basic data flow graph; After the addition, obtain the enhanced data flow graph.

7. The method according to claim 1, wherein The determination of the analysis target and obtaining the call path of the analysis target based on the enhanced data flow graph includes: Determine the key variables and structures related to the formation of vulnerabilities, apply the alias analysis method to determine the aliases corresponding to the key variables and the structures, and record the key variables, the structures, and their corresponding aliases in a linked list in memory; Traverse the key variables, the structures, or the aliases in the memory, and trace the call paths of the key variables, the structures, or the aliases in the enhanced data flow graph; Among them, during the tracing process, search the code segments before the variable is assigned or after it is used.

8. The method according to claim 1, wherein The potential risk judgment of the call path of the analysis target, printing the paths with potential vulnerability risks and warning the user includes: Determine the potential risk instructions according to the risk measurement indicators, and use the Elist linked list to record all the risk instructions; Traverse the call path and judge whether each instruction in the call path appears in the Elist linked list; If at least one instruction in the call path appears in the Elist linked list, determine that the call path has potential vulnerability risks, print the call path, and send a vulnerability warning to the user.

9. An automated protocol vulnerability mining device based on an enhanced data flow graph, characterized in that, Include: An initialization module for determining the analysis scope of the source code and compiling the source code into an intermediate language file based on the analysis scope; A control flow analysis module for performing control flow analysis on the intermediate language file to obtain a total control flow graph; A data flow analysis module for performing data flow analysis and improvement on the intermediate language based on the total control flow graph to generate an enhanced data flow graph; A path tracing module for determining the analysis target and obtaining the call path of the analysis target based on the enhanced data flow graph; A security check module for performing potential risk judgment on the call path of the analysis target, printing the paths with potential vulnerability risks, and warning the user.

10. An electronic device, characterized in that, Include: A processor and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Java program static analysis method based on control flow analysis and data flow analysis

    CN105608003A

  • Data control flow chart generation method and system and integrated circuit design method

    CN108170957A

  • Source code defect detection method and device

    CN112579469A

  • Program vulnerability detection method and device, electronic equipment and medium

    CN113312618A

  • MIPS architecture vulnerability mining method based on control flow and data flow analysis

    CN113497809A