Document risk detection method and device, storage medium and electronic equipment
By performing structural analysis and execution of PDF documents in the abstract syntax tree, the problem of low accuracy of risk detection of PDF documents in the prior art is solved, and efficient risk detection of PDF documents is achieved.
Patent Information
- Application Number
- CN202311516534.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-14
- Publication Date
- 2025-05-16
AI Technical Summary
In the prior art, risk detection is achieved by detecting JS code in PDF documents, resulting in a low accuracy of risk detection of PDF documents.
By obtaining the target document to be detected, performing structure analysis to obtain an abstract syntax tree, executing the code of the target document based on the abstract syntax tree, obtaining the target execution log, and predicting the risk of the target document based on the target execution log.
It achieves full coverage of suspicious code execution paths, and can accurately detect the risk of target documents using adversarial means such as shelling, encryption, and obfuscation, and improves the accuracy of risk detection of PDF documents.
Smart Images

Figure CN120012089A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of system security technology, and in particular to a document risk detection method and device, a storage medium and an electronic device. Background Art
[0002] Malware exploits vulnerabilities in document files to spread malicious code. Document-type malware is not an executable file itself, so it is easy to bypass existing security programs, and security programs have a high risk of false positives when detecting document-type malware. The most common malicious document type is PDF (Portable Document Format). JavaScript is one of the most convenient flexibility supported by PDF. JavaScript in PDF is used to change the content of the document in response to certain events and restrict the reader's actions. For example, pre-fill certain form fields or validate the entered form field values. However, attackers can exploit JavaScript in PDF to inject malware into PDF documents.
[0003] The earliest malicious PDF documents were designed to target the parsing function defects of PDF parsing software such as Adobe. For example, CVE-2009-4324 exploited the UAF vulnerability of Adobe's zlib decompression function to achieve code execution. Therefore, early malicious PDF documents often embedded their attack payloads directly in PDF documents or performed simple encoding processing. However, with the development of js obfuscation technology, the number of open source js obfuscators has continued to increase. The solution of directly matching PDF content to detect malicious PDF document types such as vulnerability exploits can easily be bypassed by code obfuscation behavior. Moreover, with the widespread use of js obfuscators, directly detecting the obfuscation behavior of js code in PDF documents as the basis for judging PDF malicious documents can easily lead to false positives.
[0004] With regard to the problem in the related art that risk detection is implemented by detecting js codes in PDF documents, resulting in a relatively low accuracy rate in risk detection of PDF documents, no effective solution has been proposed so far. Summary of the invention
[0005] The embodiments of the present invention provide a document risk detection method and device, a storage medium and an electronic device, so as to at least solve the technical problem in the related art that risk detection is implemented by detecting js code in PDF documents, resulting in a relatively low accuracy rate of risk detection of PDF documents.
[0006] According to one aspect of an embodiment of the present invention, a document risk detection method is provided, comprising: obtaining a target document to be detected, and performing structural analysis on the target document to obtain an abstract syntax tree; executing a target code corresponding to the target document according to the abstract syntax tree to obtain a target execution log, wherein the target code is used to display the target document in a target device; predicting the risk of the target document according to the target execution log to obtain a prediction result, wherein the prediction result is used to characterize whether the target document has a risk.
[0007] Furthermore, executing the target code corresponding to the target document according to the abstract syntax tree to obtain a target execution log includes: when executing the target code corresponding to the target document according to the abstract syntax tree, classifying the target code according to the abstract syntax tree to obtain multiple categories, wherein the multiple categories include at least: a first category, a second category and a third category, the risk of the target code in the first category is lower than the risk of the target code in the second category, and the risk of the target code in the second category is lower than the risk of the target code in the third category; determining the execution mode of the code according to the category corresponding to the code in the target code, wherein the execution mode is one of the following: real execution, simulated execution and inference execution, the inference execution refers to executing the branch code in the target code according to a preset inference strategy, and the inference strategy is used to determine the branch code to be executed; executing the target code corresponding to the target document according to the execution mode of the code and the abstract syntax tree to obtain the target execution log.
[0008] Furthermore, based on the category corresponding to the code, determining the execution mode of the code includes: if the category corresponding to the first code in the target code is the first category, determining that the execution mode of the first code is real execution; if the category corresponding to the second code in the target code is the second category, determining that the execution mode of the second code is simulated execution; if the category corresponding to the third code in the target code is the third category, determining that the execution mode of the third code is inferred execution.
[0009] Furthermore, the target code corresponding to the target document is executed according to the execution mode of the code and the abstract syntax tree to obtain the target execution log, including: executing the first code through the execution mode of the real execution and the abstract syntax tree to obtain a first execution log; executing the second code through the execution mode of the simulated execution and the abstract syntax tree to obtain a second execution log; executing the code in the third code through the execution mode of the inference execution and the abstract syntax tree to obtain a third execution log; obtaining the target execution log based on the first execution log, the second execution log and the third execution log.
[0010] Furthermore, before executing the code in the third code through the execution mode of the inference execution and the abstract syntax tree to obtain the third execution log, the method also includes: building a context execution environment for the target code based on the implementation environment of the functions and classes in the target code; building a grammatical environment for the target code based on the language grammar of the abstract syntax tree, wherein the target code is executed under the context execution environment and the grammatical environment.
[0011] Furthermore, executing the code in the third code through the execution mode of the inference execution and the abstract syntax tree to obtain a third execution log includes: determining the branch code and non-branch code in the third code according to the abstract syntax tree; executing the code in the branch code according to the execution mode of the inference execution to obtain the execution log corresponding to the branch code; executing the non-branch code according to the abstract syntax tree to obtain the execution log corresponding to the non-branch code; obtaining the third execution log based on the execution log corresponding to the branch code and the execution log corresponding to the non-branch code.
[0012] Furthermore, executing the code in the branch code according to the execution mode of the inference execution to obtain the execution log corresponding to the branch code includes: determining the target branch code to be executed in the branch code according to the preset inference strategy; determining the trigger condition corresponding to the target branch code; setting the trigger condition to true according to the preset inference strategy to execute the target branch code, and obtaining the execution log corresponding to the branch code.
[0013] Furthermore, before executing the target code corresponding to the target document according to the abstract syntax tree and obtaining the target execution log, the method also includes: determining multiple functions and multiple classes involved in the target code; determining a target function from the multiple functions and determining a target class from the multiple classes, wherein a risk value of the target function is higher than a preset threshold, and a risk value of the target class is higher than the preset threshold; rewriting the target function and the target class to obtain a rewritten target function and a rewritten target class, wherein, when executing the target code corresponding to the target document according to the abstract syntax tree, the target code implements the function corresponding to the target code by calling the rewritten target function and the rewritten target class.
[0014] Furthermore, performing structural parsing on the target document to obtain an abstract syntax tree includes: performing structural parsing on the target document to obtain a target code corresponding to the target document; parsing the target code through a target engine to obtain operation code information corresponding to the target code; and generating the abstract syntax tree based on the operation code information.
[0015] Furthermore, the risk of the target document is predicted based on the target execution log, and the prediction result includes: performing structural analysis on the target document to obtain version information corresponding to the target document and flow information corresponding to the target document; performing feature extraction on the target execution log, the version information and the flow information based on the target neural network model to obtain target feature information; performing risk prediction based on the target feature information through the target neural network model to obtain the prediction result.
[0016] According to another aspect of an embodiment of the present invention, a document risk detection method is also provided, including: obtaining a target document to be detected uploaded by a client; performing structural analysis on the target document in a cloud server to obtain an abstract syntax tree; executing a target code corresponding to the target document according to the abstract syntax tree to obtain a target execution log, wherein the target code is used to display the target document in a target device; predicting the risk of the target document according to the target execution log to obtain a prediction result, wherein the prediction result is used to characterize whether the target document has a risk; and returning the prediction result to the client.
[0017] According to another aspect of an embodiment of the present invention, a document risk detection device is also provided, including: an acquisition unit, used to acquire a target document to be detected, and perform structural analysis on the target document to obtain an abstract syntax tree; an execution unit, used to execute a target code corresponding to the target document according to the abstract syntax tree to obtain a target execution log, wherein the target code is used to display the target document in a target device; a prediction unit, used to predict the risk of the target document according to the target execution log to obtain a prediction result, wherein the prediction result is used to characterize whether the target document has a risk.
[0018] Furthermore, the execution unit includes: a classification subunit, which is used to classify the target code according to the abstract syntax tree when executing the target code corresponding to the target document according to the abstract syntax tree, so as to obtain multiple categories, wherein the multiple categories include at least: a first category, a second category and a third category, the risk of the target code in the first category is lower than the risk of the target code in the second category, and the risk of the target code in the second category is lower than the risk of the target code in the third category; a determination subunit, which is used to determine the execution mode of the code according to the category corresponding to the code in the target code, wherein the execution mode is one of the following: real execution, simulated execution and inference execution, wherein the inference execution refers to executing the branch code in the target code according to a preset inference strategy, and the inference strategy is used to determine the branch code to be executed; an execution subunit, which is used to execute the target code corresponding to the target document according to the execution mode of the code and the abstract syntax tree, so as to obtain the target execution log.
[0019] Furthermore, the determination subunit includes: a first determination module, used to determine that the execution mode of the first code in the target code is real execution if the category corresponding to the first code is the first category; a second determination module, used to determine that the execution mode of the second code in the target code is simulated execution if the category corresponding to the second code in the target code is the second category; and a third determination module, used to determine that the execution mode of the third code in the target code is inferred execution if the category corresponding to the third code is the third category.
[0020] Furthermore, the execution subunit includes: a first execution module, used to execute the first code through the execution mode of the real execution and the abstract syntax tree to obtain a first execution log; a second execution module, used to execute the second code through the execution mode of the simulated execution and the abstract syntax tree to obtain a second execution log; a third execution module, used to execute the code in the third code through the execution mode of the inference execution and the abstract syntax tree to obtain a third execution log; a fourth determination module, used to obtain the target execution log based on the first execution log, the second execution log and the third execution log.
[0021] Furthermore, the device also includes: a first building unit, which is used to build a context execution environment of the target code according to the implementation environment of the functions and classes in the target code before executing the code in the third code through the execution mode of the inference execution and the abstract syntax tree to obtain the third execution log; a second building unit, which is used to build a grammatical environment of the target code according to the language grammar of the abstract syntax tree, wherein the third code is executed under the context execution environment and the grammatical environment.
[0022] Furthermore, the third execution module includes: a first determination submodule, used to determine the branch code and non-branch code in the third code based on the abstract syntax tree; a first execution submodule, used to execute the code in the branch code according to the execution mode of the inference execution, and obtain the execution log corresponding to the branch code; a second execution submodule, used to execute the non-branch code according to the abstract syntax tree, and obtain the execution log corresponding to the non-branch code; the second determination submodule, used to obtain the third execution log based on the execution log corresponding to the branch code and the execution log corresponding to the non-branch code.
[0023] Furthermore, the first execution sub-module includes: a first determination sub-module, used to determine the target branch code to be executed in the branch code according to the preset reasoning strategy; a second determination sub-module, used to determine the trigger condition corresponding to the target branch code; a setting sub-module, used to set the trigger condition to true according to the preset reasoning strategy to execute the target branch code and obtain the execution log corresponding to the branch code.
[0024] Furthermore, the device also includes: a first determination unit, used to determine multiple functions and multiple classes involved in the target code before executing the target code corresponding to the target document according to the abstract syntax tree and obtaining the target execution log; a second determination unit, used to determine the target function from the multiple functions and to determine the target class from the multiple classes, wherein the risk value of the target function is higher than a preset threshold, and the risk value of the target class is higher than the preset threshold; a rewriting unit, used to rewrite the target function and the target class to obtain the rewritten target function and the rewritten target class, wherein, when the target code corresponding to the target document is executed according to the abstract syntax tree, the target code implements the function corresponding to the target code by calling the rewritten target function and the rewritten target class.
[0025] Furthermore, the acquisition unit includes: a first parsing subunit, used to perform structural parsing on the target document to obtain the target code corresponding to the target document; a second parsing subunit, used to parse the target code through a target engine to obtain operation code information corresponding to the target code; and a generating subunit, used to generate the abstract syntax tree based on the operation code information.
[0026] Furthermore, the prediction unit includes: a third parsing sub-unit, used to perform structural analysis on the target document to obtain version information corresponding to the target document and flow information corresponding to the target document; an extraction sub-unit, used to perform feature extraction on the target execution log, the version information and the flow information according to the target neural network model to obtain target feature information; a prediction sub-unit, used to perform risk prediction based on the target feature information through the target neural network model to obtain the prediction result.
[0027] According to another aspect of an embodiment of the present invention, a computer-readable storage medium is provided, wherein the storage medium stores a program, wherein when the program is executed, the device where the storage medium is located is controlled to execute any one of the document risk detection methods described above.
[0028] According to another aspect of an embodiment of the present invention, an electronic device is provided, including a memory storing an executable program; and a processor for running the program, wherein the program executes any one of the document risk detection methods described above when running.
[0029] In an embodiment of the present invention, a target document to be detected is obtained, and a structural analysis is performed on the target document to obtain an abstract syntax tree; a target code corresponding to the target document is executed according to the abstract syntax tree to obtain a target execution log, wherein the target code is used to display the target document in a target device; the risk of the target document is predicted according to the target execution log to obtain a prediction result, wherein the prediction result is used to characterize whether the target document has a risk, which solves the technical problem that the accuracy of risk detection of PDF documents is relatively low in the related art by detecting js code in PDF documents to achieve risk detection. In this solution, when it is determined that risk detection is required, the target document is structurally analyzed to obtain an abstract syntax tree, and then the code of the target document is executed according to the abstract syntax tree, which can achieve full coverage of the execution path of suspicious code, and can accurately detect the risk of target documents using countermeasures such as shelling, encryption, and obfuscation, thereby achieving the effect of improving the accuracy of risk detection of PDF documents. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0031] Figure 1 is a schematic diagram of a computer terminal provided according to Embodiment 1 of the present application;
[0032] Figure 2 is a flowchart of a document risk detection method provided in accordance with Embodiment 1 of the present invention;
[0033] Figure 3 is a schematic diagram of the structure of a document provided according to Embodiment 1 of the present invention;
[0034] Figure 4 is a code execution framework provided according to the first embodiment of the present invention;
[0035] Figure 5 is a flow chart of a document risk detection method provided in Embodiment 2 of the present invention;
[0036] Figure 6 is a schematic diagram of a document risk detection device provided in Embodiment 3 of the present invention;
[0037] Figure 7 It is a schematic diagram of a computer terminal provided according to Embodiment 4 of the present invention. DETAILED DESCRIPTION
[0038] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.
[0039] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0040] First, some nouns or terms that appear in the description of the embodiments of the present application are subject to the following explanations:
[0041] QuickJS: is a small-sized JavaScript engine that includes basic JS code parsing, JS context, JS Runtime, and JS bytecode execution functions.
[0042] Abstract Syntax Tree (AST): A tree representation of the abstract syntax structure of a text (usually source code) written in a formal language.
[0043] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0044] Example 1
[0045] According to an embodiment of the present invention, an embodiment of a document risk detection method is also provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0046] The method embodiment provided in the first embodiment of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1 The hardware structure block diagram of a computer terminal (or mobile device) for implementing a document risk detection method is shown. Figure 1 As shown, the computer terminal 10 (or mobile device 10) may include one or more ( Figure 1 102a, 102b, ..., 102n are used to illustrate) processor 102 (processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission module 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply and / or a camera. A person skilled in the art can understand that Figure 1 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 1 More or fewer components as shown, or with Figure 1 Different configurations are shown.
[0047] It should be noted that the one or more processors 102 and / or other data processing circuits described above may generally be referred to herein as "data processing circuits". The data processing circuits may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. In addition, the data processing circuit may be a single independent processing module, or may be incorporated in whole or in part into any of the other components in the computer terminal 10 (or mobile device). As described in the embodiments of the present application, the data processing circuit acts as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).
[0048] The memory 104 can be used to store software programs and modules of application software, such as program instructions / data storage devices corresponding to the risk detection method of documents in the embodiment of the present invention. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, to implement the vulnerability detection method of the above-mentioned application program. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 104 may further include a memory remotely arranged relative to the processor 102, and these remote memories may be connected to the computer terminal 10 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0049] The transmission module 106 is used to receive or send data via a network. The specific example of the above network may include a wireless network provided by a communication provider of the computer terminal 10. In one example, the transmission module 106 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission module 106 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0050] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the computer terminal 10 (or mobile device).
[0051] Under the above operating environment, this application provides Figure 2 The risk detection method of the document shown. Figure 2 4 is a flowchart of a document risk detection method according to Embodiment 1 of the present invention.
[0052] Step S201, obtaining a target document to be detected, and performing structural analysis on the target document to obtain an abstract syntax tree.
[0053] Optionally, the target document to be detected is clearly defined. It should be noted that the target document can be a PDF document. The risk detection of the target document is mainly to detect whether the target document is a malicious document. After obtaining the target document, the target document is structurally parsed to extract the encoded content, stream information, JS code, etc. in the PDF document. The typical structure of a PDF document is as follows: Figure 3 As shown, it consists of four parts:
[0054] 1. Header: Contains the PDF version information under the ISO standard.
[0055] 2. Body: It contains many objects that define the operations to be performed by the file, as well as embedded data, which can be images, text, script code, etc. Content displayed to the user is usually contained in this section. Operations such as data decryption or decompression can be defined in objects and are usually performed when the file is rendered. PDF files can be updated after initial creation, which will prompt a new body and cross-reference sections to be appended at the end.
[0056] 3. Reference table (x-ref table le): The reference table provides a list of offsets for each object in the file, while allowing random access to any object in the file. Since the PDF standard allows incremental updates to documents, this can be achieved through the presence of an x-ref table. When a document is updated, additional x-ref tables and Tracker information are appended to the end of the document.
[0057] 4.Trai ler: It shows how the pdf parser should render the file by pointing it to the object identified by the / Root tag, which is the first object to be rendered. It also contains the offset where the x-ref table starts. At the end of the Trainer, there is an end of file string "%%EOF" which is the last line of the file.
[0058] The JS code is obtained by parsing the structure of the target document, and the above-mentioned abstract syntax tree is constructed based on the JS code.
[0059] Step S202: Execute the target code corresponding to the target document according to the abstract syntax tree to obtain a target execution log, wherein the target code is used to display the target document in the target device.
[0060] Optionally, after obtaining the above-mentioned abstract syntax tree, the target code corresponding to the target document is executed through the abstract syntax tree to obtain the target execution log. It should be noted that the target execution log includes but is not limited to information such as assigned variables, used variables, the number of function calls and string concatenations such as numbers and strings, sensitive behavior content, etc.
[0061] It should be noted that when executing the target code, different execution methods can be used according to the type of code. For example, conventional variable assignment, calculation, function call and other codes do not have the risk of malicious attacks, so the real execution method can be used to execute such codes directly. For example, IO operations, database operations, network operations, thread creation and other codes that may implicitly transmit information or obtain external resources in disguised logic have a certain risk of malicious attacks, so the simulated execution method can be used to execute such codes. For example, externally controllable loops, branches, recursion and other statements have a high risk of malicious attacks, so the inference execution method can be used to execute such codes.
[0062] Step S203, predicting the risk of the target document according to the target execution log to obtain a prediction result, wherein the prediction result is used to characterize whether the target document has a risk.
[0063] Optionally, after obtaining the above target execution log, the target execution log can be feature extracted by using an AI algorithm, and risk prediction can be performed based on the extracted feature vector to obtain the above prediction result. The prediction result is used to characterize whether the target document has risks. If the target document has risks, the target document can be considered to be a malicious document.
[0064] To sum up, when it is determined that risk detection is needed, the target document is structurally parsed to obtain an abstract syntax tree, and then the code of the target document is executed according to the abstract syntax tree. This can achieve full coverage of the execution path of suspicious code, and can accurately detect the risk of target documents that use conventional countermeasures such as packing, encryption, and obfuscation, thereby achieving the effect of improving the accuracy of risk detection of PDF documents.
[0065] When executing the target code, in order to avoid the problem that some codes escape and are not executed, in the risk detection method for documents provided in Example 1 of the present application, the target code corresponding to the target document is executed according to the abstract syntax tree, and the target execution log is obtained, including: when executing the target code corresponding to the target document according to the abstract syntax tree, the target code is classified according to the abstract syntax tree to obtain multiple categories, wherein the multiple categories at least include: a first category, a second category and a third category, the risk of the target code in the first category is lower than the risk of the target code in the second category, and the risk of the target code in the second category is lower than the risk of the target code in the third category; according to the category corresponding to the code in the target code, the execution mode of the code is determined, wherein the execution mode is one of the following: real execution, simulated execution and inference execution, inference execution refers to executing branch code in the target code according to a preset inference strategy; according to the execution mode of the code and the abstract syntax tree, the target code corresponding to the target document is executed to obtain the target execution log.
[0066] In an optional embodiment, determining the execution mode of the code based on the category corresponding to the code includes: if the category corresponding to the first code in the target code is the first category, determining that the execution mode of the first code is real execution; if the category corresponding to the second code in the target code is the second category, determining that the execution mode of the second code is simulated execution; if the category corresponding to the third code in the target code is the third category, determining that the execution mode of the third code is inferred execution.
[0067] Optionally, when executing the target code corresponding to the target document according to the abstract syntax tree, it is necessary to classify the target code, that is, to classify the target code according to risk.
[0068] In an optional embodiment, the target code can be divided into three categories, namely the first category, the second category and the third category. It should be noted that the risk of the target code of the first category is lower than the risk of the target code of the second category, and the risk of the target code of the second category is lower than the risk of the target code of the third category. For example, conventional variable assignment, operation, function call and other codes do not have the risk of malicious attacks and can be classified into the first category. For example, IO operations, database operations, network operations, thread creation and other logical codes that may implicitly transmit information or obtain external resources in disguise have a certain risk of malicious attacks and can be classified into the second category. For example, externally controllable loops, branches, recursion and other statements have a higher risk of malicious attacks and can be classified into the third category.
[0069] After determining the category corresponding to the code, the execution mode of the code is determined according to the category corresponding to the code. In an optional embodiment, if the category corresponding to the first code is the first category, the execution mode of the first code is real execution; if the category corresponding to the second code is the second category, the execution mode of the second code is simulated execution; if the category corresponding to the third code is the third category, the execution mode of the third code is inferred execution.
[0070] It should be noted that inference execution refers to executing branch codes in the target code according to a preset inference strategy, that is, using some heuristic inference strategies to avoid the problem that some codes cannot be executed due to escape.
[0071] It should be noted that the above code corresponding category is the operation type corresponding to the code, for example, assignment operation. The operation type corresponding to the code is written in the syntax tree (i.e. the above abstract syntax tree) parsing engine. During the execution of the code, the execution mode is determined according to the above operation type.
[0072] After determining the execution modes of different types of codes, the target code corresponding to the target document is executed according to the execution mode of the codes and the abstract syntax tree to obtain the above-mentioned target execution log.
[0073] By adopting different execution methods according to the code categories, it is ensured that the test system will not be attacked due to the actual execution of malicious code during the test execution, and at the same time, the technical effect of full coverage of the execution path of suspicious code is guaranteed.
[0074] Optionally, in the document risk detection method provided in Example 1 of the present application, the target code corresponding to the target document is executed according to the code execution mode and the abstract syntax tree to obtain the target execution log, including: executing the first code through the real execution mode and the abstract syntax tree to obtain the first execution log; executing the second code through the simulated execution mode and the abstract syntax tree to obtain the second execution log; executing the code in the third code through the inference execution mode and the abstract syntax tree to obtain the third execution log; obtaining the target execution log based on the first execution log, the second execution log and the third execution log.
[0075] Optionally, the target code corresponding to the target document executed according to the code execution mode and the abstract syntax tree includes: executing the first code through the real execution mode and the abstract syntax tree, executing the second code through the simulated execution mode and the abstract syntax tree, and executing the third code through the inference execution mode and the abstract syntax tree, and then obtaining the above-mentioned target execution log according to the code execution process and execution result. It should be noted that when the target code is executed, it is executed as a whole, but some code operations are directly implemented (i.e., the real execution mentioned above), some code operations need to be simulated (i.e., the simulated execution mentioned above), and some code operations need to be inferred (i.e., the inference execution mentioned above).
[0076] In order to be able to execute the third code through the execution mode of inferential execution, therefore, in the risk detection method of the document provided in the first embodiment of the present application, before executing the code in the third code through the execution mode of inferential execution and the abstract syntax tree to obtain the third execution log, the method also includes: building a context execution environment of the target code based on the implementation environment of the functions and classes in the target code; building a grammatical environment of the target code based on the language grammar of the abstract syntax tree, wherein the target code is executed under the context execution environment and the grammatical environment.
[0077] Optionally, in order to be able to execute the target code through the inferential execution method, it is necessary to build a context execution environment for the target code according to the implementation environment of the functions and classes in the target code, and store the information required for the implementation of various functions and classes in the context execution environment, such as variable names and current values; current function stack, parameters; current code block; currently visible symbol list, etc. This information can be managed in different ways to form a reasonable hierarchical relationship.
[0078] After obtaining the above-mentioned context execution environment, the grammatical environment of the target code is built according to the language grammar of the abstract syntax tree. For example, the language grammar of the abstract syntax tree can be generally divided into the following categories:
[0079] 1. Variables, including variable types, creation methods, default values, etc.
[0080] 2. Expression calculation, including unary expressions (++,--,!), binary expressions (+-* / ), ternary expressions (?:), etc.;
[0081] 3. Process control, including loop, branch, goto, etc.;
[0082] 4. Functions, including ordinary functions, closure functions, variable parameter functions, yield functions, etc.;
[0083] 5. Classes and objects, including construction and destruction, inheritance, polymorphism, etc.
[0084] 6. Exception handling, including try, throw, catch, etc.
[0085] After obtaining the above-mentioned context execution environment and grammatical environment, the target code is executed in the context execution environment and grammatical environment.
[0086] The above steps can ensure the normal execution of the target code, making it easier to perform risk detection through execution logs later.
[0087] How to execute the third code by using the inference execution method is crucial. Therefore, in the risk detection method of the document provided in Example 1 of the present application, the code in the third code is executed by using the inference execution method and the abstract syntax tree to obtain the third execution log including: determining the branch code and non-branch code in the third code according to the abstract syntax tree; executing the code in the branch code according to the inference execution method to obtain the execution log corresponding to the branch code, and the inference strategy is used to determine the branch code to be executed; executing the non-branch code according to the abstract syntax tree to obtain the execution log corresponding to the non-branch code; obtaining the third execution log according to the execution log corresponding to the branch code and the execution log corresponding to the non-branch code.
[0088] Optionally, branch code and non-branch code are determined based on the above-mentioned abstract syntax tree. In an optional embodiment, branch perception capability can be used to make malicious code that attempts to use branch countermeasures to bypass dynamic detection nowhere to hide. Since branch code may require specific data and specific conditions to be executed, it is necessary to distinguish between branch code and non-branch code. Because the branch code may or may not be entered in actual execution, it is necessary to execute the branch code through the execution method of inference execution.
[0089] Therefore, after determining the branch code and the non-branch code, the branch code is executed through the trigger condition corresponding to the branch code and the execution mode of the inference execution, and the non-branch code is executed through the abstract syntax tree to obtain the above-mentioned third execution log.
[0090] It should be noted that when executing non-branch code, in order to avoid the problem of code escape and not being executed, the heuristic reasoning strategy under inferential execution can also be used to execute the non-branch code to ensure comprehensive coverage of the non-branch code.
[0091] In order to be able to execute branch codes more comprehensively, in the risk detection method for documents provided in Example 1 of the present application, the code in the branch code is executed according to the execution mode of reasoning execution, and the execution log corresponding to the branch code is obtained, including: determining the target branch code to be executed in the branch code according to a preset reasoning strategy; determining the trigger condition corresponding to the target branch code; setting the trigger condition to true according to the preset reasoning strategy to execute the target branch code, and obtaining the execution log corresponding to the branch code.
[0092] Optionally, whether the branch code needs to be executed is determined by a preset reasoning strategy in the reasoning execution, that is, the target branch code that needs to be executed in the branch code is determined by the preset reasoning strategy.
[0093] For example, a=1+1; if(a==3){do something} / / This is to determine the branch that does not need to be entered.
[0094] For example, a = get_from_internet() / / Get a from the Internet
[0095] b = base64_decode (a) / / decode a
[0096] if(b==“aaa”){do_someth ing}
[0097] The preset reasoning strategy here (for example, variable reasoning strategy) will know that a is a taint that requires a network request, and b is obtained from the pollution propagation of a, so b is obtained as a tainted value in the conditional judgment. At this time, it is necessary to control the execution of the corresponding do_something branch.
[0098] After determining the target branch code to be executed, the execution of the target branch code is achieved by setting the trigger condition corresponding to the target branch code to true, and the execution log corresponding to the above branch code is obtained based on the execution process, execution path and execution result of the target branch code.
[0099] It should be noted that the above-mentioned reasoning strategies include but are not limited to variable reasoning strategies, conditional expression reasoning strategies, environment reasoning strategies, and uncertain value reasoning strategies.
[0100] The heuristic reasoning strategy can better cover the use of branch escape, environment verification and other strategies to prevent malicious behavior from running out during sandbox and simulation execution, resulting in low risk detection accuracy.
[0101] In order to avoid the problem that malicious code in the target code launches malicious attacks on the test system when executing the target code, in the document risk detection method provided in Example 1 of the present application, before executing the target code corresponding to the target document according to the abstract syntax tree and obtaining the target execution log, the method also includes: determining multiple functions and multiple classes involved in the target code; determining a target function from the multiple functions and determining a target class from the multiple classes, wherein the risk value of the target function is higher than a preset threshold, and the risk value of the target class is higher than the preset threshold; rewriting the target function and the target class to obtain a rewritten target function and a rewritten target class, wherein, when the target code corresponding to the target document is executed according to the abstract syntax tree, the target code implements the function corresponding to the target code by calling the rewritten target function and the rewritten target class.
[0102] Optionally, multiple functions and multiple classes involved in the target code are determined. It should be noted that the above-mentioned multiple functions and multiple classes are the functions and classes that need to be called when the target code is executed. It should be noted that most of the functions and classes do not need to be rewritten. In the embodiment of the present application, only special functions and classes (i.e., the above-mentioned target functions and target classes) are rewritten, and other functions will directly call the functions in the standard virtual machine while ensuring safety. It should be noted that the risk value of the above-mentioned target function is higher than the preset threshold, and the risk value of the target class is higher than the preset threshold. For example, these special functions and classes include: dangerous functions such as eval, io system, network, database and other functions, and global functions. Safe functions that do not need to be rewritten include string operation functions, such as split; digital operation functions log, etc.; encoding and decoding operations unescape, etc., as well as other safe functions and classes that will not cause cross-contamination of samples and will not escape.
[0103] By rewriting functions and classes, it is possible to avoid malicious attacks on the test system by malicious code in the target code when the target code is executed. It is also possible to track the execution path and execution results of functions and classes, and more accurately detect whether the target document is a malicious document.
[0104] In order to improve the accuracy of the abstract syntax tree, in the document risk detection method provided in Example 1 of the present application, the target document is structurally parsed to obtain an abstract syntax tree, including: structurally parsing the target document to obtain a target code corresponding to the target document; parsing the target code through a target engine to obtain operation code information corresponding to the target code; and generating an abstract syntax tree based on the operation code information.
[0105] Optionally, the target document is structurally parsed to obtain the target code corresponding to the target document, i.e., the JS code. After obtaining the JS code, the open source small JS engine QuickJS (i.e., the above-mentioned target engine) is used to parse the JS code. QuickJS is a small JS engine that includes basic JS code parsing, JS context, JS Runtime, and JS bytecode execution functions. The target code is parsed by QuickJS to obtain the corresponding opcode information, and finally the above-mentioned abstract syntax tree is generated based on the opcode information. It should be noted that the above-mentioned abstract syntax tree can also be pruned to more accurately detect risks.
[0106] In an optional embodiment, the Figure 4 The framework diagram shown implements code speculation execution. Figure 4The framework in includes the basic support layer, runtime simulation layer, syntax simulation layer, correction reasoning capability and capability provision layer. The basic support layer includes QuickJS, configuration management, performance analysis and standard third-party libraries (Zipfile, Json, etc.). In the basic support layer, it is necessary to complete the parsing of JS code and generate an abstract syntax tree (AST). At the same time, since some standard functions have been built into QuickJS, these functions and classes can be called when the code is executed. This solution only rewrites special functions and classes, and other functions will directly call functions in the standard virtual machine while ensuring safety. These special functions and classes include: dangerous functions such as eval need to be rewritten for special processing; functions such as io system, network, database, etc. need to be rewritten to simulate simulation; global functions need to be rewritten to host their behavior logic. Safe functions that do not need to be rewritten include string operation functions, such as split; digital operation functions such as log; encoding and decoding operations such as unescape, as well as other safe functions and classes that will not cause cross-contamination of samples and will not escape.
[0107] The runtime simulation layer includes variable / scope management, branch / layer management, built-in symbol management, and user symbol management. The runtime simulation layer is used to implement functions and classes called by code, that is, the above-mentioned context execution environment.
[0108] The syntax simulation layer includes expression calculation support, loop / branch support, function support, and class support, and supports the language syntax of AST through the syntax simulation layer.
[0109] The correction reasoning capabilities include variable reasoning mechanism, conditional expression reasoning mechanism, environment reasoning mechanism and uncertain value reasoning mechanism. The reasoning correction layer has built-in various heuristic reasoning strategies. The built-in strategies are mainly used to cover some malicious codes that use branch escape, environment verification and other strategies.
[0110] The capability provision layer includes AST generation, behavior log export, and exception information export.
[0111] The above framework can cover malicious documents that use conventional countermeasures such as packing, encryption, and obfuscation, which cannot be accurately detected by static detection; compared with the malicious PDF document detection solution based on dynamic sandbox, this solution can increase the coverage of malicious documents that use environmental verification, branch bypass and other means, and achieve the technical effect of high detection and low false alarms.
[0112] In order to improve the accuracy of risk detection, in the document risk detection method provided in Example 1 of the present application, the risk of the target document is predicted based on the target execution log, and the prediction result includes: structural analysis of the target document to obtain version information corresponding to the target document and flow information corresponding to the target document; feature extraction of the target execution log, version information and flow information based on the target neural network model to obtain target feature information; risk prediction is performed based on the target feature information through the target neural network model to obtain a prediction result.
[0113] Optionally, after obtaining the target execution log, the target execution log and the version information and flow information corresponding to the target document are input into the target neural network model. It should be noted that the above-mentioned flow information is stream information, and stream is used to store binary data such as pictures and fonts.
[0114] The target neural network model is used to extract features from the target execution log, version information, and flow information, and risk prediction is performed based on the feature information to obtain the prediction results.
[0115] In the risk detection method of the document provided in the embodiment of the present application, the target document to be detected is obtained, and the target document is structurally parsed to obtain an abstract syntax tree; the target code corresponding to the target document is executed according to the abstract syntax tree to obtain a target execution log, wherein the target code is used to display the target document in the target device; the risk of the target document is predicted according to the target execution log to obtain a prediction result, wherein the prediction result is used to characterize whether the target document has a risk step, which solves the technical problem that the accuracy of risk detection of PDF documents is relatively low by detecting js code in PDF documents in the related art. In this scheme, when it is determined that risk detection is required, the target document is structurally parsed to obtain an abstract syntax tree, and then the code of the target document is executed according to the abstract syntax tree, which can achieve full coverage of the execution path of suspicious code, and can accurately detect the risk of target documents using conventional countermeasures such as shelling, encryption, and obfuscation, thereby achieving the effect of improving the accuracy of risk detection of PDF documents.
[0116] It should be noted that, for the above-mentioned method embodiments, for the sake of simplicity, they are all described as a series of action combinations, but those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present invention.
[0117] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus a necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, a magnetic disk, or an optical disk), and includes a number of instructions for a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods of various embodiments of the present invention.
[0118] Example 2
[0119] According to an embodiment of the present invention, an embodiment of a document risk detection method is also provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0120] This application provides Figure 5 The risk detection method of the document shown. Figure 5 4 is a flowchart of a document risk detection method according to the second embodiment of the present invention.
[0121] Step S501, obtaining the target document to be detected uploaded by the client;
[0122] Step S502: perform structural analysis on the target document in the cloud server to obtain an abstract syntax tree; execute the target code corresponding to the target document according to the abstract syntax tree to obtain a target execution log, wherein the target code is used to display the target document in the target device; predict the risk of the target document according to the target execution log to obtain a prediction result, wherein the prediction result is used to characterize whether the target document has risks;
[0123] Step S503: Return the prediction result to the client.
[0124] In the cloud server, the specific method of document risk detection is the same as the method in Example 1, and will not be repeated here.
[0125] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the described order of actions, because according to the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present application.
[0126] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods of each embodiment of the present application.
[0127] Example 3
[0128] According to an embodiment of the present invention, a document risk detection device for implementing the above document risk detection method is also provided. Figure 6 As shown, the device includes: an acquisition unit 601, an execution unit 602 and a prediction unit 603.
[0129] The acquisition unit 601 is used to acquire a target document to be detected, and perform structural analysis on the target document to obtain an abstract syntax tree;
[0130] An execution unit 602 is used to execute a target code corresponding to a target document according to the abstract syntax tree to obtain a target execution log, wherein the target code is used to display the target document in a target device;
[0131] The prediction unit 603 is used to predict the risk of the target document according to the target execution log to obtain a prediction result, wherein the prediction result is used to characterize whether the target document has a risk.
[0132] In the risk detection device for documents provided in the third embodiment of the present application, the target document to be detected is obtained by the acquisition unit 601, and the target document is structurally parsed to obtain an abstract syntax tree; the execution unit 602 executes the target code corresponding to the target document according to the abstract syntax tree to obtain a target execution log, wherein the target code is used to display the target document in the target device; the prediction unit 603 predicts the risk of the target document according to the target execution log to obtain a prediction result, wherein the prediction result is used to characterize whether the target document has a risk, and solves the technical problem that the accuracy of risk detection of PDF documents is relatively low by detecting the js code in the PDF document in the related art. In this scheme, when it is determined that risk detection is required, the target document is structurally parsed to obtain an abstract syntax tree, and then the code of the target document is executed according to the abstract syntax tree, which can achieve full coverage of the execution path of the suspicious code, and can accurately detect the risk of the target document using conventional countermeasures such as shelling, encryption, and obfuscation, thereby achieving the effect of improving the accuracy of risk detection of PDF documents.
[0133] Optionally, in the document risk detection device provided in Example 3 of the present application, the execution unit includes: a classification subunit, which is used to classify the target code according to the abstract syntax tree when executing the target code corresponding to the target document according to the abstract syntax tree, so as to obtain multiple categories, wherein the multiple categories include at least: a first category, a second category and a third category, the risk of the target code in the first category is lower than the risk of the target code in the second category, and the risk of the target code in the second category is lower than the risk of the target code in the third category; a determination subunit, which is used to determine the execution mode of the code according to the category corresponding to the code in the target code, wherein the execution mode is one of the following: real execution, simulation execution and inference execution, inference execution refers to executing branch code in the target code according to a preset inference strategy, and the inference strategy is used to determine the branch code to be executed; an execution subunit, which is used to execute the target code corresponding to the target document according to the execution mode of the code and the abstract syntax tree, so as to obtain a target execution log.
[0134] Optionally, in the risk detection device for a document provided in Example 3 of the present application, the determination subunit includes: a first determination module, used to determine that the execution mode of the first code in the target code is real execution if the category corresponding to the first code in the target code is the first category; a second determination module, used to determine that the execution mode of the second code in the target code is simulated execution if the category corresponding to the second code in the target code is the second category; and a third determination module, used to determine that the execution mode of the third code in the target code is inferred execution if the category corresponding to the third code in the target code is the third category.
[0135] Optionally, in the risk detection device for a document provided in Example 3 of the present application, the execution subunit includes: a first execution module, used to execute the first code through the execution mode of real execution and an abstract syntax tree to obtain a first execution log; a second execution module, used to execute the second code through the execution mode of simulated execution and an abstract syntax tree to obtain a second execution log; a third execution module, used to execute the code in the third code through the execution mode of inference execution and an abstract syntax tree to obtain a third execution log; and a fourth determination module, used to obtain a target execution log based on the first execution log, the second execution log and the third execution log.
[0136] Optionally, in the document risk detection device provided in Example 3 of the present application, the device also includes: a first building unit, used to build a context execution environment of the target code according to the implementation environment of functions and classes in the target code before executing the code in the third code through the execution mode of inference execution and the abstract syntax tree to obtain the third execution log; a second building unit, used to build a grammatical environment of the target code based on the language grammar of the abstract syntax tree, wherein the target code is executed under the context execution environment and the grammatical environment.
[0137] Optionally, in the risk detection device for a document provided in Example 3 of the present application, the third execution module includes: a first determination sub-module, used to determine the branch code and non-branch code in the third code based on the abstract syntax tree; a first execution sub-module, used to execute the code in the branch code according to the execution mode of inference execution, and obtain the execution log corresponding to the branch code; a second execution sub-module, used to execute the non-branch code according to the abstract syntax tree, and obtain the execution log corresponding to the non-branch code; the second determination sub-module, used to obtain the third execution log based on the execution log corresponding to the branch code and the execution log corresponding to the non-branch code.
[0138] Optionally, in the document risk detection device provided in Example 3 of the present application, the first execution sub-module includes: a first determination sub-module, used to determine the target branch code to be executed in the branch code based on a preset reasoning strategy; a second determination sub-module, used to determine the trigger condition corresponding to the target branch code; a setting sub-module, used to set the trigger condition to true according to the preset reasoning strategy to execute the target branch code and obtain the execution log corresponding to the branch code.
[0139] Optionally, in the document risk detection device provided in Example 3 of the present application, the device also includes: a first determination unit, used to determine multiple functions and multiple classes involved in the target code before executing the target code corresponding to the target document according to the abstract syntax tree and obtaining the target execution log; a second determination unit, used to determine the target function from the multiple functions and to determine the target class from the multiple classes, wherein the risk value of the target function is higher than a preset threshold, and the risk value of the target class is higher than the preset threshold; a rewriting unit, used to rewrite the target function and the target class to obtain a rewritten target function and a rewritten target class, wherein, when the target code corresponding to the target document is executed according to the abstract syntax tree, the target code implements the function corresponding to the target code by calling the rewritten target function and the rewritten target class.
[0140] Optionally, in the document risk detection device provided in Example 3 of the present application, the acquisition unit includes: a first parsing sub-unit, used to perform structural parsing on the target document to obtain a target code corresponding to the target document; a second parsing sub-unit, used to parse the target code through a target engine to obtain operation code information corresponding to the target code; and a generation sub-unit, used to generate an abstract syntax tree based on the operation code information.
[0141] Optionally, in the document risk detection device provided in Example 3 of the present application, the prediction unit includes: a third parsing sub-unit, used to perform structural analysis on the target document to obtain version information corresponding to the target document and flow information corresponding to the target document; an extraction sub-unit, used to perform feature extraction on the target execution log, version information and flow information according to the target neural network model to obtain target feature information; a prediction sub-unit, used to perform risk prediction based on the target feature information through the target neural network model to obtain a prediction result.
[0142] It should be noted that the acquisition unit 601, execution unit 602 and prediction unit 603 described above correspond to steps S201 to S203 in Embodiment 1, and the three units and corresponding steps implement the same examples and application scenarios, but are not limited to the contents disclosed in Embodiment 1 described above. It should be noted that the above modules, as part of the device, can be run in the computer terminal 10 provided in Embodiment 1.
[0143] Example 4
[0144] The embodiment of the present invention can provide a computer terminal, which can be any computer terminal device in a computer terminal group. Optionally, in this embodiment, the computer terminal can also be replaced by a terminal device such as a mobile terminal.
[0145] Optionally, in this embodiment, the computer terminal may be located in at least one network device among a plurality of network devices of the computer network.
[0146] In this embodiment, the above-mentioned computer terminal can execute the program code of the following steps in the vulnerability detection method of the application: obtain the target document to be detected, and perform structural analysis on the target document to obtain an abstract syntax tree; execute the target code corresponding to the target document according to the abstract syntax tree to obtain a target execution log, wherein the target code is used to display the target document in the target device; predict the risk of the target document according to the target execution log to obtain a prediction result, wherein the prediction result is used to characterize whether there is a risk in the target document.
[0147] Optionally, the above-mentioned computer terminal can execute the program code of the following steps in the vulnerability detection method of the application: executing the target code corresponding to the target document according to the abstract syntax tree, and obtaining the target execution log includes: when executing the target code corresponding to the target document according to the abstract syntax tree, classifying the target code according to the abstract syntax tree to obtain multiple categories, wherein the multiple categories at least include: a first category, a second category and a third category, the risk of the target code in the first category is lower than the risk of the target code in the second category, and the risk of the target code in the second category is lower than the risk of the target code in the third category; determining the execution mode of the code according to the category corresponding to the code in the target code, wherein the execution mode is one of the following: real execution, simulated execution and inference execution, inference execution refers to executing the branch code in the target code according to a preset inference strategy, and the inference strategy is used to determine the branch code to be executed; executing the target code corresponding to the target document according to the execution mode of the code and the abstract syntax tree to obtain the target execution log.
[0148] Optionally, the above-mentioned computer terminal can execute the program code of the following steps in the vulnerability detection method of the application program: determining the execution mode of the code based on the category corresponding to the code includes: if the category corresponding to the first code in the target code is the first category, then determining that the execution mode of the first code is real execution; if the category corresponding to the second code in the target code is the second category, then determining that the execution mode of the second code is simulated execution; if the category corresponding to the third code in the target code is the third category, then determining that the execution mode of the third code is inference execution.
[0149] Optionally, the above-mentioned computer terminal can execute the program code of the following steps in the vulnerability detection method of the application: executing the target code corresponding to the target document according to the execution mode of the code and the abstract syntax tree, and obtaining the target execution log includes: executing the first code through the execution mode of real execution and the abstract syntax tree to obtain the first execution log; executing the second code through the execution mode of simulated execution and the abstract syntax tree to obtain the second execution log; executing the code in the third code through the execution mode of inference execution and the abstract syntax tree to obtain the third execution log; obtaining the target execution log based on the first execution log, the second execution log and the third execution log.
[0150] Optionally, the above-mentioned computer terminal can execute the program code of the following steps in the vulnerability detection method of the application: before executing the code in the third code through the execution mode of inference execution and the abstract syntax tree to obtain the third execution log, building the context execution environment of the target code according to the implementation environment of the functions and classes in the target code; building the grammatical environment of the target code according to the language grammar of the abstract syntax tree, wherein the target code is executed in the context execution environment and the grammatical environment.
[0151] Optionally, the above-mentioned computer terminal can execute the program code of the following steps in the vulnerability detection method of the application program: execute the code in the third code through the execution mode of inference execution and the abstract syntax tree, and obtain the third execution log including: determine the branch code and non-branch code in the third code according to the abstract syntax tree; execute the code in the branch code according to the execution mode of inference execution to obtain the execution log corresponding to the branch code; execute the non-branch code according to the abstract syntax tree to obtain the execution log corresponding to the non-branch code; obtain the third execution log according to the execution log corresponding to the branch code and the execution log corresponding to the non-branch code.
[0152] Optionally, the above-mentioned computer terminal can execute the program code of the following steps in the vulnerability detection method of the application: execute the code in the branch code according to the execution mode of inference execution, and obtain the execution log corresponding to the branch code, including: determining the target branch code to be executed in the branch code according to the preset reasoning strategy; determining the trigger condition corresponding to the target branch code; setting the trigger condition to true according to the preset reasoning strategy to execute the target branch code, and obtaining the execution log corresponding to the branch code.
[0153] Optionally, the above-mentioned computer terminal can execute the program code of the following steps in the vulnerability detection method of the application: before executing the target code corresponding to the target document according to the abstract syntax tree and obtaining the target execution log, determine the multiple functions and multiple classes involved in the target code; determine the target function from the multiple functions and determine the target class from the multiple classes, wherein the risk value of the target function is higher than a preset threshold, and the risk value of the target class is higher than the preset threshold; rewrite the target function and the target class to obtain the rewritten target function and the rewritten target class, wherein, when the target code corresponding to the target document is executed according to the abstract syntax tree, the target code implements the function corresponding to the target code by calling the rewritten target function and the rewritten target class.
[0154] Optionally, the above-mentioned computer terminal can execute the program code of the following steps in the vulnerability detection method of the application: structurally parsing the target document to obtain an abstract syntax tree includes: structurally parsing the target document to obtain a target code corresponding to the target document; parsing the target code through a target engine to obtain operation code information corresponding to the target code; generating an abstract syntax tree based on the operation code information.
[0155] Optionally, the above-mentioned computer terminal can execute the program code of the following steps in the vulnerability detection method of the application: predicting the risk of the target document based on the target execution log, and obtaining the prediction result includes: performing structural analysis on the target document to obtain version information corresponding to the target document and flow information corresponding to the target document; performing feature extraction on the target execution log, version information and flow information based on the target neural network model to obtain target feature information; performing risk prediction based on the target feature information through the target neural network model to obtain a prediction result.
[0156] Optionally, Figure 7 is a structural block diagram of a computer terminal according to an embodiment of the present invention. Figure 7 As shown, the computer terminal 10 may include: one or more ( Figure 7 (only one is shown) processor 102 and memory 104.
[0157] Among them, the memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the security vulnerability detection method and device in the embodiment of the present invention. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, the detection method of the above-mentioned system vulnerability attack is realized. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include a memory remotely arranged relative to the processor, and these remote memories may be connected to the terminal A via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.
[0158] The processor can call the information and application stored in the memory through the transmission device to execute the following steps: obtain the target document to be detected, and perform structural analysis on the target document to obtain an abstract syntax tree; execute the target code corresponding to the target document according to the abstract syntax tree to obtain a target execution log, wherein the target code is used to display the target document in the target device; predict the risk of the target document according to the target execution log to obtain a prediction result, wherein the prediction result is used to characterize whether there is a risk in the target document.
[0159] Optionally, the processor may also execute program code of the following steps: executing the target code corresponding to the target document according to the abstract syntax tree, and obtaining a target execution log including: when executing the target code corresponding to the target document according to the abstract syntax tree, classifying the target code according to the abstract syntax tree to obtain multiple categories, wherein the multiple categories include at least: a first category, a second category and a third category, the risk of the target code in the first category is lower than the risk of the target code in the second category, and the risk of the target code in the second category is lower than the risk of the target code in the third category; determining the execution mode of the code according to the category corresponding to the code in the target code, wherein the execution mode is one of the following: real execution, simulated execution and inference execution, inference execution refers to executing branch code in the target code according to a preset inference strategy, and the inference strategy is used to determine the branch code to be executed; executing the target code corresponding to the target document according to the execution mode of the code and the abstract syntax tree to obtain a target execution log.
[0160] Optionally, the processor may also execute the following steps of program code: determining the execution mode of the code based on the category corresponding to the code includes: if the category corresponding to the first code in the target code is the first category, determining that the execution mode of the first code is real execution; if the category corresponding to the second code in the target code is the second category, determining that the execution mode of the second code is simulated execution; if the category corresponding to the third code in the target code is the third category, determining that the execution mode of the third code is inferred execution.
[0161] Optionally, the processor may also execute the following steps of program code: executing the target code corresponding to the target document according to the code execution method and the abstract syntax tree, and obtaining the target execution log including: executing the first code through the real execution method and the abstract syntax tree to obtain the first execution log; executing the second code through the simulated execution method and the abstract syntax tree to obtain the second execution log; executing the code in the third code through the inference execution method and the abstract syntax tree to obtain the third execution log; obtaining the target execution log based on the first execution log, the second execution log and the third execution log.
[0162] Optionally, the processor may also execute the following program code: before executing the code in the third code through the inference execution mode and the abstract syntax tree to obtain the third execution log, building the context execution environment of the target code according to the implementation environment of the functions and classes in the target code; building the grammatical environment of the target code according to the language grammar of the abstract syntax tree, wherein the target code is executed under the context execution environment and the grammatical environment.
[0163] Optionally, the processor may also execute the program code of the following steps: executing the code in the third code through the execution mode of inferential execution and the abstract syntax tree, and obtaining the third execution log including: determining the branch code and non-branch code in the third code according to the abstract syntax tree; executing the code in the branch code according to the execution mode of inferential execution, and obtaining the execution log corresponding to the branch code; executing the non-branch code according to the abstract syntax tree, and obtaining the execution log corresponding to the non-branch code; obtaining the third execution log according to the execution log corresponding to the branch code and the execution log corresponding to the non-branch code.
[0164] Optionally, the processor may also execute the following program code: executing the code in the branch code according to the execution mode of inference execution, and obtaining the execution log corresponding to the branch code, including: determining the target branch code to be executed in the branch code according to a preset reasoning strategy; determining the trigger condition corresponding to the target branch code; setting the trigger condition to true according to the preset reasoning strategy to execute the target branch code, and obtaining the execution log corresponding to the branch code.
[0165] Optionally, the processor may also execute program code of the following steps: before executing the target code corresponding to the target document according to the abstract syntax tree and obtaining the target execution log, the method further includes: determining multiple functions and multiple classes involved in the target code; determining a target function from multiple functions and determining a target class from multiple classes, wherein a risk value of the target function is higher than a preset threshold, and a risk value of the target class is higher than a preset threshold; rewriting the target function and the target class to obtain a rewritten target function and a rewritten target class, wherein, when executing the target code corresponding to the target document according to the abstract syntax tree, the target code implements the function corresponding to the target code by calling the rewritten target function and the rewritten target class.
[0166] Optionally, the processor may also execute the program code of the following steps: structurally parsing the target document to obtain an abstract syntax tree, including: structurally parsing the target document to obtain a target code corresponding to the target document, wherein an inference strategy is used to determine a branch code to be executed; parsing the target code through a target engine to obtain operation code information corresponding to the target code; and generating an abstract syntax tree based on the operation code information.
[0167] Optionally, the processor may also execute the program code of the following steps: predicting the risk of the target document based on the target execution log, and obtaining the prediction results including: performing structural analysis on the target document to obtain version information corresponding to the target document and flow information corresponding to the target document; performing feature extraction on the target execution log, version information and flow information based on the target neural network model to obtain target feature information; performing risk prediction based on the target feature information through the target neural network model to obtain a prediction result.
[0168] It can be understood by those skilled in the art that Figure 7 The structure shown is for illustration only, and the computer terminal may also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a handheld computer, a mobile Internet device (Mobile Internet Devices, MID), a PAD, or other terminal devices. Figure 7 The structure of the electronic device is not limited. For example, the computer terminal 10 may also include Figure 7 More or fewer components (such as network interfaces, display devices, etc.) shown in, or having Figure 7 Different configurations shown.
[0169] A person of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware related to the terminal device through a program, and the program can be stored in a computer-readable storage medium, and the storage medium may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0170] Example 5
[0171] The embodiment of the present invention further provides a computer-readable storage medium. Optionally, in this embodiment, the computer-readable storage medium can be used to store the program code executed by the document risk detection method provided in the first embodiment.
[0172] Optionally, in this embodiment, the above storage medium may be located in any computer terminal in a computer terminal group in a computer network, or in any mobile terminal in a mobile terminal group.
[0173] Optionally, in this embodiment, the above-mentioned storage medium is configured to store program codes for executing the following steps: obtaining a target document to be detected, and performing structural analysis on the target document to obtain an abstract syntax tree; executing a target code corresponding to the target document according to the abstract syntax tree to obtain a target execution log, wherein the target code is used to display the target document in a target device; predicting the risk of the target document based on the target execution log to obtain a prediction result, wherein the prediction result is used to characterize whether there is a risk in the target document.
[0174] Optionally, the storage medium is configured to store program code for executing the following steps: executing the target code corresponding to the target document according to the abstract syntax tree, and obtaining a target execution log including: when executing the target code corresponding to the target document according to the abstract syntax tree, classifying the target code according to the abstract syntax tree to obtain multiple categories, wherein the multiple categories include at least: a first category, a second category and a third category, the risk of the target code in the first category is lower than the risk of the target code in the second category, and the risk of the target code in the second category is lower than the risk of the target code in the third category; determining the execution mode of the code according to the category corresponding to the code in the target code, wherein the execution mode is one of the following: real execution, simulated execution and inference execution, inference execution refers to executing branch code in the target code according to a preset inference strategy, and the inference strategy is used to determine the branch code to be executed; executing the target code corresponding to the target document according to the execution mode of the code and the abstract syntax tree to obtain a target execution log.
[0175] Optionally, the above-mentioned storage medium is configured to store program code for executing the following steps: determining the execution mode of the code based on the category corresponding to the code includes: if the category corresponding to the first code in the target code is the first category, then determining that the execution mode of the first code is real execution; if the category corresponding to the second code in the target code is the second category, then determining that the execution mode of the second code is simulated execution; if the category corresponding to the third code in the target code is the third category, then determining that the execution mode of the third code is inference execution.
[0176] Optionally, the above-mentioned storage medium is configured to store program code for executing the following steps: executing the target code corresponding to the target document according to the execution mode of the code and the abstract syntax tree, and obtaining the target execution log includes: executing the first code through the execution mode of real execution and the abstract syntax tree to obtain the first execution log; executing the second code through the execution mode of simulated execution and the abstract syntax tree to obtain the second execution log; executing the code in the third code through the execution mode of inference execution and the abstract syntax tree to obtain the third execution log; obtaining the target execution log based on the first execution log, the second execution log and the third execution log.
[0177] Optionally, the above-mentioned storage medium is configured to store program code for executing the following steps: before executing the code in the third code through the execution mode of inference execution and the abstract syntax tree to obtain the third execution log, the method also includes: building a context execution environment for the target code based on the implementation environment of the functions and classes in the target code; building a grammatical environment for the target code based on the language grammar of the abstract syntax tree, wherein the target code is executed under the context execution environment and the grammatical environment.
[0178] Optionally, the above-mentioned storage medium is configured to store program code for executing the following steps: executing code in the third code through the execution mode of inferential execution and the abstract syntax tree, and obtaining a third execution log including: determining branch code and non-branch code in the third code according to the abstract syntax tree; executing code in the branch code according to the execution mode of inferential execution to obtain the execution log corresponding to the branch code; executing the non-branch code according to the abstract syntax tree to obtain the execution log corresponding to the non-branch code; obtaining the third execution log based on the execution log corresponding to the branch code and the execution log corresponding to the non-branch code.
[0179] Optionally, the above-mentioned storage medium is configured to store program code for performing the following steps: executing the code in the branch code according to the execution mode of inference execution, and obtaining the execution log corresponding to the branch code includes: determining the target branch code to be executed in the branch code according to a preset reasoning strategy; determining the trigger condition corresponding to the target branch code; setting the trigger condition to true according to the preset reasoning strategy to execute the target branch code, and obtaining the execution log corresponding to the branch code.
[0180] Optionally, the above-mentioned storage medium is configured to store program code for executing the following steps: before executing the target code corresponding to the target document according to the abstract syntax tree and obtaining the target execution log, the method also includes: determining multiple functions and multiple classes involved in the target code; determining the target function from the multiple functions and determining the target class from the multiple classes, wherein the risk value of the target function is higher than a preset threshold, and the risk value of the target class is higher than a preset threshold; rewriting the target function and the target class to obtain a rewritten target function and a rewritten target class, wherein, when executing the target code corresponding to the target document according to the abstract syntax tree, the target code implements the function corresponding to the target code by calling the rewritten target function and the rewritten target class.
[0181] Optionally, the above-mentioned storage medium is configured to store program codes for executing the following steps: structurally parsing the target document to obtain an abstract syntax tree including: structurally parsing the target document to obtain a target code corresponding to the target document; parsing the target code through a target engine to obtain operation code information corresponding to the target code; and generating an abstract syntax tree based on the operation code information.
[0182] Optionally, the above-mentioned storage medium is configured to store program codes for executing the following steps: predicting the risk of the target document based on the target execution log, and obtaining the prediction results including: performing structural analysis on the target document to obtain version information corresponding to the target document and flow information corresponding to the target document; performing feature extraction on the target execution log, version information and flow information based on the target neural network model to obtain target feature information; performing risk prediction based on the target feature information through the target neural network model to obtain a prediction result.
[0183] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.
[0184] In the above embodiments of the present invention, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0185] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0186] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0187] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0188] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, magnetic disk or optical disk and other media that can store program codes.
[0189] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A document risk detection method, characterized in that: include: Obtaining a target document to be detected, and performing structural analysis on the target document to obtain an abstract syntax tree; Executing a target code corresponding to the target document according to the abstract syntax tree to obtain a target execution log, wherein the target code is used to display the target document in a target device; The risk of the target document is predicted according to the target execution log to obtain a prediction result, wherein the prediction result is used to characterize whether the target document has a risk.
2. The method according to claim 1, characterized in that Executing the target code corresponding to the target document according to the abstract syntax tree to obtain a target execution log includes: When executing the target code corresponding to the target document according to the abstract syntax tree, classifying the target code according to the abstract syntax tree to obtain a plurality of categories, wherein the plurality of categories at least include: a first category, a second category, and a third category, the risk of the target code of the first category is lower than the risk of the target code of the second category, and the risk of the target code of the second category is lower than the risk of the target code of the third category; Determining an execution mode of the code according to the category corresponding to the code in the target code, wherein the execution mode is one of the following: real execution, simulated execution and inference execution, wherein the inference execution refers to executing the branch code in the target code according to a preset inference strategy, wherein the inference strategy is used to determine the branch code to be executed; The target code corresponding to the target document is executed according to the execution mode of the code and the abstract syntax tree to obtain the target execution log.
3. The method according to claim 2, characterized in that Determining the execution mode of the code according to the category corresponding to the code includes: If the category corresponding to the first code in the target code is the first category, determining that the execution mode of the first code is real execution; If the category corresponding to the second code in the target code is the second category, determining that the execution mode of the second code is simulated execution; If the category corresponding to the third code in the target code is the third category, it is determined that the execution mode of the third code is inference execution.
4. The method according to claim 3, characterized in that Executing the target code corresponding to the target document according to the execution mode of the code and the abstract syntax tree, obtaining the target execution log includes: Execute the first code by using the actual execution mode and the abstract syntax tree to obtain a first execution log; Execute the second code through the execution mode of the simulated execution and the abstract syntax tree to obtain a second execution log; Execute the code in the third code by using the execution mode of the inference execution and the abstract syntax tree to obtain a third execution log; The target execution log is obtained according to the first execution log, the second execution log and the third execution log.
5. The method according to claim 1, characterized in that Before executing the target code corresponding to the target document according to the abstract syntax tree to obtain the target execution log, the method further includes: Building a context execution environment for the target code according to the implementation environment of functions and classes in the target code; A grammatical environment of the target code is constructed according to the language grammar of the abstract syntax tree, wherein the target code is executed in the context execution environment and the grammatical environment.
6. The method according to claim 4, characterized in that Executing the code in the third code through the execution mode of the inference execution and the abstract syntax tree, obtaining a third execution log includes: Determine branch codes and non-branch codes in the third code according to the abstract syntax tree; Execute the code in the branch code according to the execution mode of the inference execution to obtain the execution log corresponding to the branch code; Execute the non-branch code according to the abstract syntax tree to obtain an execution log corresponding to the non-branch code; The third execution log is obtained according to the execution log corresponding to the branch code and the execution log corresponding to the non-branch code.
7. The method according to claim 4, characterized in that Executing the code in the branch code according to the execution mode of the inference execution, obtaining the execution log corresponding to the branch code includes: Determining a target branch code to be executed in the branch code according to the preset reasoning strategy; Determine the trigger condition corresponding to the target branch code; The trigger condition is set to true according to the preset reasoning strategy to execute the target branch code and obtain the execution log corresponding to the branch code.
8. The method according to claim 1, characterized in that Before executing the target code corresponding to the target document according to the abstract syntax tree to obtain the target execution log, the method further includes: Determine multiple functions and multiple classes involved in the target code; Determining a target function from the multiple functions and determining a target class from the multiple classes, wherein a risk value of the target function is higher than a preset threshold, and a risk value of the target class is higher than the preset threshold; The target function and the target class are rewritten to obtain a rewritten target function and a rewritten target class, wherein when the target code corresponding to the target document is executed according to the abstract syntax tree, the target code implements the function corresponding to the target code by calling the rewritten target function and the rewritten target class.
9. The method according to claim 1, characterized in that The target document is structurally parsed to obtain an abstract syntax tree including: Performing structural analysis on the target document to obtain a target code corresponding to the target document; Parsing the target code through a target engine to obtain operation code information corresponding to the target code; The abstract syntax tree is generated according to the operation code information.
10. The method according to claim 1, characterized in that The risk of the target document is predicted based on the target execution log, and the prediction results include: Performing structural analysis on the target document to obtain version information corresponding to the target document and stream information corresponding to the target document; Extracting features from the target execution log, the version information, and the flow information according to the target neural network model to obtain target feature information; The target neural network model performs risk prediction based on the target feature information to obtain the prediction result.
11. A document risk detection method, characterized in that: include: Get the target document to be detected uploaded by the client; Performing structural analysis on the target document in the cloud server to obtain an abstract syntax tree; Executing a target code corresponding to the target document according to the abstract syntax tree to obtain a target execution log, wherein the target code is used to display the target document in a target device; predicting the risk of the target document according to the target execution log to obtain a prediction result, wherein the prediction result is used to characterize whether the target document has a risk; The prediction result is returned to the client.
12. A document risk detection device, characterized in that: include: An acquisition unit, used for acquiring a target document to be detected, and performing structural analysis on the target document to obtain an abstract syntax tree; an execution unit, configured to execute a target code corresponding to the target document according to the abstract syntax tree to obtain a target execution log, wherein the target code is used to display the target document in a target device; The prediction unit is used to predict the risk of the target document according to the target execution log to obtain a prediction result, wherein the prediction result is used to characterize whether the target document has a risk.
13. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein when the program is executed, the device where the storage medium is located is controlled to execute the document risk detection method according to any one of claims 1 to 11.
14. An electronic device, characterized in that: include: A memory storing an executable program; A processor is used to run the program, wherein the program executes the document risk detection method described in any one of claims 1 to 11 when running.