Extraction device, extraction method, and extraction program

The extraction device addresses the challenges of dynamic bytecode instrumentation by analyzing execution traces to extract symbol tables, enabling efficient and accurate instrumentation of script engines with unknown specifications without manual analysis or support functions.

JP2025070867APending Publication Date: 2025-05-02NIPPON TELEGRAPH & TELEPHONE CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023181456
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-10-20
Publication Date
2025-05-02

AI Technical Summary

Technical Problem

Existing dynamic bytecode instrumentation techniques face challenges in accurately and efficiently providing the instrumentation function, especially for script engines with unknown internal specifications, and often require manual reverse engineering or support functions like debuggers.

Method used

The proposed solution involves an extraction device that includes a symbol table analysis unit and a symbol table extraction unit. This device analyzes the execution trace of a program to acquire structural information of the symbol table and then extracts symbol tables based on this information, enabling dynamic bytecode instrumentation without the need for manual analysis or support functions.

Benefits of technology

This approach allows for accurate and efficient provision of dynamic bytecode instrumentation, enabling instrumentation of script engines with unknown internal specifications without manual reverse engineering or support functions, and ensures accurate reference and search of symbol tables.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025070867000001_ABST
    Figure 2025070867000001_ABST
Patent Text Reader

Abstract

To provide an extraction device, an extraction method, and an extraction program that accurately and efficiently provide a dynamic bytecode instrumentation function.SOLUTION: An extraction device (analysis device 10) includes a symbol table analysis unit 1223 and a symbol table extraction unit 1232. A symbol table analysis unit 1223 analyzes an execution trace obtained by having a program execute a first script, and acquires structure information of a symbol table generated by the program. The symbol table extraction unit 1232 extracts, on the basis of the structure information, the symbol table from a memory area used by the program to execute a second script.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to an extraction device, an extraction method, and an extraction program. [Background technology]

[0002] [Script Analysis] Script analysis techniques are used for a variety of purposes, including compiler optimization for just-in-time (JIT) compilation, software testing and debugging, fuzzing, and malware analysis.

[0003] [Program Instrumentation] One of the techniques for dynamically analyzing programs is instrumentation, which is a technique for obtaining information about the execution state of a program during execution by adding code with analytical functions to the program to be analyzed and then executing the program.

[0004] For example, instrumentation could insert logging code after each execution of a program to determine the total number of instructions executed, and similarly insert logging code after each execution of each branch to determine the control flow that was followed.

[0005] This type of instrumentation is an important technology that is widely used for software testing as well as cybersecurity purposes such as malware analysis and vulnerability discovery.

[0006] Here, instrumentation can be divided into dynamic instrumentation and static instrumentation. Dynamic instrumentation is a technique for dynamically adding analysis code by using a technique for dynamically changing the behavior of a program during execution. Static instrumentation is a technique for statically adding analysis code by using a technique for rewriting a program before execution.

[0007] Furthermore, the targets of instrumentation are diverse, including source code, scripts, executable binaries (hereafter referred to as binaries), bytecode, etc. Nowadays, not only is testing for scripts important, but the opportunities for malicious scripts to be used in attacks are also increasing, making instrumentation for scripts important.

[0008] One of the most representative instrumentation techniques is dynamic binary instrumentation. Dynamic binary instrumentation is a technique for obtaining information about the execution state by dynamically adding analytical code to the binary program to be analyzed at runtime. In dynamic binary instrumentation, the injection of analytical code is mainly achieved by using hooks.

[0009] Specifically, the flow of injecting analysis code is as follows: First, the code to be injected is placed in memory. Then, when a specific, pre-defined command or function is executed, a hook is placed so that a branch is made to the code to be injected. The code is then executed. At the end of the code, execution is returned to the branch source and the original processing is resumed.

[0010] In current dynamic binary instrumentation techniques, this is typically achieved using a technique called inline hooks, or by hooking into a conversion called dynamic binary translation in a virtual machine (VM).

[0011] [Necessity to deal with obfuscation] One technique for hindering analysis is obfuscation. Obfuscation involves applying conversions to programs that make them difficult to interpret, primarily to hinder static analysis. In the case of script obfuscation, for example, parts of the script are encoded or encrypted, and then dynamically decoded or decrypted at run time before execution. In such cases, the type of script that will be executed is not clear until it is executed. This makes static analysis of the script difficult.

[0012] Malicious scripts and protected scripts are generally obfuscated, making static analysis difficult. Software protection is an important technique for protecting the intellectual property rights of creators. Reverse engineering is a technique for clarifying the technology used to implement software. This is a technique for analyzing a program to understand its structure and specifications. Software protection is a technique for protecting software from program analysis for such reverse engineering.

[0013] Therefore, there are two main methods for instrumenting scripts: statically adding analysis scripts to scripts before they are executed, and dynamically adding analysis bytecodes when the script is converted to bytecode by a script engine at execution time.

[0014] [Script execution method] Scripts are executed by a script engine (also called an interpreter). Scripts are generally converted into bytecode at runtime, and the bytecode is interpreted and executed by a Virtual Machine (VM). For this reason, scripts are analyzed before execution, and the bytecode is analyzed at execution.

[0015] As mentioned above, if the script is obfuscated, static analysis before execution becomes difficult. Therefore, it is necessary to dynamically analyze the bytecode and obtain variable information from the information obtained at execution time.

[0016] From the above, in order to implement instrumentation for scripts, dynamic instrumentation of bytecode, i.e., dynamic bytecode instrumentation, is required.

[0017] Here, a method for implementing dynamic instrumentation for Java bytecode has been proposed (Non-Patent Document 1). Also, a method for implementing dynamic instrumentation for ActionScript 3 bytecode has been proposed (Non-Patent Document 2). [Prior art documents] [Non-patent literature]

[0018] [Non-Patent Document 1] W. Binder, J. Hulaas, and P. Moret, “Advanced Java Bytecode Instrumentation”, In Proceedings of the 5th International Symposium on Principles and Practice of Programming in Java, pp. 135-144, 2007. [Non-Patent Document 2] Jarkko Turkulainen, “Reflash: practical ActionScript3 instrumentation with RABCDAsm”, 2016 Summary of the Invention [Problem to be solved by the invention]

[0019] However, with conventional techniques, it can be difficult to provide dynamic bytecode instrumentation accurately and efficiently.

[0020] For example, the techniques described in Non-Patent Documents 1 and 2 use information specific to each VM, and have a problem in that they require individual design and implementation for various script languages.

[0021] Dynamic bytecode instrumentation in scripts generally requires the use of support functions such as a debugger provided by the script engine. This is because the internal specifications of the VM in the script engine that controls the execution of the script are often not made public, making it difficult to monitor the execution state or change the execution state required for instrumentation without a support function.

[0022] However, if such support functions are not provided, it is necessary to reverse engineer the VM to reveal its internal specifications, and then independently observe the execution state and analyze the bytecode to obtain information for instrumentation.

[0023] It is not realistic to manually analyze, design, and implement this individually for a script engine, in terms of the amount of work required.

[0024] In addition, conventional techniques do not provide a detailed analysis of the structure of symbol tables or references between symbol tables, and therefore may not be able to accurately reference or search symbol tables, which may affect the accuracy of bytecode porting.

[0025] The present invention has been made in view of the above, and aims to provide an extraction device, an extraction method, and an extraction program that can provide dynamic bytecode instrumentation functionality accurately and efficiently, even for script engines that do not have a support function such as a debugger and whose internal specifications are unknown, without requiring manual individual analysis, design, and implementation, and based on accurate bytecode transplantation. [Means for solving the problem]

[0026] In order to solve the above-mentioned problems and achieve the object, the extraction device of the present invention is characterized by having a symbol table analysis unit that analyzes an execution trace obtained by executing a first script in a program and obtains structural information of a symbol table generated by the program, and a symbol table extraction unit that extracts a symbol table from a memory area used by the program to execute a second script based on the structural information. Effect of the Invention

[0027] According to the present invention, dynamic bytecode instrumentation can be implemented accurately and efficiently. [Brief description of the drawings]

[0028] [Figure 1] FIG. 1 is a diagram illustrating an example of a configuration of a script engine. [Diagram 2] FIG. 2 is a diagram showing pseudo code of the VM of the script engine. [Diagram 3] FIG. 3 is a diagram for explaining an example of instrumentation by bytecode porting. [Figure 4] FIG. 4 is a diagram illustrating an example of the configuration of the analysis device according to the embodiment. [Diagram 5] FIG. 5 is a diagram showing an example of a first test script used for detecting a virtual program counter (VPC). [Figure 6] FIG. 6 is a diagram illustrating an example of the second test script. [Figure 7] FIG. 7 is a diagram illustrating an example of the third test script. [Figure 8] FIG. 8 is a diagram illustrating an example of the fourth test script. [Figure 9] FIG. 9 is a diagram showing an example of a porting script. [Figure 10] FIG. 10 is a diagram showing the bytecode and symbol table generated by executing the porting script shown in FIG. [Figure 11] FIG. 11 is a diagram illustrating an example of a bytecode. [Figure 12] FIG. 12 is a diagram illustrating an example of an execution trace. [Figure 13] FIG. 13 is a diagram illustrating an example of a VM execution trace. [Figure 14] FIG. 14 is a diagram for explaining the structure of the symbol table. [Figure 15] FIG. 15 is a diagram for explaining the interpretation execution function. [Figure 16] FIG. 16 is a diagram for explaining a taint analysis in which a taint is propagated when a pointer is referenced. [Figure 17] FIG. 17 is a diagram for explaining the reference analysis. [Figure 18] FIG. 18 is a diagram for explaining the structural analysis. [Figure 19]FIG. 19 is a diagram for explaining the structural analysis. [Figure 20] FIG. 20 is a diagram for explaining the process of the virtual program counter detection unit. [Figure 21] FIG. 21 is a diagram illustrating the process of the bytecode cache detector. [Figure 22] FIG. 22 is a diagram showing an example of the configuration of the injection unit. [Figure 23] FIG. 23 is a flowchart illustrating the processing procedure of the analysis process according to the embodiment. [Figure 24] FIG. 24 is a flowchart showing the procedure of the execution trace acquisition process. [Diagram 25] FIG. 25 is a flowchart showing the procedure of the virtual program counter detection process. [Figure 26] FIG. 26 is a flowchart illustrating a procedure for the bytecode cache detection process. [Figure 27] FIG. 27 is a flowchart showing the procedure of the interpretation execution function detection process. [Figure 28] FIG. 28 is a flowchart showing the processing procedure of the symbol table detection process. [Figure 29] FIG. 29 is a flowchart showing the processing procedure of the symbol table analysis process. [Diagram 30] FIG. 30 is a flowchart showing the processing procedure of the reference analysis process. [Diagram 31] FIG. 31 is a flowchart showing the procedure of the structure analysis process. [Diagram 32] FIG. 32 is a flowchart showing the procedure of the value object analysis process. [Diagram 33] FIG. 33 is a flowchart showing the procedure of the bytecode extraction process. [Diagram 34] FIG. 34 is a flowchart showing the processing procedure of the symbol table extraction process. [Diagram 35] FIG. 35 is a flowchart showing the procedure of the injection process. [Diagram 36]FIG. 36 is a flowchart showing the procedure of the bytecode injection process. [Figure 37] FIG. 37 is a flowchart showing the processing procedure of the symbol table injection process. [Figure 38] FIG. 38 is a diagram illustrating an example of a computer that realizes the analysis device by executing a program. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0029] Hereinafter, an embodiment of an analysis device, an analysis method, and an analysis program according to the present application will be described in detail with reference to the drawings. The present invention is not limited to the embodiment described below. The analysis device also functions as an extraction device. The analysis method and the analysis program include an extraction method and an extraction program, respectively.

[0030] [Embodiment] An analysis device according to an embodiment first analyzes the VM of a script engine. The analysis device executes a test script while monitoring the binary of the script engine, and acquires a branch trace and a memory access trace as an execution trace. The analysis device analyzes a virtual machine (VM) based on the execution trace, and acquires, as architecture information, a virtual program counter (VPC) and a bytecode cache in which VM instructions to be executed are stored. Then, based on the analysis result of the VM of the script engine, the analysis device detects a symbol table that holds information about variables. Furthermore, the analysis device analyzes the symbol table.

[0031] The analysis device may output the results of the analysis of the symbol table. The analysis device may also use the results of the analysis of the symbol table for providing dynamic bytecode instrumentation.

[0032] The analysis device provides the script engine with dynamic bytecode instrumentation functionality based on the analysis results of the VM and the analysis results of the symbol table. Based on the acquired architecture information, the analysis device executes a porting script to port the bytecode and symbol table to the bytecode and symbol table generated by the script engine in its memory space based on the porting destination script, thereby providing the script engine with dynamic bytecode instrumentation. The porting destination script is the subject of analysis by the script engine with dynamic bytecode instrumentation functionality. In the following description, the process of porting the bytecode or symbol table may be referred to as injection.

[0033] The configuration of a typical script engine and its function will be described with reference to Figs. 1 and 2. Fig. 1 is a diagram for explaining an example of the configuration of a script engine. As shown in Fig. 1, the script engine 1 has a bytecode compiler 2 and a VM 3. The bytecode compiler 2 also has a syntax analysis unit 4 and a bytecode generation unit 5. The VM 3 also has a code cache unit 6, a fetch unit 7, a decode unit 8, and an execution unit 9. The fetch unit 7, the decode unit 8, and the execution unit 9 are executed repeatedly and are called an interpreter loop. The script engine 1 then accepts the input of a script.

[0034] The syntax analysis unit 4 receives a script as input, and generates an Abstract Syntax Tree (AST) through lexical analysis and syntax analysis, and outputs it to the bytecode generation unit 5. The bytecode generation unit 5 receives the AST as input, converts it into bytecode, and stores it in the code cache unit 6.

[0035] The fetch unit 7 fetches the VM opcode from the code cache unit 6 and outputs it to the decode unit 8. Here, the VM opcode refers to the opcode part of the VM instruction. The decode unit 8 receives the VM opcode as input, interprets the VM opcode using a decoder dispatcher, and dispatches it to the corresponding program. The execution unit 9 executes the program corresponding to the VM instruction. The contents written in the script are executed by executing the VM instructions one after another through a repetition of the interpreter loop.

[0036] The function of the components of the script engine will be described with reference to FIG. 2. FIG. 2 is a diagram showing pseudocode of the VM of the script engine. As shown in FIG. 2, the pseudocode first initializes the VPC (line 1). In the pseudocode, a while loop is an interpreter loop (lines 2 to 7). In the pseudocode, the fetch unit 7 acquires the VM opcode of the VM instruction at the position pointed to by the VPC in the bytecode cache that holds the bytecode (line 3). The decoder uses a Switch statement to interpret the VM instruction (line 4), and the dispatcher calls the instruction handler based on the VM opcode (lines 5 and after). In the pseudocode, the instruction handler performs an operation corresponding to the instruction. Input and output are performed using a virtual stack and virtual registers (line 6), and constants and variables are referenced via a symbol table (line 7).

[0037] [An example of instrumentation using bytecode porting] The analysis device according to the embodiment analyzes a script engine whose internal specifications are unknown, acquires information for dynamic bytecode instrumentation, and provides the script engine with a dynamic bytecode instrumentation function.

[0038] The analysis device provides the script engine with a dynamic bytecode instrumentation function, for example, an instrumentation function based on bytecode porting.

[0039] Fig. 3 is a diagram explaining an example of instrumentation by bytecode porting. In this instrumentation, the script engine executes the ported script, and if there is a process that meets the hook condition, it saves and saves the process that meets the hook condition (Fig. 3 (1)).

[0040] In this instrumentation, the processing part corresponding to the hook condition is overwritten with a branch instruction to the stub, and the stub is executed (Fig. 3 (2) and (3)). The stub includes a process to obtain a specific variable from the symbol table using a symbol table VM instruction, a process to call the ported bytecode with a branch VM instruction after passing the obtained variable, and a process to branch back to the original location after the ported bytecode is executed.

[0041] Next, the stub calls the ported bytecode, executes it (Figure 3 (4)), and branches back to the original location after the ported bytecode is executed. Then, after branching back to the original location by executing the stub, the saved processing is restored and processing is resumed (Figure 3 (5)).

[0042] Before the process of providing the analysis function, the analysis device executes the test script while monitoring the binary of the script engine, and obtains architecture information including the VPC, the bytecode cache, and the symbol table.

[0043] In this way, the analysis device can detect various architectural information through analysis based on the execution trace and VM execution trace, even for script engines whose VM internal specifications are unknown, and can thus provide dynamic bytecode instrumentation functionality to the script engine without the need for manual reverse engineering or the use of support functions such as a debugger.

[0044] Furthermore, the analysis device can accurately refer to and search for symbol tables by analyzing the structure of the symbol tables and the references between the symbol tables in detail. This allows the analysis device to inject the symbol tables based on the analysis results. As a result, the analysis device can accurately and efficiently provide dynamic bytecode instrumentation functions.

[0045] [Analysis equipment configuration] Next, the configuration of the analysis device 10 according to the embodiment will be specifically described with reference to Fig. 4. Fig. 4 is a diagram illustrating an example of the configuration of the analysis device 10 according to the embodiment.

[0046] As shown in Fig. 4, the analysis device 10 includes an input unit 11, a control unit 12, a storage unit 13, and an output unit 14. The analysis device 10 receives input of a test script and a script engine binary. Note that, although the present embodiment describes a script engine as an example of a target, the present invention is not limited to this as long as it is an execution environment using a virtual machine. Therefore, the script engine binary may be referred to as a virtual machine binary.

[0047] The input unit 11 is composed of input devices such as a keyboard and a mouse, and receives input of information from the outside and inputs it to the control unit 12. The input unit 11 also has a communication interface for transmitting and receiving various information to and from other devices connected via a wired connection or a network, and receives input of information transmitted from other devices. The input unit 11 receives input of test scripts, script engine binaries, porting scripts, porting destination scripts, and hook settings, and outputs them to the control unit 12.

[0048] The test script is a script that is input when dynamically analyzing a script engine to obtain an execution trace. The test script includes a first test script for VM analysis, a second test script for symbol table detection, and a third test script for symbol table detection. The script engine binary is an executable file that constitutes the script engine. The script engine binary may be composed of multiple executable files.

[0049] [Test script configuration] Let us explain about test scripts. A test script is a script that is input when dynamically analyzing a script engine. This test script focuses on the number of branch instruction executions and memory reads and writes, and is used to capture the difference in the behavior of the script engine that occurs when the test script is executed a different number of times. This test script is prepared before the analysis and is created manually. Creating it requires knowledge of the specifications of the target script language.

[0050] 5 is a diagram showing an example of a first test script used for detecting VPCs. The first test script uses a repetitive process (line 2). The first test script changes the execution conditions and generates differences by increasing or decreasing the number of repetitions (line 2) and the number of repeated statements (lines 3 to 5) in the test script. The first test script is subject to processes from the execution trace acquisition process to the bytecode cache detection process, which will be described later.

[0051] FIG. 6 is a diagram showing an example of a second test script. In the second test script, a characteristic value variable is used to enable matching of values ​​to detect a symbol table. For example, in the example of FIG. 6, the characteristic value is "1234". By including the characteristic value in the second test script, it becomes possible to detect the storage location of this characteristic value from the memory access trace by matching the characteristic value at the time of symbol table detection. The second test script is subject to the execution trace acquisition process and the symbol table detection process, which will be described later. The second test script is used to detect a symbol table having a value object.

[0052] FIG. 7 is a diagram showing an example of a third test script. In the third test script, multiple variables of characteristic values ​​are used. For example, in the example of FIG. 7, the characteristic values ​​are "1234", "5.678", and "9012". The third test script is used in conjunction with the second test script to take differences and detect a symbol table that holds value objects in a linked list.

[0053] Fig. 8 is a diagram showing an example of a fourth test script. Variable g is a global variable. Variable l is a local variable. In a script engine, multiple symbol tables may be held for each value scope, such as a global scope and a local scope. Such a symbol table is detected by a test script using values ​​of multiple scopes, as shown in Fig. 8.

[0054] [Porting script] Fig. 9 is a diagram showing an example of a porting script. By executing the porting script, bytecode and a symbol table to be ported during dynamic bytecode instrumentation are extracted. Fig. 10 is a diagram showing the bytecode and the symbol table generated by executing the porting script shown in Fig. 9.

[0055] As shown in Fig. 9, the porting script includes a function to be ported 202 and a call section 204 of the function to be ported. The function to be ported 202 includes the process content to be ported 203. When the porting script is executed, bytecode 205 (see Fig. 10) corresponding to the processing portion of the process content to be ported 203 is extracted and ported to the bytecode generated by the script engine in its own memory space based on the porting destination script. By describing various processes for analyzing the behavior of the script in this section, the porting destination script can be analyzed.

[0056] Then, the caller 204 of the function to be ported calls the function to be ported 202, generating bytecode and a symbol table 210 (see FIG. 10), which makes it possible to extract the function. The symbol table 210 is ported to a symbol table generated by the script engine in its own memory space based on the porting destination script.

[0057] 11 is a diagram showing an example of a bytecode. Bytecode 205 is an example, and a bytecode exists for each routine. As shown in bytecode 205, for example, the bytecode has a data structure including an opcode 2051 and an operand 2052. The opcode 2051 is a value that specifies the type of operation, and is actually specified by a corresponding hexadecimal value. The operand 2052 is a value that is the target of the operation, and is generally specified by an ordinal number in a symbol table.

[0058] The storage unit 13 is realized by a semiconductor memory element such as a RAM (Random Access Memory) or a flash memory, or a storage device such as a hard disk or an optical disk, and stores a processing program that operates the analysis device 10, data used during execution of the processing program, etc. The storage unit 13 has an execution trace database (DB) 131, and an architecture information DB 132 that stores architecture information acquired by the virtual machine analysis unit 121 and the symbol table analysis unit 122 (described later).

[0059] The execution trace DB 131 stores the execution trace acquired by the execution trace acquisition unit 1211. The execution trace DB 131 is managed by the analysis device 10. Of course, the execution trace DB 131 may be managed by another device (such as a server). In this case, the execution trace acquisition unit 1211 (described later) outputs the acquired execution trace to a management server of the execution trace DB 131 or the like via a communication interface of the output unit 14, and stores the execution trace in the execution trace DB 131.

[0060] Configure Execution Tracing Next, the execution trace will be explained. Fig. 12 is a diagram showing an example of an execution trace. As described above, an execution trace is composed of a branch trace and a memory access trace. Fig. 12 shows an excerpt of a part of an execution trace. The structure of an execution trace will be explained below with reference to Fig. 12.

[0061] An execution trace has an element called "trace." In "trace," it is indicated whether the log line is a branch trace or a memory access trace.

[0062] A branch trace log line has the format shown in lines 1 to 10 of Figure 12, and consists of three elements: type, src, and dst. type indicates whether the executed branch instruction was a call instruction, a jmp instruction, or a ret instruction. src indicates the address of the branch source, and dst indicates the address of the branch destination.

[0063] A log line of a memory access trace has the format shown, for example, in lines 11 to 13 of Figure 12, and consists of six elements: type, target, base, index, size, and value. Type indicates whether the memory access is a read or write. Target indicates the memory address that is the target of the memory access. Base indicates the value of the base register. Index indicates the value of the index register. Size indicates the size of the memory access. Base and index only exist if the base register or index register is used during memory access. Additionally, value stores the value resulting from the memory access.

[0064] Returning to Fig. 4, the control unit 12 will be described. The control unit 12 has an internal memory for storing programs that define various processing procedures and required data, and executes various processes using these. For example, the control unit 12 is an electronic circuit such as a CPU (Central Processing Unit) or an MPU (Micro Processing Unit). The control unit 12 has a virtual machine analysis unit 121, a symbol table analysis unit 122, an extraction unit 123, and an injection unit 124.

[0065] The virtual machine analysis unit 121 analyzes the VM of the script engine. The virtual machine analysis unit 121 obtains a plurality of execution traces by changing the conditions at the time of execution, analyzes the plurality of execution traces using differential execution analysis, and obtains the VPC. The virtual machine analysis unit 121 detects a bytecode cache from the execution traces. The bytecode cache stores the VM instructions to be executed.

[0066] The virtual machine analysis unit 121 includes an execution trace acquisition unit 1211 , a virtual program counter detection unit 1212 , and a bytecode cache detection unit 1213 .

[0067] The execution trace acquisition unit 1211 receives the first to third test scripts and the script engine binary as input. The execution trace acquisition unit 1211 acquires an execution trace by executing the first to third test scripts while monitoring the execution of the script engine binary.

[0068] The execution trace is composed of a branch trace and a memory access trace. The branch trace records the type of branch instruction at the time of execution, the branch source address, and the branch destination address. The memory access trace can record all or part of the type of memory operation at the time of execution (read / write), the memory address and value of the operation target, the base register, the index register, and the size of the memory access. It is known that the branch trace and memory access trace can be acquired by hooking the branch instruction and the memory operation instruction, inserting a code for log output, and executing it. The execution trace acquired by the execution trace acquisition unit 1211 is stored in the execution trace DB 131.

[0069] Furthermore, when acquiring the execution trace, the execution trace acquisition unit 1211 acquires an API (Application Programming Interface) trace and stores it in the execution trace DB 131. The API trace is a record of the system API called during execution and its arguments.

[0070] The virtual program counter detection unit 1212 extracts and analyzes the execution trace for the first test script stored in the execution trace DB 131 to detect a VPC. The virtual program counter detection unit 1212 analyzes a plurality of execution traces using differential execution analysis focusing on the number of memory reads to detect a VPC. The virtual program counter detection unit 1212 utilizes the fact that a read into a memory that holds a VPC always occurs after the execution of each VM instruction, and detects the VPC by discovering the read destination.

[0071] For this reason, the virtual program counter detection unit 1212 uses differential execution analysis focusing on the number of memory reads to detect VPCs. The virtual program counter detection unit 1212 compares execution traces of multiple test scripts acquired using the first test script, and finds memories whose memory read counts change in proportion to both the increase and decrease of the number of repetitions and the number of repeated statements. The virtual program counter detection unit 1212 then refers to the boundaries of each VM instruction and narrows down the memory values ​​read to those whose values ​​always point to the start points of the VM instructions. The virtual program counter detection unit 1212 detects these memories as VPCs.

[0072] (Configuring VM Execution Tracing) Here, the VM execution trace will be explained. Fig. 13 is a diagram showing an example of a VM execution trace. As described above, the VM execution trace is a record of the VM opcode and the VPC. Fig. 13 shows a part of the VM execution trace. The configuration of the VM execution trace will be described below with reference to Fig. 13.

[0073] A log line of a VM execution trace is, for example, in the format shown in FIG. 13 and consists of two elements, vpc and vmop (vm opcode). vpc indicates the value of the VPC. Also, vmop indicates the value of the VM opcode virtually assigned to each pointer that points to the beginning of the VM instruction handler to be executed, obtained from the pointer cache.

[0074] The bytecode cache detection unit 1213 detects a bytecode cache, which is a cache in which virtual machine instructions to be executed are stored, from the VM execution trace based on the execution trace, the VPC and the VM execution trace.

[0075] The bytecode cache detection unit 1213 detects the memory area pointed to by the VPC as a bytecode cache from the VM execution trace. The bytecode cache detection unit 1213 detects the code location of the caller of the memory allocation function that allocated this bytecode cache from the execution trace. The bytecode cache detection unit 1213 detects all memory areas allocated at this code location from the VM execution trace as code caches.

[0076] The bytecode cache detection unit 1213 detects a code location that writes to the bytecode cache from the execution trace. The bytecode cache detection unit 1213 detects writing by this code location from the VM execution trace as an update of the bytecode cache.

[0077] The symbol table analysis unit 122 detects and analyzes a symbol table and includes an interpretation execution function detection unit 1221, a symbol table detection unit 1222, a symbol table analysis unit 1223, and a value object analysis unit 1224.

[0078] Here, the structure of the symbol table will be explained with reference to Fig. 14. Fig. 14 is a diagram for explaining the structure of the symbol table. As shown in Fig. 14, the symbol table includes multiple data structures (structures or arrays). Also, in the symbol table, references between data structures are defined by pointers. Note that structures and arrays can be said to be variables of a type that includes multiple values.

[0079] 14, the symbol table has a structure 251, a structure 252, an array 253, and a structure 254. Furthermore, structure 251 includes a pointer to structure 252. Furthermore, structure 252 includes a pointer to array 253. Furthermore, array 253 includes a pointer to structure 254.

[0080] The end of the symbol table, structure 254, contains a value object. A value object is data related to a value used in a script. For example, a value object includes a variable name, a variable type, and a value. "var_name", "int", and "1234" correspond to the variable name, variable type, and value, respectively.

[0081] There are structures that hold information about the execution state of a script, and structures that hold information about a function that is being executed. Such structures are called management structures. Structure 251 is a management structure.

[0082] A function that interprets and executes a script is called an interpretation execution function, and the interpretation execution function has a pointer to a management structure as an argument. The interpretation execution function will be explained with reference to Figure 15. Figure 15 is a diagram for explaining the interpretation execution function.

[0083] The function "interp" in Figure 15 is an interpretation execution function. Additionally, "script_ctx_info" and "func_info" are arguments to the function "interp." "script_ctx_info" is a pointer to a structure that holds information about the execution state of the script. Additionally, "func_info" is a pointer to a structure that holds information about the function being executed. In other words, both "script_ctx_info" and "func_info" are pointers to management structures.

[0084] Note that the management structure and the symbol table do not have to correspond one-to-one. The management structure can have multiple different symbol tables for each scope within the script of the values ​​held by the symbol table.

[0085] The symbol table analyzer 122 obtains the execution trace, the VPC, and the bytecode cache from the virtual machine analyzer 121 .

[0086] The interpretive execution function detection unit 1221 detects an interpretive execution function from an execution trace. First, the interpretive execution function detection unit 1221 receives input of an execution trace, a VPC, and a bytecode cache.

[0087] The interpretive execution function detection unit 1221 detects functions from call and return branches in a branch trace included in the execution trace. Then, the interpretive execution function detection unit 1221 detects, among the detected functions, a function that performs memory access to a VPC or a bytecode cache as an interpretive execution function.

[0088] The symbol table detection unit 1222 detects a symbol table from the execution trace. First, the symbol table detection unit 1222 receives an input of the execution trace and an interpretation execution function. The symbol table detection unit 1222 executes the interpretation execution function while monitoring it, and obtains arguments of the interpretation execution function.

[0089] Symbol table detection unit 1222 assigns a unique taint tag to the acquired arguments that are pointers. Furthermore, symbol table detection unit 1222 executes the interpretation execution function while propagating the taint tag. This method of propagating taint tags is called taint analysis.

[0090] Fig. 16 is a diagram for explaining taint analysis in which a taint is propagated when a pointer is referenced. In the taint analysis in this embodiment, in addition to general taint propagation in accordance with data movement, symbol table detection unit 1222 also propagates a taint when a pointer is referenced. As shown in Fig. 16, first, symbol table detection unit 1222 assigns taint tag 273 to an address of memory area 271. Then, when data to which taint tag 273 has been assigned is referenced as a pointer during execution of a test script, symbol table detection unit 1222 propagates taint tag 273 to an address that holds a value referenced by the pointer.

[0091] Propagating a taint tag may involve, for example, adding information to a taint map, which is a data structure that holds a correspondence between memory addresses and taint tags assigned to each address.

[0092] Here, the symbol table detection unit 1222 searches the memory access trace for a characteristic value included in the test script and identifies a matching memory access (matching of characteristic values). The symbol table detection unit 1222 stores the identified target of memory access as a value in the symbol table.

[0093] The symbol table detection unit 1222 detects a memory area to which a taint tag that is the same as a value in the symbol table is assigned as a symbol table, and outputs a taint map.

[0094] 14, symbol table detection unit 1222 identifies a memory access to value "1234" of structure 254. In this case, "1234" is a characteristic value. Here, each pointer from structure 251, which is a management structure, to value "1234" of structure 254 is assigned the same taint tag as value "1234."

[0095] As a result, symbol table detection unit 1222 can detect a symbol table including a pointer to structure 252, a pointer to array 253, a pointer to structure 254, and the value of structure 254, "1234".

[0096] On the other hand, at this point, the structure of the structure or array including the value and the pointer is unknown. For example, it is not known at this point that the pointer to array 253 that structure 252 has is a member variable of structure 252, and the offset of the member variable from the start address of structure 252 is also unknown. Therefore, symbol table analyzer 1223 analyzes the structures of the structures and arrays included in the symbol table. Symbol table analyzer 1223 performs reference analysis and structure analysis.

[0097] (Reference analysis) Reference analysis will be described with reference to Fig. 17. Fig. 17 is a diagram for explaining reference analysis. Symbol table analysis unit 1223 performs reference analysis by utilizing the fact that the effective address of a structure is specified by offset reference from a starting address (base register) and that the effective address of an array is specified by index reference from a starting address (base register).

[0098] If a memory access with a taint tag uses a base register and does not use an index register, the symbol table analyzer 1223 determines that the access destination of the memory access is a member variable of a structure.Then, the symbol table analyzer 1223 calculates an offset from the difference between the effective address and the base register.The symbol table analyzer 1223 also obtains the start address of the structure from the base register.

[0099] 17, before the reference analysis is performed, it is unknown whether data structure 282 is a structure or an array. The memory access to which taint tag 286 of data structure 282 is assigned uses a base register and does not use an index register. As a result, symbol table analyzer 1223 regards the access destination of the memory access to which taint tag 286 of data structure 282 is assigned as a member variable of the structure. Here, symbol table analyzer 1223 sets the value of the base register as the start address of data structure 282, which is a structure. Then, symbol table analyzer 1223 calculates an offset from the start address of data structure 282 from the difference between the effective address of the member variable and the base register.

[0100] If a memory access with a taint tag uses a base register and an index register, the symbol table analyzer 1223 determines that the access destination of the memory access is an element of an array. The symbol table analyzer 1223 then obtains the head and subscript of the array from the base register and index register.

[0101] 17, before the reference analysis is performed, it is unknown whether data structure 283 is a structure or an array. The memory access to which taint tag 287 of data structure 283 is assigned uses a base register and an index register. Thus, symbol table analyzer 1223 regards the access destination of memory access to which taint tag 287 of data structure 283 is assigned as an element of an array, and obtains the head and subscript of the array from the base register and index register. Note that "C[i]" means "element with subscript i of array C."

[0102] (Structural analysis) Structural analysis will be described with reference to Figures 18 and 19. Figures 18 and 19 are diagrams for explaining structural analysis. Symbol table analyzer 1223 analyzes whether the detected symbol table holds value objects in an array or in a linked list, and if it holds them in a linked list, analyzes the structure. For example, in the symbol table ending in structure 254 in Figure 14 (where structure 251 is a management structure), array 253 holds a reference to structure 254, which is a value object, so it is determined that the detected symbol table holds value objects in an array.

[0103] On the other hand, if a structure holds a reference to a structure that is a value object, the symbol table analyzer 1223 determines that the detected symbol table holds the value object in a linked list. In this case, the symbol table analyzer 1223 regards the symbol table as a data structure that combines a linked list and a data structure other than a linked list. For example, the linked list includes a reference to a value object and a reference to the next element.

[0104] When the symbol table holds value objects in an array, the length of the reference from the management structure to each value object is constant regardless of the number of value objects.

[0105] For example, in the symbol table ending at structure 254 in Figure 14 (where structure 251 is a management structure), it takes three references to reach the value object, so the reference length is 3. Since value objects are managed as an array, the reference length is constant regardless of the number of value objects. Note that a management structure cannot become a value object, so the minimum reference length is 1.

[0106] On the other hand, when the symbol table holds value objects in a linked list, the length of the reference from the management structure to each value object may vary depending on the number of value objects.

[0107] When the symbol table analyzer 1223 determines that the symbol table holds value objects in a linked list, it compares the references from the management structures of the symbol tables for multiple test scripts with different numbers of variables to the value objects, and determines (estimates) whether the data structure is a linked list or not, by taking the difference. This enables the symbol table analyzer 1223 to detect the structures of the value objects and the linked list.

[0108] If the symbol table analysis unit 1223 determines that the symbol table holds value objects in a linked list, it can detect the structure of the value objects and the linked list by comparing references to value objects from the symbol table management structure for multiple test scripts with different numbers of variables and taking the differences.

[0109] For example, if the obtained reference lengths are 3, 5, and 4, at least three references are required from the management structure to the value object, and two references are not enough to reach the value object. From this, symbol table analyzer 1223 estimates that the three data structures including the common part, the management structure (corresponding to two references), are data structures other than a linked list. Then, symbol table analyzer 1223 estimates that the subsequent references, which are the differences, are linked lists.

[0110] 18 includes data structures other than the linked list from structure 291, which is a management structure, to the structure immediately preceding structure 2921, and linked list 292. Linked list 292 includes structure 2921 and structure 2922. Structure 2922 is a value object.

[0111] 19 includes data structures other than the linked list from structure 296, which is a management structure, to the structure immediately preceding structure 2971, and linked list 297. Linked list 297 includes structure 2971, structure 2972, structure 2973, structure 2974, structure 2975, and structure 2976. Structures 2974, structure 2975, and structure 2976 are value objects.

[0112] Here, structures such as value objects (structure 2974, etc.) shown in Figs. 18 and 19 may be called tagged unions. For example, when a test script such as "i=1234" is held, the tagged union has a part that holds a value such as "1234" and another part that holds a tag such as "0x1". However, it is not easy to distinguish between the tag part of the tagged union and the part of the union that holds the value of the test script. Note that in Figs. 18 and 19, the value part is expressed in decimal such as "1234" to facilitate understanding, but the value exists in memory as a binary number (or hexadecimal number).

[0113] The value object analysis unit 1224 determines that the portion hit by matching "1234" corresponding to the test script is the portion of the tagged union where a difference occurs depending on the value of the test script (the portion that holds the value of the test script). In this way, the value object analysis unit 1224 detects the portion excluding the portion where a difference occurs depending on the value of the test script. The value object analysis unit 1224 also considers variable name portions such as "i", "j", and "k" as portions where a difference occurs depending on the value of the test script. For example, the value object analysis unit 1224 detects the value portion "1234" and the variable name portion "i" from the structure 2974.

[0114] Similarly, when "j=5.678" and "k="9012"" are stored as test scripts, the value object analysis unit 1224 detects the portion excluding the portion where the difference occurs depending on the value of the test script. The value object analysis unit 1224 compares the portions of the structures storing the respective test scripts excluding the portion where the difference occurs depending on the value of the test script, and detects the member variables where the difference occurs depending on the type of the test script as the portion where the tag is stored (the portion where the difference occurs depending on the type). For example, for "j=5.678" of the test script, the value object analysis unit 1224 detects the portion "5.678" that stores the value and the portion "0x2" that stores the other tag. The type of "j=5.678" is, for example, a floating point. For example, for "k="9012"" of the test script, the value object analysis unit 1224 detects the portion ""9012"" that stores the value and the portion "0x3" that stores the other tag. Note that the type of "k="9012"" is, for example, a string.

[0115] In this way, the value object analysis unit 1224 can appropriately extract the tag part of a tagged union by detecting the part where a difference occurs depending on the type of the test script from the part of the tagged union that is detected by comparing the test script values ​​and excluding the part where a difference occurs depending on the test script value. Note that in the above-mentioned process, a search for the test script value is performed to detect the part where a difference occurs depending on the test script value, but this is not limited to this, and for example, the detection process can be performed by using the number of occurrences of the value instead of searching for the value.

[0116] The extraction unit 123 receives a porting script as an input. The extraction unit 123 executes the porting script to extract porting bytecode and porting symbol table based on information about the architecture acquired by analysis by the virtual machine analysis unit 121 and information about the symbol table acquired by analysis by the symbol table analysis unit 122. The extraction unit 123 includes a bytecode extraction unit 1231 and a symbol table extraction unit 1232.

[0117] The bytecode extraction unit 1231 executes the porting script while monitoring writing and execution of the bytecode cache detected by the bytecode cache detection unit 1213, and extracts the data written in the bytecode cache as the porting bytecode.

[0118] The symbol table extraction unit 1232 executes the porting script, obtains a pointer to the management structure from the argument of the interpretation execution function, and further extracts the porting symbol table by tracing back from the management structure using the symbol table reference and structure information.

[0119] The injection unit 124 injects a porting bytecode and a porting script into the bytecode and the symbol table generated in the memory space of the script engine by the execution of the porting destination script, respectively.

[0120] The configuration of the injection unit 124 will be described with reference to Fig. 22. Fig. 22 is a diagram showing an example of the configuration of the injection unit. As shown in Fig. 22, the injection unit 124 has an execution operation unit 1241, a bytecode injection unit 1242, and a symbol table injection unit 1243.

[0121] The execution control unit 1241 controls the execution of the porting destination script by the script engine. The execution control unit 1241 includes an execution unit 12411, an execution suspension unit 12412, a memory enumeration unit 12413, and an execution resumption unit 12414.

[0122] The execution unit 12411 causes the script engine to start executing the porting destination script. The execution pause unit 12412 causes the script engine to pause the execution of the porting destination script. The memory enumeration unit 12413 extracts and enumerates the memory areas of the bytecode and symbol table from the process of the script engine that is executing the porting destination script. The execution resumption unit 12414 causes the script engine to start executing the porting destination script.

[0123] The bytecode injection unit 1242 injects (writes) the porting bytecode into the memory area of ​​the enumerated bytecode. The symbol table injection unit 1243 injects (writes) the porting bytecode into the memory area of ​​the enumerated symbol table.

[0124] The output unit 14 is, for example, a liquid crystal display or a printer, and outputs various information including information related to the analysis device 10. The output unit 14 may also be an interface that controls input and output of various data between an external device, and may output various information to an external device.

[0125] [Virtual program counter detection processing] Next, the processing of the virtual program counter detection unit 1212 will be described. The virtual program counter detection unit 1212 detects VPC and pointer cache. The detection of the virtual program counter is realized by analyzing the log of the memory access trace of the acquired execution trace. The virtual program counter detection unit 1212 uses differential execution analysis focusing on the number of times memory is read. FIG. 20 is a diagram for explaining the processing of the virtual program counter detection unit 1212.

[0126] The virtual program counter detection unit 1212 extracts one execution trace by the first test script from the execution trace DB 131. The number of times a VPC is read is proportional to the number of repetitions in the test script and the number of statements in the repetitive process. When the number of repetitions is N and the number of repeated statements is M, approximately MN VPCs are read. Therefore, the virtual program counter detection unit 1212 extracts memory that has increased by 4MN and 9MN in the execution trace for the first test script in which N and M have been increased to 2N and 2M, respectively, and 3N and 3M. Specifically, as shown in FIG. 20, the virtual program counter detection unit 1212 extracts a memory area that has a monotonically increasing Read / Write for each VM instruction execution ((1) in FIG. 20).

[0127] Then, the virtual program counter detection unit 1212 detects, as a VPC, a memory value that always points to the start point of a VM instruction. Specifically, the virtual program counter detection unit 1212 compares the VPC's pointing destination with the address of the VM instruction handler, and narrows down the memory area to the matching memory area ((2) in FIG. 20).

[0128] [Bytecode cache detection part] Next, a description will be given of the processing of the bytecode cache detection unit 1213. FIG.

[0129] The bytecode cache detection unit 1213 detects a memory area pointed to by a VPC as a bytecode cache from a VM execution trace ((1) in FIG. 21).

[0130] The bytecode cache detection unit 1213 detects the code location of the caller of the memory allocation function that allocated this bytecode cache from the execution trace ((2) in FIG. 21). The bytecode cache detection unit 1213 detects all memory areas allocated at this code location from the VM execution trace as bytecode caches ((3) in FIG. 21).

[0131] The bytecode cache detection unit 1213 detects a code location that writes to the bytecode cache from the execution trace ((4) in FIG. 21). The bytecode cache detection unit 1213 detects writing by this code location from the VM execution trace as an update of the bytecode cache ((5) in FIG. 21).

[0132] [Analysis equipment processing procedure] The processing procedure of the analysis device 10 will be described with reference to Fig. 23. Fig. 23 is a flowchart showing the processing procedure of the analysis process according to the embodiment.

[0133] First, the input unit 11 accepts input of a test script and a script engine binary (step S11).

[0134] A script engine binary is an executable file that makes up a script engine. A script engine binary may consist of multiple executable files.

[0135] Next, the execution trace acquisition unit 1211 executes an execution trace acquisition process (step S12). The virtual program counter detection unit 1212 executes a virtual program counter detection process (step S13). The bytecode cache detection unit 1213 executes a bytecode cache detection process (step S14).

[0136] Next, the interpretation execution function detection unit 1221 executes an interpretation execution function detection process (step S15). The symbol table detection unit 1222 executes a symbol table detection process (step S16). The symbol table analysis unit 1223 executes a symbol table analysis process (step S17).

[0137] Here, if the extraction process is not executed (step S18, No), the output unit 14 outputs the bytecode, the symbol table, and the analysis result of the symbol table (step S19). On the other hand, if the extraction process is executed (step S18, Yes), the analysis device 10 proceeds to step S20. For example, whether or not the extraction process is executed is designated in advance by the user. Furthermore, when proceeding to step S20, the analysis device 10 functions as an extraction device.

[0138] The process from step S20 onwards will be described. The value object analysis unit 1224 executes a value object table analysis process (step S20). The bytecode extraction unit 1231 executes a bytecode extraction process (step S21). The symbol table extraction unit 1232 executes a symbol table extraction process (step S22).

[0139] If the injection process is not executed (step S23, No), the output unit 14 outputs the bytecode, the symbol table, and the structural information of the symbol table (step S24). On the other hand, if the injection process is executed (step S23, Yes), the analysis device 10 proceeds to step S25.

[0140] The process after step S25 will be described. The injection unit 124 executes the injection process (step S25). The output unit 14 outputs the result of the injection process (step S26).

[0141] The details of the processing procedure of the execution trace acquisition processing (step S12 in FIG. 23) will be described with reference to FIG 24. FIG 24 is a flowchart showing the processing procedure of the execution trace acquisition processing.

[0142] 24, first, the execution trace acquisition unit 1211 receives a test script and a script engine binary as input (step S1201). Next, the execution trace acquisition unit 1211 hooks the virtual machine to acquire a branch trace (step S1202). In addition, the execution trace acquisition unit 1211 hooks the virtual machine to acquire a memory access trace (step S1203).

[0143] Here, the execution trace acquiring unit 1211 inputs the test script into the virtual machine and executes it (step S1204), and then the execution trace acquiring unit 1211 stores the acquired execution trace in the execution trace DB 131 (step S1205).

[0144] If the execution trace acquisition unit 1211 has not executed all the input test scripts (step S1206, No), the process returns to step S1204 and repeats the process. On the other hand, if the execution trace acquisition unit 1211 has executed all the input test scripts (step S1206, Yes), the process ends.

[0145] The virtual program counter detection process (step S13 in FIG. 23) will be described in detail with reference to FIG 25. FIG 25 is a flowchart showing the procedure of the virtual program counter detection process.

[0146] As shown in FIG. 25, first, the virtual program counter detection unit 1212 extracts one execution trace by the first test script from the execution trace DB (step S1301).

[0147] Next, the virtual program counter detection unit 1212 focuses on the memory access trace and counts the number of reads for each memory read destination (step S1302). The virtual program counter detection unit 1212 receives as input the first test script used to acquire the execution trace (step S1303). The virtual program counter detection unit 1212 analyzes the test script to acquire the number of repetitions and the number of repeated statements (step S1304).

[0148] Next, the virtual program counter detection unit 1212 extracts another execution trace by the first test script having a different number of repetitions or number of repeated statements from the execution trace DB (step S1305). The virtual program counter detection unit 1212 focuses on the memory access trace and counts up the number of reads for each memory read destination (step S1306). The virtual program counter detection unit 1212 receives as input the first test script used to obtain the execution trace (step S1307).

[0149] Then, the virtual program counter detection unit 1212 analyzes the test script to obtain the number of repetitions and the number of repeated statements (step S1308).The virtual program counter detection unit 1212 narrows down the memory read destinations to only those whose read counts change in proportion to the increase or decrease in the number of repetitions or the number of repeated statements (step S1309).The virtual program counter detection unit 1212 narrows down the memory read destinations to those whose read values ​​always point to the start point of the VM instruction (step S1310).

[0150] If the virtual program counter detection unit 1212 narrows down the memory read destinations to only one (step S1311, Yes), it stores the narrowed down read destination as a virtual program counter in the architecture information DB 132 (step S1312). On the other hand, if the virtual program counter detection unit 1212 cannot narrow down the memory read destinations to only one (step S1311, No), it returns to step S1305 and repeats the process.

[0151] The bytecode cache detection process (step S14 in FIG. 23) will be described in detail with reference to FIG. 26. FIG. 26 is a flowchart showing the procedure of the bytecode cache detection process.

[0152] As shown in FIG. 26, first, the bytecode cache detection unit 1213 receives an execution trace and a VM execution trace as input (step S1401).

[0153] Next, the bytecode cache detection unit 1213 acquires the memory area pointed to by the VPC from the VM execution trace (step S1402). The bytecode cache detection unit 1213 acquires the code location of the caller of the memory allocation function that has allocated the memory area from the execution trace (step S1403).

[0154] Next, the bytecode cache detection unit 1213 detects all areas secured at the code location as bytecode cache (step S1404). The bytecode cache detection unit 1213 acquires the code location that writes to the bytecode cache from the execution trace (step S1405). The bytecode cache detection unit 1213 detects all areas written at the code location as updates to the bytecode cache (step S1406).

[0155] The bytecode cache detection unit 1213 returns the bytecode cache and its updated location (step S1407).

[0156] The interpretation execution function detection process (step S15 in FIG. 23) will be described in detail with reference to FIG 27. FIG 27 is a flowchart showing the procedure of the interpretation execution function detection process.

[0157] 27, first, the interpretation execution function detection unit 1221 receives an execution trace as an input (step S1501). The interpretation execution function detection unit 1221 receives a VPC and a bytecode cache as an input (step S1502).

[0158] Next, the interpretive execution function detection unit 1221 detects each function from the call and return branches in the branch trace (step S1503). Here, the interpretive execution function detection unit 1221 extracts one function (step S1504). The interpretive execution function detection unit 1221 scans the memory access trace within the function (step S1505).

[0159] If memory access to the VPC and the bytecode cache is detected (Yes at step S1506), the interpretive execution function detection unit 1221 outputs the function as an interpretive execution function (step S1507).

[0160] If no memory access to the VPC or bytecode cache is found (step S1506, No), the interpreted execution function detection unit 1221 extracts the next function (step S1508) and returns to step S1505.

[0161] In this way, the interpretive execution function detection unit 1221 detects interpretive execution functions in which memory access is seen to either or both of the virtual program counter, which is a variable that points to the instructions of the script engine that executes the script, and the bytecode cache, which is the memory area pointed to by the virtual program counter, in the execution trace.

[0162] The symbol detection process (step S16 in FIG. 23) will be described in detail with reference to FIG 28. FIG 28 is a flowchart showing the processing procedure of the symbol table detection process.

[0163] 28, first, the symbol table detection unit 1222 receives an execution trace as an input (step S1601). The symbol table detection unit 1222 receives an interpretation execution function as an input (step S1602).

[0164] Next, symbol table detection unit 1222 executes the interpretation execution function while monitoring it (step S1603). Symbol table detection unit 1222 acquires arguments of the interpretation execution function (step S1604). If the arguments are pointers, symbol table detection unit 1222 assigns each of the arguments a unique taint tag (step S1605). Symbol table detection unit 1222 executes while propagating the taint tag according to the pointer reference (step S1606).

[0165] Here, the symbol table detection unit 1222 extracts one characteristic value in the test script (step S1607).The symbol table detection unit 1222 searches the memory access trace for the characteristic value (step S1608).

[0166] When a memory access with a characteristic value is detected (step S1609, Yes), the symbol table detection unit 1222 detects the target of the memory access as a value in the symbol table (step S1611).

[0167] If no memory access of a characteristic value is found (step S1609, No), the symbol table detection unit 1222 extracts the next characteristic value (step S1610).

[0168] If all characteristic values ​​have not been processed (step S1612, No), the symbol table detection unit 1222 extracts the next characteristic value (step S1610).

[0169] When all characteristic values ​​have been processed (Yes at step S1612), the symbol table detection unit 1222 checks the taint tags assigned to the values ​​in the symbol table (step S1613).

[0170] The symbol table detection unit 1222 determines that the taint tag is related to a symbol table (step S1614). The symbol table detection unit 1222 determines that the memory area to which the taint tag is assigned is a symbol table (step S1615). A taint map related to the taint tag is output (step S1616).

[0171] In this way, the symbol table detection unit 1222 detects, as a symbol table, multiple structures that contain addresses on a reference path from a first address to a second address that contains a value specified in the script in a memory area used by the script engine (an example of a program) to execute a script.

[0172] For example, the symbol table detection unit 1222 detects, as a symbol table, multiple structures including addresses on a reference path from a first address (the start address of the management structure) indicated by a pointer of an argument of the interpretation execution function to a second address (the address of the characteristic value).

[0173] Furthermore, symbol table detection unit 1222 assigns taint tags to addresses on a reference path from the first address to the second address, and detects, as a symbol table, a plurality of structures including addresses to which taint tags have been assigned.

[0174] The symbol table analysis process (step S17 in FIG. 23) will be described in detail with reference to FIG 29. FIG 29 is a flowchart showing the processing procedure of the symbol table detection process.

[0175] 29, first, the symbol table analyzer 1223 receives an execution trace as an input (step S1701). The symbol table analyzer 1223 receives a taint map of the symbol table as an input (step S1702).

[0176] Next, the symbol table analyzer 1223 extracts one value with a taint tag from the taint map (step S1703).The symbol table analyzer 1223 searches the memory access trace with the extracted value (step S1704).

[0177] If no memory access using the value is found (step S1705, No), symbol table analyzer 1223 extracts the next value (step S1706) and returns to step S1704.

[0178] If a memory access using a value is found (step S1705, Yes), the symbol table analysis unit 1223 executes a reference analysis process (step S1707) and a structure analysis process (step S1708) and outputs information on the reference and structure of the symbol table (step S1709).

[0179] In this way, the symbol table analysis unit 1223 analyzes the execution trace obtained by having a script engine (an example of a program) execute a test script (an example of a first script), and obtains structural information of the symbol table generated by the script engine.

[0180] This enables the symbol table analysis unit 1223 to analyze the script engine, detect the symbol table, and perform detailed analysis of its structure and the references between symbol tables, thereby making it possible to obtain information about the internal structure of the symbol table even for script engines of a wide variety of scripting languages ​​whose specifications are unknown.

[0181] For example, the structural information is the top address of each element (structure or array) of the symbol table shown in Fig. 14, the address and value of a pointer included in each element of the symbol table, etc. Therefore, the analysis device 10 can identify the structure of each element and the reference relationship between elements based on the structural information.

[0182] The reference analysis process (step S1707 in FIG. 29) will be described in detail with reference to Fig. 30. Fig. 30 is a flowchart showing the processing procedure of the reference analysis process.

[0183] As shown in FIG. 30, first, symbol table analyzer 1223 receives as input a specific memory access seen in a memory access trace (step S1741).

[0184] The symbol table analyzer 1223 determines whether the memory access uses a base register (step S1742) and whether the memory access uses an index register (step S1743).

[0185] If the memory access does not use the base register (step S1742, No), the symbol table analyzer 1223 determines that the memory access destination is neither an array nor a structure (step S1755). If the memory access uses the base register but not the index register (step S1742: Yes and step S1743: No), the symbol table analyzer 1223 determines that the memory access destination is a member variable of a structure (step S1744). If the memory access uses both the base register and the index register (step S1742 and step S1743: Yes), the symbol table analyzer 1223 determines that the memory access destination is an element of an array (step S1749).

[0186] Following step S1744, symbol table analyzer 1223 executes the following process. Symbol table analyzer 1223 acquires the base register and effective address at the time of the memory access (step S1745). Symbol table analyzer 1223 detects the value of the base register as the start address of the structure at the memory access destination (step S1746).

[0187] The symbol table analyzer 1223 calculates the offset of the member variable of the structure of the memory access destination from the difference between the effective address and the base register (step S1747). The symbol table analyzer 1223 outputs the start address and offset of the structure (step S1748).

[0188] Following step S1749, symbol table analyzer 1223 executes the following process. Symbol table analyzer 1223 acquires the base register and index register at the time of the memory access (step S1750). Symbol table analyzer 1223 detects the value of the base register as the starting address of the array at the memory access destination (step S1751).

[0189] The symbol table analyzer 1223 detects the value of the index register as a subscript of the array of the memory access destination (step S1752).The symbol table analyzer 1223 calculates the offset of the element of the array of the memory access destination from the difference between the effective address and the base register (step S1753).The symbol table analyzer 1223 outputs the top address, subscript, and offset of the array (step S1754).

[0190] In this way, when the memory access to an address on the reference path from the first address to the second address shown in the execution trace uses both the base register and the index register, the symbol table analyzer 1223 determines that the target of the memory access is an element of an array, and obtains the first address of the array and the subscript of the element from the base register and the index register. Also, when the memory access to an address on the reference path from the first address to the second address shown in the execution trace uses the base register but does not use the index register, the symbol table analyzer 1223 determines that the target of the memory access is a member variable of a structure, and calculates the offset of the member variable from the first address of the structure based on the effective address and the base register. This allows the symbol table analyzer 1223 to perform analysis regardless of whether the element of the symbol table is a structure or an array.

[0191] The structure analysis process (step S1708 in FIG. 29) will be described in detail with reference to FIG 31. FIG 31 is a flowchart showing the procedure of the structure analysis process.

[0192] As shown in FIG. 31, first, symbol table analyzer 1223 receives as input a specific memory access seen in a memory access trace (step S1771).

[0193] If the value objects of the symbol table are managed by an array (Yes at step S1772), the symbol table analysis unit 1223 outputs the structure of the array obtained by the reference analysis (step S1773).

[0194] If the value objects of the symbol table are not managed in an array (step S1772, No), the symbol table analyzer 1223 accepts multiple test scripts as input (step S1774). The symbol table analyzer 1223 performs processing up to the reference analysis processing for each test script (step S1775). The symbol table analyzer 1223 extracts references from the management structure to value objects from the references of each test script (step S1776).

[0195] Here, symbol table analysis unit 1223 obtains the difference in the references (step S1777). If no difference in the length of the references is found in accordance with the number of values ​​in the test script (step S1778, No), symbol table analysis unit 1223 ends the process.

[0196] If a difference in the length of the reference is found according to the number of values ​​in the test script (Yes at step S1778), the symbol table analyzer 1223 determines that the part where the difference is found is a linked list (step S1779).

[0197] The symbol table analyzer 1223 detects, from the reference, a member variable having a reference that connects elements of the linked list (step S1780). The symbol table analyzer 1223 detects, from the reference, a value object portion that an element of the linked list has (step S1781). The symbol table analyzer 1223 outputs the structure of the linked list (step S1782).

[0198] Symbol table detection unit 1222 can detect, as a symbol table, a plurality of structures including addresses on a reference path for each of a plurality of test scripts. Symbol table analysis unit 1223 detects a structure of a part (unlinked list) common to a plurality of symbol tables and a structure of a part (linked list) not common to a plurality of symbol tables based on the length of each reference path of a plurality of symbol tables detected by symbol table detection unit 1222. In this way, symbol table analysis unit 1223 can analyze the structure of a symbol table with higher accuracy by using a plurality of test scripts.

[0199] The value object analysis process (step S20 in FIG. 23) will be described in detail with reference to FIG 32. FIG 32 is a flowchart showing the procedure of the value object analysis process.

[0200] 32, first, the value object analysis unit 1224 receives a plurality of test scripts as input (step S2021). The value object analysis unit 1224 receives a plurality of structures as input (step S2022).

[0201] The value object analysis unit 1224 detects member variables that hold values ​​of the test script from the structure by matching (step S2023).The value object analysis unit 1224 extracts a portion of the structure that does not depend on the value of the test script (step S2024).

[0202] The value object analysis unit 1224 finds differences between structures while associating them with test scripts (step S2025). The value object analysis unit 1224 detects member variables for which differences are found depending on the type of the test script as parts that hold tags (step S2026).

[0203] The value object analysis unit 1224 detects the correspondence between the tag value and the type by examining the dependency between the test script type and the tag value from the difference (step S2027).The value object analysis unit 1224 detects the type held by the structure from the type of the tag value for the test script value (step S2028).

[0204] The bytecode extraction process (step S21 in FIG. 23) will be described in detail with reference to FIG 33. FIG 33 is a flowchart showing the procedure of the bytecode extraction process.

[0205] 33, first, the bytecode extraction unit 1231 accepts a VPC as an input (step S2101). The bytecode extraction unit 1231 accepts a bytecode cache as an input (step S2102). The bytecode extraction unit 1231 accepts a script to be extracted (e.g., a porting destination script) as an input (step S2103). The bytecode extraction unit 1231 accepts a script engine binary as an input (step S2104).

[0206] The bytecode extraction unit 1231 executes the script to be extracted (step S2105). The bytecode extraction unit 1231 stops the execution at the interpretation execution function (step S2106). The bytecode extraction unit 1231 extracts the bytecode from the bytecode cache (step S2107). The bytecode extraction unit 1231 resumes the execution, and stops the execution again at the time of the execution of the first VM instruction (step S2108).

[0207] The bytecode extraction unit 1231 extracts the value of the VPC (step S2109), and outputs the extracted bytecode and the value of the VPC (step S2110).

[0208] The symbol table extraction process (step S22 in FIG. 23) will be described in detail with reference to Fig. 34. Fig. 34 is a flowchart showing the processing procedure of the symbol table extraction process.

[0209] 34, first, symbol table extraction unit 1232 accepts symbol table reference and structure information as input (step S2201). Symbol table extraction unit 1232 accepts a script to be extracted as input (step S2202). Symbol table extraction unit 1232 accepts a script engine binary as input (step S2203).

[0210] The symbol table extraction unit 1232 executes the script to be extracted (step S2204). The execution is stopped by the interpretation execution function (step S2205). The symbol table extraction unit 1232 acquires a pointer to the management structure from the argument of the interpretation execution function (step S2206). The symbol table extraction unit 1232 extracts the symbol table by tracing the management structure from the symbol table reference and structure information (step S2207). The symbol table extraction unit 1232 outputs the extracted symbol table (step S2208).

[0211] In this way, symbol table extraction unit 1232 extracts a symbol table from the memory area used by the script engine to execute the porting script (an example of a second script) based on the structural information. By using the structural information, which is detailed information on the structure of the symbol table and the references between the symbol tables, symbol table extraction unit 1232 can detect the symbol table more accurately. As a result, more accurate porting (injection) of the bytecode and the symbol table becomes possible.

[0212] Symbol table extraction unit 1232 extracts, as a symbol table, a number of structures including the addresses obtained by repeatedly calculating the addresses of the pointers included in the structures and the addresses of the beginnings of the structures pointed to by the pointers, based on the structure information, starting from a specific structure. Repeated address calculation corresponds to the process of tracing the reference from the management structure, as shown in step S2207.

[0213] The symbol table extraction unit 1232 can extract a symbol table based on the interpretation execution function. That is, the symbol table extraction unit 1232 extracts, as a symbol table, a plurality of structures including addresses obtained by repeatedly calculating the address of a pointer included in a structure and the address of the beginning of a structure pointed to by a pointer based on structural information, starting from a management execution entity pointed to by a pointer of an argument of an interpretation execution function in which memory access to either or both of a virtual program counter, which is a variable pointing to an instruction of a program executing a porting script, and a bytecode cache, which is a memory area pointed to by the virtual program counter, is observed in an execution trace, as a symbol table.

[0214] By determining the starting point in this manner, the symbol table extraction unit 1232 can accurately detect the symbol table.

[0215] The injection process (step S25 in FIG. 23) will be described in detail with reference to FIG 35. FIG 35 is a flowchart showing the procedure of the injection process.

[0216] 35, first, the injection unit 124 receives a test script and a script engine binary at the input unit 11 (step S2501). The injection unit 124 inputs a target script to a target script engine at the execution unit 2411 and executes it (step S2502).

[0217] The target script engine is the script engine that corresponds to the script engine binary. The target script is, for example, the script to be ported to.

[0218] The injection unit 124 pauses the target script engine at the beginning of the interpretation execution function in the execution pause unit 2412 (step S2503). The injection unit 124 receives the bytecode and VPC value extracted by the extraction unit 123 as input (step S2504). The injection unit 124 receives the symbol table extracted by the extraction unit 123 as input (step S2505).

[0219] The injection unit 124 executes a bytecode injection process (step S2506) and a symbol table injection process (step S2507). The injection unit 124 resumes the execution of the target script engine in the execution resumption unit 2414 (step S2508).

[0220] The bytecode injection process (step S2506 in FIG. 35) will be described in detail with reference to FIG. 36. FIG. 36 is a flowchart showing the procedure of the bytecode injection process.

[0221] 36, first, the bytecode injection unit 1242 accepts a process of a running target script engine as an input (step S2541). The bytecode injection unit 1242 accepts a VPC as an input (step S2542). The bytecode injection unit 1242 accepts a bytecode cache as an input (step S2543). The bytecode injection unit 1242 accepts an extracted bytecode as an input (step S2544).

[0222] The bytecode injection unit 1242 writes the extracted bytecode to the bytecode cache of the process of the target script engine (step S2545). The bytecode injection unit 1242 sets the memory address where the extracted bytecode is written to the VPC (step S2546).

[0223] The symbol table injection process (step S2507 in FIG. 35) will be described in detail with reference to FIG. 37. FIG. 37 is a flowchart showing the processing procedure of the symbol table injection process.

[0224] 37, first, the symbol table injection unit 1243 accepts a process of a running target script engine as an input (step S2571). The symbol table injection unit 1243 accepts structural information of a symbol table as an input (step S2572). The symbol table injection unit 1243 accepts an extracted symbol table as an input (step S2573).

[0225] The symbol table injection unit 1243 obtains the start address of the management structure from the argument of the interpretation execution function (step S2574). The symbol table injection unit 1243 extracts a symbol table of one scope from the extracted symbol table (step S2575). The symbol table injection unit 1243 refers to the start address of the management structure based on the structural information of the symbol table, and traces to the symbol table of the corresponding scope (step S2576).

[0226] The symbol table injection unit 1243 writes the symbol table of the extracted symbol table of the scope to the traced destination (step S2577). When the symbol table injection unit 1243 has injected all the scopes of the extracted symbol table (step S2578, Yes), the process ends.

[0227] If all the scopes of the extracted symbol table have not been injected (step S2578, No), symbol table injection unit 1243 extracts the symbol table of the next scope from the extracted symbol table (step S2579), and returns to step S2576.

[0228] In this way, symbol table injection unit 1243 injects the symbol table extracted by symbol table extraction unit 1232 into the position specified based on the structural information in the memory area used by the program to execute the porting destination script (third script example). Symbol table injection unit 1243 can accurately inject the symbol table by using the structural information.

[0229] [System configuration of the embodiment] Each component of analysis device 10 shown in Fig. 4 is a functional concept, and does not necessarily have to be physically configured as shown in the figure. In other words, the specific form of distribution and integration of the functions of analysis device 10 is not limited to that shown in the figure, and all or part of it can be functionally or physically distributed or integrated in any unit depending on various loads, usage conditions, etc.

[0230] Furthermore, all or any part of each process performed in analysis device 10 may be realized by a CPU and a program analyzed and executed by the CPU. Furthermore, each process performed in analysis device 10 may be realized as hardware using wired logic.

[0231] Furthermore, among the processes described in the embodiments, all or part of the processes described as being performed automatically can be performed manually. Alternatively, all or part of the processes described as being performed manually can be performed automatically by a known method. In addition, the process procedures, control procedures, specific names, and information including various data and parameters described above and shown in the drawings can be changed as appropriate unless otherwise specified.

[0232] [program] 38 is a diagram showing an example of a computer in which analysis device 10 is realized by executing a program. Computer 1000 has, for example, memory 1010 and CPU 1020. Computer 1000 also has hard disk drive interface 1030, disk drive interface 1040, serial port interface 1050, video adapter 1060, and network interface 1070. Each of these components is connected by bus 1080.

[0233] The memory 1010 includes a ROM 1011 and a RAM 1012. The ROM 1011 stores a boot program such as a BIOS (Basic Input Output System). The hard disk drive interface 1030 is connected to a hard disk drive 1090. The disk drive interface 1040 is connected to a disk drive 1100. A removable storage medium such as a magnetic disk or an optical disk is inserted into the disk drive 1100. The serial port interface 1050 is connected to a mouse 1110 and a keyboard 1120, for example. The video adapter 1060 is connected to a display 1130, for example.

[0234] The hard disk drive 1090 stores, for example, an OS 1091, an application program 1092, a program module 1093, and program data 1094. That is, a program that defines each process of the analysis device 10 is implemented as a program module 1093 in which code executable by the computer 1000 is written. The program module 1093 is stored, for example, in the hard disk drive 1090. For example, the program module 1093 for executing the same process as the functional configuration of the analysis device 10 is stored in the hard disk drive 1090. The hard disk drive 1090 may be replaced by an SSD (Solid State Drive).

[0235] Furthermore, setting data used in the processing of the above-described embodiment is stored as program data 1094, for example, in the memory 1010 or the hard disk drive 1090. Then, the CPU 1020 reads out the program module 1093 or the program data 1094 stored in the memory 1010 or the hard disk drive 1090 into the RAM 1012 as necessary, and executes them.

[0236] Note that the program module 1093 and the program data 1094 are not limited to being stored in the hard disk drive 1090, but may be stored in, for example, a removable storage medium and read by the CPU 1020 via the disk drive 1100 or the like. Alternatively, the program module 1093 and the program data 1094 may be stored in another computer connected via a network (such as a local area network (LAN) or wide area network (WAN)). The program module 1093 and the program data 1094 may then be read by the CPU 1020 from the other computer via the network interface 1070.

[0237] Although the embodiments of the present invention have been described above, the present invention is not limited by the descriptions and drawings that form part of the disclosure of the present invention according to the present embodiments. In other words, other embodiments, examples, and operation techniques made by those skilled in the art based on the present embodiments are all included in the scope of the present invention. [Explanation of symbols]

[0238] 10 Analysis device 11 Input section 12 Control section 13 Storage section 14 Output section 121 Virtual Machine Analysis Department 122 Symbol Table Analysis Unit 123 Extraction part 124 Injection part 131 Execution Trace DB 132 Architecture Information DB 1211 Execution trace acquisition unit 1212 Virtual Program Counter Detector 1213 Bytecode Cache Detector 1221 Interpretation Execution Function Detection Unit 1222 Symbol Table Detector 1223 Symbol Table Analysis Unit 1224 Value Object Parser 1231 Bytecode Extraction Unit 1232 Symbol Table Extraction Unit 1241 Execution Operation Unit 1242 Bytecode Injector 1243 Symbol Table Injection Section 12411 Executive Department 12412 Execution pause section 12413 Memory Enumeration Unit 12414 Execution restart section

Claims

1. a symbol table analysis unit that analyzes an execution trace obtained by executing a first script by a program and acquires structural information of a symbol table generated by the program; a symbol table extraction unit that extracts a symbol table from a memory area used by the program to execute a second script based on the structural information; An extraction apparatus comprising:

2. The symbol table extraction unit extracts, as a symbol table, a plurality of structures including addresses obtained by repeatedly calculating, based on the structural information, an address of a pointer included in a structure and a start address of the structure indicated by the pointer, starting from a specific structure.

2. The extraction device according to claim 1 .

3. The symbol table extraction unit extracts, as a symbol table, a plurality of structures including addresses obtained by repeatedly calculating, based on the structural information, the addresses of pointers included in structures and the addresses of the tops of structures pointed to by the pointers, starting from a management execution entity pointed to by a pointer of an argument of an interpretation execution function in which memory access to either or both of a virtual program counter, which is a variable pointing to an instruction of a program executing the second script, and a bytecode cache, which is a memory area pointed to by the virtual program counter, is seen in the execution trace, and extracts, based on the structural information, a plurality of structures including addresses obtained by repeatedly calculating the addresses of pointers included in structures and the top addresses of structures pointed to by the pointers, 2. The extraction device according to claim 1 .

4. a symbol table injection unit that injects the symbol table extracted by the symbol table extraction unit into a position in the memory area used by the program to execute the third script, the position being specified based on the structural information; 2. The extraction device of claim 1 further comprising:

5. A method of extraction performed by an extraction device, comprising: a symbol table analysis step of analyzing an execution trace obtained by executing a first script by a program and acquiring structural information of a symbol table generated by the program; a symbol table extraction step of extracting a symbol table from a memory area used by the program to execute a second script based on the structural information; An extraction method comprising the steps of:

6. a symbol table analysis step of analyzing an execution trace obtained by executing a first script by a program and acquiring structural information of a symbol table generated by the program; a symbol table extraction step of extracting a symbol table from a memory area used by the program to execute a second script based on the structural information; An extraction program that causes a computer to execute the above steps.