A method and device for collecting instruction streams in a hybrid mode
By enabling only the basic block translation callback in the initial stage of target program simulation, and dynamically enabling instruction execution callbacks and memory usage callbacks in the instruction acquisition area, the problems of large system overhead and low instruction flow acquisition efficiency caused by callback strategies in the prior art are solved, and more efficient instruction flow acquisition is achieved.
Patent Information
- Application Number
- CN202411132416.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-16
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2044-08-16
AI Technical Summary
In the prior art, callback strategies lead to large system overhead and low instruction flow acquisition efficiency, especially in the non-instruction acquisition area information acquisition.
The instruction stream acquisition method adopts a hybrid mode. Only the basic block translation callback is enabled in the initial stage of the target program simulation. The callback strategy is dynamically adjusted through the translation block execution callback and the instruction execution callback, distinguishing between the instruction acquisition area and the non-instruction acquisition area, and only enabling the instruction execution callback and memory usage callback in the instruction acquisition area.
By dynamically adjusting the callback strategy, the information acquisition overhead in the non-instruction acquisition area is reduced, the efficiency of instruction flow acquisition is improved, and the system operation pressure is reduced.
Smart Images

Figure CN119088657B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technologies, and in particular, to a method and apparatus for collecting instruction streams in a hybrid mode. Background Art
[0002] The performance simulation of a processor microarchitecture is one of the important links in the development of a CPU processor. The performance simulation of a cycle-accurate microarchitecture is usually about 1000 - 10000 times slower than the actual machine operation. At the same time, since some benchmark programs contain a large number of instructions, for example, a single test of SpecCpu2017 contains trillions of instructions, it is necessary to collect fragments of the instruction stream. The collection of instruction stream fragments is usually implemented by a functional simulator, and then the performance simulation of the instruction stream fragments is performed by a cycle-accurate microarchitecture simulator. The instruction stream information required for performance simulation includes: the initial states of all registers of the CPU processor at the beginning of the sampling fragment and the initial state of the target program memory, and detailed information such as the address, instruction code, code length, values of registers used by the instruction, memory address and data value of the memory operation performed by the instruction within the sampling fragment. Specifically, an instruction stream collector is loaded in the functional simulator in the form of a plug-in, and the instruction stream collector is used to register callbacks. The functional simulator translates the collected instructions into instructions (translation blocks) on the host machine. To avoid multiple translations, the functional simulator puts the translation blocks into a code cache and connects multiple translation blocks through jumps on the host machine. If a callback is registered, the call of the callback will also be added to the translation block.
[0003] Functionally speaking, if the functional simulator wants to collect the above-mentioned instruction stream information, it needs to register callbacks (such as instruction execution callbacks and memory usage callbacks) in the instruction stream collector at the initial stage of simulation, so as to embed the call of the callback in the translation block when translating the target instruction, thereby realizing the monitoring of the instruction execution and memory usage conditions. Summary of the Invention
[0004] In the prior art, a callback strategy is to register a memory usage callback in the instruction stream collector at the initial stage of simulation to obtain the memory information of the instruction stream fragment. Specifically, the memory information of the instruction is determined by enabling the memory access callback after each instruction execution. Since the callback is a piece of program, executing the callback every time after each instruction execution starting from the starting point of the program will consume running time. Therefore, the above-mentioned memory usage callback delays the running speed of the instruction simulation. In addition to the additional overhead of the memory usage callback, the instruction stream collector also needs to maintain the information of the memory usage callback, which will also increase the simulation overhead.
[0005] Other callback strategies in the prior art register multiple callbacks in the instruction stream collector during the simulation initialization phase, and during the simulation, the callback strategy is not dynamically adjusted (i.e., the registered callbacks are not adjusted), resulting in the use of unnecessary callbacks in the system and further reducing the speed of instruction stream collection.
[0006] The present invention is used to solve the problems of large system overhead and low instruction stream collection efficiency caused by the existing callback strategy. To solve the above technical problems, a first aspect of the present invention provides a method for collecting instruction streams in a hybrid mode, including:
[0007] Only enable the basic block translation callback during the simulation initialization phase of the target program;
[0008] After storing the translation blocks obtained by translating the basic blocks of the target program into the code cache, call the basic block translation callback to determine whether it is in the instruction collection area. If not, enable the translation block execution callback. If so, enable the instruction execution callback and the memory usage callback;
[0009] When calling the translation block execution callback, determine whether it enters the entry point of the instruction collection area after determining the relevant basic block. If so, obtain the memory information and register information of the target program based on the memory information of the simulation process and record them in the instruction stream collection file, and clear the code cache of the translation block;
[0010] When calling the instruction execution callback, obtain the instruction information and record it in the instruction stream collection file, and determine whether it is at the end point of the instruction collection area after determining the relevant basic block. If so, clear the code cache of the translation block.
[0011] In some embodiments of the present invention, obtaining the memory information and register information of the target program based on the memory information of the simulation process includes:
[0012] Obtain the memory information of the simulation process;
[0013] Identify the memory addresses in the memory information and obtain the page table and page table attributes corresponding to the memory addresses;
[0014] Judge whether the page table attributes contain the identifier of the target program. If so, determine that the memory address corresponding to the page table is the memory address of the target program, and record the page table and its corresponding memory address in the instruction stream collection file;
[0015] Obtain the register information and record it in the instruction stream collection file.
[0016] In some embodiments of the present invention, before enabling the instruction execution callback and the memory usage callback, the following steps are further included: traversing the instruction machine codes of each instruction in the translation block, and performing the following operations on the instruction machine code of each instruction: checking whether the instruction machine code is in the hash table, if not, parsing the instruction machine code to obtain instruction register information, and inserting the instruction machine code and the instruction register information into the hash table, if so, continuing to traverse the next instruction in the translation block;
[0017] After all instructions in the translation block have been traversed, enable the instruction execution callback and the memory usage callback.
[0018] In some embodiments of the present invention, before the basic block translation callback determines whether to enter the instruction collection area, the following steps are further included: querying the number of instructions in the translation block;
[0019] When enabling the translation block execution callback, the number of instructions in the translation block is also passed into the translation block execution callback;
[0020] When enabling the instruction execution callback, the instruction machine code of each instruction in the translation block is also passed into the instruction execution callback.
[0021] In some embodiments of the present invention, obtaining instruction information includes:
[0022] Searching for the register information of the instruction machine code of the current instruction in the hash table, and recording the register information and the instruction machine code of the current instruction in the instruction stream collection file;
[0023] Obtaining the register value in the current register state of the processor, and recording the register value in the instruction stream collection file;
[0024] Checking whether the current instruction crosses the page table, if so, recording the additional page table information in the instruction stream collection file;
[0025] Checking whether the current instruction is a jump instruction and whether it is a system call;
[0026] If it is a jump instruction, record the page table information of the destination address in the instruction stream collection file;
[0027] If it is not a jump instruction and is a system call, compare the memory before and after the instruction call, and record the comparison result in the instruction stream collection file;
[0028] If it is not a system call, check whether there is an exception, if so, record the exception information in the instruction stream collection file.
[0029] In some embodiments of the present invention, the execution process of the memory usage callback includes:
[0030] Obtain instruction memory information, where the instruction memory information includes: the data length, type, data virtual address, and physical address of the memory access;
[0031] Record the instruction memory information in the instruction stream collection file;
[0032] Determine whether the data access crosses pages. If it does, record additional page table information in the instruction stream collection file.
[0033] In some embodiments of the present invention, the translation block execution callback determines the entry point for entering the instruction collection area after determining the relevant basic block, including:
[0034] Add the number of instructions in the translation block to the historical instruction accumulation value to obtain a new instruction accumulation value;
[0035] Determine whether to enter the instruction collection area based on the new instruction accumulation value.
[0036] The second aspect of the present invention provides an instruction stream collection device in a hybrid mode. The instruction stream collection device includes:
[0037] A first configuration unit for enabling only the basic block translation callback in the initial stage of target program simulation;
[0038] A second configuration unit for storing the translation blocks obtained by translating the basic blocks of the target program into the code cache and then calling the basic block translation callback to determine whether it is in the instruction collection area. If not, enable the translation block execution callback. If so, enable the instruction execution callback and the memory usage callback;
[0039] An instruction collection area entry point monitoring unit for determining the entry point for entering the instruction collection area after determining the relevant basic block when calling the translation block execution callback. If so, obtain the memory information and register information of the target program based on the memory information of the simulation process and record them in the instruction stream collection file, and clear the code cache of the translation block;
[0040] An instruction collection area end point monitoring unit for obtaining instruction information and recording it in the instruction stream collection file when calling the instruction execution callback, and determining whether it is at the end point of the instruction collection area after determining the relevant basic block. If so, clear the code cache of the translation block.
[0041] The third aspect of the present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the method described in any of the foregoing embodiments.
[0042] In a fourth aspect of the present invention, there is provided a computer storage medium having a computer program stored thereon, and when the computer program is executed by a processor of a computer device, the method described in any of the foregoing embodiments is implemented.
[0043] In a fifth aspect of the present invention, there is provided a computer program product, which includes a computer program, and when the computer program is executed by a processor of a computer device, the method described in any of the foregoing embodiments is implemented.
[0044] The method and apparatus for collecting instruction streams in a hybrid mode provided by the present invention enable only basic block translation callbacks in the initial stage of target program simulation; store the translation blocks obtained by translating the basic blocks of the target program into a code cache, and then call the basic block translation callback to determine whether it is in the instruction collection area. If not, enable the translation block execution callback. If so, enable the instruction execution callback and the memory usage callback; when calling the translation block execution callback, determine whether it enters the entry point of the instruction collection area after determining the relevant basic block. If so, obtain the memory information and register information of the target program based on the memory information of the simulation process and record them in the instruction stream collection file, clear the code cache of the translation block. When calling the instruction execution callback, obtain the instruction information and record it in the instruction stream collection file, and determine whether it is at the end point of the instruction collection area after determining the relevant basic block. If so, clear the code cache of the translation block. The present invention distinguishes between the instruction collection area and the non-instruction collection area. In the non-instruction collection area, only the translation block execution callback is enabled. In the instruction collection area, the instruction execution callback and the memory usage callback are enabled, which can reduce information acquisition in the non-instruction collection area and improve the sampling efficiency. By timely clearing the translation blocks in the code cache, the present invention can re-judge whether it is in the instruction collection area, and thus adjust the enabling strategy of the callback according to the judgment result, improving the sampling efficiency.
[0045] Furthermore, when entering the entry point of the instruction collection area, the memory information and register information of the target program are obtained based on the memory information of the simulation process and recorded in the instruction stream collection file. Specifically, when executing to the entry point of the instruction collection area, the memory information of the simulation process is obtained; the memory addresses in the memory information are identified, and the page table corresponding to the memory address is obtained; according to the page table corresponding to the memory address, the page table attributes are determined; it is judged whether the page table attributes contain the identifier of the target program. If so, the memory address corresponding to the page table is determined as the memory address of the target program, and the page table and its corresponding memory address are recorded in the instruction stream collection file. Through the above technical solutions, the memory of the target program and the memory of the function simulator can be distinguished, thereby realizing the collection of memory information at the entry point of the instruction collection area. At the same time, through the above technical solutions, the operation of the memory usage callback and the maintenance of the memory usage callback information are omitted, which can reduce the system operation pressure and improve the instruction stream collection efficiency.
[0046] To make the above and other objects, features, and advantages of the present invention more obvious and understandable, the following provides preferred embodiments in conjunction with the accompanying drawings and describes them in detail as follows. Brief Description of the Drawings
[0047] To more clearly illustrate the technical solutions in the embodiments of the present invention or in the prior art, the following briefly introduces the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0048] Figure 1 Shows the structural diagram of the instruction stream acquisition system in the hybrid mode of the embodiment of the present invention;
[0049] Figure 2A Shows the schematic diagram of the correspondence between the basic block and the translation block in the embodiment of the present invention;
[0050] Figure 2B Shows the schematic diagram of the relationship between the translation block and the callback in the embodiment of the present invention;
[0051] Figure 3 Shows the schematic diagram of the sampling segment of the benchmark program in the embodiment of the present invention;
[0052] Figure 4 Shows the flowchart of the instruction stream acquisition method in the hybrid mode of the embodiment of the present invention;
[0053] Figure 5 Shows the flowchart of the execution process of the basic block translation callback in the embodiment of the present invention;
[0054] Figure 6 Shows the flowchart of the execution process of the translation block execution callback in the embodiment of the present invention;
[0055] Figure 7 Shows the flowchart of the execution process of the instruction execution callback in the embodiment of the present invention;
[0056] Figure 8 Shows the flowchart of the execution process of the memory usage callback in the embodiment of the present invention;
[0057] Figure 9 Shows the structural diagram of the instruction stream acquisition device in the hybrid mode of the embodiment of the present invention;
[0058] Figure 10 Shows the structural diagram of the computer device in the embodiment of the present invention.
[0059] Description of the Reference Numerals in the Drawings:
[0060] 101, Functional Simulator;
[0061] 102. Instruction stream collector;
[0062] 103. Processor;
[0063] 104. Code cache;
[0064] 901. First configuration unit;
[0065] 902. Second configuration unit;
[0066] 903. Instruction collection area entry point monitoring unit;
[0067] 904. Instruction collection area end point monitoring unit;
[0068] 1002. Computer device;
[0069] 1004. Processor;
[0070] 1006. Memory;
[0071] 1008. Driving mechanism;
[0072] 1010. Input / output module;
[0073] 1012. Input device;
[0074] 1014. Output device;
[0075] 1016. Rendering device;
[0076] 1018. Graphical user interface;
[0077] 1020. Network interface;
[0078] 1022. Communication link;
[0079] 1024. Communication bus. Detailed implementation manners
[0080] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0081] It should be noted that the terms "first", "second", etc. in the specification, claims and above-mentioned drawings of the present invention are used to distinguish similar objects and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, product or equipment comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or equipment.
[0082] This specification provides the method operation steps as described in the embodiments or flowcharts, but may include more or fewer operation steps based on routine or non-creative labor. The step sequences listed in the embodiments are only one way among the execution sequences of numerous steps and do not represent the only execution sequence. When the actual system or device product is executed, it can be executed in the order shown in the embodiments or drawings or in parallel.
[0083] It should be noted that the data involved in this application (including but not limited to the data for analysis, stored data, displayed data, etc.) are all information and data authorized by the user or fully authorized by all parties.
[0084] Take Figure 3 as an example, Figure 3 The B-C segment and D-E segment in Figure 3 are the instruction sampling areas. In order to collect the memory information at the starting point of the instruction stream segment ( Figure 3 point B in
[0085] ), the instruction stream collector needs to collect the memory operations of all instructions from the starting point of the program ( Figure 2AAs shown, the functional simulator T10 translates the basic block T20 in the target program code into a translation block of the host machine. Each basic block T20 contains multiple instructions T30, and the instructions T30 in the basic block T20 correspond to the translations T60 in the translation block T50. To avoid repeated translations, the functional simulator places the translation blocks into the host machine code cache T40 and connects multiple translation blocks through the jumps T70 of the host machine. If the callback function is enabled, the call of the callback function will also be added to the translation block, as Figure 2B shown. In a typical instruction stream collector, instruction callbacks and memory usage callbacks are enabled in the initial stage of simulation, and during the simulation, the callback policy is not dynamically adjusted, which also causes the speed of the instruction stream collection process to decrease.
[0086] In an embodiment of the present invention, a hybrid-mode instruction stream collection system is provided, as Figure 1 shown, including: a functional simulator 101, an instruction stream collector 102, a processor 103, and a code cache 104.
[0087] The functional simulator 101 refers to a Functional simulator that simulates the instruction set of a specified architecture, such as open-source QEMU, GEM5, and closed-source Intel SDE, and is used to simulate the target program. The instruction stream collector 102 is loaded into the functional simulator 101 in the form of a plug-in, and the instruction stream collector 102 is used to register callbacks (functions), as Figure 2B shown, and the callback functions include:
[0088] Basic block translation callback: This function is called when a basic block is translated;
[0089] Translation block execution callback: This function is called after a translation block is executed;
[0090] Instruction execution callback: This function is called when a target instruction is executed;
[0091] Memory usage callback: This function is called when an instruction uses a memory operation.
[0092] The functional simulator 101 is used to translate the basic blocks of the target program into translation blocks and store them in the code cache 104.
[0093] The processor 103 is used to register only the basic block translation callback in the instruction stream collector 102 during the initial stage of target program simulation; after storing the translated block into the code cache 104, it calls the basic block translation callback to determine whether it is in the instruction collection area. If not, it registers the translated block execution callback in the instruction stream collector 102. If so, it registers the instruction execution callback and the memory usage callback in the instruction stream collector 102; when calling the translated block execution callback, it determines whether it enters the entry point of the instruction collection area after the relevant basic block. If so, it obtains the memory information and register information of the target program based on the memory information of the simulation process and records them into the instruction stream collection file, and clears the code cache 104 of the translated block; when calling the instruction execution callback, it obtains the instruction information and records it into the instruction stream collection file, and determines whether it is at the end point of the instruction collection area after the relevant basic block. If so, it clears the code cache of the translated block.
[0094] Before implementing this embodiment, it is necessary to pre - establish the basic block translation callback, the basic block translation callback, the instruction execution callback, and the memory usage callback.
[0095] This embodiment distinguishes between the instruction collection area and the non - instruction collection area. In the non - instruction collection area, only the translated block execution callback is enabled. In the instruction collection area, the instruction execution callback and the memory usage callback are enabled, which can reduce information acquisition in the non - instruction collection area and improve the sampling efficiency. By timely clearing the translated blocks in the code cache, the present invention can re - judge whether it is in the instruction collection area, and thus adjust the enabling strategy of the callback according to the judgment result, improving the sampling efficiency.
[0096] In some embodiments of the present invention, the schematic diagram of the program operation and sampling segment is as Figure 3 shown. Figure 3 In, A - Z indicates the entire operation of a benchmark program. Since there are too many instructions, it is impossible to sample the whole process. Specifically, during implementation, according to the instruction statistics within the basic block of the program operation, representative instruction segments can be found. For example, Figure 3 in, the B - C segment and the D - E segment in are the instruction sampling areas. Use the instructions in these segments and attach the weight of each segment to represent the execution of the entire program. Usually Figure 3 the number of instructions in the irrelevant area in is much larger than the number of instructions in the sampling area. This requires collecting the instruction stream information for the B - C and D - E segments. Taking the B - C segment as an example, the following information needs to be collected:
[0097] At the initial point B of the segment, collect the complete register information at point B, register snapshot; at the initial point B of the segment, collect the current memory information of the entire program, memory snapshot; from point B to point C, the execution situation of each instruction, including the instruction address, instruction code, and the register status used by the instruction. If the instruction has a memory operation, record the memory operation information.
[0098] In an embodiment of the present invention, a method for collecting instruction streams in a hybrid mode is further provided. As Figure 4 shown, it includes:
[0099] Step 401: Only enable the basic block translation callback during the initial stage of target program simulation.
[0100] Specifically, enabling the basic block translation callback in this step means registering a basic block translation callback function in the instruction stream collector. When the basic block in the target program is translated into a translation block and the translation block is stored in the code cache, the basic block translation callback can be called. Both the basic block and the translation block contain multiple instructions, and the instructions in the basic block correspond to the instructions in the translation block.
[0101] Step 402: After storing the translation block in the code cache, call the basic block translation callback to determine whether it is in the instruction collection area. If not, enable the translation block execution callback. If so, enable the instruction execution callback and the memory usage callback.
[0102] In this step, enabling the translation block execution callback means registering the translation block execution callback in the instruction stream collector, and enabling the instruction execution callback and the memory usage callback means registering the instruction execution callback and the memory usage callback in the instruction stream collector.
[0103] When enabling the translation block execution callback, after the translation block is executed, the translation block execution callback will be called for execution.
[0104] When enabling the instruction execution callback and the memory usage callback, after the instructions in the translation block are executed, the instruction execution callback and the memory usage callback will be called for execution.
[0105] Step 403: When calling the translation block execution callback, determine whether it enters the entry point of the instruction collection area after determining the relevant basic block. If so, obtain the memory information and register information of the target program based on the memory information of the simulation process and record them in the instruction stream collection file, and clear the code cache of the translation block.
[0106] In this step, by clearing the code cache of the translation block after the entry point of the instruction collection area, it is possible to re-translate the basic block into a translation block, return to step 402 to re-determine enabling the instruction execution callback and the memory usage callback, achieving the purpose of dynamically adjusting the callback strategy.
[0107] During implementation, if it does not enter the entry point of the instruction collection area, the translation block execution callback will be enabled again after the translation block is executed.
[0108] Step 404: When calling the instruction execution callback, obtain the instruction information and record it in the instruction stream collection file, and determine whether it is at the end point of the instruction collection area after determining the relevant basic block. If so, clear the code cache of the translation block.
[0109] After the instruction is executed, the number of executed instructions is accumulated to the historical instruction accumulation value to obtain the updated instruction accumulation value, and based on the updated instruction accumulation value, it is determined whether to reach the end point of the collection area.
[0110] In this step, after reaching the end point of the instruction collection area, the code cache of the translation block is cleared, enabling the retranslation of the basic block to obtain the translation block, and returning to step 402 to re-determine the enabling of the translation block execution callback, achieving the purpose of dynamically adjusting the callback strategy.
[0111] In an embodiment of the present invention, as Figure 5 shown, the execution process inside the basic block translation callback includes:
[0112] Step 501, query the number of instructions in the translation block.
[0113] Specifically, the number of instructions in the translation block here refers to the number of instructions of the executed translation block, that is, the total number of executed instructions.
[0114] Step 502, determine whether it is in the instruction collection area based on the number of instructions in the translation block.
[0115] Specifically, after the translation block is executed, the number of instructions of the executed translation block is accumulated to the historical instruction accumulation value to obtain the total number of executed instructions (i.e., the updated historical instruction accumulation value), and based on the total number of executed instructions, it is determined whether to enter the collection area. For example, the number of instructions in the collection area is in [1,000,000 - 2,000,000], that is, after executing 1 million instructions, enter the collection area, and after executing 2 million instructions, exit the collection area. Then, when the total number of executed instructions is greater than or equal to one million, enter the collection area, and when the total number of executed instructions is equal to two million, exit the collection area.
[0116] Step 503, if not in the instruction collection area, register the translation block instruction callback function in the instruction stream collector, and use the number of instructions of the current translation block as the input parameter of the translation block instruction callback function. By passing in the number of instructions of the translation block, multiple queries can be avoided, improving the sampling efficiency.
[0117] Step 504, if in the instruction collection area, traverse each instruction in the translation block and perform the following operations:
[0118] Step 5041, obtain the instruction machine code (opcode) of the instruction.
[0119] Step 5042, check whether the instruction machine code is in the hash table. If not in the hash table, parse the instruction machine code opcode to obtain the instruction register information (the register information required for the source / destination of the instruction), and insert the instruction machine code and the instruction register information into the hash table. If in the hash table, continue to traverse the next instruction in the translation block.
[0120] Step 505, after traversing the instructions in the translation block, register the instruction execution callback and memory usage callback in the instruction stream collector, and pass the instruction machine code to the instruction callback function.
[0121] In an embodiment of the present invention, as Figure 6 shown, after the translation block is executed, the execution process of the translation block execution callback includes:
[0122] Step 601, receive the number of instructions n in the translation block.
[0123] Step 602, there is an instruction counter in the instruction stream collector. Use the instruction counter to calculate the cumulative value of the instructions. Add n instructions on the basis of the historical instruction cumulative value, that is, g_instCount += n, to obtain a new instruction cumulative value.
[0124] Among them, the historical instruction cumulative value in this step refers to the total number of executed instructions in the previous text.
[0125] Step 603, determine whether to enter the entry point of the instruction collection area after the translation block is executed according to the new instruction cumulative value.
[0126] If the entry point of the instruction collection area is not entered, directly return, and wait to continue calling the translation block execution callback after the next translation block is executed.
[0127] If the entry point of the instruction collection area is entered, execute the following process:
[0128] Step 6041, obtain the memory information of the simulated process.
[0129] Specifically, when the simulated process is loaded into memory by the operating system, the operating system will provide a memory mapping file to express the memory mapping information. In this step, the memory information can be determined through the memory mapping file.
[0130] Among them, the memory information of the simulated process usually includes multiple lines of memory information. Each line of memory information includes the memory address range, range usage permissions, offset, device number, file number and file name corresponding to the memory range loaded, etc.
[0131] Step 6042, identify the memory addresses in the memory information of the simulated process, and obtain the page table and page table attributes corresponding to the memory addresses.
[0132] Specifically, the page table is a data structure used for memory management, which is used to record the mapping relationship between virtual memory addresses and physical memory addresses. In a system using paged memory management, the virtual memory space is divided into fixed-size pages. The page table is located in the page table area of the system space, and each process has its own page table. The page table attributes include access permission attributes, memory type attributes and mapping attributes.
[0133] Step 6043: Determine whether the page table attribute contains the identifier of the target program. If so, determine that the memory address corresponding to the page table is the memory address of the target program, and record the page table and its corresponding memory address in the instruction stream collection file. If not, execute Step 6044.
[0134] Step 6044: Determine whether all the memory addresses in the memory information of the simulated process have been analyzed. If so, execute Step 6045. If not, return to Step 6041 to continue execution.
[0135] Step 6045: Obtain the register information and record it in the instruction stream collection file, and clear the code cache of the translation block. Specifically, the register information obtained in this step refers to the information of all the registers of the current CPU processor, including integer, floating-point, and system registers, etc. Specifically, this step can obtain the register information from the functional simulator.
[0136] Among them, obtaining the register information is a part of the initial state of instruction collection. Clearing the translation block code cache can prepare for subsequent retranslation and enabling the instruction execution callback.
[0137] This embodiment can distinguish the memory of the target program and the memory of the functional simulator, so as to realize the collection of the memory information at the entry point of the instruction collection area. At the same time, through the above technical solutions, the operation of the memory usage callback and the maintenance of the memory usage callback information are omitted, which can reduce the system operation pressure and improve the instruction stream collection efficiency.
[0138] In an embodiment of the present invention, as Figure 7 shown, the execution process of the instruction execution callback includes:
[0139] Step 701: Receive the incoming parameters, where the incoming parameters include the instruction machine code of the current instruction.
[0140] Step 702: Search for the register information of the instruction machine code of the current instruction in the hash table, and record the register information and the instruction machine code of the current instruction in the instruction stream collection file.
[0141] This step can avoid multiple instruction parses by searching for the register information of the instruction machine code in the hash table, and improve the acquisition efficiency of the register information of the instruction.
[0142] Step 703: Obtain the register value in the current register state of the processor, and record the register value in the instruction stream collection file.
[0143] Step 704: Check whether the current instruction crosses page tables. If so, record the additional page table information in the instruction stream collection file. The page table translation is maintained by the operating system. Whether it crosses pages is determined based on the starting address and length of the instruction, as well as the length of the current page table. If it crosses pages, a page walk is required and it is traversed level by level according to the number of levels of the multi-level page table. The specific determination process of the additional page table information can refer to the prior art and will not be elaborated here.
[0144] Step 705: Check whether the current instruction is a jump instruction.
[0145] Step 706: If it is a jump instruction, record the page table information of the destination address in the instruction stream collection file. Specifically, the page table information of the destination address is obtained through address translation.
[0146] Step 707: If it is not a jump instruction, check whether it is a system call.
[0147] Step 708: If it is a system call, compare the memory before and after the instruction call and record the comparison result in the instruction stream collection file.
[0148] Step 709: If it is not a system call, check whether there is an exception. If there is an exception, record the exception information in the instruction stream collection file. Specifically, the exceptions here include but are not limited to instructions causing overflow or page table missing, etc. After having an exception handling mechanism, come back and execute this instruction again. The status of the exception is usually saved in the program status register of the computer system. This step can determine whether there is an exception by querying the status register.
[0149] Step 710: If there is no exception, record the instruction machine code in the instruction stream collection file.
[0150] Step 711: Increment the instruction count by 1, i.e., g_instCont += 1, and check whether it is at the end point of the instruction collection area.
[0151] If it is at the end point of the instruction collection area, clear the code cache of the translation block, update the information of the next instruction collection area, and determine whether it is the last instruction collection area. If so, end the entire simulation instruction collection. If not, continue to translate the basic block and execute the Figure 4 shown process.
[0152] If it is not at the end point of the collection area, continue to execute the next instruction and repeat the above steps 601 to 711.
[0153] In an embodiment of the present invention, as Figure 8 shown, the memory usage callback execution process includes the following process:
[0154] Step 801: Obtain instruction memory information, where the instruction memory information includes: the data length of the memory access, the type of the memory access, the virtual address of the data accessed in the memory, and the physical address.
[0155] Specifically, the type of the memory access includes but is not limited to read (load), write (store), prefetch, etc.
[0156] Step 802: Record the instruction memory information in the instruction stream collection file.
[0157] Step 803: Determine whether the data access crosses pages. If it does, record additional page table information in the instruction stream collection file.
[0158] The information obtained in this embodiment can provide a data basis for the input of other projects, such as microarchitecture performance simulators, etc.
[0159] In some specific embodiments of the present invention, using the instruction stream sampling method of the hybrid mode of the above-mentioned multiple callback modes to perform instruction stream collection for SPECCPU 2017 only takes 1-2 days. Compared with the 3-4 weeks taken by the existing instruction stream sampling method (such as Intel SDE pinball) before the present invention, the sampling efficiency can be greatly improved.
[0160] Based on the same inventive concept, the present invention also provides a hybrid-mode instruction stream collection device as described in the following embodiments. Since the principle of the hybrid-mode instruction stream collection device for solving problems is similar to that of the hybrid-mode instruction stream collection method, the implementation of the hybrid-mode instruction stream collection device can refer to the hybrid-mode instruction stream collection method, and the repeated parts will not be elaborated. Specifically, as Figure 9 shown, the hybrid-mode instruction stream collection device includes:
[0161] The first configuration unit 901 is used to enable only the basic block translation callback in the initial stage of the target program simulation;
[0162] The second configuration unit 902 is used to store the translation blocks obtained by translating the basic blocks of the target program into the code cache and then call the basic block translation callback to determine whether it is in the instruction collection area. If not, enable the translation block execution callback. If so, enable the instruction execution callback and the memory usage callback;
[0163] The instruction collection area entry point monitoring unit 903 is used to determine whether the relevant basic block enters the entry point of the instruction collection area when calling the translation block execution callback. If so, obtain the memory information and register information of the target program based on the memory information of the simulation process and record them in the instruction stream collection file, and clear the code cache of the translation block;
[0164] The instruction acquisition area end point monitoring unit 904 is used to call the instruction execution callback to obtain instruction information and record it in the instruction stream acquisition file, and determine whether it is at the end point of the instruction acquisition area after determining the relevant basic block. If so, the code cache of the translation block is cleared.
[0165] In specific implementation, the execution processes of the basic block translation callback, the translation block execution callback, the instruction execution callback, and the memory usage callback can respectively refer to Figures 5 to 8 the embodiments shown, and this embodiment will not be elaborated herein.
[0166] This embodiment distinguishes between the instruction acquisition area and the non-instruction acquisition area. In the non-instruction acquisition area, only the translation block execution callback is enabled, and in the instruction acquisition area, the instruction execution callback and the memory usage callback are enabled, which can reduce information acquisition in the non-instruction acquisition area and improve the sampling efficiency. By timely clearing the translation blocks in the code cache, the present invention can re-execute the judgment of whether it is in the instruction acquisition area, thereby adjusting the enabling strategy of the callback according to the judgment result and improving the sampling efficiency.
[0167] In an embodiment of the present invention, a computer device is further provided, as Figure 10 shown. The computer device 1002 may include one or more processors 1004, such as one or more central processing units (CPUs), and each processing unit may implement one or more hardware threads. The computer device 1002 may also include any memory 1006 for storing any kind of information such as code, settings, data, etc. In specific implementation, the method described in any of the above embodiments is stored in the memory 1006. Non-limitingly, for example, the memory 1006 may include any one or a combination of the following: any type of RAM, any type of ROM, flash memory devices, hard disks, optical discs, etc. More generally, any memory can store information using any technology. Further, any memory can provide volatile or non-volatile retention of information. Further, any memory can represent a fixed or removable component of the computer device 1002. In one case, when the processor 1004 executes the associated instructions stored in any memory or combination of memories, the computer device 1002 can perform any operation of the associated instructions. The computer device 1002 also includes one or more drive mechanisms 1008 for interacting with any memory, such as a hard disk drive mechanism, an optical disc drive mechanism, etc.
[0168] The computer device 1002 may further include an input / output module 1010 (I / O) for receiving various inputs (via the input device 1012) and for providing various outputs (via the output device 1014). A specific output mechanism may include a presentation device 1016 and an associated graphical user interface 1018 (GUI). In other embodiments, the input / output module 1010 (I / O), the input device 1012, and the output device 1014 may not be included, and it may only be a computer device in the network. The computer device 1002 may further include one or more network interfaces 1020 for exchanging data with other devices via one or more communication links 1022. One or more communication buses 1024 couple the components described above together.
[0169] The communication link 1022 may be implemented in any manner, for example, through a local area network, a wide area network (e.g., the Internet), a point-to-point connection, etc., or any combination thereof. The communication link 1022 may include any combination of hardwired links, wireless links, routers, gateway functions, name servers, etc. governed by any protocol or combination of protocols.
[0170] An embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it executes the steps of the above method.
[0171] An embodiment of the present invention also provides a computer-readable instruction. When the processor executes the instruction, the program therein causes the processor to execute the method described in any of the foregoing embodiments.
[0172] An embodiment of the present invention also provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor of a computer device, it implements the method described in any of the foregoing embodiments.
[0173] It should be understood that in various embodiments of the present invention, the magnitudes of the serial numbers of the above processes do not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0174] It should also be understood that in the embodiments of the present invention, the term "and / or" is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, in the present invention, the character " / " generally represents an "or" relationship between the associated objects before and after.
[0175] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in the present invention can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0176] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0177] In several embodiments provided by the present invention, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed couplings, direct couplings, or communication connections to each other can be indirect couplings or communication connections through some interfaces, devices, or units, and can also be in electrical, mechanical, or other forms of connection.
[0178] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of the embodiments of the present invention.
[0179] In addition, the functional units in each embodiment of the present invention can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0180] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs.
[0181] Specific embodiments are applied in the present invention to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A mixed mode instruction stream acquisition method, characterized in that: include: Only basic block translation callbacks are enabled during the initial stage of target program simulation; After storing the translation block obtained by translating the basic block of the target program into the code cache, the basic block translation callback is called to determine whether it is located in the instruction acquisition area, if not, the translation block execution callback is enabled, and if so, the instruction execution callback and the memory usage callback are enabled; When calling the translation block to execute the callback, determine whether the relevant basic block enters the entry point of the instruction collection area, if so, obtain the memory information and register information of the target program based on the memory information of the simulation process and record them in the instruction stream collection file, and clear the code cache of the translation block; When the instruction execution callback is called, the instruction information is obtained and recorded in the instruction stream acquisition file, and it is determined whether the relevant basic block is at the end point of the instruction acquisition area, and if so, the code cache of the translation block is cleared; When an instruction uses memory operations, the memory usage callback is called.
2. The method according to claim 1, characterized in that The memory information and register information of the target program are obtained based on the memory information of the simulation process, including: Get the memory information of the simulation process; Identify a memory address in the memory information, and obtain a page table and page table attributes corresponding to the memory address; Determine whether the page table attribute contains the identifier of the target program, and if so, determine that the memory address corresponding to the page table is the memory address of the target program, and record the page table and its corresponding memory address in the instruction stream acquisition file; Get register information and record it in the instruction stream capture file.
3. The method according to claim 1, characterized in that Before enabling instruction execution callback and memory usage callback, also include: Traversing the instruction machine code of each instruction in the translation block, and performing the following operations on the instruction machine code of each instruction: checking whether the instruction machine code is in the hash table, if not, parsing the instruction machine code to obtain instruction register information, and inserting the instruction machine code and instruction register information into the hash table, if yes, continuing to traverse the next instruction in the translation block; After all instructions in the translation block have been traversed, the instruction execution callback and the memory usage callback are enabled.
4. The method according to claim 3, characterized in that Before the basic block translation callback determines whether to enter the instruction collection area, it also includes: querying the number of instructions in the translation block; When enabling the translation block to execute the callback, the number of instructions in the translation block is also passed into the translation block execution callback; When the instruction execution callback is enabled, the instruction machine code of each instruction in the translation block is also passed into the instruction execution callback.
5. The method according to claim 3, characterized in that The command information includes: Searching the register information of the instruction machine code of the current instruction in the hash table, and recording the register information and the instruction machine code of the current instruction in the instruction stream acquisition file; Acquire a register value in the current register state of the processor, and record the register value in the instruction stream acquisition file; Check whether the current instruction crosses the page table, and if so, record the additional page table information in the instruction stream acquisition file; Check whether the current instruction is a jump instruction and whether it is a system call; If it is a jump instruction, the page table information of the destination address is recorded in the instruction stream acquisition file; If it is not a jump instruction but a system call, the memory before and after the instruction call is compared, and the comparison result is recorded in the instruction stream acquisition file; If it is not called by the system, check whether there is an exception. If there is an exception, record the exception information in the instruction stream collection file.
6. The method according to claim 1, characterized in that The memory usage callback execution process includes: Obtaining instruction memory information, wherein the instruction memory information includes: data length, type, data virtual address and logistics address of memory access; Recording the instruction memory information in the instruction stream acquisition file; Determine whether the data access crosses pages. If so, record additional page table information in the instruction stream acquisition file.
7. The method according to claim 1, characterized in that The translation block executes the callback to determine whether to enter the entry point of the instruction collection area after the relevant basic block enters, including: Adding the number of instructions in the translation block to the historical instruction accumulation value to obtain a new instruction accumulation value; Determine whether to enter the instruction collection area based on the new instruction accumulation value.
8. A mixed mode instruction stream acquisition device, characterized in that: The instruction stream acquisition device comprises: A first configuration unit is used to enable only basic block translation callbacks in an initial stage of target program simulation; A second configuration unit is used to store the translation block obtained by translating the basic block of the target program into the code cache and then call the basic block translation callback to determine whether it is located in the instruction acquisition area, if not, enable the translation block execution callback, and if so, enable the instruction execution callback and the memory usage callback; An instruction collection area entry point monitoring unit is used to determine whether the relevant basic block enters the entry point of the instruction collection area after calling the translation block execution callback, and if so, obtain the memory information and register information of the target program based on the memory information of the simulation process and record them in the instruction stream collection file, and clear the code cache of the translation block; An instruction collection area end point monitoring unit is used to obtain instruction information and record it to an instruction stream collection file when calling the instruction execution callback, determine whether the relevant basic block is at the end point of the instruction collection area, and if so, clear the code cache of the translation block; When an instruction uses memory operations, the memory usage callback is called.
9. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.
10. A computer storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor of a computer device, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Data acquisition method and device, and computer readable storage medium
CN115145806A
Instruction processing method and device
CN117742791A