Method and apparatus for performing region formation of a dynamic binary translation processor
By introducing a region formation manager and delay formation strategy into the dynamic binary conversion processor, the problem that the existing technology is difficult to capture the internal loop in the outer loop is solved, and more efficient dynamic binary conversion and performance improvement is achieved.
Patent Information
- Application Number
- CN201810995122.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2017-09-29
- Filing Date
- 2018-08-29
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2038-08-29
AI Technical Summary
The prior art is difficult to effectively capture multiple internal loops in the outer loop in the dynamic binary conversion processor, resulting in the inability to achieve optimization such as loop expansion and parallelization, which affects performance improvement.
By introducing a region formation manager in the dynamic binary conversion processor, a delay region formation strategy is adopted, and multiple hot code blocks are accumulated and area formation is triggered, ensuring sufficient analysis information is required to capture the inner loop in the outer loop, and the final area is optimized through technologies such as region expansion and loop selection.
The ability to capture loop nesting as a single unit is achieved more comprehensively, improving the code performance after dynamic binary conversion, specifically manifested in improving loop coverage and processing speed.
Smart Images

Figure CN109582444B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates generally to dynamic binary translation processors, and more particularly, to methods and apparatus for performing region formation in a dynamic binary translation processor. Background Art
[0002] Binary translation is the process of translating / converting code into a functionally equivalent version. Many types of binary translation also include attempts to optimize the translated code in order to achieve improved performance when the translated code is executed. Binary translation includes selecting multiple portions of the code to be translated as a single unit. These portions are collectively referred to as "regions". The improvement in performance depends at least in part on the manner in which the multiple portions of the code are selected to be included in the regions to be translated. BRIEF DESCRIPTION OF THE DRAWINGS
[0003] Figure 1 is a block diagram of an example implementation of a first example computing system having an example first dynamic binary translation system including an example region formation manager.
[0004] Figure 2 yes Figure 1 Block diagram of an example implementation of an example region formation manager.
[0005] Figure 3 The diagram shows Figure 2 An example control flow graph of the operation of an example region formation manager.
[0006] Figure 4 is a block diagram of an example implementation of a second example computing system having an example second dynamic binary translation system including an example super region formation manager.
[0007] Figure 5 yes Figure 4 A block diagram of an example implementation of an example super-region formation manager.
[0008] Figure 6 The diagram shows Figure 1 A table of example contents of an example queue buffer included in the super-region formation manager.
[0009] Figure 7 Is Figure 1 An example of a control flow graph generated by the super region formation manager.
[0010] Figure 8 yes Figure 5 Block diagram of an example merge candidate selector and checker for an example super region formation manager.
[0011] Fig. 9It means that it can be executed to achieve Figure 2 A flow chart of machine readable instructions for an example region formation manager.
[0012] Fig.10 It means that it can be executed to achieve Figure 2 A flow diagram of machine readable instructions for an example first execution profiler and an example queue manager.
[0013] Fig.11 It means that it can be executed to achieve Figure 2 A flow chart of machine readable instructions for an example initializer of an example region formation manager.
[0014] Fig.12 It means that it can be executed to achieve Figure 2 A flowchart of machine readable instructions for an example initial region former.
[0015] Fig.13 It means that it can be executed to achieve Figure 2 A flow chart of machine readable instructions for an example area expander.
[0016] Fig.14 It means that it can be executed to achieve Figure 2 A flow chart of machine readable instructions for an example region analyzer, an example loop selector, and an example trimmer.
[0017] Fig.15 It means that it can be executed to achieve Figure 5 A flow diagram of machine readable instructions for an example super region formation manager.
[0018] Fig.16 It means that it can be executed to achieve Figure 5 A flowchart of machine readable instructions for an example merge candidate selector and checker, an example entry point selector, and an example super region creator.
[0019] Fig.17 is constructed to execute Figure 9-14 Instructions to achieve Figure 1 and Figure 2 Block diagram of an example processing platform of an example region formation manager.
[0020] Fig.18 is constructed to execute Figure 15-16 Instructions to achieve Figure 4 and Figure 5 Block diagram of an example processing platform of an example super-region formation manager.
[0021] The drawings are not drawn to scale. Wherever possible, the same reference numbers will be used throughout the drawings and accompanying written description to refer to the same or like parts. DETAILED DESCRIPTION
[0022] A dynamic binary converter is used to convert code (machine-readable instructions) from a first instruction set (also referred to as a "native instruction set") to a second instruction set (also referred to as a "target instruction set") for execution on a processor or hardware accelerator. In some examples, the dynamic binary converter not only converts the code but also attempts to optimize the converted code. Optimization refers to the process of trying to fine-tune the code during the conversion process in a way that the converted code is executed more efficiently than the original native (unconverted) code. Some dynamic binary converters perform code conversion by converting a sequence of codes (also referred to as a block of instructions or instruction blocks) and then performing a conversion of the code. The conversion of the code is referred to as "conversion" in this article. The result of performing the code conversion is cached and used when the code block is to be re-executed. Some dynamic binary converters are capable of performing an optimization process, selecting an area of the code to be optimized. For example, a sequence / block of instructions containing a loop is typically converted to be executed as a single unit, which typically reduces the number of clock cycles required to execute the code.
[0023] Region formation is the process of selecting blocks / sequences of code instructions to be converted into a single unit to form a single conversion. The performance benefit obtained by executing the converted code (relative to the unconverted code) depends on the degree to which the converted code is optimized, which in turn depends at least in part on the region formation. Some conventional methods of performing region formation focus on collecting code blocks that are equally hot or almost equally hot into a single region, while excluding cold code blocks. Hot code blocks (also called "hot code") are code blocks that are executed for a long time relative to other code blocks. Code that forms a loop that is usually executed iteratively is usually hot code. Conventional methods of such region formation also focus on limiting the size of the region. Another conventional method emphasizes the formation of regions including "super blocks". Super blocks include code blocks that have a single entry point and multiple exit points when represented as a control flow graph. Another conventional method for performing region formation involves constructing regions using observed execution traces. However, none of the above methods attempts to ensure that multiple inner loops of an outer loop (e.g., loop nesting) within the code are captured in a single region rather than being segmented into multiple regions. Unfortunately, when region formation results in a loop (or loop nest) being segmented into multiple regions, optimizations such as loop unrolling and / or parallelization are not possible. As a result, the above methods for performing region formation generally reduce the performance benefits that would be achieved if such a loop (and / or loop nest) were captured in a single region formation.
[0024] Figure 11 is a block diagram of an example implementation of a first example computing system 100 having an example first dynamic binary translation system 102. The first dynamic binary translation system 102 includes an example region formation manager 104 and an example dynamic binary code converter 106. Figure 1 In the example of , the example native code memory 108 stores a computer program to be translated by the first dynamic binary translation system 102. The dynamic binary code converter 106 of the first dynamic binary translation system 102 converts the computer program from a native instruction set to a target instruction set using any dynamic binary translation technology. In some examples, the first dynamic binary translation system 102 converts the code in a manner that includes performing any of a variety of optimization techniques such as parallelization and loop unrolling. When executed by the example processor 110, the translated computer program and the untranslated computer program have equivalent functionality.
[0025] In some examples, the hardware of the example processor 110 is designed to execute a target instruction set, and the processor 110 uses an example software emulator 112 to execute a native instruction set. In some examples, the processor 110 begins executing native code stored in the example native code memory 108 while running the example software emulator 112. The example region formation manager 104 monitors the execution of the native code by the software emulator 112 of the processor 110 and identifies instruction sets / blocks to be placed in regions. The instruction sets / blocks to be placed in regions are provided by the region formation manager 104 to the example dynamic binary translator 106 for conversion as a single unit (so as to produce a single conversion). The dynamic binary translator 106 stores the converted code in the example target code memory 114 for execution by the processor 110. As described above, the code conversion performed by the first dynamic binary translation system includes optimizations that improve the processing speed and efficiency achieved when the target code is executed by the processor 110.
[0026] Figure 2 yes Figure 1 2. Block diagram of an example implementation of an example region formation manager 104. In some examples, the region formation manager 104 includes an example region memory 202, an example first execution profiler 204, an example queue manager 206, an example queue 208, an example initializer 210, an example initial region former 212, an example region expander 214, an example region analyzer 216, an example loop feature memory 218, an example loop selector 220, an example criteria memory 222, and an example region pruner 224.
[0027] In some examples, the example software emulator 112 of the example processor 110 executes native code stored in the native code memory 108 (see Figure 1). When the native code is being executed, the example first execution profiler 204 monitors the execution of the native code. When the first execution profiler 204 determines that the instruction or instruction / code block is "hot" (e.g., takes a long time to execute, is executed most frequently, etc.), the first execution profiler 204 notifies the example queue manager 206. The queue manager 206 responds by adding information that can be used to identify the hot code block to the example queue 208. In some examples, the information identifying the hot code block includes the IP address of the start instruction of the hot code block. In some examples, the queue manager 206 also places context information about the hot code block (such as a thread identifier, an execution count, a branch identifier that identifies the branch statement that leads to the hot code block, etc.) in the queue 208. In some examples, the queue manager 206 only adds the hot code block to the queue that is the target of the backward branch (e.g., the IP address of the branch statement leading to the hot code block is higher than the IP address of the hot code block).
[0028] In some examples, aspects of the example queue manager 206 and / or the example queue 208 are implemented as one or more hardware structures or in-memory structures. In some examples, the first execution profiler 204 is implemented as a hardware-based profiler and an interrupt is triggered when a code block is determined to be hot. In some such examples, the queue manager 206 is implemented in software (e.g., as an interrupt service routine) and responds to the interrupt by adding information identifying the hot code block to the queue 208. In some examples, the first execution profiler 204 and the queue manager 206 are implemented as software-based profilers, and the queue 208 is implemented by hardware. In some such examples, the execution of one or more special instructions causes the hot code block to be added to the queue 208. In some examples, the first execution profiler 204, the queue manager 206, and the queue 208 are all implemented in hardware. In some examples, the first execution profiler 204, the queue manager 206, and the queue 208 are all implemented in software without any special instructions.
[0029] In some examples, the example queue 208 is implemented as a least recently used (LRU) queue with limited capacity that discards the least recently used hot code blocks first (when capacity is reached). In some examples, when a hot code block that has been identified by the first execution profiler 204 to be added to the queue 208 is already resident in the queue 208, the queue manager 206 causes the hot code block to be moved to the head of the queue 208 and no new hot code blocks are added to the queue 208. In this way, the identified blocks are considered to be recently used and are not discarded before other less used blocks are discarded.
[0030] Still see Figure 2, the example initializer 210 monitors the length of the example queue 208. When the length of the queue 208 reaches a threshold queue length value, the initializer 210 generates a trigger at the example initial region former 212. In some examples, the threshold queue length value that will cause the trigger to be generated is greater than 1. Therefore, unlike conventional region formation methods that immediately begin forming regions based on the identification of a single hot code block, the region formation manager 104 operates after a threshold number of hot code blocks have been detected. Delaying region formation until a threshold number of hot code blocks are added to the queue 208 ensures that profiling information about a relatively large portion of the native unconverted code can be collected. Since loop nests tend to span a larger number of code blocks than just the innermost loop, obtaining profiling information about a larger number of code blocks results in a richer data set, which typically includes information about loop branches. Obtaining profiling information about a larger number of code blocks also reduces the likelihood of unnecessarily truncating region growth due to a lack of sufficient amount of profiling information.
[0031] Still see Figure 2 , in response to a trigger generated by the example initializer 210, the example initial region former 212 begins to form the initial region. In some examples, the initial region former 212 begins to form the initial region by selecting the first hot code block to be included in the initial region from the queue 208. As described above, each hot code block may include one or more native code instructions. In some examples, the selected first hot code block corresponds to the last hot code block added to the queue 208. Then, the initial region former 212 grows the initial region by generating a control flow graph of the code being executed. In some examples, the initial region former 212 generates a control flow graph by acquiring and decoding the native instructions included in the hot code block (e.g., performing a static decoding technique). In some examples, the initial region former 212 determines that the first hot code block leads to one or more other code blocks. In some such examples, the initial region former 212 determines which hot code block is the hottest hot code block, and then adds the hottest hot code block to the initial region. In some examples, the initial region former 212 determines which hot code block is the hottest based on the profiling information provided by the first execution profiler 204. The initial region former 212 continues to grow the initial region in this manner (e.g., adding blocks placed along the hottest path) until the initial region former 212 determines that other code blocks are not hot enough to be added to the initial region, or until the path leads to the next hot code block previously added to the initial region.
[0032] Figure 3is an example control flow graph 300 and includes example code blocks (block A, block B, block C, block D, block E, block F, block G) connected by example paths (path 1, path 2, path 3, path 4, path 5, path 6, path 7, path 8, path 9, and path 10). The blocks (blocks AG) represent blocks of instructions included in the code to be converted, and the paths (paths 1-10) represent the direction of flow between the code blocks (blocks AG) when executed by the example processor 110 (see Figure 1 ). In some examples, blocks AG may be identified using the IP address corresponding to the first instruction of each such block. In the control flow graph 300, block A is connected to block B via path 1, indicating that after executing the instructions of block A, the processor 110 continues to execute the instructions of block B. Block B is coupled to block C via path 2, and is also coupled to block E via path 4. Therefore, after executing the instructions of block B, the processor 110 branches to execute the instructions of block C or block E. Block C is also connected back to block B via path 3, and is also connected to block D via path 5. Therefore, after the processor 110 executes the instructions of block C, the processor branches to execute the instructions of block B or block D. In addition to block C, block E is also coupled to block D via path 6. Block D loops back to itself via path 8. As a result, the instructions of block D can be executed after any instructions of block E, block C, or block D. In addition, block D is connected to block F via path 7. After executing the instructions of block D, the processor 110 executes the instructions of block F or executes the instructions of block D again. Block G is connected to itself via path 9 and is also connected to block A via path 10, so that after executing instructions of block G, processor 110 re-executes instructions of block G, or executes instructions of block A.
[0033] exist Figure 3 In FIG. 1 , block B and block F are shaded to indicate that both block B and block F are hot code blocks. As described above, hot code is used to indicate code that is executed in more clock cycles than other code blocks, or is frequently re-executed compared to other code blocks. In contrast, cold code blocks (blocks A, C, D, E, and G) may be executed infrequently and / or use relatively few clock cycles.
[0034] Reference now Figure 2 and Figure 3, when block B is the last code block added to the example queue 208, the example initial region former 212, in response to the trigger, adds the example block B to the initial region. As the last code block added to the queue 208, block B is the most recently executed code block included in the queue 208. In addition, the initial region former 212 determines that block B leads to example blocks C and example E. In this example, the initial region former 212 determines that the hottest path from block B leads to block C and proceeds to add block C to the initial region. In addition, the initial region former 212 determines that block D is not hot enough to be added to the initial region, and therefore terminates the initial region at block C. In some examples, the initial region former 212 stores identification information identifying the initial regions (e.g., block B and block C) in the region memory 202. Using a static decoding method to support region formation rather than relying on observed traces (as is often done in conventional region formation methods), the region formation manager 212 disclosed herein is able to explore cooler code paths that may not have been executed in the window during which the trace was collected.
[0035] See also Figure 2 , after forming the initial region, the example initial region former 212 notifies the example region expander 214 that the initial region is ready to be expanded. The region expander 214 responds by expanding the initial region to form an extended region including the initial region. In some examples, the region expander 214 accesses the region memory 202 to obtain the initial region (e.g., to determine the code blocks included in the initial region and the paths connecting the code blocks). The region expander 214 analyzes a set of exits of the initial region and determines which exit is the hottest exit. Then, the region expander 214 evaluates the code blocks reachable via the hottest exit, and based on the result of the evaluation, adds the path including the code blocks reachable via the hottest exit to the extended region. In some examples, the evaluation includes determining whether the path extended from the evaluated exit: 1) reaches a hot code block that is already connected to the initial region (or reaches the extended region if the initial region has been expanded), or 2) reaches a hot code block within a threshold number of code blocks (even if the code block between the exit and the hot code block is a cold code block). When performing this evaluation, the region expander 214 ignores any back edges (e.g., Figure 3 Block D has a back edge). Ignoring the back edge helps the region formation manager 104 capture the outer loop when there are multiple inner loops in the body of the outer loop but only some of the code blocks containing the inner loop headers are present in the queue 208. See also Figure 3, the example control graph 300 contains three loop heads, example block B, example block D, and example block F, but only block B and block F are present in the queue 208. As a result, if the back edges are followed (not ignored), there will be no non-loop path starting from block C and leading to block F, so that expansion along path 5 will not be possible.
[0036] If the evaluated path meets any of the two criteria provided above, the example region expander 214 adds the code block located on the path starting from the hottest exit to the expanded region. Figure 3 In the example control flow diagram 300, when determining whether the path extended from the hottest exit (e.g., block C) meets the above criteria, the path goes from block C to block D, which is a cold code block. However, block D reaches block F, which is a hot code block. Assuming that the threshold (e.g., maximum) number of blocks that can be located in the path between the exit (e.g., block C) and the hot code block (e.g., block F) is two blocks, the path from block C to block F meets the evaluation criteria, so that the path extended from the hottest exit (block C) will be added to the initial region by the region expander 214, thereby forming an extended region.
[0037] See also Figure 2 and Figure 3 , the example region expander 214 next determines which code blocks on the hottest path starting from the hottest exit (e.g., block C) to add to the initial region. In some examples, the region expander 214 adds code blocks along the hottest path extending from the exit (e.g., example block C) until a threshold number of code blocks have been added or until blocks associated with back edges have been added. Figure 3 When looking at the example control flow graph 300 of FIG. 1 , the region expander 214 starts at the block with the hottest exit (block C) and continues to add the next block (e.g., example block D) until a threshold number of blocks is reached or a block with a back edge is reached. After adding block D, the region expander 214 terminates the path because block D is associated with a back edge. Therefore, even if a path extending from block C and including block D is added to the extended region because it can reach example block F from block C, the added path is limited to block D because block D has a back edge.
[0038] After adding the example block D to the initial region to form the extended region, the example region expander 214 continues to grow the extended region by iteratively selecting the hottest exit from all exits of the extended region, determining whether to add a path associated with the hottest exit, and then adding the path based on the determination. When the last hottest exit has been evaluated, the region expander 214 completes the formation of the extended region. In some examples, the region expander 214 stores information identifying the code blocks included in the extended region and any other desired information in the example region memory 202.
[0039] See also Figure 2 , after the example region expander 214 completes the growth of the expansion region, the example region analyzer 216 analyzes the example expansion region stored in the example region memory 202. In some examples, the analysis performed by the region analyzer 216 includes identifying all loops (and loop nests) included in the expansion region and determining the characteristics (interestingness metrics) of each loop. In some examples, the region analyzer 216 stores information identifying loops and loop characteristics in the example feature memory 218. In some examples, the characteristics may include a number of instructions included in the loop nest, information identifying code blocks included in the loop nest, the depth of the loop nest, whether the loop nest can be shrunk, etc. In some examples, the region analyzer 216 obtains the characteristics from the example first execution profiler 204.
[0040] When loop nest features are stored in the example feature memory 218, the example loop selector 220 selects one of the loop nests based on the stored features. In some examples, the loop selector 210 uses the loop nest criteria stored in the criteria memory 222 to select one of the loop nests. In some such examples, the loop selector 220 examines the loop nest features stored in the loop feature memory 218 and determines which loop nest meets the criteria. In some examples, the criteria specify 1) the selected loop nest includes less than a threshold number of instructions and / or blocks, 2) the selected loop nest has a nesting depth less than or equal to a threshold nesting depth, and / or 3) the selected loop nest is reducible, etc. In some examples, the execution region is formed so that the processing of the code block included in the region can be offloaded to the hardware accelerator. In some such examples, the criteria may include any constraints imposed by the architecture of the hardware accelerator. In some examples, the architecture of the hardware accelerator may not be able to handle loops with more than a threshold number of certain types of instructions. In some such examples, a restriction on the number of such instructions may be added as a criterion to be met when selecting a loop nest. In some examples, loop selector 220 selects the largest loop nest that meets the criteria.
[0041] After the example loop selector 220 selects one of the loop nests, the loop selector 220 notifies the example region pruner 224 of the identity of the selected loop nest. The region pruner 224 uses the identity of the selected loop nest to prune all other loops / loop nests from the expansion region. In some examples, the region pruner 224 also performs queue cleaning actions. Example queue cleaning actions may include: 1) removing all hot code blocks that are part of the selected loop nest from the example queue 208, 2) clearing all hot code blocks from the queue 208, regardless of whether the hot code blocks are included in the selected loop nest, and so on. In some examples, the region pruner 224 does not perform any queue cleaning actions, so that the hot code blocks contained in the queue 208 after the region formation process remain in the queue 208.
[0042] After pruning is complete, the selected loop nest is the only remaining portion of the expanded region and becomes the final region. In some examples, the example region pruner 224 causes information identifying code blocks included in the final region to be stored in the region memory 202. In addition, the region pruner 224 provides the information identifying the final region to the example dynamic binary translator 102 (see Figure 1 ). The dynamic binary translator 102 converts the code included in the final region into a single unit. The resulting converted code is referred to herein as a "translation". In addition, the dynamic binary translator 102 causes the translation to be stored in the example target code memory 114 (see Figure 1 ) for access and execution by processor 110 and / or hardware accelerator.
[0043] The example region formation manager 104 disclosed herein provides several advantages over conventional region formation tools, including the ability to more fully capture loop nests for conversion as a single unit. To illustrate one or more advantages obtained when using the region formation manager 104 disclosed herein over conventional methods, a strawman region former is modeled and loop-related metrics of loop coverage and high-resident loop coverage are obtained. The workload used includes a mixture of benchmarks from a suite representing a mixture of client and server workloads. In addition, the region formation manager 104 is modeled and similar loop metrics are obtained. For both models, metrics are obtained for measuring the ability of the region former to capture loops (called "loop coverage"). Loop coverage is defined as the ratio of 1) dynamic instructions from loops and loop nests captured in a region to 2) the total number of dynamic instructions executed. For example, if in a workload of 10 million instructions, the converted region constitutes 5 million instructions, and 3 million of the 5 million instructions are from loops, then the loop coverage is 3 / 10 (30%). Comparison of loop coverage obtained using the strawman region former relative to the region former manager 104 shows a 20% improvement obtained by using the region former manager 104 disclosed herein. The improved coverage is obtained due to the region former manager's ability to: 1) accumulate multiple hot blocks in a queue (or pool) and trigger the region former process / method after the queue / pool is filled, rather than triggering the region former process immediately after each hot block is detected; 2) use the region expander to perform a secondary growing phase (in addition to the main growing phase performed by the initial region former) to identify and grow cold paths (assuming the identified cold paths lead to hot code blocks), and / or 3) select a loop nest having characteristics that are superior to other loop nests and prune the other loop nests from the final region.
[0044] Figure 4 is a block diagram of an example implementation of a second example computing system 400 having an example second dynamic binary translation system 402, including an example super region formation manager 404 in addition to the example region formation manager 104 and the example dynamic binary code converter 106. Figure 4 In the example of , the computer system 400 also includes an example native code memory 108, an example processor 110 that can execute a software emulator 112, an example target code memory 114, an example compiler 406, an example conversion result and metadata memory 408, and an example second execution profiler 410. In some examples, the target code stored in the target code memory 114 includes a converted native code block (also referred to as a conversion). Figure 1-Figure 3As described, the conversion is generated by the dynamic binary code converter 106, and at least some of the conversions are generated based on the regions formed by the region formation manager 104. In some examples, the processor 110 causes the compiler 406 to compile the conversion. The compiler 406 causes the compiled conversion to be placed in the conversion result and metadata storage 408. Then, the processor 110 obtains the conversion result from the conversion result and metadata storage 408 for execution. In some examples, the compiler 406 and / or the second execution profiler 410 generates conversion metadata corresponding to the conversion when executing the conversion. The conversion metadata is stored together with the corresponding conversion in the conversion result and metadata storage 408. In some examples, the conversion metadata may indicate the number of times the corresponding conversion is executed, the list of conversions to exit to the conversion, and the list of conversions for the conversion to exit, etc.
[0045] Figure 5 yes Figure 4 4 is a block diagram of an example implementation of an example super region formation manager 404. In some examples, the super region formation manager 404 includes an example transformation queue 502, an example transformation queue sampler 504, an example policy memory 506, an example transformation queue buffer 508, an example graph generator 510, an example merge candidate selector and checker 512, an example entry point selector 514, and an example super region creator 516.
[0046] In some examples, the example conversion queue 502 implemented using a hardware, circular buffer HERE queue monitors the example processor 110 (see Figure 1 ) execution of a conversion. In some such examples, conversion queue 502 periodically interrupts execution of a conversion by processor 110 and collects information about the conversion being executed at the time of the interruption (referred to as the "current conversion"). In some examples, the collected information includes information identifying the current conversion and further identifying the number of times the current conversion has been sampled by conversion queue 502. In some examples, the size of conversion queue 502 is configurable.
[0047] Still see Figure 5, the example conversion queue sampler 504 periodically reads the entire contents of the conversion queue and records the sampling information contained therein as an array in the example conversion queue buffer 508. As a result, the conversion queue buffer 508 includes a snapshot of the information contained in the example conversion queue 502. In some examples, the conversion queue buffer 508 includes M×N entries, where M is the number of entries in the conversion queue 502 and N is the number of snapshots saved for each conversion included in the conversion queue 502. In some examples, when a sampling interruption occurs and the conversion queue 502 determines that the sample count of the current conversion meets the sample count threshold, the current conversion is eligible for merging with other conversions. In some examples, the sample count threshold is statically defined, and in some examples, the sample count threshold can be dynamically adjusted based on the most recently executed conversion and / or the policy stored in the example policy memory 506. The first conversion eligible for merging is the conversion with a sample count exceeding the sample count threshold, and is also referred to as a "seed" in this article.
[0048] In some examples, when the example transformation queue buffer 508 has identified a seed (e.g., the current transformation sampled satisfies a sample count threshold), the example graph generator 510 builds a transformation graph based on the contents of the transformation queue buffer 508. In some examples, the transformation graph includes a set of nodes coupled by edges. The nodes of the graph represent transformations, and an edge connecting two nodes (e.g., node A and node B) indicates that one of the corresponding transformations can be reached through the other corresponding transformation (e.g., when executed, the transformation of node A at least sometimes branches to the transformation of node B). In some examples, the graph generator 510 begins building the graph by representing the seed as the first node of the graph, finding an entry in the transformation queue buffer 508 that contains the seed, and using the entry to determine edges between the seed and other transformations.
[0049] Figure 6 Included are example contents of an example conversion queue buffer 508. Figure 6 In the example of , the example contents of the conversion queue buffer 508 include four example entries (example first entry 602, example second entry 604, example third entry 606, and example fourth entry 608). The entries include information about the execution order of the example conversions (example conversion 10, example conversion 11, example conversion 12, example conversion 13, example conversion 14, and example conversion 15).
[0050] refer to Figure 6 and Figure 7 To prepare the graph, the example graph generator 510 represents the seed as the first node of the graph. Figure 6In the illustrated example, transition 10 corresponds to the seed, causing the graph generator 510 to generate a node 10 corresponding to transition 10. Next, the graph generator 510 examines the contents of the transition queue buffer 508 to identify the entry containing the seed. Thus, the graph generator 510 identifies example entry 1 602, example entry 2 604, example entry 3 606, and example entry 4 608. Next, the graph generator 510 examines the entries (e.g., entry 1 602, entry 2 604, entry 3 606, entry 4 608) to determine which transitions (e.g., transition 10) were performed before the seed. Figure 6 In the example of , transition 11 is executed before transition 10, as reflected in three entries (entry 1 602, entry 3 606, and entry 4 608). To capture this information, graph generator 510 adds an example node 11 representing transition 11 to graph 700, and connects node 11 to node 10 using an example first edge 702. Due to the queue buffer contents (see Figure 6 ) indicates that transition 11 is executed before transition 10, so first edge 702 is assigned a weight of 3. Graph generator 510 also determines based on the entries in transition queue buffer 508 that transition 15 is executed before transition 10, as reflected in one entry (entry 2 604) of the transition queue buffer contents (see Figure 6 ). To capture this information, graph generator 510 adds example node 15 to the graph and connects node 15 to node 10 using an example second edge 704 having a weight of 1. The contents of queue buffer 600 assigned a weight of 1 indicate that transition 15 is executed once before transition 10.
[0051] Next, the example graph generator 510 uses the contents of the example transformation queue buffer 508 to identify the transformations to be performed after the seed. Figure 6 As illustrated, example transition 12 is performed after example transition 10 in three of the entries (example entry 1 602, example entry 2 604, example entry 3 606) and example transition 13 is performed after transition 10 in one of the entries (e.g., entry 4 608). Graph generator 510 captures this information by adding example node 12 representing transition 12 to graph 700 and coupling node 10 to node 12 with example third edge 706 assigned weight 3. Additionally, graph generator 510 adds example node 13 and couples node 10 to node 13 with example fourth edge 708 assigned weight 1.
[0052] After adding node 12 and node 13, example graph generator 510 queries the contents of example transformation queue buffer 508 to determine whether any transformations have been performed after example transformation 12 or example transformation 13. Figure 6, example transition 14 is executed twice after transition 12 (entry 1 602 and entry 3 606), causing graph generator 510 to add example node 14 representing transition 14 to graph 700, and further add example fifth edge 710 with edge weight 2 connecting node 12 to node 14. In some examples, graph generator 510 stops adding additional nodes to the graph based on a heuristic limit that can be set using one or more policies.
[0053] Figure 8 The block diagram of includes an example implementation of an example merge candidate selector and checker 512. Now referring to Figure 5 and Figure 8 , when the example graph generator 510 has completed generating the graph 700, the example merge candidate selector 512 operates to select a set of transformations that are candidates to be merged with the seeds. In some examples, the merge candidate selector 512 includes an example priority queue 802 that uses the edge weights of the graph 700 as a priority key, an example merge candidate memory, an example stopping criteria evaluator 806, an example chain evaluator 808, an example edge weight evaluator 810, and an example controller 812. In some examples, the controller 812 of the merge candidate selector 512 accesses the graph 700 generated by the graph generator 510 and causes the seeds (e.g., the example transformation 10) identified in the graph 700 to be placed in the priority queue 802.
[0054] Then, the example controller 812 of the merge candidate selector and checker 512 dequeues the seed from the priority queue 802 and stores information identifying the seed in the merge candidate memory 804. When the seed information is placed in the merge candidate memory 804, the controller 812 checks the graph generated by the graph generator 510 to identify the predecessor node and the successor node of the seed. The controller 812 then provides information identifying the transformation represented by the predecessor node and the successor node to the stopping criteria evaluator 806, which evaluates the transformation according to a set of stopping criteria. In some examples, the stopping criteria evaluator 806 adds the transformation represented by the successor node and the predecessor node to the priority queue 802 unless the transformation meets any stopping criteria. In some examples, the set of stopping criteria includes: 1) the node / transformation is a current transformation (and therefore has already been added), 2) the node / transformation is no longer valid (which may occur if the node / transformation is no longer executed), 3) the metadata of the node / transformation is the same as the metadata of another transformation already included in the merge candidate store, 4) the node / transformation is illegal due to, for example, SMC, 5) the node / transformation has mismatched transformation options (e.g., when the node / transformation has already been transformed using a different optimization technique), and / or 6) the node / transformation is based on a previously generated super region. After evaluating the nodes / transformations according to the stopping criteria, the remaining nodes / transformations in the priority queue 802 are evaluated by the example chain evaluator 808.
[0055] In some examples, the example chain evaluator 808 determines whether the conversion / node included in the example priority queue 802 is directly linked to the seed (via one or more edges) and / or is linked to the seed by one or more of the other conversions / nodes included in the priority queue 802. In some examples, the example chain evaluator causes the conversion that is not included in at least one node chain including the seed to be eliminated / removed from the priority queue 802. (As used herein, a node / conversion chain refers to a set of nodes / conversions linked via one or more edges.) The conversion removed from the priority queue 802 is no longer considered eligible to be merged with the seed. When determining whether the conversion / node included in the priority queue 802 is also included in at least one node chain including the seed, the chain evaluator 808 also obtains information about the extended chain including the seed. In some examples, the extended chain is an extension of the chain included in the graph 700. In some examples, the extended chain is identified using information stored in the conversion queue buffer 508 and / or by accessing and scanning the conversion-related metadata stored in the conversion results and metadata storage 408. In some examples, the chain evaluator 808 causes the conversion / node included in the extended chain to be added to the priority queue 802, and the added conversion / node is identified to the example stopping criteria evaluator 806 to be evaluated according to the stopping criteria. As described above, the stopping criteria evaluator 806 removes any conversion / node that meets the stopping criteria from the priority queue 802. In some examples, the chain evaluator 808 only extends the chain to include nodes / conversions that result in chains with less than a threshold number of instructions, less than a threshold number of blocks, and / or less than a threshold number of exits. The threshold number of instructions, blocks, and exits can be selected based on any of a plurality of factors including the available register resources that the converter is configured to use.
[0056] In some examples, the example edge weight evaluator 810 also evaluates the nodes / transformations included in the priority queue 802. In some examples, the edge weight evaluator 810 evaluates the nodes / transformations in the priority queue 802 by identifying which edge of the seed has the heaviest weight. As described above, in this context, the edge weight refers to the number of times a first transformation represented by a first node is executed before (or after) a second transformation represented by a second node. In some examples, the edge weight evaluator 810 selects a transformation coupled to the seed by an edge having the heaviest weight to merge with the seed. In some examples, the edge weight evaluator 810 determines whether the weight of the heaviest edge is at least a threshold portion of the total weight of all edges coupled to the seed. When the edge weight evaluator 810 determines that the weight of the heaviest edge is at least a threshold portion of the total weight of all edges coupled to the seed, the edge weight evaluator 810 selects the corresponding transformation to merge with the seed. When the weight of the heaviest edge is not at least a threshold portion of the total weight of all edges coupled to the seed, the corresponding transformation is not selected for merging with the seed. In some examples, the threshold portion of the total weight is four eighths or five eighths of the total weight of all edges of the seed. In some examples, the value of the threshold portion of the total weight is modifiable and can be based on the example processor 110 (see Figure 1 ) to adjust the dynamic execution of the program to achieve higher performance gains and loop coverage.
[0057] Still see Figure 8 In some examples, when a transformation located in the priority queue is selected for merging (e.g., after evaluation by the example stopping criteria evaluator 806, the example chain evaluator 808, and the example edge weight evaluator 810), the controller 812 dequeues the selected transformation from the example priority queue 802 and adds it to the example merge candidate memory 804. Thereafter, the example merge candidate selector and checker 512 repeats the above evaluation process for the dequeued transformation (e.g., using the stopping criteria evaluator 806, the chain evaluator 808, and the edge weight evaluator 810) to determine which (if any) other transformations are to be included in the merge candidate memory 804. In some examples, the evaluation process is repeated until no other candidates remain in the priority queue 802. In some examples, the number of transformations selected for merging is at least two. In some examples, the number of candidate transformations selected for merging is modifiable.
[0058] Still reference Figure 5 and Figure 8, when there are no remaining transformations in the priority queue, the example controller 812 of the example merge candidate selector and checker 512 notifies the example entry point selector 514 that a set of transformations has been selected for merging and that the set of transformations is stored in the example merge candidate memory 804. The entry point selector 514 performs an entry point selection process to identify entry points for transformations. In some examples, the entry point selector 514 begins the selection process by making all transformations to be included in the merge (referred to as "merge transformations") unreachable. The entry point selector 514 then uses a depth-first search algorithm to traverse each transformation to search for loops. In some examples, the entry point selector 514 obtains transformations and associated instructions from the example transformation results and metadata memory 408 (see Figure 4 ). For each loop detected, the entry point selector 514 selects the best loop header. In some examples, for each conversion that includes the loop, the entry point selector 514 initially identifies the entry point of the conversion with the lowest extended instruction pointer (EIP) register value as the loop header. In some examples, the entry point selector then uses a trace graph to identify the loop entry. In some examples, the trace graph is used to identify the following loop entries: 1) most frequently appearing before a conversion that is not one of the merged conversions, and 2) after the conversion that most often exits the loop. In some examples, if the entry found using the trace graph also has the most frequent edge count (based on a heuristic), then the entry found using the trace graph is selected as the entry point. If the entry found using the trace graph does not have the most frequent edge count, the entry with the lowest execution address is selected as the loop header. In some examples, the entry point selector 514 counts the entry points of the conversions to be merged and compares the count with a threshold entry count value. If the threshold entry count value is exceeded, the merging process fails.
[0059] In some examples, when the transformation selection performed by the example entry point selector 514 fails (eg, fewer than two merging transformations are identified, a threshold entry count value is exceeded, etc.), the example processor 110 (see Figure 4 ) continues executing the current conversion (for example, the conversion that was executing when the interrupt timer expires).
[0060] In some examples, when the conversion selection process performed by the example entry point selector 514 is successful, the entry point selector 514 provides information identifying the conversion to be merged and the selected entry point to the example super region creator 520. The super region creator 520 topologically sorts the entry point EIP register values of the candidate conversions and accumulates the basic blocks from the merged conversions. Assuming that a set of block and instruction thresholds are not exceeded, the super region creator 520 merges the basic blocks together to form a super region. If the block and instruction thresholds are exceeded, the merging process fails and the super region creation process starts again with the current conversion. Super region creation cannot be started immediately because it is likely to cause another failure. In some examples, information related to the failed super conversion attempt can be saved and tried again at a later time. In some examples, the instruction limit is 500 and the block limit is 64. In some examples, the super region creator 520 provides the information identifying the super region to the dynamic binary translator to reconvert as a single unit. If the re-conversion is successful, the translation used as the seed is removed from the translation results and metadata memory 408 and / or the hardware translation lookup table, and the newly created super-translation is inserted in its place. If the re-conversion is unsuccessful due to limited resources, the dynamic binary transcoder 106 saves the super-region information in the target code memory 114 and attempts to perform the re-conversion of the super-region at a later time when more resources are available.
[0061] The super region created using the second dynamic binary translation system outperforms conventional region maker techniques in terms of loop coverage, region size, and instructions per cycle (IPC) gain. The loop coverage used here refers to the ratio of 1) dynamic instructions from loops and loop nests captured in a region to 2) the total number of dynamic instructions executed. Region size refers to the average number of native instructions in the region where the conversion is performed. The performance gain obtained when executing the optimized conversion generated by the super region is used to determine the IPC gain. Using a mixture of benchmarks from various tool suites representing a mixture of client and server workloads, the second dynamic binary translation system obtains 6% more loop coverage, 1.62 times larger regions, and 1.4% overall IPC gain than the baseline system using the conventional region maker.
[0062] Although in Figure 1 and Figure 2 An example manner of implementing an example first dynamic binary translation system 102 having an example region formation manager 104 is illustrated in Figure 4 , Figure 5 and Figure 8 An example manner of implementing an example second dynamic binary translation system 402 having an example super region formation manager 404 is illustrated in FIG. Figure 1 , Figure 2 , Figure 4 , Figure 5 and Figure 8 One or more of the illustrated elements, processes and / or devices may be combined, divided, rearranged, omitted, eliminated and / or implemented in any other manner. In addition, the example dynamic binary code converter 106, the region formation manager 104, the example region memory 203, the example first execution analyzer 204, the example queue manager 206, the example queue 208, the example initializer 210, the example initial region former 212, the example region expander 214, the example loop feature memory 218, the example loop selector 220, the example standard memory 222, the example region pruner 224, the example super region formation manager 404, the example conversion result and metadata memory 408, the example example second execution profiler 410, example transition queue 502, example transition queue sampler 504, example policy memory 506, example transition queue buffer 508, example graph generator 510, example merge candidate selector and checker 512, example entry point selector 514, example super region creator 520, example priority queue 802, example merge candidate memory 804, example stopping criteria evaluator 806, example chain evaluator 808, example edge weight evaluator 810, example controller 812, and / or more generally, Figure 1 An example of a first dynamic binary translation system and Figure 4 The example second dynamic binary translation system may be implemented by hardware, software, firmware and / or any combination of hardware, software and / or firmware. Thus, for example, the example dynamic binary code converter 106, the region formation manager 104, the example region memory 203, the example first execution profiler 204, the example queue manager 206, the example queue 208, the example initializer 210, the example initial region former 212, the example region expander 214, the example loop feature memory 218, the example loop selector 220, the example standard memory 222, the example region pruner 224, the example super region formation manager 404, the example conversion result and metadata storage 408, example second execution analyzer 410, example transition queue 502, example transition queue sampler 504, example policy memory 506, example transition queue buffer 508, example graph generator 510, example merge candidate selector and checker 512, example entry point selector 514, example super region creator 520, example priority queue 802, example merge candidate memory 804, example stopping criteria evaluator 806, example chain evaluator 808, example edge weight evaluator 810, example controller 812, and / or more generally, Figure 1 An example of a first dynamic binary translation system and Figure 5Any of the example second dynamic binary translation systems can be implemented by one or more analog or digital circuits, logic circuits, (multiple) programmable processors, (multiple) application specific integrated circuits (ASICs), (multiple) programmable logic devices (PLDs), and / or (multiple) field programmable logic devices (FPLDs). When any apparatus or system claim of this patent is read to cover pure software and / or firmware implementations, the example dynamic binary code converter 106, the region formation manager 104, the example region memory 203, the example first execution profiler 204, the example queue manager 206, the example queue 208, the example initializer 210, the example initial region former 212, the example region expander 214, the example loop feature memory 218, the example loop selector 220, the example standard memory 222, the example region pruner 224, the example super region formation manager 404, the example conversion result and metadata memory 408, the example second execution profiler 410, the example conversion queue At least one of the columns 502, the example transition queue sampler 504, the example policy memory 506, the example transition queue buffer 508, the example graph generator 510, the example merge candidate selector and checker 512, the example entry point selector 514, the example super region creator 520, the example priority queue 802, the example merge candidate memory 804, the example stopping criteria evaluator 806, the example chain evaluator 808, the example edge weight evaluator 810, and / or the example controller 812 is expressly defined herein to include a non-transitory computer readable storage device or storage disk, such as a memory, a digital versatile disk (DVD), a compact disk (CD), a Blu-ray disk, etc., which includes software and / or firmware. In addition, Figure 1 An example of a first dynamic binary translation system 102 and / or Figure 4 An example second dynamic binary translation system 402 may include as Figure 2 , Figure 5 and Figure 8 One or more elements, processes and / or devices in addition to or instead of the elements, processes and / or devices illustrated in the drawings, and / or may include one or more of any or all of the elements, processes and devices illustrated in the drawings.
[0063] Indicates that it is used to implement Figure 1 A flowchart of example machine readable instructions for the region formation manager 104 is provided in Figure 9-14 is shown in and is used to implement Figure 4 A flowchart of example machine readable instructions for the super region formation manager 404 is provided in Fig.15 and Fig.16In these examples, the machine-readable instructions include instructions for execution by a processor (such as the processor 1712 or the processor 1812 shown in the example processor platform 1700 and the example processor platform 1800) and in conjunction with the following Fig.17 and Fig.18 The program may be embodied in software stored on a non-transitory computer-readable storage medium such as a CD-ROM, floppy disk, hard drive, digital versatile disk (DVD), Blu-ray disk, or memory associated with processor 1712 (or processor 1812), but all and / or or a portion thereof may alternatively be executed by a device other than processor 1712 (or processor 1812) and / or may be embodied in firmware or dedicated hardware. In addition, although reference is made to Figure 9-Figure 16 The flowcharts illustrated in describe example procedures, but many other methods of implementing the example region formation manager 104 and / or the example super region formation manager 404 may be used instead. For example, the order of execution of the various blocks may be changed, and / or some of the blocks described may be changed, eliminated, or combined. Additionally or alternatively, any or all of the blocks may be implemented by one or more hardware circuits (e.g., discrete and / or integrated analog and / or digital circuits, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), comparators, operational amplifiers (op-amps), logic circuits, etc.) that are structured to perform corresponding operations without executing software or firmware.
[0064] As mentioned above, the present invention may be implemented using coded instructions (eg, computer and / or machine readable instructions) stored on a non-transitory computer and / or machine readable medium. Figure 9-Figure 16In the example process of, non-transient computer and / or machine readable media such as: hard disk drive, flash memory, read-only memory (ROM), compact disk (CD), digital versatile disk (DVD), cache, random access memory (RAM) and / or any other storage device or storage disk that stores information therein for any length of time (e.g.: for an extended period of time, permanently, during a brief instance, during temporary buffering and / or information caching). As used herein, the term "non-transient computer readable storage medium" is explicitly defined to include any type of computer readable storage device and / or storage disk, and excludes propagated signals and excludes transmission media. "Include" and "include" (and all forms and tenses thereof) are used as open terms in this article. Therefore, whenever a claim lists any content following any form of "include" or "includes" (e.g., including, including, etc.), it is understood that additional elements, items, etc. may exist without exceeding the scope of the corresponding claim. As used herein, when the phrase "at least" is used as a transition term used synchronously with the claim, it is open like the terms "include" and "include".
[0065] Fig. 9 The program 900 begins at block 902, where the example queue manager 206 causes a code block to be placed in the example queue 208 in response to a notification from the example first execution profiler 204. The example initializer 210 monitors the queue length and determines whether the queue contains a threshold number of hot code blocks (block 904). If the threshold number of hot code blocks has not been met, the initializer 210 continues to monitor and evaluate the queue length (block 904). If the initializer 210 determines that the threshold number of hot code blocks is included in the queue 208 (block 904), the initializer 210 generates a trigger (block 906). The example initial region former 212 responds to the trigger by growing the initial region during the main region growth phase (block 908). The example region expander 214 expands the initial region in the secondary region growth phase (block 910). The example region analyzer 216 analyzes the loop nest of the expanded region generated during the secondary region growth phase and generates a set of features about the loop (block 912). The example loop selector 220 then applies criteria to the feature to identify and select loop nests to be included in the final region (block 914). The example region pruner 224 prunes the expanded region until only the selected loop nests remain to form the final region (block 916). In some examples, the region pruner 224 also performs queue cleanup actions to remove one or more of the hot code blocks from the queue 208. The region pruner stores information identifying / specifying the final region in the region memory 202 (block 918), and the procedure 900 ends.
[0066] Fig.10The program 1000 of 1000 begins at block 1002, where the example first execution profiler 204 monitors and profiles the execution of program code by the example processor 110. When the execution profiler 204 identifies the hot code block to the example queue manager 206 (block 1004), the queue manager 206 determines whether the hot code block is already contained in the queue 208 (block 1006). If the hot code block is already in the queue 208, the queue manager 206 moves the hot code block to the head of the queue 208 (block 1008). If the hot code block is not yet in the queue 208, the queue manager 206 adds the hot code block to the end of the queue 208 (block 1010). After the hot code block is moved to the head of the queue 208 or added to the end of the queue 208, as long as the processor 110 is executing the code, the execution profiler 204 continues to profile the execution of the code by the processor 110 (block 1002) and the program continues to operate in the described manner.
[0067] Fig.11 The example process 1100 begins at block 1102, where the example initializer 210 monitors the queue length (block 1102) and then determines whether the queue length meets a threshold (block 1104). If the queue length does not meet the threshold, the initializer 210 continues to monitor the queue length (block 1102). If the queue length does meet the threshold, the initializer 210 triggers the initial region former to begin executing the initial region growth phase (block 1106). After generating the trigger, the initializer 210 continues to monitor the queue length (block 1102) to identify when the next region will be generated.
[0068] Fig.12 The program 1200 begins at block 1202, where, in response to a trigger generated by the example initializer 210, the example initial region former 212 selects a hot code block from the queue 208 that was most recently placed in the queue 208. The initial region former 212 then generates a control flow graph that starts at the hot code block and continues to include code blocks placed along the hottest path of the control flow graph until a threshold number of blocks have been added to the initial region or the hottest path leads to a code block that is too cold to be placed in the initial region (block 1204), at which point the program 1200 ends.
[0069] Fig.13The program 1300 of begins at block 1302, where the example region expander 214 identifies an exit included in the initial region. In some examples, the region expander 214 begins identifying exits in response to a notification from the example initial region shaper indicating that the initial growth phase has been completed. The region expander 214 identifies and selects the hottest exit (block 1304). Then, the region expander 214 determines whether to expand the initial region from the hottest exit (block 1306). In some examples, the region expander 214 makes a decision by evaluating the code blocks reachable via the hottest exit and, based on the result of the evaluation, adding the path including the code blocks reachable via the hottest exit to the expanded region. In some examples, the evaluation includes determining whether the path expanded from the evaluated exit: 1) reaches a hot code block that is already connected to the initial region (or reaches the expanded region if the initial region has been expanded), or 2) reaches a hot code block within a threshold number of code blocks (even if the code blocks between the exit and the hot code block are cold code blocks). In making this evaluation, region expander 214 ignores any back edges included in the code blocks that help region formation manager 104 capture an outer loop when only a subset of its inner loop is represented by the code blocks included in queue 208 .
[0070] When the region expander 214 determines that the path starting from the hottest path will not be added to the initial region, the region expander 214 removes the previously evaluated exits from the exit list (box 1312) and determines whether there are any exits to be evaluated (box 1314). When there are no exits to be evaluated, the program ends. When there are still exits to be evaluated, the region expander selects the hottest exit from the list of remaining exits again (box 1304). When the region expander 214 determines that the path starting from the hottest path will be added to the initial region, the region expander 214 adds the code blocks placed along the path (box 1308). In some examples, the region expander 214 adds blocks placed along the path until encountering (and including) a code block associated with a back edge or until a threshold number of code blocks have been added. After adding the path, the region expander 214 identifies additional exits included in the extended region as a result of adding the path (box 1310), and then removes any exits that have been evaluated from the exit list to be evaluated (box 1312). If there are more exits to be evaluated, the zone expander continues to select and evaluate the next exit (block 1306) and add paths based on the evaluation (block 1308). Alternatively, if there are no exits remaining in the exit list to be evaluated (determined at block 1314), the process 1300 ends.
[0071] Fig.14The process 1400 begins at block 1402, where the example region analyzer 216 analyzes the expansion region to identify all loop nests included in the expansion region. The region analyzer 216 also characterizes the loop nests by identifying a set of properties / features of the loop nests (block 1404). In some examples, the region analyzer 216 applies criteria to the loop nest features to select loops to be included in the final region (block 1406). The region analyzer 216 identifies the selected loops to the region pruner 224, which responds to the information by pruning all loop nests except the selected loop nests from the expansion region (block 1408). The region pruner 224 stores the final region and provides the final region to the dynamic binary translator 104 (see Figure 4 ) to be converted into a single unit. Thereafter, the process 1400 ends.
[0072] Fig.15 The process 1500 begins at block 1502 where the example conversion queue 502 (see Figure 5 ) in the example processor 110 (see Figure 4 ) generates an interrupt at block 1502 and collects information identifying the conversion currently being executed for inclusion in the conversion queue 502. Also at block 1502, the example conversion queue sampler 504 periodically captures the contents of the conversion queue 502 for storage in the example conversion queue buffer 508 (see Figure 5 ). The conversion queue also determines whether the count value associated with the currently executing conversion meets the sample count threshold (block 1504). The count value corresponds to the number of times the conversion queue 502 has sampled the currently executing conversion. When the count value meets the sample count threshold, the currently executing conversion is determined to be "hot" and is identified as a "seed" (block 1506). When the count value does not meet the sample count threshold, the currently executing conversion is not determined to be "hot" and the program returns to periodically interrupting the processor 110 (block 1502). The graph generator 510 generates a graph based on the seed, as described above with respect to Figure 5 Next, the example merges the candidate selector and the checker 512 (see Figure 5 ) uses the graph to identify a set of transformations that are candidates for merging with the seed (block 1510). The example entry point selector identifies a set of entry points associated with the loops contained in the transformation (block 1512). Finally, the transformation to be merged with the seed and the entry points identified for the transformation are provided to the example super region creator 520 (see Figure 5 ), which uses this information to merge the transformation into the super region (block 1514). In addition, the super region creator provides the super region to the transformation for re-transformation (also at block 1514). Thereafter, the process 1500 ends.
[0073] Fig.16 The process 1600 begins at block 1602 where the example merge candidate selector and checker 512 (see Figure 5 ) of an example controller 812 (see Figure 8 ) dequeues the seed instance priority queue 802 (see Figure 8 ). The controller then causes the seed's predecessor and successor transformations to be added to the priority queue 802. In some examples, the controller 812 uses the example graph generator 510 (see Figure 5 ) to identify the seed and the predecessor and successor transitions. The stopping criteria evaluator then evaluates the predecessor and successor transitions in the priority queue according to a set of stopping criteria and removes any transitions that meet any of the stopping criteria from the priority queue 802 (box 1606). In some examples, the example chain evaluator 808 identifies a set of chains and extended chains that contain the seed and adds the transitions included in the chains to the priority queue 802 (box 1608). In some examples, the stopping criteria evaluator 806 operates again to evaluate the newly added transitions according to the stopping criteria and again removes any transitions that meet any of the criteria from the priority queue 802 (box 1610). The example edge weight evaluator 810 (see Figure 8 ) selects the transformation contained in the priority queue that is coupled to the seed by the edge with the heaviest weight as a candidate for merging with the seed (box 1612). In some examples, the edge weight evaluator 810 causes the selected transformation to be placed in the merge candidate memory 804 and causes the selected transformation to become a new "seed" (box 1614). The controller 812 then determines whether the priority queue 802 is empty (box 1616). If the priority queue 802 is not empty, more transformations will be evaluated for possible merging with the seed, and control returns to box 1602 and subsequent boxes as described above. If the priority queue 802 is empty, all transformations that are candidates for merging have been identified, and the program 1600 ends.
[0074] Fig.17 Is able to execute Figure 9-14 Instructions to achieve Figure 2 The processor platform 1700 can be, for example, a server, a personal computer, a mobile device (e.g., a cell phone, a smartphone, an iPad, etc.). TM tablet device such as a ), or any other type of computing device.
[0075] The processor platform 1700 of the illustrated example includes a processor 1712. The processor 1712 in the illustrated example is hardware. For example, the processor 1712 can be implemented by one or more integrated circuits, logic circuits, microprocessors or controllers from any desired family or manufacturer. The hardware processor can be a semiconductor-based (e.g., silicon-based) device. In this example, the processor 1712 implements the example first execution profiler 204, the example queue manager 206, the example initializer 210, the example initial region former 212, the example region expander 214, the example region analyzer 216, the example loop selector 220, and the example region trimmer 224.
[0076] The processor 1712 of the illustrated example includes a local memory 1713 (e.g., a cache). The processor 1712 in the illustrated example communicates with a main memory including a volatile memory 1714 and a non-volatile memory 1716 via a bus 1718. The volatile memory 1714 may be implemented by a synchronous dynamic random access memory (SDRAM), a dynamic random access memory (DRAM), a RAMBUS dynamic random access memory (RDRAM), and / or any other type of random access memory device. The non-volatile storage 1716 may be implemented by a flash memory and / or any other desired type of memory device. A memory controller controls access to the main memory 1714, 1716. The local memory 1713, the volatile memory 1714, and the non-volatile memory 1716 may be used to implement any one or all of the example regional memory 202, the example queue 208, the example feature memory 218, and / or the example standard memory 222.
[0077] The processor platform 1700 in the illustrated example also includes an interface circuit 1720. The interface circuit 1720 may be implemented by any type of interface standard, such as an Ethernet interface, a Universal Serial Bus (USB), and / or a PCI Express interface.
[0078] In the example shown, one or more input devices 1722 are connected to the interface circuit 1720. The input device(s) 1722 permit a user to input data and / or commands into the processor 1712. The input device(s) may be implemented, for example, by an audio sensor, microphone, camera (still or video), keyboard, button, mouse, touch screen, track pad, track ball, iso-point mouse, and / or voice recognition system. In some examples, the input device 1722 may be used to input any standard stored in the standard memory 222.
[0079] One or more output devices 1724 are also connected to the interface circuit 1720 in the illustrated example. The output device 1724 may be implemented, for example, by a display device (e.g., a light emitting diode (LED), an organic light emitting diode (OLED), a liquid crystal display, a cathode ray tube display (CRT), a touch screen, a tactile output device, a printer, and / or a speaker). Therefore, the interface circuit 1720 in the illustrated example typically includes a graphics driver card, a graphics driver chip, and / or a graphics driver processor.
[0080] The interface circuitry 1720 in the illustrated example also includes communication devices such as transmitters, receivers, transceivers, modems, and / or network interface cards to facilitate the exchange of data with an external machine (e.g., any kind of computing device) via a network 1726 (e.g., an Ethernet connection, a digital subscriber line (DSL), a telephone line, a coaxial cable, a cellular telephone system, etc.).
[0081] The processor platform 1700 of the illustrated example also includes one or more mass storage devices 1728 for storing software and / or data. Examples of such mass storage devices 1728 include floppy disk drives, hard disk drives, compact disk drives, Blu-ray disk drives, RAID systems, and digital versatile disk (DVD) drives.
[0082] Figure 9-Figure 16 The encoded instructions 1732 may be stored in the mass storage device 1728, in the volatile memory 1714, in the non-volatile memory 1716, and / or on a removable tangible computer-readable storage medium such as a CD or DVD.
[0083] Fig.18 Is able to execute Fig.15 and Fig.16 Instructions to achieve Figure 4 The processor platform 1800 can be, for example, a server, a personal computer, a mobile device (e.g., a cell phone, a smart phone, an iPad, etc.). TM tablet device such as a ), or any other type of computing device.
[0084] The processor platform 1800 of the illustrated example includes a processor 1812. The processor 1812 in the illustrated example is hardware. For example, the processor 1812 can be implemented by one or more integrated circuits, logic circuits, microprocessors or controllers from any desired family or manufacturer. The hardware processor can be a semiconductor-based (e.g., silicon-based) device. In this example, the processor 1812 can be used to implement the example second execution profiler 410, the example conversion queue 502, the example conversion queue sampler 504, the example graph generator 510, the example merge candidate selector and checker 512, the example entry point selector 514, the example super region creator 520, the example controller 812, the example priority queue 802, the example stop standard evaluator 806, the example chain evaluator 808, the example edge weight evaluator 810 and / or the example super region creator 520.
[0085] The processor 1812 of the illustrated example includes a local memory 1813 (e.g., a cache). The processor 1812 in the illustrated example communicates with a main memory including a volatile memory 1814 and a non-volatile memory 1816 via a bus 1818. The volatile memory 1814 may be implemented by a synchronous dynamic random access memory (SDRAM), a dynamic random access memory (DRAM), a RAMBUS dynamic random access memory (RDRAM), and / or any other type of random access memory device. The non-volatile memory 1816 may be implemented by a flash memory and / or any other desired type of memory device. A memory controller controls access to the main memory 1814, 1816. The local memory 1813, the volatile memory 1814, and the non-volatile memory 1816 may be used to implement any or all of the example policy memory 506, the example conversion queue buffer 508, and / or the example merge candidate memory 804.
[0086] The processor platform 1800 in the illustrated example also includes an interface circuit 1820. The interface circuit 1820 may be implemented by any type of interface standard, such as an Ethernet interface, a Universal Serial Bus (USB), and / or a PCI Express interface.
[0087] In the example shown, one or more input devices 1822 are connected to the interface circuit 1820. The input device(s) 1822 permit a user to input data and / or commands into the processor 1812. The input device(s) may be implemented, for example, by an audio sensor, microphone, camera (still or video), keyboard, button, mouse, touch screen, track pad, track ball, iso-point mouse, and / or voice recognition system. In some examples, the input device 1822 may be used to input any policy information stored in the policy memory 506.
[0088] One or more output devices 1824 are also connected to the interface circuit 1820 in the illustrated example. The output device 1824 may be implemented, for example, by a display device (e.g., a light emitting diode (LED), an organic light emitting diode (OLED), a liquid crystal display, a cathode ray tube display (CRT), a touch screen, a tactile output device, a printer, and / or a speaker). Therefore, the interface circuit 1820 in the illustrated example typically includes a graphics driver card, a graphics driver chip, and / or a graphics driver processor.
[0089] The interface circuitry 1820 in the illustrated example also includes communication devices such as transmitters, receivers, transceivers, modems, and / or network interface cards to facilitate the exchange of data with an external machine (e.g., any kind of computing device) via a network 1826 (e.g., an Ethernet connection, a digital subscriber line (DSL), a telephone line, a coaxial cable, a cellular telephone system, etc.).
[0090] The processor platform 1800 of the illustrated example also includes one or more mass storage devices 1828 for storing software and / or data. Examples of such mass storage devices 1828 include floppy disk drives, hard disk drives, compact disk drives, Blu-ray disk drives, RAID systems, and digital versatile disk (DVD) drives.
[0091] Fig.15 and Fig.16 The encoded instructions 1832 may be stored in the mass storage device 1828, in the volatile memory 1814, in the non-volatile memory 1816, and / or on a removable tangible computer-readable storage medium such as a CD or DVD.
[0092] As can be appreciated from the foregoing, example methods, apparatuses, and products for performing region formation to identify multiple code blocks of a computer program that are to be converted by a dynamic binary translator into a single unit have been disclosed. The region formation manager disclosed herein provides enhanced loop coverage of computer code, which allows the dynamic binary translator to generate converted code that executes faster and more efficiently. Some such example region formation managers disclosed herein perform two region formation phases (a primary growth phase and a secondary growth phase). The initial region formation is the primary growth phase, and the extended region formation is the secondary growth phase. In addition, some such region formation managers disclosed herein initiate the formation of the initial growth phase only after multiple hot code blocks are identified using an execution profiler (rather than just a single hot code block). In addition, the disclosed region formation manager analyzes an extended region containing multiple loops to identify the best loop nest to be included in the final region. These operational aspects of such region formation managers enable the generation of a final region that includes superior loop coverage compared to conventional techniques and results in improved dynamic binary conversion. Some super region formation managers disclosed herein provide enhanced loop coverage of computer code to be converted by a dynamic binary translator by combining two or more regions of converted computer code to form super regions. The super region is provided to the dynamic binary converter for reconversion as a single unit. As a result, the dynamic binary converter can generate converted code that is executed more quickly and efficiently. The super region formation manager creates a super region by first identifying a hot conversion as a seed and then identifying a predecessor conversion and / or a successor conversion of the seed. The predecessor conversion and the successor conversion of the seed are used to identify other candidate conversions that are included in the same chain as the seed and can be combined with the seed to create a super region. The candidate conversions to be included in the super region are subject to a set of evaluations that are designed to improve loop coverage and region size. The candidate conversions are then combined to form a super region, which is then reconverted by the dynamic binary converter. In some examples, the entire process of forming a super region results in greater loop coverage, greater region, and greater IPC gain.
[0093] The following further examples are disclosed herein.
[0094] Example 1 is a device for performing region formation of a control flow graph. In Example 1, the device includes an initial region former for forming an initial region starting from the first hot code block of the control flow graph. The initial region former adds a hot code block located on the first hottest path of the control flow graph to the initial region until reaching one of the following: a hot code block previously added to the initial region, or a cold code block. The device of Example 1 also includes a region expander for expanding the initial region to form an extended region. The extended region includes the initial region, and the region expander begins to expand the initial region at the hottest exit of the initial region. The region expander adds a hot code block located on the second hottest path to the extended region until one of the following situations: a threshold path length is met, or a back edge of the control flow graph is added to the extended region. The device of Example 1 also includes a region pruner for pruning all loops except the selected loop from the extended region. The selected loop forms the region.
[0095] Example 2 includes the apparatus of Example 1, and further includes a region analyzer for identifying and generating features of a plurality of loops and loop nests included in the extended region. Additionally, the apparatus of Example 2 includes a loop selector for selecting one of the loops identified by the loop analyzer based on a set of criteria. The loop selector applies the set of criteria to the features generated by the region analyzer.
[0096] Example 3 includes the apparatus of Example 1, and further includes: a queue for containing block identifiers of a plurality of code blocks of the control flow graph identified as hot; and an initializer for generating a trigger when the queue includes a threshold number of the code blocks. In Example 3, the initial region former initiates the formation of the initial region in response to the trigger. In addition, the threshold number is greater than 1.
[0097] Example 4 includes the apparatus of Example 3. In Example 4, the first hot code block is the last code block added to the queue before the trigger is generated.
[0098] Example 5 includes the apparatus of Example 3 and / or Example 4. In Example 5, the initial region former and the region expander use the queue to determine whether a code block is hot.
[0099] Example 6 includes the apparatus of any one of Examples 1, 2, and 3. In Example 6, the threshold path length is a first threshold path length. In Example 6, the region expander further 1) determines whether the second hottest path has at least one of the following situations before adding the second hottest path to the extended region: connects back to a block in the extended region or leads to a second hot code block within a second threshold path length, and 2) based on the determination, adds the code block on the second hottest path to the extended region.
[0100] Example 7 includes the apparatus of Example 6. In Example 7, the region expander determines that the second hottest path leads to the second hot code block within the second threshold path length, and the second hottest path includes a cold code block between the hottest exit of the initial region and the second hot code block.
[0101] Example 8 includes the apparatus of any one of Examples 1, 2, and 3. In Example 8, the region expander further 1) iterates over the corresponding hottest exits of the initial region and the expanded region in descending order of heat, and 2) adds code blocks placed along the corresponding hottest path to the expanded region when the corresponding hottest path meets the criteria.
[0102] Example 9 includes the apparatus of any one of Examples 1, 2, and 3. In Example 9, the region is a first region, and the apparatus further comprises a super region former for merging a first transformation generated based on the first region and a second transformation generated based on the second region. The merging of the first transformation and the second transformation forms a super region.
[0103] Example 10 includes the apparatus of Example 9. The apparatus of Example 10 also includes: a queue for identifying a hot transition; a graph generator for generating a flow graph including the hot transition, a predecessor transition and a successor transition of the hot transition; and a transition selector for selecting at least one of the predecessor transition and the successor transition for merging with the hot transition. In Example 10, the hot transition is the first transition, and the at least one of the predecessor transition and the successor transition is the second transition.
[0104] Example 11 is one or more non-transient machine-readable storage media including machine-readable instructions. The instructions, when executed, cause at least one processor to form at least an initial region. The initial region starts at the first hot code block of a control flow graph and includes a hot code block located on the first hottest path extending from the first hot code block of the control flow graph until one of the following situations: a hot code block previously added to the initial region is reached, or a cold code block is reached. The instructions of Example 11 also cause the at least one processor to extend the initial region to form an extended region. The extended region includes the initial region. Extending the initial region starts at the hottest exit of the initial region and includes adding a hot code block located on the second hottest path of the control flow graph until a threshold path length has been met or a hot code block associated with a back edge of the control flow graph is added to the extended region. The instructions of Example 11 also cause the at least one processor to prune a set of loops other than the selected loop from the extended region. The selected loop forms the final region.
[0105] Example 12 includes the one or more non-transitory machine-readable storage media of Example 11, and further includes instructions for causing the at least one processor to identify loops and loop nests included in the extended region and generate features of the loops and loop nests. Additionally, the instructions of Example 11 cause the processor to select one of the loops based on the features.
[0106] Example 13 includes one or more non-transitory machine-readable storage media of Example 11, and further includes instructions that cause the at least one processor to identify a code block included in the control flow graph as hot based on profiling, and add the name of the code block identified as hot to a queue. The instructions of the example also cause the processor to generate a trigger when the queue includes a threshold number of code blocks. In Example 11, the forming of the initial region is initiated in response to the trigger, and the threshold number is greater than 1.
[0107] Example 14 includes the one or more non-transitory machine-readable storage media of Example 13. In Example 14, the first hot code block is a last code block added to the queue before generating the trigger.
[0108] Example 15 includes one or more non-transitory machine-readable storage media of Example 11. In Example 15, the threshold path length is a first threshold path length. Example 15 also includes instructions that cause the at least one processor to determine whether the second hottest path has at least one of the following situations before adding the code block located on the second hottest path to the region: connecting back to a previously formed portion of the initial region or the extended region, or leading to a second hot code block within a second threshold path length. Based on this determination, the instructions of Example 15 cause the processor to add the code block located on the second hottest path to the extended region.
[0109] Example 16 includes the one or more non-transitory machine-readable storage media of Example 15. In Example 16, the second hottest path is determined to lead to the second hot code block within the second threshold path length, and the second hottest path includes a cold code block between the hottest exit of the initial region and the second hot code block.
[0110] Example 17 includes one or more non-transitory machine-readable storage media of Example 14, and also includes instructions that cause the at least one processor to iteratively select the corresponding hottest exits of the region in descending order of heat, and add code blocks along the corresponding hottest path to the extended region based on whether the corresponding hottest path meets a criterion.
[0111] Example 18 includes the one or more non-transitory machine-readable storage media of any one of Examples 11, 12, 13, 14, 15, 16, and 17. In Example 18, the final region is a first region. The instructions further cause the at least one processor to merge a first transformation generated based on the first region with a second transformation generated based on the second region. In Example 18, the merging of the first transformation and the second transformation forms a super region.
[0112] Example 19 includes the one or more non-transitory machine-readable storage media of Example 18. In Example 19, the instructions further cause the one or more processors to identify a hot transition, generate a flow graph including the hot transition, a predecessor transition, and a successor transition of the hot transition, and select at least one of the predecessor transition and the successor transition for merging with the hot transition. In Example 19, the hot transition is the first transition, and the at least one of the predecessor transition and the successor transition is the second transition.
[0113] Example 20 is a method for forming a region, the method comprising forming an initial region by executing instructions using a processor. The initial region starts at a first hot code block of a control flow graph, and includes a hot code block located on the first hottest path extending from the first hot code block until one of the following situations: a hot code block previously added to the initial region is reached, or a cold code block is reached. The method of Example 20 also includes extending the initial region by executing instructions using the processor to form an extended region. The extended region includes the initial region. Extending the region starts at the hottest exit of the initial region, and includes adding a hot code block located on the second hottest path of the control flow graph until a threshold path length has been met or a hot code block associated with a back edge of the control flow graph is added to the extended region. The method of Example 20 also includes pruning the extended region by executing instructions using the processor to form a single loop. In Example 20, the single loop is a first region.
[0114] Example 21 includes the method of Example 20, and further includes: identifying loops and loop nests included in the expansion region; generating features of the loops and loop nests; and selecting one of the loops as the single loop based on the features.
[0115] Example 22 includes the method of Example 20, and further includes: identifying a code block included in the control flow graph as hot based on profiling; adding the name of the code block identified as hot to a queue; and generating a trigger when the queue includes a threshold number of code blocks. In Example 22, the forming of the initial region is initiated in response to the trigger, and the threshold number is greater than 1.
[0116] Example 23 includes the method of Example 21. In the method of Example 23, the first hot code block is the last code block added to the queue before the trigger is generated.
[0117] Example 24 includes the method of any one of Examples 20 to 23. In Example 24, the threshold path length is a first threshold path length, and the method of Example 24 further includes: before adding the code block located on the second hottest path to the region, determining whether at least one of the following situations exists in the second hottest path: connecting back to a previously formed portion of the initial region or the extended region, or leading to a second hot code block within a second threshold path length. The method of Example 24 also includes adding the code block located on the second hottest path to the extended region based on the determination.
[0118] Example 25 includes the method of any one of Examples 20 to 23. The method of Example 25 further includes merging a first transformation generated based on the first region with a second transformation generated based on the second region. In Example 25, the merging of the first transformation and the second transformation forms a super region.
[0119] Example 26 includes the method of Example 25, and further includes: identifying a thermal transition; generating a flow graph including the thermal transition, a predecessor transition, and a successor transition of the thermal transition; and selecting at least one of the predecessor transition and the successor transition for merging with the thermal transition. In Example 26, the thermal transition is the first transition, and the at least one of the predecessor transition and the successor transition is the second transition.
[0120] Example 27 includes the method of any one of Examples 20 and 21, and further includes: identifying a code block included in the control flow graph as hot based on profiling; adding the name of the code block identified as hot to a queue; and generating a trigger when the queue includes a threshold number of code blocks. In Example 27, the forming of the initial region is initiated in response to the trigger, and the threshold number is greater than 1.
[0121] Example 28 includes the method of any one of Examples 20-23 and 27. In Example 28, the threshold path length is a first threshold path length, and the method of Example 28 further includes: before adding the code block located on the second hottest path to the region, determining whether the second hottest path has at least one of the following situations: connecting back to a previously formed portion of the initial region or the extended region, or leading to a second hot code block within a second threshold path length. The method of Example 28 also includes adding the code block located on the second hottest path to the extended region based on the determination.
[0122] Example 29 includes the method of any of Examples 20-23 and 28, and further includes merging a first transformation generated based on the first region with a second transformation generated based on the second region. In Example 29, the merging of the first transformation and the second transformation forms a super region.
[0123] Example 30 includes the method of Example 29, and further includes: identifying a thermal transition; generating a flow graph including the thermal transition, a predecessor transition, and a successor transition of the thermal transition; and selecting at least one of the predecessor transition and the successor transition for merging with the thermal transition. In Example 30, the thermal transition is the first transition, and the at least one of the predecessor transition and the successor transition is the second transition.
[0124] Example 31 is an apparatus comprising means for performing the method of any of Examples 20-23 and 27-30.
[0125] Example 32 is a machine readable storage comprising machine readable instructions. The instructions, when executed, implement the method of any one of Examples 20-23 and 27-30.
[0126] Example 33 is a device for performing region formation of a control flow graph. The device of Example 33 includes a device for forming an initial region starting from the first hot code block of the control flow graph. The device for forming the initial region adds a hot code block located on the first hottest path of the control flow graph to the initial region until reaching one of the following: a hot code block previously added to the initial region, or a cold code block. The device of Example 33 also includes a device for extending the initial region to form an extended region including the initial region. The device for extending the initial region begins to extend the initial region at the hottest exit of the initial region. In addition, the device for extending the initial region adds a hot code block located on the second hottest path to the extended region until one of the following situations: a threshold path length is met, or a back edge of the control flow graph is added to the extended region. The device of Example 33 also includes a device for pruning all loops except the selected loop from the extended region. In Example 33, the selected loop forms the region.
[0127] Example 34 includes the apparatus of Example 33, and further includes: means for identifying and generating characteristics of a plurality of loops and loop nests included in the extended region; and means for selecting one of the loops identified by the loop analyzer based on a set of criteria. In Example 34, the set of criteria is applied to the characteristics by the means for selecting one of the loops.
[0128] Example 35 includes the apparatus of Example 33, and further includes means for storing block identifiers of a plurality of code blocks of the control flow graph identified as hot. In addition, the apparatus of Example 35 includes means for generating a trigger when the means for storing the block identifiers includes a threshold number of the code blocks. In Example 46, the means for forming the initial region initiates formation of the initial region in response to the trigger. In addition, the threshold number is greater than 1.
[0129] Example 36 includes the apparatus of Example 35. In Example 36, the first hot code block is a last code block added to the means for storing the block identifier before generating the trigger.
[0130] Example 37 includes the apparatus of any one of Examples 35 and 36. In Example 37, the means for forming the initial region and the means for extending the initial region use the means for storing a block identifier to determine whether a code block is hot.
[0131] Example 38 includes the apparatus of any one of Examples 33, 34, and 35. In Example 38, the threshold path length is a first threshold path length, and before adding the second hottest path to the extended region, the means for extending the initial region determines whether the second hottest path at least one of: connects back to a block in the extended region, or leads to a second hot code block within a second threshold path length. In Example 38, based on the determination, the means for extending the initial region adds the code block on the second hottest path to the extended region.
[0132] Example 39 includes the apparatus of Example 38. In Example 39, the apparatus for expanding the initial region determines whether the second hottest path leads to the second hot code block within the second threshold path length, and whether the second hottest path includes a cold code block between the hottest exit of the initial region and the second hot code block.
[0133] Example 40 includes the apparatus of any one of Examples 33, 34, and 35. In Example 40, the means for extending the initial region further iterates over the corresponding hottest exits of the initial region and the extended region in descending order of heat. Additionally, when the corresponding hottest path meets a criterion, the means for extending the initial region adds a code block placed along the corresponding hottest path to the extended region.
[0134] Example 41 includes the apparatus of any one of Examples 33, 34, and 35. In Example 41, the region is a first region, and the apparatus further comprises means for merging a first transformation generated based on the first region and a second transformation generated based on the second region. In Example 41, the merging of the first transformation and the second transformation forms a super region.
[0135] Example 42 includes the apparatus of Example 41. The apparatus of Example 42 also includes: a queue for identifying a hot transition; means for generating a flow graph including the hot transition, a predecessor transition and a successor transition of the hot transition; and means for selecting at least one of the predecessor transition and the successor transition for merging with the hot transition. In Example 42, the hot transition is the first transition, and at least one of the predecessor transition and the successor transition is the second transition.
[0136] Example 43 is a machine-readable medium comprising code, which when executed causes the machine to perform the method of any one of Examples 20 to 26.
[0137] Although certain example methods, apparatus, and articles of manufacture have been disclosed herein, the scope of coverage of this patent is not limited thereto. On the contrary, this patent covers all methods, apparatus, and articles of manufacture that fall within the scope of the claims of this patent.
Claims
1. An apparatus for performing region formation of a control flow graph, the apparatus comprising: an initial region former for forming an initial region starting from a first hot code block of the control flow graph, the initial region former for adding hot code blocks located on a first hottest path of the control flow graph to the initial region until reaching one of: a hot code block previously added to the initial region, or a cold code block; a region expander, the region expander being used to expand the initial region to form an extended region including the initial region, the region expander being used to start expanding the initial region at the hottest exit of the initial region, the region expander being used to add hot code blocks located on the second hottest path to the extended region until one of the following situations occurs: a threshold path length is met, or a back edge of the control flow graph is added to the extended region; as well as A region pruner is used to prune all loops except selected loops from the extended region, the selected loops forming a final region.
2. The device according to claim 1, characterized in that Also includes: a region analyzer for identifying and generating features of a plurality of loops and loop nests contained in the extended region; as well as A cycle selector for selecting one of the cycles identified by the region analyzer based on a set of criteria, the cycle selector applying the set of criteria to the features generated by the region analyzer.
3. The device according to any one of claims 1 and 2, characterized in that The final region is a first region, and the apparatus further includes a super region former, which is used to merge a first transformation generated based on the first region and a second transformation generated based on the second region, wherein the merger of the first transformation and the second transformation is used to form a super region.
4. The device according to claim 3, characterized in that Also includes: a queue, the queue being used to identify a hot transition; A graph generator, the graph generator being used to generate a flow graph including the thermal transition, a predecessor transition and a successor transition of the thermal transition; as well as A transition selector for selecting at least one of the predecessor transition and the successor transition for merging with the thermal transition, the thermal transition being the first transition, and the at least one of the predecessor transition and the successor transition being the second transition.
5. A method for forming a region, the method comprising: An initial region is formed by executing instructions with a processor, the initial region starting at a first hot code block of a control flow graph and including hot code blocks located on a first hottest path extending from the first hot code block until one of the following situations is reached: a hot code block previously added to the initial region is reached, or a cold code block is reached; expanding the initial region by executing instructions with the processor to form an extended region, the extended region including the initial region, the expansion of the initial region starting at a hottest exit of the initial region and including adding hot code blocks located on a second hottest path of the control flow graph until one of the following situations: a threshold path length is met, or a hot code block associated with a back edge of the control flow graph is added to the extended region; as well as The extended region is pruned into a single loop by executing instructions with the processor, the single loop being a final region.
6. The method according to claim 5, characterized in that Also includes: identifying loops and loop nests included in the extension region; generating characteristics of the loops and loop nests; as well as One of the loops is selected based on the feature to become the single loop.
7. The method according to claim 5, characterized in that Also includes: identifying, based on profiling, whether a code block included in the control flow graph is hot; adding the name of the code block identified as hot to a queue; as well as A trigger is generated when the queue includes a threshold number of code blocks, the forming of the initial region is initiated in response to the trigger, and the threshold number is greater than one.
8. The method according to claim 7, characterized in that The first hot code block is the last code block added to the queue before the trigger is generated.
9. The method according to any one of claims 5 to 8, characterized in that The threshold path length is a first threshold path length, the method further comprising: Before adding the code block located on the second hottest path to form the extended region, determining whether the second hottest path has at least one of the following situations: connecting back to a previously formed portion of the initial region or the extended region, or leading to a second hot code block within a second threshold path length; and Based on the determination, the code block located on the second hottest path is added to the extension area.
10. The method according to any one of claims 5 to 8, characterized in that Also includes: The final region is a first region, a first transformation generated based on the first region is merged with a second transformation generated based on the second region, and the merger of the first transformation and the second transformation is used to form a super region.
11. The method according to claim 10, characterized in that Also includes: Identify thermal transitions; generating a flow graph including the thermal transition, a preceding transition and a succeeding transition of the thermal transition; as well as At least one of the predecessor and successor transitions is selected for merging with the thermal transition, the thermal transition being the first transition, and the at least one of the predecessor and successor transitions being the second transition.
12. A machine-readable memory comprising machine-readable instructions which, when executed, implement the method of any one of claims 5 to 8.
13. An apparatus for performing region formation of a control flow graph, the apparatus comprising: means for forming an initial region starting at a first hot code block of the control flow graph, the means for forming the initial region being used to add hot code blocks located on a first hottest path of the control flow graph to the initial region until reaching one of: a hot code block previously added to the initial region, or a cold code block; means for extending the initial region to form an extended region including the initial region, the means for extending the initial region being used to start extending the initial region at the hottest exit of the initial region, the means for extending the initial region being used to add hot code blocks located on the second hottest path to the extended region until one of the following situations occurs: a threshold path length is met, or a back edge of the control flow graph is added to the extended region; as well as Means for pruning all loops from the expanded region except for selected loops, the selected loops forming a final region.
14. The device according to claim 13, characterized in that Also includes: means for identifying and generating features of a plurality of loops and loop nests contained in said extended region; as well as Means for selecting the identified one of the plurality of cycles based on a set of criteria, the means for selecting the one of the plurality of cycles applying the set of criteria to the feature.
15. The device according to claim 13, characterized in that Also includes: means for storing block identifiers of a plurality of code blocks of the control flow graph identified as hot; means for generating a trigger when the means for storing the block identifiers includes a threshold number of the code blocks, the means for forming the initial region being configured to initiate formation of the initial region in response to the trigger, and the threshold number is greater than one.
16. The device according to claim 15, characterized in that The first hot code block is the last code block added to the means for storing the block identifier before generating the trigger.
17. The device according to any one of claims 15 and 16, characterized in that The means for forming the initial region and the means for extending the initial region determine whether a code block is hot using the means for storing a block identifier.
18. The device according to any one of claims 13 to 15, characterized in that The threshold path length is a first threshold path length, and before adding the second hottest path to form the extended region, the means for extending the initial region is further configured to: determining whether the second hottest path at least one of the following: connects back to a block in the extended region, or leads to a second hot code block within a second threshold path length; as well as Based on the determination, the code block on the second hottest path is added to the extension region.
19. The device according to claim 18, characterized in that The means for expanding the initial region determines that the second hottest path leads to the second hot code block within the second threshold path length, and the second hottest path includes a cold code block between the hottest exit of the initial region and the second hot code block.
20. The device according to any one of claims 13, 14 and 15, characterized in that The device for expanding the initial area is further used for: Iterate over the corresponding hottest outlets of the initial region and the expansion region in descending order of heat; as well as When the path starting from the hottest exit meets the criteria, a code block placed along the path starting from the hottest exit is added to the extension area.
21. The device according to any one of claims 13, 14 and 15, characterized in that The final region is a first region, and the device further comprises a device for merging a first transformation generated based on the first region and a second transformation generated based on the second region, wherein the merging of the first transformation and the second transformation is used to form a super region.
Citation Information
Patent Citations
A lightweight service based dynamic binary rewriter framework
CN102483700A
Hot route searching method in assembly code hot function
CN1783009A