Missing program filling processing method, electronic device and computer-readable storage medium
By analyzing the dominator tree and post-dominator tree of the program, splitting it into super blocks and rewriting jump instructions into DMA calls, the problem of low hardware filling efficiency is solved and the program execution efficiency and resource utilization are improved.
Patent Information
- Application Number
- CN202511105959.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-08
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-08-08
AI Technical Summary
In the prior art, the hardware filling efficiency of missing pages is low, resulting in low program execution efficiency.
By building dominator trees and post-dominator trees, the program is split and merged to form super blocks, and jump instructions are rewritten into instruction-filled function call sequences to trigger DMA and optimize the program execution path.
It improves the execution efficiency and resource utilization of the program, reduces branch jumps, and realizes the parallelization of memory access and calculation.
Smart Images

Figure CN120596152B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of software security technology, and in particular to a missing program filling processing method, electronic equipment and computer-readable storage medium. Background Art
[0002] Modern operating systems provide each process with an independent virtual address space through virtual memory management. Virtual addresses must be mapped to physical memory pages before they can be accessed.
[0003] When the computer's processing core executes a program, if the page containing the virtual address accessed by the program has not yet been loaded into physical memory (for example, the page table entry is marked as "not present"), or if there is insufficient permission (for example, an attempt to write to a read-only page), a page fault will be thrown and an exception will be triggered. The missing page needs to be processed.
[0004] In the prior art, when a large program is running, the missing pages in the memory are usually filled by hardware, which is inefficient. Summary of the Invention
[0005] Embodiments of the present invention provide a missing program filling processing method, an electronic device, and a computer-readable storage medium to solve the problem of low efficiency of missing program filling using hardware in the prior art.
[0006] In a first aspect, an embodiment of the present invention provides a method for filling a missing program, comprising:
[0007] Establish a dominator tree and a post-dominator tree according to the initial procedure;
[0008] The initial program is split into multiple super blocks according to the post-dominator tree;
[0009] Splitting and / or merging the multiple super blocks according to the post-dominator tree and the dominator tree to obtain multiple updated super blocks; wherein the size of each super block is not greater than a preset size;
[0010] Rewrite the jump instructions between each super block into a function call sequence filled with instructions; wherein the function call sequence filled with instructions is used to trigger DMA;
[0011] Determine whether each super block after rewriting is larger than a preset size;
[0012] If there is a super block larger than the preset size, the process jumps to the step of splitting and / or merging the multiple super blocks according to the post-dominator tree and the dominator tree to obtain the updated multiple super blocks;
[0013] Otherwise, for any super block, adjust the order of the basic blocks in the super block in the program so that the basic blocks in the super block are connected in memory;
[0014] The size of each super block is obtained and stored in a first preset area in the memory.
[0015] In a second aspect, an embodiment of the present invention provides an electronic device comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, it implements the missing program filling processing method as described in the first aspect or any possible implementation of the first aspect.
[0016] In a third aspect, an embodiment of the present invention provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the missing program filling processing method in the first aspect or any possible implementation of the first aspect.
[0017] The embodiment of the present invention provides a missing program filling processing method, electronic device and computer-readable storage medium. The above method includes: establishing a dominator tree and a post-dominator tree according to an initial program; splitting the initial program according to the post-dominator tree to obtain multiple super blocks; splitting and / or merging the multiple super blocks according to the post-dominator tree and the dominator tree to obtain multiple updated super blocks; wherein the size of each super block is not greater than a preset size; rewriting the jump instructions between each super block into a function call sequence filled with instructions; wherein the function call sequence filled with instructions is used to trigger DMA; determining whether each rewritten super block is greater than a preset size; if there is a super block greater than the preset size, jumping to the step of splitting and / or merging the multiple super blocks according to the post-dominator tree and the dominator tree to obtain multiple updated super blocks; otherwise, for any super block, adjusting the order of the basic blocks in the super block in the program so that the basic blocks in the super block are connected in the memory; obtaining the size of each super block and storing it in a first preset area in the memory. This application reasonably divides the initial program into multiple super blocks to reduce the branch jumps in the traditional control flow graph; at the same time, it pre-fetches data through DMA before jumping, uses the instruction filling sequence to allow the CPU to continue to execute other tasks, realizes the parallelism of memory access and calculation, and effectively improves the execution efficiency of the program. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 This is a flow chart of a missing program filling processing method provided by an embodiment of the present invention;
[0019] Figure 2 This is a flowchart of another missing program filling processing method provided by an embodiment of the present invention;
[0020] Figure 3 1 is a schematic structural diagram of a missing program filling processing device provided by an embodiment of the present invention;
[0021] Figure 4is a schematic diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0022] The embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0023] See also Figure 1 and Figure 2 , which shows a flowchart of the implementation of the missing program filling processing method provided by an embodiment of the present invention, and is described in detail as follows:
[0024] The missing program filling processing method includes:
[0025] S101: Establish a dominator tree and a post-dominator tree according to the initial procedure;
[0026] Control flow analysis is a crucial technique in compiler optimization and code analysis. By understanding and representing all possible control flow paths in a program, it assists the compiler in making improvements and enhancements in various areas. This includes optimizing program execution efficiency, detecting potential errors in the code, and analyzing the behavior of concurrent programs. The core work of control flow analysis primarily involves constructing a control flow graph (CFG), deeply analyzing the structure and relationships of basic blocks, and effectively detecting and handling various loop structures in the program.
[0027] In a control flow graph (CFG), for nodes A and B, if every path from the entry node (program entry) to node B passes through node A, node A dominates node B and node A is the dominant node of node B. If node A is the dominant node of node B and there is no other dominant node, then node A is the direct dominant node of node B;
[0028] With the entry node as the root, each node's parent is its direct dominant node, forming a tree structure called a dominator tree. In a dominator tree, except for the entry node, each node has exactly one direct dominant node. The dominator tree reflects the dominance hierarchy between nodes, with the entry node as the root.
[0029] For nodes A and B, if all paths from node B to all exit nodes must pass through node A, then node A is called the post-dominator of node B. If node A is the post-dominator of node B and there are no other post-dominators, then node A is the direct post-dominator of node B.
[0030] A post-dominator tree is a tree structure with the exit node (program exit) as the root. Each node's parent is its immediate post-dominator. Except for the exit node, each node has exactly one immediate post-dominator. The post-dominator tree reflects the necessary paths from a node to the exit.
[0031] Control flow analysis and dominance relationships can accurately analyze a program's execution paths and dependencies. Within a control flow graph (CFG), control flow analysis can identify the execution order and conditions of each basic block in a program, thereby constructing an overall control flow model for the program. A dominance tree is used to characterize the required path from the entry point, while a post-dominance tree is used to characterize the required path to the exit point. Both construct a tree structure using "direct dominance / post-dominance nodes" and are key tools for understanding the control flow hierarchy in program analysis. By analyzing dominance and post-dominance relationships, the compiler can determine which blocks in the program are nodes on the critical path and which blocks are executed more frequently during program execution.
[0032] Based on this, the present application establishes the dominator tree and post-dominator tree of the initial program for control flow analysis.
[0033] S102: Split the initial program according to the post-dominator tree to obtain multiple super blocks;
[0034] A superblock is a structured code region used in program control flow analysis. It aims to simplify program structure and improve compilation optimization efficiency by merging continuous linear code fragments. It is a larger code unit than a basic block, with one and only one entry node, which serves as the starting point for the entire superblock and is controlled by all other nodes within the superblock (that is, the path from any node within the superblock to the program exit must pass through the entry node). A superblock has one exit node, and control flow flows linearly from entry to exit, with no internal branches or loops (the exit node can connect to multiple subsequent nodes, but the execution path within the superblock is unique). Nodes within a superblock are executed sequentially, with no conditional branches or jumps occurring in between (unless jumping to the exit node).
[0035] By integrating linear code segments, super blocks abstract the control flow of the program into a simple "inter-block jump" model, which can balance code locality and global optimization capabilities, making the analysis and conversion of complex programs more efficient. It is an important foundation for compiler optimization and high-performance code generation.
[0036] Based on this, the present application uses a post-dominator tree to split the initial program into multiple super blocks to optimize the overall efficiency and performance of the program.
[0037] The specific splitting process is described in detail as follows:
[0038] In one possible implementation, reference Figure 2 , S102 includes:
[0039] S1021: If the initial program is not larger than the preset size, exit;
[0040] In order to avoid complex splitting of programs that are too small, this application sets a preset threshold (preset size). If the program size (for example, the number of instructions) is not greater than the preset threshold, it means that the program is very small and there is no need to split it, thus avoiding complex splitting and saving compilation resources.
[0041] For example, all basic blocks in the program can be traversed and the number of instructions recorded to obtain the size of the initial program. Some instructions require special processing, such as skipping debugging-related instructions and counting the number of instructions contained in the string representation of the inline assembly.
[0042] S1022: Otherwise, the path from the starting block to the ending block on the dominator tree is determined to be the target path; wherein the starting block is the program entry and the ending block is the program exit;
[0043] On the post-dominator tree, there is only one path from the starting block (entry block) to the ending block (exit block) along the post-dominator relationship.
[0044] S1023: Split the initial program into multiple super blocks using each node on the target path as a boundary.
[0045] The target path typically represents the main control flow of a program (such as the main function call chain). Therefore, this application treats each node on the target path as a split point (boundary) for splitting. The split super block still maintains the continuity of this path, facilitating optimization. At the same time, compared to indiscriminate splitting, this method focuses on the critical path, avoiding the metadata overhead caused by excessive splitting, and is suitable for optimizing large-scale programs. This application splits the program by target path nodes, converting complex control flow graphs into multiple linear execution regions. This not only preserves the critical control flow of the program, but also decomposes complex problems into local optimization problems. It is an efficient compilation optimization strategy.
[0046] In a possible implementation, S1023 may include:
[0047] 1. For any node in the target path, take the node as the head basic block and the direct parent node of the node in the post-dominator tree as the tail basic block; start from the head basic block and perform a breadth-first search on the successor nodes of the head basic block until the tail basic block. The head basic block and all basic blocks searched except the tail basic block form the super block corresponding to the node.
[0048] A breadth-first search hierarchically traverses all reachable nodes, ensuring that the superblock contains all possible paths from the head basic block to the tail basic block. This systematically constructs a superblock that complies with the post-dominance constraint. Furthermore, since the tail basic block is the parent node on the post-dominator tree and serves as the entry point for the next superblock, this application ensures that the exit node of each superblock is the tail basic block, but does not include the tail basic block itself.
[0049] S103: splitting and / or merging the multiple super blocks according to the post-dominator tree and the dominator tree to obtain multiple updated super blocks; wherein the size of each super block is not greater than a preset size;
[0050] If the super block is split too large, the optimization granularity will be too coarse, while if it is too small, the number and time of program filling will increase; based on this, the present application splits and / or merges the super block through the dominator tree / post-dominator tree, so that each super block is close to the preset size, reducing the number of filling times and improving the overall execution efficiency of the program.
[0051] In a possible implementation, S103 may include:
[0052] S1031: Determine whether each super block is no larger than a preset size;
[0053] S1032: If there is a super block larger than a preset size, split the super block larger than the preset size according to the dominator tree and the post-dominator tree so that each super block after the split is no larger than the preset size;
[0054] The core of splitting the super block in S1032 is to utilize the structural characteristics of the dominator tree and the post-dominator tree to split the overly large super block into multiple sub-blocks that meet the size limit while ensuring the correctness of the control flow.
[0055] Based on this, in a possible implementation, refer to Figure 2 , S1032 may include:
[0056] 1. Use a superblock larger than a preset size as the first target superblock;
[0057] 2. Determine whether there is a jump relationship between the head basic block and the tail basic block of the first target super block;
[0058] 3. If a jump relationship exists, remove the jump relationship and determine whether the direct parent node of the head basic block in the first target super block is still the tail basic block after removing the jump relationship;
[0059] Jump relationships can cause loops or complex branches within a superblock, disrupting the linear structure. Therefore, in this application, if there is a jump relationship between the head and tail basic blocks, it indicates that the superblock may contain a loop structure, and you can try to remove it to expose more splitting opportunities.
[0060] It should be noted that the jump relationship is only temporarily removed and will be restored after the split is completed.
[0061] 4. If, after removing the jump relationship, the direct parent node of the head basic block is not the tail basic block, then perform the first operation to split the first target super block to obtain at least two super blocks;
[0062] If, after removing the jump relationship, the direct post-dominant node of the head basic block is no longer the tail basic block (i.e., the post-dominant tree structure changes), it means that the jump is a key edge of the loop and can be split using the first operation.
[0063] In a possible implementation, performing the first operation to split the first target super block to obtain at least two super blocks may include:
[0064] (1) In the first target super block, determine the path from the head basic block to the tail basic block, and split the first target super block with each node on the path as the boundary to obtain at least two super blocks.
[0065] This application can use the same method of splitting the initial program, recursively splitting.
[0066] 5. If there is no jump relationship, or if the direct parent node of the head basic block in the first target super block is still the tail basic block after removing the jump relationship, determine whether the tail basic block has any child node other than the head basic block in the post-dominator tree, and if the child node is in the first target super block.
[0067] If there is no jump relationship, or the direct parent node of the head basic block in the first target super block is still the tail basic block after removing the jump relationship, the structural characteristics of the post-dominator tree are used to find more splitting possibilities for splitting.
[0068] 6. If it exists, perform the second operation to split the first target super block to obtain at least two super blocks;
[0069] If the tail basic block has other child nodes in the post-dominator tree and is within the current super block, it indicates that a parallel path exists, and the second operation can be used for splitting.
[0070] In a possible implementation, performing the second operation to split the first target super block to obtain at least two super blocks may include:
[0071] (1) Determine the child nodes of the tail basic block in the first target super block on the post-dominator tree;
[0072] If the tail basic block has no child nodes, it may be because there are other nodes in the superblock without successor basic blocks. You can try connecting these dangling nodes to the tail node, rebuilding the post-dominator tree, and then searching for the tail basic block's children. If the tail basic block has only one child in the post-dominator tree, use the above conditions to search for the tail basic block's child's children... until you find multiple child nodes.
[0073] (2) The sub-node closest to the average split among all sub-nodes is selected as the target sub-node;
[0074] (3) Split the first target super block into two super blocks with the target child node as the boundary.
[0075] If the tail basic block has other child nodes besides the head basic block in the post-dominator tree, a breadth-first search is performed on the subsequent basic blocks starting from the candidate child nodes, and a new super block is obtained with the tail basic block as the end point. The remaining part of the original super block forms another super block, which is the complement of the other super block. Among all the splits, a split method that is closest to an average split is selected to reduce the number of splits. After the split is completed, two new super blocks are obtained, and all control flows across super blocks between the two are removed. Here, it is also preferable to keep the complement super block within the legal size, because this split may produce dangling basic blocks in the complement super block, making the tail block of the complement super block unable to post-dominate the head block, increasing the difficulty of subsequent splitting.
[0076] 7. If not, determine whether the head basic block has any child nodes other than the tail basic block on the dominator tree, and the child nodes are within the first target super block;
[0077] 8. If it exists, perform the third operation to split the first target super block to obtain at least two super blocks;
[0078] If it does not exist, the third operation is used to split it using the structural characteristics of the dominator tree.
[0079] 9. If it does not exist, a breadth-first search is performed starting from the beginning basic block. When the searched basic blocks reach the preset size, the searched basic blocks are formed into a super block, and the basic blocks that have not been searched are formed into a second super block.
[0080] 10. Determine whether each superblock is no larger than the preset size;
[0081] 11. If there is a super block larger than the preset size, jump to the step of using the super block larger than the preset size as the first target super block and continue execution.
[0082] Since the super block after the first target super block is split may be larger than the preset size, after one split is completed, it is re-determined whether each super block meets the size limit. If not, the super block larger than the preset size can be further split. Through cyclic splitting, each super block can meet the size limit.
[0083] refer to Figure 2 ,This application adopts a series of heuristic algorithms to analyze the nodes on the ,critical path and avoid jumps between high execution frequency blocks, ,and obtain a better splitting solution.
[0084] S1033: Merge each super block to obtain updated super blocks; wherein each updated super block is no larger than a preset size.
[0085] In this application, the oversized superblocks are first split so that each superblock is smaller than the preset size. However, the splitting may produce blocks that are too small, so the superblocks need to be merged again.
[0086] In a possible implementation, S1033 may include:
[0087] 1. Sort each superblock in the order in which the program is executed;
[0088] 2. Use the last superblock as the second target superblock;
[0089] 3. Finding the first number and the second number such that the sum of the sizes of the second target super block and the first number of super blocks preceding the second target super block is no greater than a predetermined size, and the sum of the sizes of the second target super block and the second number of super blocks preceding the second target super block is greater than a predetermined size; wherein the second number is equal to the first number plus 1;
[0090] 4. Merge the second target super block with the first number of super blocks preceding the second target super block to obtain a new super block.
[0091] 5. Use the previous super block of the new super block as the new second target super block, and jump to the step of searching for the first quantity and the second quantity, so that the sum of the sizes of the second target super block and the first number of super blocks before the second target super block is not greater than the preset size, and the sum of the sizes of the second target super block and the second number of super blocks before the second target super block is greater than the preset size, and continue to execute the step until the new super block includes the first super block.
[0092] Under the premise of meeting the size limit, this application merges adjacent super blocks from back to front until the first super block is also merged, resulting in multiple super blocks of a size close to the preset size. The fusion is performed from the end of the execution sequence to ensure that the code executed later is optimized first.
[0093] In this application, oversized superblocks are first split, so that each superblock is smaller than the preset size. However, this splitting may produce overly small blocks, so the superblocks need to be merged to avoid excessive splitting causing an overflow of small blocks, which increases the number of subsequent program fills and increases the fill time.
[0094] S104: rewriting the jump instructions between the super blocks into a function call sequence filled with instructions; wherein the function call sequence filled with instructions is used to trigger DMA;
[0095] Direct Memory Access (DMA) is a computer system feature that enables peripheral devices to directly access system memory without requiring data transfer through the CPU. In traditional data transfer methods, the CPU is responsible for reading data from a storage device and writing it to memory. This approach not only consumes CPU resources but also reduces overall system efficiency. However, with DMA, peripheral devices can autonomously read and write data directly from or to memory without CPU intervention. A DMA controller manages and coordinates this process. When a data transfer is required, the DMA controller first sends a request to the CPU. Once the CPU approves it, the DMA controller takes over bus control and performs the data transfer. Once completed, the DMA controller returns control to the CPU. This allows the CPU to continue performing other tasks while DMA is performing data transfers, significantly improving system efficiency and performance.
[0096] The jump instruction rewriting between super blocks is mainly used to trigger direct memory access (DMA) operations. After starting DMA in the instruction fill sequence, the CPU can continue to execute the fill instructions while DMA completes the data transfer in the background, thereby "hiding" the memory access delay and improving the execution efficiency of the program.
[0097] Jump instructions between superblocks typically come from two sources: first, jump instructions from all basic blocks within a superblock to the last basic block; and second, control flow removed from the algorithm to allow for more splits or after a split. Jump instructions are replaced by a function sequence filled with instructions, whose inputs are the current program counter (PC) value and the offset of the address to jump to relative to the current PC.
[0098] The specific process includes:
[0099] 1. When the program is initialized, the offset of the current super block relative to the entire program is initialized to D=0;
[0100] 2. The program runs normally;
[0101] 3. When a jump across super blocks is encountered, an automatic reloading procedure is triggered. The input of this procedure is the program address at the time of triggering and the offset of the target relative to the current address.
[0102] 4. Read D and calculate the absolute address of the jump A=D+x;
[0103] 5. The offsets of all superblocks relative to the program are stored in a separate memory area. The jump address is compared with the offset one by one to find the superblock where the jump address is located. The superblock is replaced in the instruction memory and D is set to the offset of the superblock.
[0104] 6. After the replacement is completed, the program counter is set to AD and the program execution is resumed;
[0105] S105: Determine whether each super block after rewriting is larger than a preset size;
[0106] The size of the super block may change due to instruction padding. To avoid overly large super blocks affecting cache efficiency or memory layout, if there are super blocks larger than the preset size, they can be split and / or merged again.
[0107] S106: If there is a super block larger than the preset size, the process jumps to the step of splitting and / or merging the multiple super blocks according to the post-dominator tree and the dominator tree to obtain multiple updated super blocks;
[0108] S107: Otherwise, for any super block, adjust the order of the basic blocks in the super block in the program so that the basic blocks in the super block are connected in the memory;
[0109] For each superblock, the order of its internal basic blocks is adjusted so that they are stored contiguously in physical memory. If the relative positions of two basic blocks without jump instructions change after the adjustment, the corresponding jump or filler instructions need to be added.
[0110] Furthermore, if the size of the super block after the supplementary jump or filling instruction is larger than the preset size, it is necessary to jump to the step of splitting and / or merging multiple super blocks according to the post-dominant tree and the dominant tree to obtain the updated multiple super blocks and continue to adjust the size of the super block.
[0111] S108: Obtain the size of each super block and store it in a first preset area in the memory.
[0112] To obtain the size of each super block, the sizes of each super block can be accumulated to form an array, representing the offset of each super block in memory relative to the program 0 address. This array is stored in the first preset area for runtime call.
[0113] In the embodiment of the present invention, the initial program is reasonably divided into multiple super blocks, reducing the branch jumps in the traditional control flow graph; at the same time, data is pre-fetched through DMA before the jump, and the instruction filling sequence is used to allow the CPU to continue to execute other tasks, realizing the parallelization of memory access and calculation, and effectively improving the execution efficiency and resource utilization of the program.
[0114] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0115] The following are device embodiments of the present invention. For details not fully described therein, reference may be made to the corresponding method embodiments described above.
[0116] Figure 3 A schematic diagram of the structure of a missing program filling processing device provided by an embodiment of the present invention is shown. For ease of explanation, only the parts related to the embodiment of the present invention are shown, which are described in detail as follows:
[0117] like Figure 3 As shown, the missing program filling processing device includes:
[0118] a dominator tree building module 21 for building a dominator tree and a post-dominator tree according to an initial program;
[0119] A super block splitting module 22 is used to split the initial program into multiple super blocks according to the post-dominator tree;
[0120] a superblock optimization module 23 for splitting and / or merging the plurality of superblocks according to the post-dominator tree and the dominator tree to obtain a plurality of updated superblocks; wherein the size of each superblock is not greater than a preset size;
[0121] The instruction rewriting module 24 is used to rewrite the jump instructions between the super blocks into a function call sequence filled with instructions; wherein the function call sequence filled with instructions is used to trigger DMA;
[0122] A super block size identification module 25 is used to determine whether each super block after rewriting is larger than a preset size;
[0123] A first determination module 26 is configured to jump to the step of splitting and / or merging multiple super blocks according to the post-dominator tree and the dominator tree to obtain multiple updated super blocks if a super block larger than a preset size exists;
[0124] A second judgment module 27 is configured to adjust the order of the basic blocks in any super block in the program so that the basic blocks in the super block are connected in the memory;
[0125] The size storage module 28 is used to obtain the size of each super block and store it in a first preset area in the memory.
[0126] In one possible implementation, the super block splitting module 22 may include:
[0127] A program size determination unit, configured to exit if the initial program is not larger than a preset size;
[0128] a target path determination unit, configured to otherwise determine the path from the start block to the end block on the dominator tree as the target path; wherein the start block is the program entry and the end block is the program exit;
[0129] The first splitting unit is used to split the initial program into multiple super blocks based on each node on the target path as a boundary.
[0130] In a possible implementation, the first splitting unit may be specifically configured to:
[0131] 1. For any node in the target path, take the node as the head basic block and the direct parent node of the node in the post-dominator tree as the tail basic block; start from the head basic block and perform a breadth-first search on the successor nodes of the head basic block until the tail basic block. The head basic block and all basic blocks searched except the tail basic block form the super block corresponding to the node.
[0132] In a possible implementation, the first determination module 26 may include:
[0133] A size determination unit, configured to determine whether each super block is no larger than a preset size;
[0134] a second splitting unit, configured to, if a super block is larger than a preset size, split the super block larger than the preset size according to the dominator tree and the post-dominator tree, so that each super block after the split is no larger than the preset size;
[0135] The fusion unit is used to fuse the super blocks to obtain updated super blocks; wherein the updated super blocks are no larger than a preset size.
[0136] In a possible implementation manner, the second splitting unit may be specifically configured to:
[0137] 1. Use a superblock larger than a preset size as the first target superblock;
[0138] 2. Determine whether there is a jump relationship between the head basic block and the tail basic block of the first target super block;
[0139] 3. If a jump relationship exists, remove the jump relationship and determine whether the direct parent node of the head basic block in the first target super block is still the tail basic block after removing the jump relationship;
[0140] 4. If, after removing the jump relationship, the direct parent node of the head basic block is not the tail basic block, then perform the first operation to split the first target super block to obtain at least two super blocks;
[0141] 5. If there is no jump relationship, or if the direct parent node of the head basic block in the first target super block is still the tail basic block after removing the jump relationship, determine whether the tail basic block has any child node other than the head basic block in the post-dominator tree, and if the child node is in the first target super block.
[0142] 6. If it exists, perform the second operation to split the first target super block to obtain at least two super blocks;
[0143] 7. If not, determine whether the head basic block has any child nodes other than the tail basic block on the dominator tree, and the child nodes are within the first target super block;
[0144] 8. If it exists, perform the third operation to split the first target super block to obtain at least two super blocks;
[0145] 9. If it does not exist, a breadth-first search is performed starting from the beginning basic block. When the searched basic blocks reach the preset size, the searched basic blocks are formed into a super block, and the basic blocks that have not been searched are formed into a second super block.
[0146] 10. Determine whether each superblock is no larger than the preset size;
[0147] 11. If there is a super block larger than the preset size, jump to the step of using the super block larger than the preset size as the first target super block and continue execution.
[0148] In a possible implementation, performing the first operation to split the first target super block to obtain at least two super blocks may include:
[0149] (1) In the first target super block, determine the path from the head basic block to the tail basic block, and split the first target super block with each node on the path as the boundary to obtain at least two super blocks.
[0150] In a possible implementation, performing the second operation to split the first target super block to obtain at least two super blocks may include:
[0151] (1) Determine the child nodes of the tail basic block in the first target super block on the post-dominator tree;
[0152] (2) The sub-node closest to the average split among all sub-nodes is selected as the target sub-node;
[0153] (3) Split the first target super block into two super blocks with the target child node as the boundary.
[0154] In one possible implementation, the fusion unit may be specifically used to:
[0155] 1. Sort each superblock in the order in which the program is executed;
[0156] 2. Use the last superblock as the second target superblock;
[0157] 3. Finding the first number and the second number such that the sum of the sizes of the second target super block and the first number of super blocks preceding the second target super block is no greater than a predetermined size, and the sum of the sizes of the second target super block and the second number of super blocks preceding the second target super block is greater than a predetermined size; wherein the second number is equal to the first number plus 1;
[0158] 4. Merge the second target super block with the first number of super blocks preceding the second target super block to obtain a new super block.
[0159] 5. Use the previous super block of the new super block as the new second target super block, and jump to the step of searching for the first quantity and the second quantity, so that the sum of the sizes of the second target super block and the first number of super blocks before the second target super block is not greater than the preset size, and the sum of the sizes of the second target super block and the second number of super blocks before the second target super block is greater than the preset size, and continue to execute the step until the new super block includes the first super block.
[0160] Figure 4 Schematic diagram of an electronic device provided by an embodiment of the present invention. Figure 4 As shown, the electronic device 3 of this embodiment includes a processor 30 and a memory 31. The memory 31 stores a computer program 32. When the processor 30 executes the computer program 32, the steps of the above-described method embodiments are implemented. Alternatively, when the processor 30 executes the computer program 32, the functions of the modules / units in the above-described device embodiments are implemented.
[0161] Exemplarily, the computer program 32 may be divided into one or more modules / units, which are stored in the memory 31 and executed by the processor 30 to implement the present invention. The one or more modules / units may be a series of computer program instruction segments capable of implementing specific functions, and the instruction segments are used to describe the execution process of the computer program 32 in the electronic device 3.
[0162] The electronic device 3 may include, but is not limited to, a processor 30 and a memory 31. Those skilled in the art will appreciate that Figure 4 It is only an example of electronic device 3 and does not constitute a limitation of electronic device 3. It may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, electronic device 3 may also include input and output devices, network access devices, buses, etc.
[0163] The processor 30 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.
[0164] The memory 31 can be an internal storage unit of the electronic device 3, such as the hard drive or memory of the electronic device 3. The memory 31 can also be an external storage device of the electronic device 3, such as a plug-in hard drive, a Smart Media Card (SMC), a Secure Digital (SD) card, a flash memory card, etc. equipped on the electronic device 3. Furthermore, the memory 31 can include both the internal storage unit of the electronic device 3 and an external storage device. The memory 31 is used to store the computer program 32 and other programs and data required by the electronic device 3. The memory 31 can also be used to temporarily store data that has been output or is about to be output.
[0165] For the sake of convenience and brevity, the division of the above functional modules / units is only used as an example. In actual applications, the above functions can be assigned to different functional modules / units as needed. The above modules / units can be implemented in the form of hardware, software, or a combination of hardware and software.
[0166] An embodiment of the present invention further provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the methods in the above-mentioned method embodiments.
[0167] An embodiment of the present invention further provides a computer program product, including a computer program, which, when executed by a processor, implements the methods in the above-mentioned method embodiments.
[0168] The term "computer program" includes computer program code, which may be in source code form, object code form, executable file, or some intermediate form. Computer-readable media may include any entity or device capable of carrying computer program code, recording media, USB flash drives, removable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunications signals, and software distribution media.
[0169] In the above embodiments, the descriptions of each embodiment have their own focus. For parts not described or recorded in detail in one embodiment, please refer to the relevant descriptions of other embodiments. Unless otherwise specified or there is a logical conflict, the terms and / or descriptions between different embodiments are consistent and can be referenced to each other. The technical features of different embodiments can be combined to form new embodiments based on their inherent logical relationships.
[0170] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included in the scope of protection of the present invention.
Claims
1. A missing program filling processing method, characterized in that: include: Establish a dominator tree and a post-dominator tree according to the initial procedure; Splitting the initial program according to the post-dominator tree to obtain multiple super blocks; Splitting and / or merging multiple super blocks according to the post-dominator tree and the dominator tree to obtain multiple updated super blocks; wherein the size of each super block is not greater than a preset size; Rewriting the jump instructions between the super blocks into a function call sequence filled with instructions; wherein the function call sequence filled with instructions is used to trigger DMA; Determining whether each super block after rewriting is larger than the preset size; If a super block larger than the preset size exists, the process jumps to the step of splitting and / or merging the multiple super blocks according to the post-dominator tree and the dominator tree to obtain the updated multiple super blocks and continues to execute; Otherwise, for any super block, adjust the order of the basic blocks in the super block in the program so that the basic blocks in the super block are connected in memory; Obtaining the size of each super block and storing it in a first preset area in the memory; The step of splitting and / or merging the multiple super blocks according to the post-dominator tree and the dominator tree to obtain the updated multiple super blocks includes: Determining whether each super block is no larger than the preset size; If a super block exists that is larger than the preset size, split the super block that is larger than the preset size according to the dominator tree and the post-dominator tree so that each super block after the split is no larger than the preset size; Merging the super blocks to obtain updated super blocks; wherein each updated super block is no larger than the preset size; The merging of the super blocks to obtain updated super blocks includes: Sort each superblock in the order in which the program is executed; Use the last superblock as the second target superblock; searching for a first number and a second number such that the sum of the sizes of the second target super block and the first number of super blocks preceding the second target super block is no greater than the predetermined size, and the sum of the sizes of the second target super block and the second number of super blocks preceding the second target super block is greater than the predetermined size; wherein the second number is equal to the first number plus 1; Merging the second target super block with the first number of super blocks preceding the second target super block to obtain a new super block; The step of using the previous super block of the new super block as the new second target super block and jumping to the step of searching for the first quantity and the second quantity so that the sum of the sizes of the second target super block and the first number of super blocks before the second target super block is not greater than the preset size and the sum of the sizes of the second target super block and the second number of super blocks before the second target super block is greater than the preset size is continued until the new super block includes the first super block.
2. The missing program filling processing method according to claim 1, characterized in that: The initial program is split according to the post-dominator tree to obtain multiple super blocks, including: If the initial program is not larger than the preset size, exit; Otherwise, determining the path from the starting block to the ending block on the post-dominator tree as the target path; wherein the starting block is the program entry and the ending block is the program exit; The initial program is split into multiple super blocks using each node on the target path as a boundary.
3. The missing program filling processing method according to claim 2, characterized in that: The initial program is split into multiple super blocks based on each node on the target path as a boundary, including: For any node in the target path, the node is taken as a head basic block, and the direct parent node of the node on the post-dominator tree is taken as a tail basic block; starting from the head basic block, a breadth-first search is performed on the successor nodes of the head basic block until the tail basic block, and the head basic block and all basic blocks searched except the tail basic block form a super block corresponding to the node.
4. The missing program filling processing method according to claim 1, characterized in that: The step of splitting the super block larger than the preset size according to the dominator tree and the post-dominator tree so that each super block after the splitting is no larger than the preset size includes: Using the super block larger than the preset size as a first target super block; Determining whether there is a jump relationship between the head basic block and the tail basic block of the first target super block; If a jump relationship exists, removing the jump relationship, and determining whether the direct parent node of the head basic block in the first target super block is still the tail basic block after removing the jump relationship; If, after removing the jump relationship, the direct parent node of the head basic block is not the tail basic block, performing a first operation to split the first target super block to obtain at least two super blocks; If no jump relationship exists, or if the direct parent node of the head basic block in the first target super block is still the tail basic block after removing the jump relationship, determining whether the tail basic block has any child node other than the head basic block on the post-dominator tree, and the child node is in the first target super block; If so, performing a second operation to split the first target super block to obtain at least two super blocks; If not, determining whether the head basic block has other child nodes on the dominator tree except the tail basic block, and whether the child nodes are in the first target super block; If so, performing a third operation to split the first target super block to obtain at least two super blocks; If it does not exist, a breadth-first search is performed starting from the head basic block. When the basic blocks searched reach the preset size, the basic blocks searched are formed into a super block, and the basic blocks not searched are formed into a second super block. Determining whether each super block is no larger than the preset size; If there is a super block larger than the preset size, the process jumps to the step of using the super block larger than the preset size as the first target super block and continues to execute.
5. The missing program filling processing method according to claim 4, characterized in that: The performing of the first operation to split the first target super block to obtain at least two super blocks includes: In the first target super block, a path from the head basic block to the tail basic block is determined, and the first target super block is split with each node on the path as a boundary to obtain at least two super blocks.
6. The missing program filling processing method according to claim 4, characterized in that: The performing of the second operation to split the first target super block to obtain at least two super blocks includes: Determine each child node of the tail basic block in the first target super block on the post-dominating tree; The child node closest to the average split among all child nodes is selected as the target child node; The first target super block is split into two super blocks with the target child node as the boundary.
7. An electronic device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and the processor implements the missing program filling processing method according to any one of claims 1 to 6 when executing the computer program.
8. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the missing program filling processing method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Software code processing method and electronic equipment
CN114003868A
Method and system for automatically filling data, electronic equipment and storage medium
CN115391248A