Floating Point Stack Processing Method, Apparatus and Processor
Through multi-threaded parallel processing and preset mapping rules, dynamically map floating point stack instructions to the target register, solving the overhead and performance losses caused by the introduction of extended registers in the prior art, and achieving efficient floating point stack translation.
Patent Information
- Application Number
- CN202011631165.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-30
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2040-12-30
AI Technical Summary
The prior art requires the introduction of additional extended registers when implementing floating-point stack translation of Intel x86 processors, resulting in additional overhead and software performance losses, and occupying floating-point register resources.
Through multi-threaded parallel processing and preset mapping rules, floating-point stack instructions are dynamically mapped to the target register, avoiding the introduction of additional extended registers and achieving efficient translation of the floating-point stack.
Improves the translation performance of floating point stacks, reduces additional overhead, and ensures efficient and accurate software migration.
Smart Images

Figure CN114691209B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technologies, and particularly to a floating-point stack processing method, apparatus, and processor. Background Art
[0002] For a new-generation architecture microprocessor, although it has more advanced design concepts and higher cost performance, the software ecosystem has become an unfavorable factor restricting its development. Adopting binary translation technology for software transplantation can enable the microprocessor to be compatible with the already mature software ecosystem. For example, the Intel x86 processor has a good software ecosystem, and software designed based on the Intel x86 architecture can be transplanted to a modern RISC (Reduced Instruction Set Computer) processor architecture through binary translation technology. However, due to special historical reasons, the organization method of the floating-point registers of the floating-point component x87 of the Intel x86 processor is different from that of modern processors. This structural difference not only brings difficulties to software transplantation using binary translation, but also causes a certain loss in software performance.
[0003] In the prior art, by additionally adding a memory stack and extended registers, fast translation of the X87 floating-point stack is achieved. This solution realizes the translation of the X87 floating-point stack through the interaction between the translator and the running environment. Among them, in the translator, the information of the actual x87 floating-point stack (80 bits) is simulated by using 8 consecutive 8-byte (64-bit) memories, and in the running environment, 16 64-bit physical floating-point registers are used for register mapping with the memory stack. Among them, the lower 8 physical floating-point registers and the memory stack have a fixed mapping method, and the higher 8 physical floating-point registers form extended registers as the simulated register stack of the running environment. When translating floating-point operation instructions, according to the fixed mapping of the registers, the operands of the floating-point instructions are directly mapped to the corresponding registers, reducing the need to verify the assumed stack top (TOP) for each execution of the translated basic block, thereby improving the efficiency of translating the x87 floating-point stack.
[0004] However, the current method needs to introduce additional extended registers to accelerate the translation process, which not only brings additional overhead, but also requires normalizing the basic blocks through the memory stack in the translator, introducing memory access operations, which will cause a loss in the performance of the transplanted software. Moreover, during the fast translation process, a relatively large number of floating-point registers are occupied, and an auxiliary translator may also be required, which will also increase the corresponding cost. Summary of the Invention
[0005] The present application provides a floating-point stack processing method, apparatus, and processor. Without introducing additional extended registers, the processing of the X87 floating-point stack can be achieved through multi-thread parallel processing and corresponding mapping strategies, effectively improving the translation performance of the floating-point stack.
[0006] In a first aspect, the present application provides a floating-point stack processing method, including:
[0007] Multiple threads process each first basic block in parallel, where the first basic block is any basic block included in the source program in the to-be-emulated register;
[0008] When there is at least one unprocessed first basic block, obtain the source floating-point stack instruction from the unprocessed first basic block, perform mapping processing according to the source floating-point stack instruction and a preset mapping rule, obtain a target register mapped to the to-be-emulated register and a second basic block, to complete the floating-point stack processing, where the second basic block is any basic block included in the source program in the target register.
[0009] In a possible design, after obtaining the second basic block, it further includes:
[0010] Save each second basic block in a basic block cache in sequence according to the next program counter of a specified floating-point stack instruction, where the specified floating-point stack instruction is the last floating-point stack instruction in the current second basic block.
[0011] In a possible design, after obtaining the second basic block, it further includes:
[0012] Check each second basic block in sequence through a single thread to determine whether there is an error in the process of the multiple threads processing each first basic block in parallel;
[0013] When it is determined that there is an error, reprocess the current second basic block through the single thread and update the real-time stack pointer corresponding to the current second basic block.
[0014] In a possible design, the step of checking each second basic block in sequence through a single thread to determine whether there is an error in the process of the multiple threads processing the multiple first basic blocks includes:
[0015] Obtain the real-time stack pointer corresponding to the current second basic block;
[0016] Judge whether the real-time stack pointer is consistent with the stack pointer corresponding to the current second basic block;
[0017] When the judgment result is yes, determine that the process of the corresponding thread obtaining the current second basic block is correct;
[0018] When the judgment result is negative, it is determined that there is an error in the process of the corresponding thread obtaining the current second basic block.
[0019] In a possible design, the mapping process according to the source floating-point stack instruction and the preset mapping rule to obtain the target register mapped to the register to be simulated includes:
[0020] Obtain the first stack pointer corresponding to the source floating-point stack instruction;
[0021] Determine the corresponding target register according to the first stack pointer, the preset mapping rule, and the register to be simulated to which the source floating-point stack instruction belongs.
[0022] In a possible design, when the source floating-point stack instruction is a push instruction or a pop instruction, the mapping process according to the source floating-point stack instruction and the preset mapping rule to obtain the target register mapped to the register to be simulated includes:
[0023] Perform adjustment processing on the first stack pointer to obtain a second stack pointer;
[0024] Determine the corresponding target register according to the second stack pointer, the preset mapping rule, and the register to be simulated to which the source floating-point stack instruction belongs.
[0025] In a possible design, the performing adjustment processing on the first stack pointer to obtain a second stack pointer includes:
[0026] When the source floating-point stack instruction is the push instruction, perform adjustment processing on the first stack pointer according to the preset push instruction to obtain the second stack pointer; or
[0027] When the source floating-point stack instruction is the pop instruction, perform adjustment processing on the first stack pointer according to the preset pop instruction to obtain the second stack pointer.
[0028] In a possible design, the multiple threads process each first basic block in parallel, including:
[0029] For each first basic block, the following operations are performed: all threads simultaneously obtain any first basic block; select the thread that preferentially ends the atomic operation from all the threads; and
[0030] Process the obtained first basic block through the selected thread; wherein each first basic block includes a preset atomic flag bit, and the preset atomic flag bit is used to indicate whether the atomic operation performed by the thread that obtains the first basic block on the first basic block has ended.
[0031] Second aspect, the present application provides a floating-point stack processing device, including:
[0032] A basic block processing module, configured to process each first basic block in parallel by multiple threads, where the first basic block is any basic block included in the source program in the register to be simulated;
[0033] A floating-point stack processing module, configured to, when there is at least one of the first basic blocks that has not been processed yet, obtain source floating-point stack instructions from the first basic block that has not been processed yet, perform mapping processing according to the source floating-point stack instructions and a preset mapping rule, obtain a target register and a second basic block that are mapped to the register to be simulated, so as to complete the floating-point stack processing, where the second basic block is any basic block included in the source program in the target register.
[0034] In a possible design, the device further includes: a storage module, configured to:
[0035] Save each second basic block in a basic block cache in sequence according to the next program counter of a specified floating-point stack instruction, where the specified floating-point stack instruction is the last floating-point stack instruction in the current second basic block.
[0036] In a possible design, the device further includes an inspection module, and the inspection module includes:
[0037] A judgment unit, configured to sequentially inspect each second basic block through a single thread to determine whether there is an error in the processing process of each first basic block by the multiple threads in parallel;
[0038] A reprocessing unit, configured to, when it is determined that there is an error, reprocess the current second basic block through the single thread and update the real-time stack pointer corresponding to the current second basic block.
[0039] In a possible design, the judgment unit is specifically configured to:
[0040] Obtain the real-time stack pointer corresponding to the current second basic block;
[0041] Judge whether the real-time stack pointer is consistent with the stack pointer corresponding to the current second basic block;
[0042] When the judgment result is yes, determine that the processing process of the corresponding thread obtaining the current second basic block is correct;
[0043] When the judgment result is no, determine that the processing process of the corresponding thread obtaining the current second basic block is incorrect.
[0044] In a possible design, the floating-point stack processing module is specifically configured to:
[0045] Obtain a first stack pointer corresponding to the source floating-point stack instruction;
[0046] Determine a corresponding target register according to the first stack pointer, the preset mapping rule, and the register to be simulated to which the source floating-point stack instruction belongs.
[0047] In a possible design, when the source floating-point stack instruction is a push instruction or a pop instruction, the floating-point stack processing module further includes an adjustment processing unit and a determination unit;
[0048] The adjustment processing unit is configured to perform adjustment processing on the first stack pointer to obtain a second stack pointer;
[0049] The determination unit is configured to determine a corresponding target register according to the second stack pointer, the preset mapping rule, and the register to be simulated to which the source floating-point stack instruction belongs.
[0050] In a possible design, the adjustment processing unit is specifically configured to:
[0051] When the source floating-point stack instruction is the push instruction, perform adjustment processing on the first stack pointer according to a preset push instruction to obtain the second stack pointer;
[0052] When the source floating-point stack instruction is the pop instruction, perform adjustment processing on the first stack pointer according to a preset pop instruction to obtain the second stack pointer.
[0053] In a possible design, the basic block processing module is specifically configured to:
[0054] For each first basic block, perform the following operations: all threads simultaneously obtain any first basic block; select a thread that preferentially ends the atomic operation from all the threads; and
[0055] Process the obtained first basic block through the selected thread; wherein each first basic block includes a preset atomic flag bit, and the preset atomic flag bit is used to indicate whether the atomic operation performed by the thread that obtains the first basic block on the first basic block ends.
[0056] In a third aspect, the present application provides a processor, including the floating-point stack processing device according to any one of the second aspects.
[0057] This application provides a floating-point stack processing method, apparatus, and processor. Multiple threads are used to process each first basic block in parallel, where the first basic block is any basic block included in the source program in the register to be simulated. When there is at least one first basic block that has not been processed yet, first obtain the source floating-point stack instructions from the unprocessed first basic blocks, and then perform mapping processing according to the source floating-point stack instructions and the preset mapping rules to obtain the target register mapped to the register to be simulated and the second basic block, thereby completing the floating-point stack processing, where the second basic block is any basic block included in the source program in the target register. Thus, based on the multi-thread parallel processing strategy, the preset mapping rules, and the floating-point stack instructions, the operands of the register to be simulated are flexibly and dynamically mapped to the target register on the basis of the original register to complete the processing of the floating-point stack. There is no need to introduce additional extended registers, and thus there is no need to perform related maintenance on the extended registers and no additional overhead is incurred. Moreover, it can effectively improve the translation performance of the translator for the floating-point stack. Description of the Drawings
[0058] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0059] Figure 1 FIG. is a schematic diagram of an application scenario provided by an embodiment of the present application;
[0060] Figure 2 FIG. is a schematic flowchart of a floating-point stack processing method provided by an embodiment of the present application;
[0061] Figure 3 FIG. is a schematic flowchart of a multi-thread parallel processing provided by an embodiment of the present application;
[0062] Figure 4 FIG. is a schematic flowchart of obtaining a target register provided by an embodiment of the present application;
[0063] Figure 5 FIG. is a schematic diagram of a dynamic mapping process provided by an embodiment of the present application;
[0064] Figure 6 FIG. is another schematic flowchart of obtaining a target register provided by an embodiment of the present application;
[0065] Figure 7 FIG. is a schematic flowchart of a verification mechanism provided by an embodiment of the present application;
[0066] Figure 8A schematic diagram of inspection and verification provided by an embodiment of the present application;
[0067] Figure 9 A schematic flowchart of an inspection process provided by an embodiment of the present application;
[0068] Figure 10 A schematic structural diagram of a floating-point stack processing device provided by an embodiment of the present application;
[0069] Figure 11 A schematic structural diagram of another floating-point stack processing device provided by an embodiment of the present application;
[0070] Figure 12 A schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0071] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present application. On the contrary, they are merely examples of methods and devices consistent with some aspects of the present application as detailed in the appended claims.
[0072] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and do not necessarily have to be used to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0073] Although the microprocessors with a new generation of architecture have more advanced design concepts and higher cost performance, their software ecosystems have become an adverse factor restricting their development. Binary translation technology can be used for software transplantation to make the microprocessors compatible with the mature software ecosystems. For example, Intel x86 processors have a good software ecosystem, and software designed based on the Intel x86 architecture can be transplanted to modern RISC (Reduced Instruction Set Computer) processor architectures through binary translation technology. However, due to special historical reasons, the organization method of the floating-point registers of the floating-point unit x87 of Intel x86 processors is different from that of modern processors. This structural difference will cause difficulties in software transplantation using binary translation and will also cause certain losses in software performance. In the prior art, the translation of the X87 floating-point stack can be achieved by additionally increasing the memory stack and extended registers. However, the current method needs to introduce additional extended registers to accelerate the translation process, which not only brings additional overhead but also causes performance loss of the transplanted software due to the need to normalize the basic blocks. Moreover, during the fast translation process, because a large number of floating-point registers are occupied, an auxiliary translator may be required, which will also increase the corresponding cost. It can be seen that for the processing of the X87 floating-point stack, the existing solutions are not the best choice.
[0074] To address the above problems, the embodiments of the present application provide a floating-point stack processing method, apparatus, and processor. Each first basic block is processed in parallel by multiple threads, where the first basic block is any basic block included in the source program in the register to be simulated. When there is at least one first basic block that has not been processed yet, first obtain the source floating-point stack instruction from the first basic blocks that have not been processed yet, and then perform mapping processing according to the source floating-point stack instruction and the preset mapping rule to obtain the target register mapped to the register to be simulated and the second basic block, thereby completing the floating-point stack processing, where the second basic block is any basic block included in the source program in the target register. Thus, based on the multi-thread parallel processing strategy, the preset mapping rule, and the floating-point stack instruction, the processing of the floating-point stack can be realized without introducing extended registers, the operands of the floating-point stack in the register to be simulated are dynamically mapped to the target register, without causing additional overhead, and the translation performance of the translator for the floating-point stack can be effectively improved.
[0075] Next, an exemplary application scenario of the embodiments of the present application will be introduced.
[0076] Figure 1A schematic diagram of an application scenario provided by an embodiment of the present application. The floating-point stack processing method provided by the embodiment of the present application can be executed by the floating-point stack processing device provided by the embodiment of the present application. The floating-point stack processing device provided by the embodiment of the present application can be a device such as a computer, a server, or a server cluster. The present application embodiment does not limit the specific type. For example, Figure 1 as shown Figure 1 In the figure, server 11 is taken as an example. Among them, server 11 can be the server to which the target register in the floating-point stack processing method provided by the embodiment of the present application belongs, and server 12 is the server to which the register to be simulated belongs. Through the floating-point stack processing method provided by the embodiment of the present application, the operands of the floating-point stack in server 12 are dynamically mapped to server 11, realizing the software transplantation of the source program from the processor of server 12 to the processor of server 11, without introducing an additional memory data structure to maintain it, without additional overhead, and effectively improving the translation performance of the floating-point stack.
[0077] Next, specific embodiments will be used to describe in detail the technical solutions of the present invention and how the technical solutions of the present application solve the above technical problems. These several specific embodiments below can be combined with each other. For the same or similar concepts or processes, they may not be repeated in some embodiments. Next, the embodiments of the present invention will be described with reference to the accompanying drawings.
[0078] Figure 2 A flowchart of a floating-point stack processing method provided by an embodiment of the present application. As Figure 2 shown, the floating-point stack processing method provided in this embodiment includes:
[0079] S101: Multiple threads process each first basic block in parallel.
[0080] Among them, the first basic block is any basic block included in the source program in the register to be simulated.
[0081] Using binary translation technology, the source program that runs maturely in the processor to which the register to be simulated belongs can be transplanted to the processor to which the target register belongs. Among them, a basic block refers to a sequence of program execution statements, and a source program can include multiple basic blocks. When performing binary translation processing, for each first basic block, multiple threads can be used to process the first basic block in parallel. Among them, the first basic block is any basic block in which the source program runs in the processor to which the register to be simulated belongs. Multiple threads processing each first basic block in parallel means that for any one of the first basic blocks in all the first basic blocks, multiple threads process the any one of the first basic blocks at the same time. Multiple threads processing each first basic block in parallel can enable multiple threads to perform binary translation of the basic block at the same time. It can be understood that the threads in the embodiment of the present application refer to the corresponding threads with binary translation functions.
[0082] For example, a possible implementation of step S101 can be as Figure 3 shown Figure 3 A schematic diagram of a multi-thread parallel processing flow provided by an embodiment of the present application. Taking any one of the first basic blocks in all the first basic blocks as an example, this implementation includes:
[0083] S1011: All threads simultaneously obtain any one of the first basic blocks.
[0084] To improve the binary translation processing efficiency of the basic block, all threads simultaneously obtain any one of the multiple basic blocks included in the source program, where any one of the basic blocks is the first basic block. When the number of basic blocks included in the source program is one, that is, the number of first basic blocks is one, then for this one first basic block, the binary translation processing is performed by a single thread. Through this step, all threads can run simultaneously.
[0085] S1012: Select a thread that preferentially ends the atomic operation from all threads, and process the obtained first basic block through the selected thread.
[0086] Wherein, each first basic block includes a preset atomic flag bit, and the preset atomic flag is used to indicate whether the atomic operation performed by the thread that obtains the first basic block on the first basic block is ended.
[0087] After step S1011, for a first basic block, multiple threads will simultaneously obtain the first basic block, and then this step can ensure that each first basic block is processed by a single thread.
[0088] Specifically, after all threads obtain any first basic block, each thread will query the preset atomic flag bit included in the first basic block and perform an atomic operation on the first basic block. Among them, the atomic operation refers to the complete step of reading the preset atomic flag bit, then modifying the preset atomic flag bit, and then writing the preset atomic flag bit. When the thread finishes writing the preset atomic flag bit, it indicates that the atomic operation of the thread ends. Therefore, the preset atomic flag bit can represent whether the atomic operation of the thread on the first basic block ends. For any first basic block, all threads are performing atomic operations. Since the speeds at which different threads execute atomic operations may be different, there will inevitably be threads that finish the atomic operation first. Therefore, a thread that finishes the atomic operation first can be selected from all the threads performing atomic operations on the first basic block, and then the first basic block obtained by the selected thread is subjected to binary translation processing. Correspondingly, the remaining unselected threads will also obtain other first basic blocks for corresponding processing. Thus, it is ensured that all threads perform binary translation processing on the first basic block simultaneously, but for each first basic block, in fact, it is the single thread with the fastest processing speed among all threads that performs binary translation processing on it, thereby improving the efficiency of binary translation processing of the basic block.
[0089] S102: When there is at least one first basic block that has not been processed to completion, obtain the source floating-point stack instruction from the first basic blocks that have not been processed to completion, and perform mapping processing according to the source floating-point stack instruction and the preset mapping rule to obtain the target register mapped to the register to be simulated and the second basic block, so as to complete the floating-point stack processing.
[0090] Among them, the second basic block is any basic block included in the source program.
[0091] In step S101, multiple threads process each first basic block in parallel. During the processing, there may be first basic blocks that have not been processed to completion. When there is at least one first basic block that has not been processed to completion, obtain the source floating-point instruction from the first basic blocks that have not been processed to completion. Among them, the form in which the first basic block has not been processed to completion can be that the first basic block cannot be normally executed for binary translation processing due to reasons such as accidental interruption or program error, or the processing of the first basic block has not been completed within the preset time period. This embodiment does not make any limitations in this regard. When it is determined that there is at least one first basic block that has not been processed to completion, obtain the source floating-point stack instruction from the first basic block that has not been processed to completion. After that, perform mapping processing according to the source floating-point stack instruction and the preset mapping rule to obtain the target register mapped to the register to be simulated and the second basic block, thereby completing the floating-point stack processing.
[0092] The mapping process can be understood as follows: based on the source floating-point stack and the preset mapping rules, the operands of the floating-point registers involved in the first basic block that has not been processed yet are mapped from the registers to be simulated to the target registers, so as to obtain the target registers mapped to the registers to be simulated. Then, according to the target registers, binary translation processing can be normally performed on each first basic block. Each first basic block after binary translation processing is the second basic block. Therefore, the second basic block is any basic block included in the source program in the target registers.
[0093] When performing the mapping process, the source floating-point stack instructions in the first basic block corresponding to the registers to be simulated and the preset mapping rules are adopted. The preset mapping rules can be expressed as the functional relationship of the floating-point stack of the registers to be simulated. When the floating-point stacks are different, the results determined by the functional relationship are different. Therefore, the mapping process of the operands from the registers to be simulated to the target simulated registers can become a dynamic mapping mechanism, without the need to maintain an additional memory data structure, making the mapping process more flexible.
[0094] After obtaining the second basic blocks, in a possible implementation manner, each second basic block can be sequentially stored in the basic block cache according to the next program counter (Program Counter, abbreviated as PC) of the specified floating-point stack instruction. For each second basic block, the specified floating-point stack instruction refers to the last floating-point stack instruction in the current second basic block. For example, if the basic block is represented by TB, then each second basic block can be sequentially stored in the basic block cache (TB cache) in the format of "TB -> PC". In the format "TB -> PC", "TB" represents the current second basic block to be saved, and "PC" in the format "TB -> PC" represents the next program counter of the last floating-point stack instruction in the current second basic block. In addition, the sequential storage order can be set according to the actual working conditions. For example, the storage order can be the size of the "PC" space, and this embodiment does not make any limitations in this regard.
[0095] The floating-point stack processing method provided by the embodiments of the present application first processes each first basic block in parallel by multiple threads, where the first basic block is any basic block included in the source program in the register to be simulated. Then, when there is at least one first basic block that has not been processed yet, first obtain the source floating-point stack instruction from the first basic blocks that have not been processed yet, and then perform mapping processing according to the source floating-point stack instruction and the preset mapping rule to obtain the target register mapped to the register to be simulated and the second basic block, where the second basic block is any basic block included in the source program in the target register. The floating-point stack processing method provided by this embodiment effectively improves the translation efficiency by processing the first basic block in parallel by multiple threads. In addition, dynamic mapping of operands is realized based on the source floating-point stack instruction and the preset mapping rule, without introducing additional extended registers, and thus there is no need to perform related maintenance on the extended registers and no additional overhead is brought. Moreover, it can effectively improve the translation performance of the translator for the floating-point stack.
[0096] In a possible design, the possible implementation manner of obtaining the target register mapped to the register to be simulated by performing mapping processing according to the source floating-point stack instruction and the preset mapping rule in step S102 is as Figure 4 shown Figure 4 is a schematic flowchart of a process for obtaining a target register provided by an embodiment of the present application, and this implementation manner includes:
[0097] S201: Obtain the first stack pointer corresponding to the source floating-point stack instruction.
[0098] When there is at least one first basic block that has not been processed yet, obtain the source floating-point stack instruction from the first basic blocks that have not been processed yet, and further obtain the first stack pointer corresponding to the source floating-point stack instruction.
[0099] S202: Determine the corresponding target register according to the first stack pointer, the preset mapping rule, and the register to be simulated to which the source floating-point stack instruction belongs.
[0100] The first stack pointer corresponding to the source floating-point stack instruction is actually the stack pointer of the floating-point stack of the register to be simulated stored in the translator. For example, X87_top can be used to represent the stack pointer. Further, actual_top / temp_top are respectively used to represent the stack pointers corresponding to the first basic block during multi-threaded translation of the basic block and single-threaded binary translation of the first basic block. ST(i) is used to represent the floating-point register of the corresponding floating-point stack, that is, the register to be simulated, and F(j) is used to represent the floating-point register mapped to the register to be simulated, that is, the target register. The preset mapping rule provided by this embodiment can be expressed by the following expression:
[0101]
[0102] Among them, i represents any floating-point register in the register to be simulated, and j represents the floating-point register corresponding to i in the target register after being processed by the preset mapping rule. Both i and j are integers greater than zero. During the mapping process, in order not to occupy too many floating-point registers, usually the maximum values of i and j are set to 8, that is, only 8 floating-point registers are needed to implement the dynamic mapping of operands through the floating-point stack processing method provided by the embodiments of the present application.
[0103] Specifically, ST(i) in the preset mapping rule represents any floating-point register to be simulated, such as ST(1), ST(2), ST(3), etc. After ST(i) is determined, since the floating-point register performs relative addressing with respect to TOP, ST(i) can perform relative addressing with respect to temp_top to directly address the target register and obtain F(j).
[0104] It should be noted that the above expression is for one basic block. For multiple basic blocks included in the source program, actual_top can be used to replace temp_top in the above expression.
[0105] Through the above expression of the preset mapping rule, based on the first stack pointer corresponding to the source floating-point stack and the register to be simulated to which it belongs, the operands of the floating-point stack can be dynamically mapped from the register to be simulated to the target register, thereby obtaining the mapped target register.
[0106] During the actual mapping process, during translation initialization, the value of X87_top can be updated to actual_top / temp_top, that is, X87_top -> actual_top / temp_top. Therefore, only the preset mapping rule needs to be set according to actual_top / temp_top for dynamic mapping.
[0107] In a possible design, since actual_top is frequently used for mapping, a general-purpose register can be introduced for the transition during the dynamic mapping process, such as Figure 5 shown Figure 5 is a schematic diagram of a dynamic mapping process provided by the embodiments of the present application. For temp_top, since it is used separately by each single thread during the parallel processing of multiple first basic blocks by multiple threads, there is no need to introduce a general-purpose register for mapping processing.
[0108] It should be noted that the register to be simulated performs relative addressing with respect to the corresponding stack pointer, while the target register performs direct addressing with respect to the stack pointer. Therefore, the code representing the basic block after translation processing can be directly mapped to the target register, and when processing the basic block to obtain the second basic block, it is no longer necessary to adjust the state of the target register, and there is no need to perform normalization processing after maintaining additional memory structure data.
[0109] It should be understood that when using 8 target registers for mapping processing, it is necessary to save or restore these 8 target registers according to the actual context, but it is not necessary to perform context switching during the process of obtaining the second basic block from each first basic block, and it can be performed when necessary according to the actual situation.
[0110] The floating-point stack processing method provided by the embodiments of the present application, after obtaining the source floating-point stack instruction, first obtains the first stack pointer corresponding to the source floating-point stack instruction, and then obtains the target register mapped to the register to be simulated according to the first stack pointer, the preset mapping rule, and the register to be simulated to which the source floating-point stack instruction belongs. Among them, a preset mapping rule is established based on the stack pointer, so that the mapping process is a dynamic mapping process, effectively improving the translation processing performance.
[0111] In Figure 4 Based on the shown embodiment, further, the obtained source floating-point stack instruction may be a push instruction or a pop instruction. When the source floating-point stack instruction is a push instruction or a pop instruction, in step S102, mapping processing is performed according to the source floating-point stack instruction and the preset mapping rule to obtain a target register mapped to the register to be simulated. A possible implementation manner is as Figure 6 shown Figure 6 which is a schematic flowchart of another method for obtaining a target register provided by the embodiments of the present application. This implementation manner includes:
[0112] S301: Adjust the first stack pointer to obtain a second stack pointer.
[0113] After obtaining the source floating-point stack instruction, it is possible to determine whether the source floating-point stack instruction is a push instruction or a pop instruction. When the determination result is that the source floating-point stack instruction is a push instruction or a pop instruction, the obtained first stack pointer should be adjusted to obtain the corresponding second stack pointer.
[0114] For example, when the source floating-point stack instruction is a push instruction, the first stack pointer is adjusted according to the preset push instruction to obtain the corresponding second stack pointer. Specifically, the preset push instruction may be "++", that is, for a basic block, temp_top is adjusted according to the principle of temp_top++.
[0115] Correspondingly, when the source floating-point stack instruction is a pop instruction, the first stack pointer is adjusted according to a preset pop instruction to obtain a corresponding second stack pointer. Specifically, the preset pop instruction can be "--", that is, for a first basic block, temp_top is adjusted according to the principle of temp_top--.
[0116] Among them, in the process of adjusting the first stack pointer to obtain the second stack pointer, in order to ensure that the current_top of the second stack pointer does not overflow, according to the principle that negative numbers in the processor are in two's complement, it can be restored to 7 after adjustment.
[0117] S302: Determine a corresponding target register according to the second stack pointer, a preset mapping rule, and a to-be-simulated register to which the source floating-point stack instruction belongs.
[0118] For push instructions or pop instructions, the mapping process is based on the stack pointer after adjustment processing.
[0119] Among them, the implementation principle and technical effect of step S302 are similar to those of step S202. The difference is that the current stack pointer is the second stack pointer after adjustment processing, and the specific process will not be elaborated here.
[0120] In the floating-point stack processing method provided by the embodiments of the present application, when the source floating-point stack instruction is a push instruction or a pop instruction, first, the obtained first stack pointer is adjusted to obtain a corresponding second stack pointer. Then, according to the second stack pointer, a preset mapping rule, and a to-be-simulated register to which the source floating-point stack instruction belongs, a corresponding target register is determined. For floating-point registers containing push / pop instructions, the dynamic mapping of operands can still be achieved through a preset mapping rule, making the floating-point stack processing process more flexible and effectively improving the translation processing performance.
[0121] Through the description of the above embodiments, during the mapping process to obtain the target register, and after the first basic block is translated accordingly to obtain the second basic block, the second basic block is sequentially stored in the basic block cache according to the next program counter of the specified floating-point stack instruction. To ensure the correctness of the parallel processing process of multiple threads for multiple first basic blocks and improve the accuracy of translation processing. Therefore, after obtaining the second basic block, the X87 floating-point stack processing method provided by the embodiments of the present application further includes an inspection and verification step.
[0122] Figure 7 It is a flowchart of a verification mechanism provided by the embodiments of the present application. As Figure 7 shown, the steps provided in this embodiment include:
[0123] S401: Check each second basic block one by one through a single thread to determine whether there is an error in the processing of each first basic block in parallel by multiple threads.
[0124] Use a single thread to check each second basic block stored in the basic block cache one by one, so as to determine whether there is an error in the translation processing of step S101, thereby further ensuring the accuracy of the translation processing.
[0125] S402: When it is determined that there is an error, reprocess the current second basic block through a single thread and update the real-time stack pointer corresponding to the current second basic block.
[0126] After inspection and judgment, when it is determined that there is an error, reprocess the second basic block with an error in the current processing through a single thread and update the real-time stack pointer corresponding to the current second basic block.
[0127] After obtaining the second basic block, save the second basic block in the basic block cache one by one according to the next program counter of the specified floating-point stack instruction. At the same time, for this second basic block, the stack pointer of the basic block corresponding to it before translation processing, that is, the stack pointer corresponding to the first basic block, and the stack pointer corresponding to the second basic block after the first basic block is translated and processed are also saved. For example, when saving the obtained second basic block in the basic block cache according to the next program counter of the specified floating-point stack instruction, the temp_top corresponding to the first basic block is saved in TB->top_in, and the temp_top corresponding to the second basic block is saved in TB->top_out. After it is determined that there is an error, reprocess the current second basic block through a single thread. At the same time, update the temp_top corresponding to the current second basic block saved in TB->top_out, that is, update the real-time stack pointer of the second basic block that needs to be reprocessed.
[0128] Among them, the inspection and verification process of the single thread is carried out in sequence, that is, for each second basic block in the basic block cache, it is carried out one by one according to its storage order until all the basic blocks are inspected and verified. The inspection and verification process is as Figure 8 shown Figure 8 is a schematic diagram of inspection and verification provided by an embodiment of the present application.
[0129] The floating-point stack processing method provided by the embodiment of the present application further performs an inspection and verification step after obtaining the second basic block. First, each current second basic block is sequentially inspected through a single thread to determine whether there is an error in the processing process of multiple threads for multiple first basic blocks. When it is determined that there is an error, the current second basic block is reprocessed through the single thread, and at the same time, the corresponding real-time stack pointer is updated. Thus, the correctness of the parallel processing of multiple threads for multiple first basic blocks is ensured, and the accuracy of the translation processing is improved.
[0130] Based on the Figure 7 embodiment, a possible implementation of step S401 is as Figure 9 shown Figure 9 and is a schematic flowchart of an inspection process provided by the embodiment of the present application. This implementation includes:
[0131] S4011: Obtain the real-time stack pointer corresponding to the current second basic block;
[0132] S4012: Determine whether the real-time stack pointer is consistent with the stack pointer corresponding to the current second basic block;
[0133] S4013: When the judgment result is yes, determine that the processing process of the corresponding thread for obtaining the current second basic block is error-free;
[0134] S4014: When the judgment result is no, determine that the processing process of the corresponding thread for obtaining the current second basic block is incorrect.
[0135] The real-time stack pointer corresponding to the current second basic block in the basic block cache is obtained through a single thread, and then this real-time stack pointer is compared with the stack pointer corresponding to the current second basic block that has been saved in the basic block cache, that is, compared with the temp_top corresponding to the current second basic block saved in TB->top_out, to determine whether the two are consistent. If the judgment result is yes, it is determined that the processing process of obtaining the current second basic block is error-free. On the contrary, if the judgment result is no, it is determined that the processing process of obtaining the current second basic block is incorrect. Among them, the inspection process provided in this embodiment sequentially inspects each current second basic block, that is, steps S4011 to S4014 are performed on each current second basic block until all current second basic blocks are inspected, so as to complete the judgment on whether there is an error in the processing process of multiple threads for multiple first basic blocks.
[0136] The floating-point stack processing method provided by the embodiments of the present application checks each second basic block sequentially through a single thread. Among them, first, the real-time stack pointer corresponding to the current second basic block is obtained, and then it is determined whether the real-time stack pointer is consistent with the stack pointer corresponding to the current second basic block that has been saved in the basic block cache. If they are consistent, that is, the judgment result is yes, it is determined that the processing process of the corresponding thread for the current second basic block is correct. If they are inconsistent, that is, the judgment result is no, it is determined that the processing process of the corresponding thread for the current second basic block is incorrect. By checking the processing process of obtaining the current second basic block through a single thread, the incorrect processing process can be determined, and then it can be reprocessed through a single thread to achieve verification, ensuring the correctness of the parallel processing of multiple first basic blocks by multiple threads and improving the accuracy of translation processing.
[0137] The following is an embodiment of the device of the present application, which can be used to execute the steps of the floating-point stack processing method provided by the corresponding method embodiment above. For the details not disclosed in the device embodiment of the present application, please refer to the method embodiment of the present application.
[0138] Figure 10 It is a schematic structural diagram of a floating-point stack processing device provided by an embodiment of the present application, as Figure 10 shown. The floating-point stack processing device 100 provided in this embodiment includes:
[0139] A basic block processing module 101, which is used for multiple threads to process each first basic block in parallel.
[0140] Among them, the first basic block is any basic block included in the source program in the register to be simulated.
[0141] A floating-point stack processing module 102, which is used to obtain the source floating-point stack instruction from the first basic blocks that have not been processed when there is at least one first basic block that has not been processed, and perform mapping processing according to the source floating-point stack instruction and the preset mapping rule to obtain the target register and the second basic block that are mapped to the register to be simulated, so as to complete the floating-point stack processing.
[0142] Among them, the second basic block is any basic block included in the source program in the target register.
[0143] In a possible design, the floating-point stack processing device 100 further includes: a storage module, which is used for:
[0144] Save each second basic block in the basic block cache in sequence according to the next program counter of the specified floating-point stack instruction, and the specified floating-point stack instruction is the last floating-point stack instruction in the current second basic block.
[0145] In Figure 10 Based on the shown embodiment, Figure 11A structural schematic diagram of another floating-point stack processing device provided by an embodiment of the present application, as Figure 11 shown, the floating-point stack processing device 100 provided in this embodiment further includes a verification module 103. Among them, the verification module 103 includes:
[0146] A judgment unit, configured to sequentially verify each second basic block through a single thread to determine whether there is an error in the processing process of each first basic block by multiple threads in parallel;
[0147] A reprocessing unit, configured to reprocess the current second basic block through a single thread and update the real-time stack pointer corresponding to the current second basic block when it is determined that there is an error.
[0148] In a possible design, the judgment unit is specifically configured to:
[0149] Obtain the real-time stack pointer corresponding to the current second basic block;
[0150] Judge whether the real-time stack pointer is consistent with the stack pointer corresponding to the current second basic block;
[0151] When the judgment result is yes, it is determined that the processing process of the corresponding thread for obtaining the current second basic block is correct;
[0152] When the judgment result is no, it is determined that the processing process of the corresponding thread for obtaining the current second basic block is incorrect.
[0153] In a possible design, the floating-point stack processing module 102 is specifically configured to:
[0154] Obtain the first stack pointer corresponding to the source floating-point stack instruction;
[0155] Determine the corresponding target register according to the first stack pointer, the preset mapping rule, and the register to be simulated to which the source floating-point stack instruction belongs.
[0156] In a possible design, when the source floating-point stack instruction is a push instruction or a pop instruction, the floating-point stack processing module 102 further includes an adjustment processing unit and a determination unit.
[0157] An adjustment processing unit, configured to perform adjustment processing on the first stack pointer to obtain a second stack pointer;
[0158] A determination unit, configured to determine the corresponding target register according to the second stack pointer, the preset mapping rule, and the register to be simulated to which the source floating-point stack instruction belongs.
[0159] In a possible design, the adjustment processing unit is specifically configured to:
[0160] When the source floating-point stack instruction is a push instruction, perform adjustment processing on the first stack pointer according to the preset push instruction to obtain a second stack pointer;
[0161] When the source floating-point stack instruction is a stack pop instruction, the first stack pointer is adjusted according to a preset stack pop instruction to obtain a second stack pointer.
[0162] In a possible design, the basic block processing module is specifically configured to:
[0163] For each first basic block, the following operations are performed: all threads simultaneously obtain any first basic block; select a thread that preferentially ends the atomic operation from all threads; and
[0164] process the obtained first basic block through the selected thread; wherein each first basic block includes a preset atomic flag bit, and the preset atomic flag bit is used to indicate whether the atomic operation performed by the thread that obtains the first basic block on the first basic block is ended.
[0165] It should be noted that the device embodiments provided in this application are only illustrative. The module division in the above device embodiments is only a logical function division, and there may be other division methods in actual implementation. For example, multiple modules can be combined or integrated. The coupling between each module can be implemented through some interfaces, and these interfaces are usually electrical communication interfaces, but it does not exclude the possibility of being a mechanical interface or other forms of interfaces. Therefore, the modules described as separate components may or may not be physically separated, and can be located in one place or distributed to different positions of any one or different devices.
[0166] The above device embodiments can be used to execute the steps provided in the corresponding method embodiments above. The specific implementation manners and technical effects are similar to the foregoing, and will not be elaborated herein.
[0167] This application also provides a processor, wherein the processor includes the above-mentioned floating-point stack processing device provided in the embodiments of this application.
[0168] Furthermore, this application also provides an electronic device, as Figure 12 shown, Figure 12 is a schematic structural diagram of an electronic device provided in the embodiments of this application. The electronic device 700 provided in this embodiment includes:
[0169] at least one processor 701; and
[0170] a memory 702 communicatively connected to at least one processor 701; wherein,
[0171] The memory 702 stores instructions that can be executed by at least one processor 701. The instructions are executed by at least one processor 701 so that at least one processor 701 can execute each step of the floating-point stack processing method in the above method embodiments. For details, reference can be made to the relevant descriptions in the foregoing method embodiments.
[0172] Optionally, the memory 702 can be either independent or integrated with the processor 701.
[0173] When the memory 702 is a device independent of the processor 701, the electronic device 700 may further include:
[0174] A bus 703 for connecting the processor 701 and the memory 702.
[0175] In a possible design, Figure 12 The processor 701 in the illustrated embodiment can be a processor.
[0176] In addition, an embodiment of the present application further provides a non-transitory computer-readable storage medium storing computer instructions, and the computer instructions are used to cause a computer to execute each step of the floating-point stack processing method in the above embodiments. For example, the readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0177] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present application. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include common general knowledge or conventional technical means in the technical field not disclosed in the present application. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present application are pointed out by the claims.
[0178] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.
Claims
1. A floating-point stack processing method, characterized in that, Including: Processing each first basic block in parallel by multiple threads, where the first basic block is any basic block included in the source program in the register to be simulated; When there is at least one of the first basic blocks that has not been processed yet, obtain the source floating-point stack instruction from the first basic block that has not been processed yet, and perform mapping processing according to the source floating-point stack instruction and the preset mapping rule to obtain the target register mapped to the register to be simulated and the second basic block to complete the floating-point stack processing, where the second basic block is any basic block included in the source program in the target register; The multiple threads processing each first basic block in parallel includes: For each first basic block, the following operations are performed: all threads simultaneously obtain any first basic block; select the thread that finishes the atomic operation first from all the threads; and process the obtained first basic block through the selected thread that finishes the atomic operation first; where each first basic block includes a preset atomic flag bit, and the preset atomic flag bit is used to indicate whether the atomic operation performed by the thread that obtains the first basic block on the first basic block is finished; the atomic operation refers to the steps of reading the preset atomic flag bit, then modifying the preset atomic flag bit, and then writing the preset atomic flag bit, and when the thread finishes writing the preset atomic flag bit, the atomic operation of the thread ends.
2. The floating-point stack processing method according to claim 1, characterized in that, After obtaining the second basic block, it further includes: Saving each second basic block in the basic block cache in sequence according to the next program counter of the specified floating-point stack instruction, where the specified floating-point stack instruction is the last floating-point stack instruction in the current second basic block.
3. The floating-point stack processing method according to claim 1, wherein After obtaining the second basic block, it further includes: Sequentially checking each second basic block through a single thread to determine whether there is an error in the process of the multiple threads processing each first basic block in parallel; When it is determined that there is an error, reprocess the current second basic block through the single thread and update the real-time stack pointer corresponding to the current second basic block.
4. The floating-point stack processing method according to claim 3, characterized in that, The sequentially checking each second basic block through a single thread includes: Obtaining the real-time stack pointer corresponding to the current second basic block; Judging whether the real-time stack pointer is consistent with the stack pointer corresponding to the current second basic block; When the judgment result is yes, it is determined that the process of the corresponding thread obtaining the current second basic block is correct; When the judgment result is no, it is determined that the process of the corresponding thread obtaining the current second basic block is incorrect.
5. The floating-point stack processing method according to any one of claims 1-4, characterized in that, When the source floating-point stack instruction is a push instruction or a pop instruction, the performing mapping processing according to the source floating-point stack instruction and the preset mapping rule to obtain the target register mapped to the register to be simulated includes: Performing adjustment processing on the first stack pointer corresponding to the source floating-point stack instruction to obtain a second stack pointer; Determining the corresponding target register according to the second stack pointer, the preset mapping rule, and the register to be simulated to which the source floating-point stack instruction belongs.
6. The floating-point stack processing method according to claim 5, characterized in that, The performing adjustment processing on the first stack pointer to obtain a second stack pointer includes: When the source floating-point stack instruction is the push instruction, the first stack pointer is adjusted according to a preset push instruction to obtain the second stack pointer; When the source floating-point stack instruction is the pop instruction, the first stack pointer is adjusted according to a preset pop instruction to obtain the second stack pointer.
7. A floating-point stack processing device, characterized in that, Comprising: A basic block processing module for processing each first basic block in parallel by multiple threads, where the first basic block is any basic block included in the source program in the register to be simulated; A floating-point stack processing module for, when there is at least one unprocessed first basic block, obtaining a source floating-point stack instruction from the unprocessed first basic block, performing mapping processing according to the source floating-point stack instruction and a preset mapping rule to obtain a target register mapped to the register to be simulated and a second basic block, so as to complete the floating-point stack processing, where the second basic block is any basic block included in the source program in the target register; When the multiple threads process each first basic block in parallel, the basic block processing module is specifically configured to, for each first basic block, perform the following operations: all threads simultaneously obtain any first basic block; select the thread that finishes the atomic operation first from all the threads; and process the obtained first basic block through the selected thread that finishes the atomic operation first; where each first basic block includes a preset atomic flag bit, and the preset atomic flag bit is used to indicate whether the atomic operation performed on the first basic block by the thread that obtains the first basic block ends; the atomic operation refers to the steps of reading the preset atomic flag bit, then modifying the preset atomic flag bit, and then writing the preset atomic flag bit, and when the thread finishes writing the preset atomic flag bit, the atomic operation of the thread ends.
8. A processor, characterized in that, Including the floating-point stack processing device according to claim 7.
Citation Information
Patent Citations
Floating-point operation process for X8b in binary translation
CN1746850A