Instruction Processing Method and Device

By canceling the merge operation of high and low operands in instruction processing, and independently processing and broadcasting the micro-operation results of high and low bits, the problem of low parallelism in the prior art is solved, and the processing efficiency and parallelism are improved.

CN114115999BActive Publication Date: 2025-06-24HYGON INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111396633.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-23
Publication Date
2025-06-24
Estimated Expiration
2041-11-23

AI Technical Summary

Technical Problem

In the prior art, the high and low bits of operands in the instruction processing method need to be merged, resulting in high-bit micro-operations relying on the results of low-bit micro-operations, resulting in low parallelism of program processing and large processor power consumption.

Method used

By canceling the merge operation of high and low bits of operands during the instruction calculation process, the high and low bits of the operands of the instruction are obtained respectively, and the high and low bits of the high and low bits of the micro-operations are performed independently, and the results are stored and broadcasted separately.

Benefits of technology

It improves the efficiency and parallelism of instruction processing, reduces the use of computer resources, and reduces the processing time of the processor.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114115999B_ABST
    Figure CN114115999B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide an instruction processing method, apparatus, device, computer program product, and computer-readable storage medium for a computer architecture. The method provided by the embodiments of the present disclosure cancels the merging operation of the high and low bits of the operand during the instruction calculation process, and processes, calculates, and broadcasts the high and low bits of the operand separately. When the instruction is processed by high and low bits, the high and low bit operations of the operand only depend on the number of bits of the register required for the operation, without waiting for the processing results of other bits. Therefore, the processing result of the instruction can be broadcast in advance, improving the efficiency and parallelism of instruction processing. At the same time, through the method provided by the embodiments of the present disclosure, the high and low bits of the operand respectively use the high and low bits of the same physical register during processing, reducing the occupation of computer resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computers and processors, and more particularly, to an instruction operation method, apparatus, device, and storage medium. Background Art

[0002] With the rapid growth of computing requirements and the increasing expansion of the total amount of data, in order to match the data processing capabilities of computers and processors, how to improve data processing efficiency is one of the main development directions of current computers and processors. Currently, continuously evolving application scenarios, such as scientific simulation, financial analysis, deep learning, etc., have also formed more and larger intensive computing loads, which pose severe challenges to the computing capabilities of processors. To address such challenges, it is necessary to optimize and improve the interior of the processor, and among them, the most crucial is the improvement of the instruction set or arithmetic unit.

[0003] In a computer microarchitecture, an instruction needs to go through processes such as fetching, decoding and conversion, distribution, renaming, scheduling, execution, and broadcasting. According to the existing instruction processing method, for the case of processing the high and low bits of an operand, the micro-operations corresponding to the instruction respectively allocate a physical register for the high and low bits of the operand. The processor first calculates the low-bit data, then calculates the high-bit data, and then combines the two parts into a physical register. When subsequent instructions operate, the combined physical register is used for processing and calculation. In this instruction processing method, the high and low parts of the operand need to be combined, and the high-bit micro-operation depends on the result of the low-bit micro-operation. The high bit needs to wait for the result of the low-bit processing to be broadcast together, so the program processing parallelism is low, the program processing is slow, and the processor power consumption is large.

[0004] However, in fact, when the program processes the high bit of the operand, the operation result of the low-bit part is not a necessary calculation prerequisite in the high-bit operation processing. Therefore, an instruction processing method is needed to reduce the dependence between the micro-operations when processing the high and low parts of the operand, so that the micro-operations only depend on the bit results they need during processing, remove the redundant data dependence, and the high-bit operation can broadcast the data in advance without waiting for the result of the low-bit operation, thereby reducing the processing time of the processor. Summary of the Invention

[0005] To solve the above problems, the present disclosure cancels the merging operation of the high and low bits of the operand during the instruction calculation process, removes the redundant data dependence in the instruction operation. The high and low-bit operations of the operand only depend on the register bits required for the operation, without waiting for the processing results of other bits. Therefore, the processing result of the instruction can be broadcast in advance, improving the efficiency and parallelism of the instruction processing. At the same time, through the method provided by the embodiments of the present disclosure, the high and low bits of the operand respectively use the high and low bits of the same physical register during processing, reducing the occupation of computer resources.

[0006] Embodiments of the present disclosure provide an instruction processing method, apparatus, device, and computer-readable storage medium.

[0007] Embodiments of the present disclosure provide an instruction processing method, the method comprising: respectively obtaining high-order bits and low-order bits of an operand of the instruction; performing a low-order micro-operation based on the obtained low-order bits of the operand, and storing a result of the low-order micro-operation to low-order bits of a physical register for the instruction; and performing a high-order micro-operation based on the obtained high-order bits of the operand, and storing a result of the high-order micro-operation to high-order bits of the physical register, wherein the micro-operation of the high-order bits is independent of the micro-operation of the low-order bits.

[0008] According to an embodiment of the present disclosure, the performing a low-order micro-operation based on the obtained low-order bits of the operand further comprises: allocating the physical register for a destination logical register of the low-order micro-operation of the instruction; the performing a high-order micro-operation based on the obtained high-order bits of the operand further comprises: mapping a destination logical register of the high-order micro-operation of the instruction to the physical register.

[0009] According to an embodiment of the present disclosure, the instruction processing method further comprises: storing the result of the low-order micro-operation to low-order bits of a physical register for the instruction; broadcasting the low-order bits stored in the physical register based on the low-order bits of the physical register; and storing the result of the high-order micro-operation to high-order bits of the physical register; broadcasting the high-order bits stored in the physical register based on the high-order bits of the physical register.

[0010] According to an embodiment of the present disclosure, the physical register is assigned an identifier for identifying that its low-order bits or high-order bits are valid bits, wherein broadcasting the low-order bits stored in the physical register comprises: broadcasting the low-order bits and high-order bits stored in the physical register, and broadcasting an identifier for identifying that the low-order bits of the physical register are valid bits; and broadcasting the high-order bits stored in the physical register comprises: broadcasting the low-order bits and high-order bits stored in the physical register, and broadcasting an identifier for identifying that the high-order bits of the physical register are valid bits.

[0011] According to an embodiment of the present disclosure, the instruction comprises a first instruction and a second instruction, and the second instruction depends on a result of the first instruction; according to a bit-width type of a destination logical register of the first instruction and a bit-width type of a source logical register of the second instruction, an association relationship between low-order bits and high-order bits of the source logical register of the second instruction and low-order bits and high-order bits of the destination logical register of the first instruction can be determined.

[0012] According to an embodiment of the present disclosure, the instructions include a first instruction and a second instruction, and the second instruction depends on the result of the first instruction. The method includes: obtaining the high-order bits and low-order bits of the operand of the first instruction and the mask number of the first instruction; based on the obtained low-order bits of the operand and the mask number, performing a low-order micro-operation and storing the result of the low-order micro-operation in the low-order bits of the physical register for the instruction; based on the obtained high-order bits of the operand and the mask number, performing a high-order micro-operation and storing the result of the high-order micro-operation in the high-order bits of the physical register; according to the bit-width type of the physical register of the first instruction and the bit-width type of the source logical register of the second instruction, determining the low-order bits of the source logical register for the low-order micro-operation of the second instruction based on the low-order bits and / or high-order bits of the physical register of the first instruction, and based on the low-order bits of the operand of the second instruction and the low-order bits of the source logical register for the low-order micro-operation of the second instruction, performing the low-order micro-operation of the second instruction, wherein the low-order bits of the source logical register for the low-order micro-operation of the second instruction are used to store the mask number of the low-order micro-operation of the second instruction; according to the bit-width type of the physical register of the first instruction and the bit-width type of the source logical register of the second instruction, determining the high-order bits of the source logical register for the high-order micro-operation of the second instruction based on the low-order bits or high-order bits of the physical register of the first instruction, and based on the high-order bits of the operand of the second instruction and the high-order bits of the source logical register of the second instruction, performing the high-order micro-operation of the second instruction, wherein the high-order bits of the source logical register for the high-order micro-operation of the second instruction are used to store the mask number of the high-order micro-operation of the second instruction, wherein the bit-width type is any one of byte, word, double word, and quad word, and different bit-width types have different effective bit-widths.

[0013] According to an embodiment of the present disclosure, the effective bit-width of the bit-width type of the physical register of the first instruction is equal to the effective bit-width of the bit-width type of the source logical register of the second instruction; the low-order bits of the source logical register for the low-order micro-operation of the second instruction depend on the low-order bits in the effective bit-width of the physical register of the first instruction; the high-order bits of the source logical register for the high-order micro-operation of the second instruction depend on the high-order bits in the effective bit-width of the physical register of the first instruction.

[0014] According to an embodiment of the present disclosure, the effective bit width of the bit width type of the physical register of the first instruction is less than the effective bit width of the bit width type of the source logical register of the second instruction; the low-order bits of the source logical register for the low-order micro-operation of the second instruction depend on the low-order bits in the effective bit width of the physical register of the first instruction; the high-order bits of the source logical register for the high-order micro-operation of the second instruction depend on the low-order bits in the effective bit width of the physical register of the first instruction.

[0015] According to an embodiment of the present disclosure, the effective bit width of the bit width type of the physical register of the first instruction is greater than the effective bit width of the bit width type of the source logical register of the second instruction; the low-order bits of the source logical register for the low-order micro-operation of the second instruction depend on the low-order bits and high-order bits in the effective bit width of the physical register of the first instruction; the high-order bits of the source logical register for the high-order micro-operation of the second instruction do not depend on the effective bit width of the physical register of the first instruction.

[0016] An embodiment of the present disclosure provides an instruction processing device, the device includes: an operand acquisition module configured to respectively acquire the high-order bits and low-order bits of the operand of the instruction; a micro-operation processing module configured to execute a low-order micro-operation based on the acquired low-order bits of the operand and store the result of the low-order micro-operation to the low-order bits of the physical register for the instruction; and execute a high-order micro-operation based on the acquired high-order bits of the operand and store the result of the high-order micro-operation to the high-order bits of the physical register, wherein the micro-operation processing of the high-order bits is independent of the micro-operation processing of the low-order bits.

[0017] An embodiment of the present disclosure provides a device for instruction processing, including: one or more processors; and one or more memories, wherein computer-executable programs are stored, and when the computer-executable programs are executed by the processors, the method according to any one of claims 1-9 is executed.

[0018] An embodiment of the present disclosure provides a computer program product, the computer program product includes computer software code, and when the computer software code is run by a processor, it is used to implement the method according to any one of claims 1-9.

[0019] Embodiments of the present disclosure provide an instruction processing method, apparatus, device, computer program product, and computer-readable storage medium for a computer architecture. The method provided by the embodiments of the present disclosure cancels the merging operation of the high and low bits of the operand during the instruction calculation process, and processes and calculates the high and low bits of the operand separately, so that when performing operations and processing on the operand, the high and low bit operations of the operand only depend on the number of bits of the register required for the operation, without waiting for the processing results of other bits. Therefore, the processing result of the instruction can be broadcast in advance, improving the efficiency and parallelism of instruction processing. At the same time, through the method provided by the embodiments of the present disclosure, the high and low bits of the operand respectively use the high and low bits of the same physical register during processing, reducing the occupation of computer resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the following will briefly introduce the drawings required for use in the description of the embodiments. Obviously, the drawings in the following description are only some exemplary embodiments of the present disclosure, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.

[0021] Figure 1 is a schematic diagram showing the computer instruction processing process according to an embodiment of the present disclosure;

[0022] Figure 2A is a flowchart showing the instruction processing method according to an embodiment of the present disclosure;

[0023] Figure 2B is a flowchart showing the instruction processing method when a second instruction depends on a first instruction according to an embodiment of the present disclosure;

[0024] Figure 3 is a schematic diagram showing the data dependency relationship in the buffer according to an embodiment of the present disclosure;

[0025] Figure 4 is a schematic diagram showing the positions of the valid bit numbers of different bit-width types in the physical register according to an embodiment of the present disclosure;

[0026] Figure 5A is a schematic diagram showing the comparison of the instruction processing timings of the solution of the present invention and the existing solution in cases of scenarios 1 and 2 according to an embodiment of the present disclosure;

[0027] Figure 5B is a schematic diagram showing the comparison of the instruction processing timings of the solution of the present invention and the existing solution in case of scenario 3 according to an embodiment of the present disclosure;

[0028] Figure 6AIt is a schematic diagram showing the corresponding relationship of bit positions between the physical register of the first instruction and the source logical register of the second instruction when the bit-width types of both the first instruction and the second instruction according to an embodiment of the present disclosure are bytes;

[0029] Figure 6B It is a schematic diagram showing the corresponding relationship of bit positions between the physical register of the first instruction and the source logical register of the second instruction when the bit-width type of the first instruction according to an embodiment of the present disclosure is a byte and the bit-width type of the second instruction is a word;

[0030] Figure 6C It is a schematic diagram showing the corresponding relationship of bit positions between the physical register of the first instruction and the source logical register of the second instruction when the bit-width type of the first instruction according to an embodiment of the present disclosure is a word and the bit-width type of the second instruction is a byte;

[0031] Figure 7 It is a schematic diagram showing an instruction processing device according to an embodiment of the present disclosure;

[0032] Figure 8 It shows a schematic diagram of an instruction processing device according to an embodiment of the present disclosure;

[0033] Figure 9 It shows a schematic diagram of the architecture of an exemplary computing device according to an embodiment of the present disclosure; and

[0034] Figure 10 It shows a schematic diagram of a storage medium according to an embodiment of the present disclosure. Detailed implementation manners

[0035] In this specification and the accompanying drawings, steps and elements that are substantially the same or similar are denoted by the same or similar reference numerals, and repeated descriptions of these steps and elements will be omitted. Meanwhile, in the description of the present disclosure, terms such as "first", "second", etc. are only used for differential description and cannot be construed as indicating or implying relative importance or order.

[0036] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present disclosure belongs. The terms used herein are only for the purpose of describing embodiments of the present invention and are not intended to limit the present invention.

[0037] To facilitate the description of the present disclosure, the following concepts related to the present disclosure are introduced.

[0038] The method of the present disclosure may be based on computer architecture. Computer architecture is a set of rules and methods that describe the functions, organization, and implementation of a computer system. This set of rules and methods is implemented through Instruction Set Architecture (ISA) and Microarchitecture. ISA is the bridge between computer hardware and software. The functions of the hardware are provided through ISA, and software uses the hardware through the instructions specified by ISA. Microarchitecture is the specific implementation of ISA, and for the same ISA, different technologies of microarchitecture can be used. Currently, common instruction sets include x86, EM64T, MMX, SSE, SSE2, SSE3, SSSE3 (Super SSE3), SSE4A, SSE4.1, SSE4.2, AVX, AVX2, AVX-512, VMX, x86-64, 3D-Now!, etc.

[0039] In summary, the solution provided by the embodiments of the present disclosure involves technologies such as computer architecture, microarchitecture, and instruction set. The embodiments of the present disclosure will be further described below with reference to the accompanying drawings.

[0040] Figure 1 is a schematic diagram showing the process of computer processing instructions according to an embodiment of the present disclosure. As Figure 1 shown, a Central Processing Unit (CPU) can be logically divided into three modules: a control unit, an arithmetic unit, and a storage unit. These three parts are connected by an internal bus of the CPU. The operation of the CPU of almost all von Neumann-type computers can be divided into five stages: fetching an instruction, decoding the instruction, executing the instruction, accessing memory to fetch data, and writing back the result.

[0041] Among them, the instruction fetch stage is the process of reading an instruction in memory into a register in the CPU. The program register is used to store the address of the next instruction. After the instruction fetch is completed, it immediately enters the instruction decoding stage. In the instruction decoding stage, the instruction decoder splits and interprets the fetched instruction according to a predetermined instruction format, and identifies and differentiates different instruction categories and various methods of obtaining operands. After decoding is completed, it enters the instruction execution stage. The task of this stage is to complete various operations specified by the instruction and specifically implement the function of the instruction. According to the needs of the instruction, it may be necessary to extract data from memory, thus entering the data access stage. The task of this stage is: according to the instruction address code, obtain the address of the operand in the main memory and read the operand from the main memory for operation. Finally, the result write-back stage writes the operation result data of the instruction execution stage back to the internal register of the CPU so that it can be quickly accessed by subsequent instructions.

[0042] According to the existing instruction processing method, in the case of processing the high and low bits of an operand, the micro-operation separately allocates a physical register for storing the masking number k for the high and low bits of the operand, that is, the k mask physical register (k PhysicalRegister Number, kprf); among them, the masking number k mask is used to perform a masking operation, which can control the corresponding position of the calculation result to be cleared, unchanged, or updated normally; the address of the k mask physical register exists in the physical register number (kprn, kPhysical Register Number). The processor first calculates the low-bit data, then calculates the high-bit data, and then combines the two parts into a physical register. When the subsequent instructions perform operations, the combined physical register is used for processing and calculation. In this instruction processing method, the high and low parts of the operand need to be combined, and the high-bit micro-operation depends on the result of the low-bit micro-operation. The high bit needs to wait for the result of the low-bit processing to be broadcast together, so the program processing parallelism is low, the program processing is slow, and the processor power consumption is large.

[0043] For example, when both the data bit width and the physical register bit width are 256 bits, when processing the AXV 256 instruction, it only needs to be decoded into one micro-operation to allocate a physical register; however, when processing AVX 512, two micro-operations need to be executed twice. At this time, the AVX 512 instruction is decoded into two micro-operations for the high and low bits, and a physical register is separately allocated for the destination logical registers of the high and low bits. The computer first performs the calculation of the lower 32 bits, then performs the calculation of the higher 32 bits, and then combines the calculation result of the higher 32 bits with the lower 32 bits. When the subsequent instructions use the result of this instruction, the combined result stored in the physical register is used for the processing of the subsequent instructions. Due to the merging operation between the high and low bits, the broadcast of the high bit depends on the result of the low-bit processing, and the high bit needs to wait for the result of the low-bit processing to be broadcast.

[0044] According to the method of the present disclosure, when implementing the above instruction processing operation, the high and low parts of the operand do not need to be combined, the high-bit micro-operation does not depend on the result of the low-bit micro-operation, and the high bit can be broadcast without waiting for the result of the low-bit processing.

[0045] For example, when processing AVX 512, the AVX 512 instruction is decoded into two micro-operations. A physical register is allocated for the lower-order destination logical register, and the lower bits of the physical register are used when the computer performs calculations on the lower 32 bits; when performing calculations on the higher 32 bits, no new physical register is allocated. Instead, the higher 32-bit destination logical register is remapped to the physical register allocated during the lower-order processing, and the higher bits of the physical register are used for higher-order calculations. The lower and higher bits are processed in parallel respectively, without a merge operation, and the higher-order processing does not depend on the result of the lower-order processing.

[0046] According to the method of the present disclosure, in the broadcast stage, the lower bits stored in the physical register are broadcast based on the lower bits of the physical register; the higher bits stored in the physical register are broadcast based on the higher bits of the physical register. The broadcast of the lower bits stored in the physical register and the broadcast of the higher bits stored in the physical register are performed independently.

[0047] For example, when processing AVX 512, when the lower-order micro-operation is completed, the lower 32 bits of the result stored in the physical register are broadcast based on the lower bits of the physical register; when the higher-order micro-operation is completed, the higher 32 bits of the result stored in the physical register are broadcast based on the higher bits of the physical register.

[0048] Optionally, the physical register may be assigned an identifier for identifying that its lower bits or higher bits are valid bits. During the broadcast stage of the lower bits stored in the physical register: broadcast the lower bits and higher bits stored in the physical register, and broadcast the identifier for identifying that the lower bits of the physical register are valid bits; and during the broadcast stage of the higher bits stored in the physical register: broadcast the lower bits and higher bits stored in the physical register, and broadcast the identifier for identifying that the higher bits of the physical register are valid bits.

[0049] For example, for the following two instructions:

[0050] Vpcmp k1,{k2},zmm1,zmm2

[0051] Vsub zmm1,{k1},zmm2,zmm3

[0052] According to the method of the present disclosure, the micro-operation sequence corresponding to the instruction is:

[0053] U1:UopK k1_lo,k2,zmm1_lo,zmm2_lo

[0054] U2:UopK k1_hi,k2,zmm1_hi,zmm2_hi

[0055] U3: Uopsub zmm1_lo, k1_lo’, zmm2_lo, zmm3_lo

[0056] U4: Uopsub zmm1_hi, k1_hi’, zmm2_hi, zmm3_hi

[0057] For this example, Vpcmp is used as the first instruction and Vsub is used as the second instruction, and the second instruction depends on the result of the first instruction. It should be understood that the second instruction depending on the result of the first instruction as described in the present disclosure means that the second instruction needs to use the register storing the result of the first instruction during processing. For various instructions of the same or different types, as long as the second instruction needs to use the register storing the result of the first instruction during processing, it can be said that the second instruction depends on the result of the first instruction.

[0058] The micro-operation sequences corresponding to the first instruction Vpcmp are U1 and U2. Among them, zmm1_lo and zmm2_lo respectively represent the source logical registers corresponding to the low-order bits of the operands zmm1 and zmm2, zmm1_hi and zmm2_hi respectively represent the source logical registers corresponding to the high-order bits of the operands zmm1 and zmm2, k2 represents the mask register, and k1_lo and k1_hi respectively represent the destination logical registers corresponding to the low-order micro-operation and the high-order micro-operation. For the sake of convenient expression, hereinafter, the logical register and the bits it stores will no longer be distinguished.

[0059] For the low-order micro-operation U1 and the high-order micro-operation U2, there is 2-bit identifier information stored in the renaming table, which is respectively used to identify whether the physical registers used by U1 and U2 are high-order valid or low-order valid. During the decoding stage, the identifier information in this renaming table is stored in the micro-operation information packets corresponding to the micro-operations U1 and U2. During the broadcast stage of the low-order bits stored in the physical register, the identifier indicating that the low-order bits of the physical register are valid bits in the micro-operation information packet of the micro-operation U1 and the address of this physical register are broadcast together. During the broadcast stage of the high-order bits stored in the physical register, the identifier indicating that the high-order bits of the physical register are valid bits in the micro-operation information packet of the micro-operation U2 and the address of this physical register are broadcast together. The broadcast of the result of the U1 micro-operation and the broadcast of the result of the U2 micro-operation are carried out independently.

[0060] For the low - order micro - operation U3 and the high - order micro - operation U4, the renaming table stores 2 - bit identifier information, which is used to identify whether the physical registers used by U3 and U4 are high - order valid or low - order valid. During the renaming stage, the identifier information in this renaming table is stored in the micro - operation information packets corresponding to the micro - operation U3 and the micro - operation U4. During the broadcast stage of the low - order bits stored in the physical register, the identifier indicating that the low - order bit of the physical register is a valid bit in the micro - operation information packet of the micro - operation U3 and the address of this physical register are broadcast together. During the broadcast stage of the high - order bits stored in the physical register, the identifier indicating that the high - order bit of the physical register is a valid bit in the micro - operation information packet of the micro - operation U4 and the address of this physical register are broadcast together. The broadcast of the result of the U3 micro - operation and the broadcast of the result of the U4 micro - operation are independent of each other.

[0061] It should be noted that the k mask mask numbers k1_lo’ and k1_hi’ in U3 and U4 are determined based on the physical register bit - width type of the instruction and k1_lo and k1_hi in U1 and U2. Although sometimes k1_lo’ and k1_hi’ in U3 and U4 are also written as k1_lo and k1_hi, in fact, the k mask mask numbers in U3 and U4 at this time have been processed by the computer and are not the same as k1_lo and k1_hi in U1 and U2.

[0062] Using the instruction processing method of the present disclosure eliminates the dependence between the high - order and low - order parts when processing them separately, improves the parallelism of program processing. When processing the high - order part, there is no need to wait for the processing result of the low - order part, and broadcasting can be advanced.

[0063] It should be understood that the method of the present disclosure is applicable to all instruction operations of various bit - widths based on the micro - architecture. For example, when processing single - instruction multiple - data (SIMD) instructions with a bit - width of 512 bits, due to the limited register bit - width, the high - order and low - order information stored in the physical register is processed and broadcast separately; when processing ordinary instructions with a bit - width of 64 bits, due to the sufficient register bit - width, the high - order and low - order information stored in the physical register may be broadcast simultaneously.

[0064] Figure 2A is a flowchart showing an instruction processing method 200 according to an embodiment of the present disclosure.

[0065] As Figure 2A shown, in step S201, the instruction starts to be processed, and the processing of this instruction in the computer is decomposed into two micro - operations based on the low - order and high - order parts of the operand respectively.

[0066] For the low - order micro - operation and high - order micro - operation of the instruction, the low - order bits of the operand of the instruction are obtained in step S202, and the high - order bits of the operand of the instruction are obtained in step S212, respectively.

[0067] Based on the low - order bits of the instruction operand obtained in step S202, in step S203, the low - order micro - operation corresponding to the instruction starts to process and execute. At the same time, based on the high - order bits of the instruction operand obtained in step S212, in step S213, the high - order micro - operation corresponding to the instruction starts to process and execute. It should be noted that at this time, step S203 and step S213 can be executed in parallel or sequentially, and there is no mutual association when the high - order micro - operation in step S213 and the low - order micro - operation in step S203 are executed.

[0068] In step S204, when the low - order micro - operation is processed, a physical register is allocated for the low - order micro - operation, and the execution result of the low - order micro - operation is stored in the low - order bits of the physical register of the instruction.

[0069] In step S214, when the high - order micro - operation is processed, no physical register is allocated for the high - order micro - operation, but the execution result of the high - order micro - operation is directly stored in the high - order bits of the physical register allocated for the low - order micro - operation.

[0070] Figure 2B It is an exemplary flowchart showing an instruction processing method when there is a sequential dependency relationship between two adjacent instructions according to an embodiment of the present disclosure. Specifically, the two adjacent instructions include a first instruction and a second instruction, and the operation of the second instruction depends on the result of the first instruction.

[0071] As Figure 2B shown, according to an embodiment of the present disclosure, when processing two instructions, a first instruction and a second instruction, and the second instruction depends on the result of the first instruction, the operations of processing the first instruction may include (as Figure 2B shown in the first - instruction processing box): obtaining the high - order bits and low - order bits of the operand of the first instruction and the mask number of the first instruction; based on the obtained low - order bits of the operand and the mask number, performing a low - order micro - operation, and storing the result of the low - order micro - operation in the low - order bits of the physical register for the instruction; based on the obtained high - order bits of the operand and the mask number, performing a high - order micro - operation, and storing the result of the high - order micro - operation in the high - order bits of the physical register.

[0072] Then, based on the operations of processing the first instruction, the operations of processing the second instruction may include (as Figure 2BAs shown in the second instruction processing box: for the low - order micro - operation and high - order micro - operation of the second instruction respectively, determine the valid bits of the corresponding source logical register, and perform operations respectively based on the determined valid bits of the source logical register.

[0073] For the low - order micro - operation of the second instruction, according to the bit - width type of the physical register of the first instruction and the bit - width type of the source logical register of the second instruction, based on the low - order bits and / or high - order bits of the physical register of the first instruction, determine the low - order bits of the source logical register for the low - order micro - operation of the second instruction, and based on the low - order bits of the operand of the second instruction and the low - order bits of the source logical register for the low - order micro - operation of the second instruction, perform the low - order micro - operation of the second instruction, where the low - order bits of the source logical register for the low - order micro - operation of the second instruction are used to store the mask number of the low - order micro - operation of the second instruction.

[0074] For the high - order micro - operation of the second instruction, according to the bit - width type of the physical register of the first instruction and the bit - width type of the source logical register of the second instruction, based on the low - order bits or high - order bits of the physical register of the first instruction, determine the high - order bits of the source logical register for the high - order micro - operation of the second instruction, and based on the high - order bits of the operand of the second instruction and the high - order bits of the source logical register of the second instruction, perform the high - order micro - operation of the second instruction, where the high - order bits of the source logical register for the high - order micro - operation of the second instruction are used to store the mask number of the high - order micro - operation of the second instruction.

[0075] For example, for the following two instructions:

[0076] Vpcmp k1,{k2},zmm1,zmm2

[0077] Vsub zmm1,{k1},zmm2,zmm3

[0078] According to the method of the present disclosure, the micro - operation sequence corresponding to the instruction is:

[0079] U1:UopK k1_lo,k2,zmm1_lo,zmm2_lo

[0080] U2:UopK k1_hi,k2,zmm1_hi,zmm2_hi

[0081] U3:Uopsub zmm1_lo,k1_lo’,zmm2_lo,zmm3_lo

[0082] U4:Uopsub zmm1_hi,k1_hi’,zmm2_hi,zmm3_hi

[0083] For this example, Vpcmp is used as the first instruction and Vsub is used as the second instruction. The micro-operation sequence corresponding to the first instruction Vpcmp is U1 and U2. Among them, zmm1_lo and zmm2_lo respectively represent the source logical registers corresponding to the low-order bits of the operands zmm1 and zmm2, zmm1_hi and zmm2_hi respectively represent the source logical registers corresponding to the high-order bits of the operands zmm1 and zmm2, k2 represents the mask register, and k1_lo and k1_hi respectively represent the destination logical registers corresponding to the low-order micro-operation and the high-order micro-operation. For the convenience of description, in the following text, the logical register and the bits it stores will no longer be distinguished.

[0084] For the first instruction, its instruction processing includes:

[0085] Obtain the high-order bits (zmm1_hi and zmm2_hi) and low-order bits (zmm1_lo and zmm2_lo) of the operands of the first instruction and the k mask number k2 of the first instruction; in U1, based on the obtained low-order bits (zmm1_lo and zmm2_lo) of the operands and the k mask number k2, perform a low-order micro-operation, and store the result of the low-order micro-operation in the low-order bit (k1_lo) of the physical register used for the instruction; in U2, based on the obtained high-order bits (zmm1_hi and zmm2_hi) of the operands and the k mask number k2, perform a high-order micro-operation, and store the result of the high-order micro-operation in the high-order bit (k1_hi) of the physical register.

[0086] The micro-operation sequence corresponding to the second instruction is U3 and U4. The micro-operation corresponding to the second instruction Vsub depends on the result of the micro-operation corresponding to the first instruction Vpcmp. Among them, zmm2_lo and zmm3_lo respectively represent the source logical registers corresponding to the low-order bits of the operands zmm2 and zmm3, zmm2_hi and zmm3_hi respectively represent the source logical registers corresponding to the high-order bits of the operands zmm2 and zmm3, k2 represents the mask register, and k1_lo’ and k1_hi’ respectively represent the source logical registers used to store the mask number corresponding to the low-order micro-operation and the high-order micro-operation. For the convenience of description, in the following text, the logical register and the bits it stores will no longer be distinguished.

[0087] Based on the bit-width type of the physical register of the first instruction and the bit-width type of the source logical register of the second instruction, determine the low-order bit (i.e., k1_lo’) of the source logical register for the high-order micro-operation of the second instruction based on the low-order bit (k1_lo) or high-order bit (k1_hi) of the physical register of the first instruction, and execute the low-order micro-operation of the second instruction based on the low-order bits (zmm2_lo and zmm3_lo) of the operands of the second instruction and the low-order bit (k1_lo’) of the source logical register of the low-order micro-operation of the second instruction, where the low-order bit of the source logical register of the low-order micro-operation of the second instruction is used to store the k mask mask number (k1_lo’) of the low-order micro-operation of the second instruction.

[0088] Based on the bit-width type of the physical register of the first instruction and the bit-width type of the source logical register of the second instruction, determine the high-order bit (i.e., k1_hi’) of the source logical register for the high-order micro-operation of the second instruction based on the low-order bit (k1_lo) or high-order bit (k1_hi) of the physical register of the first instruction, and execute the high-order micro-operation of the second instruction based on the high-order bits (zmm2_hi and zmm3_hi) of the operands of the second instruction and the high-order bit (k1_hi’) of the source logical register of the second instruction, where the high-order bit of the source logical register of the high-order micro-operation of the second instruction is used to store the k mask mask number (k1_hi’) of the high-order micro-operation of the second instruction.

[0089] It should be noted that the k mask mask numbers k1_lo’ and k1_hi’ in U3 and U4 are determined based on the bit-width type of the physical register of the instruction according to k1_lo and k1_hi in U1 and U2. Although sometimes k1_lo’ and k1_hi’ in U3 and U4 are also written as k1_lo and k1_hi, in fact, the k mask mask numbers in U3 and U4 at this time have been processed by the computer and are not equivalent to k1_lo and k1_hi in U1 and U2.

[0090] Optionally, the bit-width type is any one of byte, word, double-word, and quadruple-word, and different bit-width types have different effective bit-widths.

[0091] Figure 3 is a schematic diagram showing the data dependence relationship in the buffer according to an embodiment of the present disclosure.

[0092] As Figure 3As shown, data processing depends on producer micro-operations and consumer micro-operations. After the producer micro-operation is executed, data is written into a buffer, and the consumer micro-operation reads and uses the data from the buffer. The producer micro-operation and the consumer micro-operation share a buffer that is initially empty and has a determined size. Only when the buffer is ready to store data can the producer micro-operation put the result generated by the execution of the producer micro-operation into the buffer, otherwise it must wait. Only when the data relied on by the consumer micro-operation in the buffer is ready can the consumer micro-operation take out the data from the buffer for use, otherwise it must wait.

[0093] For two related instructions during execution, the result of the micro-operation of the previous instruction (i.e., the producer) is stored in the destination logical register of this instruction. For the micro-operation of the next instruction (i.e., the consumer), the destination logical register of the previous instruction corresponds to the source logical register of the next instruction. Among them, both the source logical register and the destination logical register are logical registers, and this logical register corresponds to a physical register.

[0094] For example, for the following two instructions:

[0095] Vpcmp k1,{k2},zmm1,zmm2

[0096] Vsub zmm1,{k1},zmm2,zmm3

[0097] In the first instruction Vpcmp, the operands zmm1 and zmm2 are masked by the mask number in the source logical register k2 and the result is stored in the destination logical register k1. In the second instruction Vsub, the operands zmm2 and zmm3 are masked by the mask number in the source logical register k1 and the result is stored in the number zmm1. The second instruction Vsub depends on the result of the first instruction Vpcmp, and the destination logical register of the first instruction corresponds to the source logical register of the second instruction.

[0098] It should be noted that the destination logical register of the first instruction corresponds to the source logical register of the second instruction, but they are not exactly the same. According to the bit-width types of the destination logical register of the first instruction and the source logical register of the second instruction, the computer calculates and determines the association relationship between the low-order bits and high-order bits of the source logical register of the second instruction and the low-order bits and high-order bits of the destination logical register of the first instruction.

[0099] Figure 4 It is a schematic diagram showing the positions of the valid bits of different bit-width types in the physical register according to an embodiment of the present disclosure.

[0100] It should be understood that the masking register is 64 bits. For 512-bit data, the 0-31 lower bits of the masking register are used to mask the lower bits 0-255 of the data, and the 32-63 higher bits of the masking register are used to mask the higher bits 256-511 of the data.

[0101] As Figure 4 shown, when the data type is Byte (B), the valid bits in the lower part of the masking register occupy bits 0-31 in the register, and the valid bits in the higher part of the data occupy bits 32-63 in the register. That is, 1 bit in the masking register is used to mask the corresponding 8 bits in the data, and thus 256 bits of the data are masked by 32 bits. As Figure 4 shown in row B of

[0102] When the data type is Word (W), the valid bits in the lower part of the masking register occupy bits 0-15 in the register, and the valid bits in the higher part of the data occupy bits 32-47 in the register. That is, 1 bit in the masking register is used to mask the corresponding 16 bits in the data, and thus 256 bits of the data are masked by 16 bits. As Figure 4 shown in row W of

[0103] When the data type is Dword (D), the valid bits in the lower part of the masking register occupy bits 0-7 in the register, and the valid bits in the higher part of the data occupy bits 32-39 in the register. That is, 1 bit in the masking register is used to mask the corresponding 32 bits in the data, and thus 256 bits of the data are masked by 8 bits. As Figure 4 shown in row D of

[0104] When the data type is Qword (Q), the valid bits in the lower part of the masking register occupy bits 0-3 in the register, and the valid bits in the higher part of the data occupy bits 32-35 in the register. That is, 1 bit in the masking register is used to mask the corresponding 64 bits in the data, and thus 256 bits of the data are masked by 4 bits. AsFigure 4 As shown in row Q, each ellipse corresponds to 4 bits. At this time, the 1 dark ellipse on the left represents that the 4 bits are valid bits, and the 256 bits of the data are masked through the 4 valid bits.

[0105] It can be seen that when the effective bit width of the bit width type of the physical register of the first instruction is equal to the effective bit width of the bit width type of the source logical register of the second instruction, the valid bits used in the source logical register of the second instruction correspond one-to-one with the valid bits in the physical register of the first instruction; when the effective bit width of the bit width type of the physical register of the first instruction is less than the effective bit width of the bit width type of the source logical register of the second instruction, the source logical register of the second instruction only needs to use some of the valid bits in the physical register of the first instruction; when the effective bit width of the bit width type of the physical register of the first instruction is greater than the effective bit width of the bit width type of the source logical register of the second instruction, the source logical register of the second instruction needs to use the high and low valid bits in the physical register of the first instruction, and fill 0 for the insufficient bits.

[0106] The following gives specific instructions and their corresponding micro-operations to illustrate the superiority of the instruction operation method of the present disclosure over the existing instruction processing method.

[0107] For example, for the following instruction example,

[0108] Vpcmp k1,{k2},zmm1,zmm2

[0109] Vsub zmm1,{k1},zmm2,zmm3

[0110] According to the existing instruction processing method, the micro-operations corresponding to this instruction are:

[0111] U1:UopK k_tmp,k2,zmm1_lo,zmm2_lo

[0112] U2:UopmK k1,k2,zmm1_hi,zmm2_hi,k_tmp

[0113] U3:Uopsub zmm1_lo,k1,zmm2_lo,zmm3_lo

[0114] U4:Uopsub zmm1_hi,k1,zmm2_hi,zmm3_hi

[0115] Taking Vpcmp as the first instruction and Vsub as the second instruction, the execution of the second instruction depends on the result of the first instruction; the first instruction Vpcmp is processed separately for high and low bits, corresponding to the micro-operations U1 and U2; the second instruction Vsub is processed separately for high and low bits, corresponding to the micro-operations U3 and U4.

[0116] In micro-operation U1, the lower bits of operand zmm1 and the lower bits of operand zmm2 are compared, and after being masked by k2, they are stored in the destination logical register k_tmp;

[0117] In micro-operation U2, the upper bits of operand zmm1 and the upper bits of operand zmm2 are compared, and after being masked by k2, they are merged with the result in k_tmp, and the merged result is stored in k1;

[0118] In micro-operation U3, the lower bits of operand zmm2 and the lower bits of operand zmm3 are subtracted, and after being masked by the source logical register k1, they are stored in zmm1_lo;

[0119] In micro-operation U4, the upper bits of operand zmm2 and the upper bits of operand zmm3 are subtracted, and after being masked by the source logical register k1, they are stored in zmm1_hi.

[0120] According to an embodiment of the present disclosure, when the bit-width type of the compare instruction as the producer and the bit-width type of the sub instruction as the consumer are the same (in this article, this situation is all referred to as Scenario 1), for the instruction processing method of the present disclosure, the micro-operations corresponding to the instruction are:

[0121] U1: UopK k1_lo, k2, zmm1_lo, zmm2_lo

[0122] U2: UopK k1_hi, k2, zmm1_hi, zmm2_hi

[0123] U3: Uopsub zmm1_lo, k1_lo, zmm2_lo, zmm3_lo

[0124] U4: Uopsub zmm1_hi, k1_hi, zmm2_hi, zmm3_hi

[0125] In micro-operation U1, the lower bits of operand zmm1 and the lower bits of operand zmm2 are compared, and after being masked by k2, they are stored in the lower bits of the destination logical register, that is, k1_lo;

[0126] In micro-operation U2, the upper bits of operand zmm1 and the upper bits of operand zmm2 are compared, and after being masked by k2, they are stored in the upper bits of the destination logical register, that is, k1_hi;

[0127] In micro-operation U3, the lower bits of operand zmm2 and the lower bits of operand zmm3 are subtracted, and after being masked by k1_lo, they are stored in zmm1_lo;

[0128] In the micro-operation U4, the high bits of the operand zmm2 and the high bits of the operand zmm3 are subtracted. After being masked by k1_hi, they are stored in zmm1_hi.

[0129] By comparing the micro-operation sequences of the existing instruction processing with those of the instruction processing of the present disclosure, it can be seen that for the existing solution, a total of two physical registers, k_tmp and k1, are allocated in the micro-operations U1 and U2 corresponding to the first instruction Vpcmp, and the processing results of the low bits and high bits are merged in the U2 micro-operation (i.e., UopmK); while for the method of the present disclosure, only one physical register, k1, is allocated in the micro-operations U1 and U2 corresponding to the first instruction Vpcmp. The low bits and high bits of k1 are respectively used for the low-bit micro-operation and the high-bit micro-operation. Only the result of the high-bit is processed in the U2 micro-operation (i.e., UopK), and there is no merging operation. Since the merging operation of the low-bit and high-bit processing results is cancelled in the method of the present disclosure, the low-bit and high-bit micro-operations are performed independently and in parallel. The high-bit operation does not depend on the low-bit operation, and the high-bit operation does not need to wait for the processing result of the low-bit. Therefore, the processing result of the high-bit can be broadcast in advance, thereby removing the data dependence redundancy between the low-bit and high-bit micro-operations, improving the parallelism of the program operation, and effectively enhancing the program operation efficiency.

[0130] It should be understood that the instructions in the examples are only for the convenience of those skilled in the art to understand or implement the content of the present disclosure, and are not for limitation. Other instructions based on the micro-architecture can all implement the method of the present disclosure.

[0131] According to the embodiments of the present disclosure, when the bit-width types of the compare instruction as the producer and the sub instruction as the consumer are the same (i.e., scenario 1), the schematic diagram of the comparison of the instruction processing timings between the solution of the present invention and the existing solution is as Figure 5A shown.

[0132] From Figure 5A it can be seen that for the situation of scenario 1, in the prior art, the micro-operation U1 is completed in the machine cycle T0, the micro-operation U2 is completed in the machine cycle T1, and the micro-operations U3 and U4 are completed in the machine cycle T2. While according to the solution of the present invention, the micro-operation U1 is completed in the machine cycle T0, the micro-operation U2 is completed in the machine cycle T0, and the micro-operations U3 and U4 are completed in the machine cycle T1. The solution of the present invention can advance the processing cycle of the micro-operation by one machine cycle compared with the existing solution.

[0133] According to an embodiment of the present disclosure, when the bit-width type of the compare instruction as the producer is less than the bit-width type of the sub instruction as the consumer (in this text, this situation is all referred to as Scenario 2), for the instruction processing method of the present disclosure, the micro-operations corresponding to the instruction are as follows:

[0134] U1: UopK k1_lo, k2, zmm1_lo, zmm2_lo

[0135] U2: UopK k1_hi, k2, zmm1_hi, zmm2_hi

[0136] U3: Uopsub zmm1_lo, k1_lo, zmm2_lo, zmm3_lo

[0137] U4: Uopsub zmm1_hi, k1_lo, zmm2_hi, zmm3_hi

[0138] In micro-operation U1, the lower bits of operand zmm1 and the lower bits of operand zmm2 are compared, and after being masked by k2, they are stored in the lower bits of the destination logical register, that is, k1_lo;

[0139] In micro-operation U2, the upper bits of operand zmm1 and the upper bits of operand zmm2 are compared, and after being masked by k2, they are stored in the upper bits of the destination logical register, that is, k1_hi;

[0140] In micro-operation U3, the lower bits of operand zmm2 and the lower bits of operand zmm3 are subtracted, and after being masked by k1_lo, they are stored in zmm1_lo;

[0141] In micro-operation U4, the upper bits of operand zmm2 and the upper bits of operand zmm3 are subtracted, and after being masked by k1_lo, they are stored in zmm1_hi.

[0142] Therefore, it can be obtained that in the case where the bit-width type of the compare instruction as the producer is less than the bit-width type of the sub instruction as the consumer (that is, Scenario 2) according to the embodiment of the present disclosure, the schematic diagram of the instruction processing timing comparison between the solution of the present disclosure and the existing solution is the same as that in the case where the bit-width type of the compare instruction as the producer is the same as the bit-width type of the sub instruction as the consumer (that is, Scenario 1), specifically as Figure 5A shown.

[0143] From Figure 5AIt can be seen that when the bit-width type of the compare instruction as the producer is less than the bit-width type of the sub instruction as the consumer (i.e., scenario 2), in the prior art, the micro-operation U1 is completed in machine cycle T0, the micro-operation U2 is completed in machine cycle T1, and the micro-operations U3 and U4 are completed in machine cycle T2. According to the solution of the present invention, the micro-operation U1 is completed in machine cycle T0, the micro-operation U2 is completed in machine cycle T0, and the micro-operations U3 and U4 are completed in machine cycle T1. The solution of the present invention can advance the processing cycle of the micro-operation by one machine cycle compared with the existing solution.

[0144] According to an embodiment of the present disclosure, when the bit-width type of the compare instruction as the producer is greater than the bit-width type of the sub instruction as the consumer (in this article, this situation is referred to as scenario 3), for the instruction processing method of the present disclosure, the micro-operations corresponding to the instruction are:

[0145] U1: UopK k1_lo, k2, zmm1_lo, zmm2_lo

[0146] U2: UopK k1_hi, k2, zmm1_hi, zmm2_hi

[0147] U3: Uopsub zmm1_lo, k1_lo, k1_hi, zmm2_lo, zmm3_lo

[0148] U4: Uopsub zmm1_hi, ZERO, zmm2_hi, zmm3_hi

[0149] At this time, the k mask value used in U4 is 0.

[0150] In the micro-operation U1, the low bits of the operand zmm1 and the low bits of the operand zmm2 are compared, and after being masked by k2, they are stored in the low bits of the destination logical register, that is, k1_lo;

[0151] In the micro-operation U2, the high bits of the operand zmm1 and the high bits of the operand zmm2 are compared, and after being masked by k2, they are stored in the high bits of the destination logical register, that is, k1_hi;

[0152] In the micro-operation U3, the low bits of the operand zmm2 and the low bits of the operand zmm3 are subtracted, and after being masked by k1_lo and k1_hi, they are stored in zmm1_lo;

[0153] In the micro-operation U4, the high bits of the operand zmm2 and the high bits of the operand zmm3 are subtracted, and after being masked by 0, they are stored in zmm1_hi.

[0154] Therefore, according to the embodiments of the present disclosure, when the bit-width type of the compare instruction as the producer is greater than the bit-width type of the sub instruction as the consumer (i.e., scenario 3), the schematic diagram of the instruction processing timing comparison between the solution of the present invention and the existing solution is as Figure 5B shown.

[0155] From Figure 5B it can be seen that for the situation of scenario 3, in the prior art, the micro-operation U1 is completed in the machine cycle T0, the micro-operation U2 is completed in the machine cycle T1, and the micro-operations U3 and U4 are completed in the machine cycle T2. According to the solution of the present invention, the micro-operation U1 is completed in the machine cycle T0, the micro-operation U2 is completed in the machine cycle T0, the micro-operation U3 is completed in the machine cycle T1, and the micro-operation U4 is completed in the machine cycle T0. The solution of the present invention can advance the processing cycle of the micro-operations by one machine cycle compared with the existing solution.

[0156] By comparing the solution of the present invention with the existing solution in the above three scenarios, it can be concluded that: the solution of the present disclosure cancels the merging operation of the high and low bits of the operands during the instruction calculation process, processes, calculates, and broadcasts the high and low bits of the operands separately, so that when the instruction is processed by high and low bits, the high and low bit operations of the operands only depend on the register bits required for the operation, without waiting for the processing results of other bits. Therefore, the processing result of the instruction can be broadcast in advance, improving the efficiency and parallelism of instruction processing.

[0157] Through Figure 5A and Figure 5B the description of the examples in it can be seen that for any one of the above three scenarios, the solution of the present disclosure can reduce the above program running time from 3 machine cycles to 2 machine cycles. For the above instructions, in an ideal situation, the method of the present disclosure can reduce the program running time by 33%.

[0158] In the case of scenario 1, taking the bit-width types of the first instruction and the second instruction both being bytes (B) as an example, Figure 6A the schematic diagram showing the corresponding relationship between the bits in the physical register of the first instruction and the source logical register of the second instruction is shown.

[0159] As Figure 6AAs shown, when the bit-width types of the first instruction and the second instruction are both B, the lower bits (i.e., bits 0 - 31) of the source logical register for the lower micro-operations of the second instruction depend on the lower bits (i.e., bits 0 - 31) of the effective bit-width of the physical register of the first instruction; the higher bits (i.e., bits 32 - 63) of the source logical register for the higher micro-operations of the second instruction depend on the higher bits (i.e., bits 32 - 63) of the effective bit-width of the physical register of the first instruction.

[0160] Similarly, when the effective bit-width of the bit-width type of the physical register of other first instructions is equal to the effective bit-width of the bit-width type of the source logical register of the second instruction, the situation is similar to the processing when the bit-width types of the first instruction and the second instruction are both B, and will not be elaborated here.

[0161] In the case of Scenario 2, taking the bit-width type of the first instruction as byte (B) and the bit-width type of the second instruction as word (W) as an example, Figure 6B A schematic diagram showing the corresponding relationship of bit positions between the physical register of the first instruction and the source logical register of the second instruction is shown.

[0162] As Figure 6B shown, when the bit-width type of the first instruction is B and the bit-width type of the second instruction is W, the lower bits (i.e., bits 0 - 15) of the source logical register for the lower micro-operations of the second instruction depend on the lower bits (i.e., bits 0 - 15) of the effective bit-width of the physical register of the first instruction; the higher bits (i.e., bits 32 - 47) of the source logical register for the higher micro-operations of the second instruction depend on the higher bits (i.e., bits 16 - 31) of the effective bit-width of the physical register of the first instruction. That is, at this time, both the lower bits and the higher bits of the source logical register of the second instruction depend on the lower bits of the physical register of the first instruction.

[0163] Similarly, when the effective bit-width of the bit-width type of the physical register of other first instructions is less than the effective bit-width of the bit-width type of the source logical register of the second instruction, the situation is similar to the processing when the bit-width type of the first instruction is B and the bit-width type of the second instruction is W, and will not be elaborated here.

[0164] In the case of Scenario 3, taking the bit-width type of the first instruction as word (W) and the bit-width type of the second instruction as byte (B) as an example, Figure 6C A schematic diagram showing the corresponding relationship of bit positions between the physical register of the first instruction and the source logical register of the second instruction is shown.

[0165] As Figure 6CAs shown, when the bit-width type of the first instruction is W and the bit-width type of the second instruction is B, in the physical register of the first instruction, the low-order valid bits and the high-order valid bits are concatenated. The low-order part of the source logical register of the second instruction depends on the concatenated valid bits for calculation; after the concatenation of the low-order valid bits and the high-order valid bits, all the high-order bits of the physical register of the first instruction are 0. At this time, the high-order bits of the source logical register of the high-order micro-operation of the second instruction no longer depend on the effective bit-width of the physical register of the first instruction.

[0166] At this time, the low-order bits (i.e., bits 0 - 31) of the source logical register for the low-order micro-operation of the second instruction depend on the low-order bits (i.e., bits 0 - 15) and the high-order bits (i.e., bits 32 - 47) of the effective bit-width of the physical register of the first instruction; the high-order bits (i.e., bits 32 - 63) of the source logical register for the high-order micro-operation of the second instruction do not depend on the effective bit-width of the physical register of the first instruction.

[0167] Similarly, when the effective bit-width of the bit-width type of the physical register of other first instructions is greater than the effective bit-width of the bit-width type of the source logical register of the second instruction, the situation is similar to the processing when the bit-width type of the first instruction is W and the bit-width type of the second instruction is B, which will not be elaborated here.

[0168] By summarizing the above scenarios, Table 1 shows the dependency relationship of the low-order part of the source logical register of the second instruction on the low-order (low) and high-order (high) parts of the destination logical register of the first instruction.

[0169] Table 1

[0170] consumer\producer BYTE WORD DWORD QWORD BYTE low low&high low&high low&high WORD low low low&high low&high DWORD low low low low&high QWORD low low low low

[0171] For Scenario 1, as shown in the content on the diagonal of Table 1, the effective bit-width of the bit-width type of the physical register of the first instruction is equal to the effective bit-width of the bit-width type of the source logical register of the second instruction. The low-order bits of the source logical register for the low-order micro-operation of the second instruction all depend on the low-order bits of the effective bit-width of the physical register of the first instruction.

[0172] For Scenario 2, as shown in the content below the diagonal of Table 1, the effective bit-width of the bit-width type of the physical register of the first instruction is less than the effective bit-width of the bit-width type of the source logical register of the second instruction. The low-order bits of the source logical register for the low-order micro-operation of the second instruction depend on the low-order bits of the effective bit-width of the physical register of the first instruction.

[0173] For scenario 3, as shown in the content above the diagonal of Table 1, the effective bit width of the bit width type of the physical register of the first instruction is greater than the effective bit width of the bit width type of the source logical register of the second instruction. The lower bits of the source logical register for the lower micro-operations of the second instruction depend on the lower bits and the higher bits of the effective bit width of the physical register of the first instruction.

[0174] By summarizing the above scenarios, Table 2 shows the dependency relationship of the higher bits of the source logical register of the second instruction on the lower bits (low) and higher bits (high) of the destination logical register of the first operation.

[0175] Table 2

[0176] consumer\producer BYTE WORD DWORD QWORD BYTE high zero zero zero WORD low high zero zero DWORD low low high zero QWORD low low low high

[0177] For scenario 1, as shown in the content of the diagonal of Table 2, the effective bit width of the bit width type of the physical register of the first instruction is equal to the effective bit width of the bit width type of the source logical register of the second instruction. The higher bits of the source logical register for the higher micro-operations of the second instruction depend on the higher bits of the effective bit width of the physical register of the first instruction.

[0178] For scenario 2, as shown in the content below the diagonal left of Table 2, the effective bit width of the bit width type of the physical register of the first instruction is less than the effective bit width of the bit width type of the source logical register of the second instruction. The higher bits of the source logical register for the higher micro-operations of the second instruction depend on the lower bits of the effective bit width of the physical register of the first instruction.

[0179] For scenario 3, as shown in the content above the diagonal right of Table 2, the effective bit width of the bit width type of the physical register of the first instruction is greater than the effective bit width of the bit width type of the source logical register of the second instruction. The higher bits of the source logical register for the higher micro-operations of the second instruction do not depend on the effective bit width of the physical register of the first instruction.

[0180] According to an embodiment of the present disclosure, Table 3 shows the correspondence between instruction operation types and type encodings.

[0181] Table 3

[0182] Operation type Type code SSE BYTE / General instruction 0 SSE WORD 1 SSE DWORD 2 SSE QWORD 3

[0183] According to the correspondence between operation types and type encodings in Table 3, for the second instruction that depends on the result of the first instruction, the lower bits (k_real_low) of the k mask actually used in its functional unit can be obtained according to the following calculation, where TD is the encoding of the operation bit width type of the first instruction according to Table 3, and TS is the encoding of the operation bit width type of the second instruction according to Table 3.

[0184] k_real_low[31:0] = low[31:0] |

[0185] {{16{~TD[0] & TD[1]}} & high[15:0], 16’b0} |

[0186] {16’b0, {8{TD[1] & ~TD[0]}} & high[7:0]}, 8’b0} |

[0187] {24’b0, {4{TD[1] & TD[0]}} & high[3:0]}, 4’b0} |

[0188] When the operation type of the second instruction is BYTE, the 31 - 0 bits of k_real_low actually used by the second instruction respectively correspond to the 31 - 0 bits of the k mask of the first instruction;

[0189] When the operation type of the second instruction is WORD, the 15 - 0 bits of k_real_low actually used by the second instruction correspond to the lower 15 - 0 bits of the k mask of the first instruction, and the 31 - 16 bits of k_real_low correspond to the upper 15 - 0 bits of the k mask of the first instruction;

[0190] When the operation type of the second instruction is DWORD, the 7 - 0 bits of k_real_low actually used by the second instruction correspond to the lower 7 - 0 bits of the k mask of the first instruction, and the 8 - 15 bits of k_real_low correspond to the upper 7 - 0 bits of the k mask of the first instruction; the 31 - 16 bits of k_real_low correspond to 16 zeros;

[0191] When the operation type of the second instruction is QWORD, the 3 - 0 bits of k_real_low actually used by the second instruction correspond to the lower 3 - 0 bits of the k mask of the first instruction, and the 7 - 4 bits of k_real_low correspond to the upper 3 - 0 bits of the k mask of the first instruction; the 31 - 8 bits of k_real_low correspond to 24 zeros.

[0192] It should be noted that the lower and upper bits of the k mask of the first instruction here are after the computer splicing of valid bits.

[0193] According to the correspondence between the operation type and the type code, for the second instruction that depends on the result of the first instruction, the upper bits (k_real_high) of the k mask actually used in its functional unit can be obtained according to the following calculation, where TD is the destination logical register of the first instruction and TS is the source logical register of the second instruction.

[0194] k_real_high[31:0] = {32{TD == TS}} & high[31:0] |

[0195] {16’b0, {{16{TD < TS}} & {{16{~TS[1] & TS[0]}} & low[31:16]}}} |

[0196] {24’b0, {{8{TD < TS}} & {{8{TS[1] & ~TS[0]}} & low[15:8]}}} |

[0197] {28’b0, {{4{TD < TS}} & {{4{TS[1] & ~TS[0]}} & low[7:4]}}}

[0198] When the operation type of the second instruction is BYTE, the 31-0 bits of k_real_high actually used by the second instruction respectively correspond to the 31-0 bits of the k mask of the first instruction;

[0199] When the operation type of the second instruction is WORD, the 15-0 bits of k_real_high actually used by the second instruction correspond to the 31-16 bits of the lower bits of the k mask of the first instruction, and the 31-16 bits of k_real_high correspond to 16 zeros;

[0200] When the operation type of the second instruction is DWORD, the 7-0 bits of k_real_high actually used by the second instruction correspond to the 15-8 bits of the lower bits of the k mask of the first instruction, and the 31-8 bits of k_real_high correspond to 24 zeros;

[0201] When the operation type of the second instruction is QWORD, the 3-0 bits of k_real_high actually used by the second instruction correspond to the 7-4 bits of the lower bits of the k mask of the first instruction, and the 31-4 bits of k_real_high correspond to 28 zeros.

[0202] It should be noted that the lower and upper bits of the k mask of the first instruction here are after the computer splicing of valid bits.

[0203] Figure 7 It is a schematic diagram showing an instruction processing apparatus 700 according to an embodiment of the present disclosure.

[0204] The instruction processing 700 may include an operand acquisition module 701 and a micro-operation processing 702.

[0205] According to an embodiment of the present disclosure, the operand acquisition module 701 may be configured to acquire the high-order bits and low-order bits of the operands of the instruction respectively. The micro-operation processing module 702 may be configured to execute low-order micro-operations based on the acquired low-order bits of the operands and store the results of the low-order micro-operations in the low-order bits of the physical register for the instruction; and execute high-order micro-operations based on the acquired high-order bits of the operands and store the results of the high-order micro-operations in the high-order bits of the physical register, wherein the micro-operation processing of the high-order bits is independent of the micro-operation processing of the low-order bits.

[0206] Optionally, in the micro-operation processing module 702, the physical register may also be allocated for the destination logical register of the low-order micro-operations of the instruction; and the destination logical register of the high-order micro-operations of the instruction may be mapped to the physical register.

[0207] Optionally, in order to broadcast the results of the low-order micro-operations and the high-order micro-operations separately, the results of the low-order micro-operations may be stored in the low-order bits of the physical register for the instruction, and the low-order bits stored in the physical register may be broadcast; and the results of the high-order micro-operations may be stored in the high-order bits of the physical register, and the high-order bits stored in the physical register may be broadcast.

[0208] Optionally, the physical register is assigned an identifier for identifying that its low-order bits or high-order bits are valid bits. Among them, broadcasting the low-order bits stored in the physical register includes: broadcasting the low-order bits and high-order bits stored in the physical register, and broadcasting the identifier for identifying that the low-order bits of the physical register are valid bits; and broadcasting the high-order bits stored in the physical register includes: broadcasting the low-order bits and high-order bits stored in the physical register, and broadcasting the identifier for identifying that the high-order bits of the physical register are valid bits.

[0209] Optionally, when the instruction processing device 700 processes the first instruction and the second instruction, and the second instruction depends on the result of the first instruction, the instruction processing device further includes: an association determination module, which determines the association relationship between the low-order bits and high-order bits of the source logical register of the second instruction and the low-order bits and high-order bits of the destination logical register of the first instruction according to the bit-width type of the physical register of the first instruction and the bit-width type of the source logical register of the second instruction, and the micro-operation processing module is further configured to: determine the low-order bits of the source logical register for the low-order micro-operation of the second instruction and the high-order bits of the source logical register for the high-order micro-operation of the second instruction according to the bit-width type of the physical register of the first instruction and the bit-width type of the source logical register of the second instruction; execute the low-order micro-operation of the second instruction based on the low-order bits of the operand of the second instruction and the low-order bits of the source logical register for the low-order micro-operation of the second instruction; and execute the high-order micro-operation of the second instruction based on the high-order bits of the operand of the second instruction and the high-order bits of the source logical register of the second instruction, where the low-order bits and high-order bits of the source logical register for the low-order micro-operation of the second instruction are respectively used to store the mask numbers for the low-order micro-operation and high-order micro-operation of the second instruction.

[0210] Optionally, when the effective bit-width of the bit-width type of the physical register of the first instruction is equal to the effective bit-width of the bit-width type of the source logical register of the second instruction, the low-order bits of the source logical register for the low-order micro-operation of the second instruction depend on the low-order bits in the effective bit-width of the physical register of the first instruction; the high-order bits of the source logical register for the high-order micro-operation of the second instruction depend on the high-order bits in the effective bit-width of the physical register of the first instruction.

[0211] Optionally, when the effective bit-width of the bit-width type of the physical register of the first instruction is less than the effective bit-width of the bit-width type of the source logical register of the second instruction, the low-order bits of the source logical register for the low-order micro-operation of the second instruction depend on the low-order bits in the effective bit-width of the physical register of the first instruction; the high-order bits of the source logical register for the high-order micro-operation of the second instruction depend on the low-order bits in the effective bit-width of the physical register of the first instruction.

[0212] Optionally, when the effective bit-width of the bit-width type of the physical register of the first instruction is greater than the effective bit-width of the bit-width type of the source logical register of the second instruction, the low-order bits of the source logical register for the low-order micro-operation of the second instruction depend on the low-order bits and high-order bits in the effective bit-width of the physical register of the first instruction; the high-order bits of the source logical register for the high-order micro-operation of the second instruction do not depend on the effective bit-width of the physical register of the first instruction.

[0213] According to another aspect of the present disclosure, an instruction processing device is also provided. Figure 8 The schematic diagram of an instruction processing device 2000 according to an embodiment of the present disclosure is shown.

[0214] As Figure 8 shown, the instruction processing device 2000 may include one or more processors 2010 and one or more memories 2020. Among them, computer-readable code is stored in the memory 2020, and when the computer-readable code is run by the one or more processors 2010, the instruction processing method described above can be executed.

[0215] The processor in the embodiment of the present disclosure may be an integrated circuit chip with the ability to process signals. The above-mentioned processor may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc., and may be of the X86 architecture or the ARM architecture.

[0216] Generally speaking, various exemplary embodiments of the present disclosure may be implemented in hardware or special circuits, software, firmware, logic, or any combination thereof. Some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that can be executed by a controller, a microprocessor or other computing devices. When aspects of the embodiments of the present disclosure are illustrated or described as block diagrams, flowcharts or using some other graphical representation, it will be understood that the blocks, devices, systems, technologies or methods described herein may be implemented as non-limiting examples in hardware, software, firmware, special circuits or logic, general hardware or controllers or other computing devices, or some combination thereof.

[0217] For example, the method or device according to an embodiment of the present disclosure may also be implemented by means of Figure 9 the architecture of the computing device 3000 shown. As Figure 9 shown, the computing device 3000 may include a bus 3010, one or more CPUs 3020, a read-only memory (ROM) 3030, a random access memory (RAM) 3040, a communication port 3050 connected to the network, an input / output component 3060, a hard disk 3070, etc. The storage device in the computing device 3000, such as the ROM 3030 or the hard disk 3070, may store various data or files of the instruction processing method provided by the present disclosure and the program instructions executed by the CPU. The computing device 3000 may also include a user interface 3080. Of course, Figure 9The architecture shown is only exemplary. When implementing different devices, one or more components in the Figure 9 Figure 9 shown computing device may be omitted according to actual needs.

[0218] According to another aspect of the present disclosure, a computer-readable storage medium is also provided. Figure 10 FIG. 4000 shows a schematic diagram of the storage medium according to the present disclosure.

[0219] As Figure 10 shown, computer-readable instructions 4010 are stored on the computer storage medium 4020. When the computer-readable instructions 4010 are run by a processor, the instruction processing method according to the embodiments of the present disclosure described with reference to the above figures can be executed. The computer-readable storage medium in the embodiments of the present disclosure may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. The non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct rambus random access memory (DR RAM). It should be noted that the memories of the methods described herein are intended to include but not be limited to these and any other suitable types of memories. It should be noted that the memories of the methods described herein are intended to include but not be limited to these and any other suitable types of memories.

[0220] Embodiments of the present disclosure also provide a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the instruction processing method according to the embodiments of the present disclosure.

[0221] Embodiments of the present disclosure provide an instruction processing method, apparatus, device, and computer-readable storage medium.

[0222] The method provided by the embodiments of the present disclosure cancels the merging operation of the high and low bits of the operands during the instruction calculation process, removes the data dependence redundancy in the instruction operation, and the high and low bit operations of the operands only depend on the number of bits of the register required for the operation, without waiting for the processing results of other bits. Therefore, the processing result of the instruction can be broadcast in advance, improving the efficiency and parallelism of the instruction processing. At the same time, by the method provided by the embodiments of the present disclosure, the high and low bits of the operand respectively use the high and low bits of the same physical register during processing, reducing the occupation of computer resources.

[0223] It should be noted that the flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains at least one executable instruction for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0224] Generally speaking, the various exemplary embodiments of the present disclosure can be implemented in hardware or dedicated circuits, software, firmware, logic, or any combination thereof. Some aspects can be implemented in hardware, while other aspects can be implemented in firmware or software that can be executed by a controller, a microprocessor, or other computing devices. When aspects of the embodiments of the present disclosure are illustrated or described as block diagrams, flowcharts, or using some other graphical representation, it will be understood that the blocks, devices, systems, techniques, or methods described herein can be implemented as non-limiting examples in hardware, software, firmware, dedicated circuits or logic, general hardware or controllers or other computing devices, or some combination thereof.

[0225] The exemplary embodiments of the present disclosure described in detail above are merely illustrative and not restrictive. Those skilled in the art should understand that various modifications and combinations can be made to these embodiments or their features without departing from the principles and spirit of the present disclosure, and such modifications should fall within the scope of the present disclosure.

Claims

1. An instruction processing method, comprising: Obtaining the high-order bits and low-order bits of the operands of the instruction respectively; Based on the obtained low-order bits of the operand, performing a low-order micro-operation, allocating a physical register for the destination logical register of the low-order micro-operation of the instruction, and storing the result of the low-order micro-operation to the low-order bits of the physical register for the instruction; And Based on the obtained high-order bits of the operand, performing a high-order micro-operation, mapping the destination logical register of the high-order micro-operation of the instruction to the physical register, and storing the result of the high-order micro-operation to the high-order bits of the physical register, wherein, the micro-operation of the high-order bits is independent of the micro-operation of the low-order bits.

2. The instruction processing method according to claim 1, further comprising: Broadcasting the low-order bits stored in the physical register based on the low-order bits of the physical register; And Broadcasting the high-order bits stored in the physical register based on the high-order bits of the physical register.

3. The instruction processing method according to claim 2, wherein, The physical register is assigned an identifier for identifying that its low-order bits or high-order bits are valid bits, wherein, broadcasting the low-order bits stored in the physical register includes: broadcasting the low-order bits and high-order bits stored in the physical register, and broadcasting an identifier for identifying that the low-order bits of the physical register are valid bits; and Broadcasting the high-order bits stored in the physical register includes: broadcasting the low-order bits and high-order bits stored in the physical register, and broadcasting an identifier for identifying that the high-order bits of the physical register are valid bits.

4. The instruction processing method according to claim 3, wherein, The instruction includes a first instruction and a second instruction, and the second instruction depends on the result of the first instruction; Determining the association relationship between the low-order bits and high-order bits of the source logical register of the second instruction and the low-order bits and high-order bits of the destination logical register of the first instruction according to the bit width type of the physical register of the first instruction and the bit width type of the source logical register of the second instruction.

5. The instruction processing method according to claim 4, wherein, The effective bit width of the bit width type of the physical register of the first instruction is equal to the effective bit width of the bit width type of the source logical register of the second instruction; The low-order bits of the source logical register for the low-order micro-operation of the second instruction depend on the low-order bits in the effective bit width of the physical register of the first instruction; The high-order bits of the source logical register for the high-order micro-operation of the second instruction depend on the high-order bits in the effective bit width of the physical register of the first instruction.

6. The instruction processing method according to claim 4, wherein, The effective bit width of the bit width type of the physical register of the first instruction is less than the effective bit width of the bit width type of the source logical register of the second instruction; The low-order bits of the source logical register for the low-order micro-operation of the second instruction depend on the low-order bits in the effective bit width of the physical register of the first instruction; The high-order bits of the source logical register for the high-order micro-operation of the second instruction depend on the low-order bits in the effective bit width of the physical register of the first instruction.

7. The instruction processing method according to claim 4, wherein, The effective bit width of the bit width type of the physical register of the first instruction is greater than the effective bit width of the bit width type of the source logical register of the second instruction; The lower bits of the source logical register for the lower micro-operations of the second instruction depend on the lower and upper bits in the effective bit-width of the physical register of the first instruction; The upper bits of the source logical register for the upper micro-operations of the second instruction do not depend on the effective bit-width of the physical register of the first instruction.

8. An instruction processing method, wherein, The instructions include a first instruction and a second instruction, and the second instruction depends on the result of the first instruction. The method includes: Obtaining the upper and lower bits of the operand of the first instruction and the mask number of the first instruction; Based on the obtained lower bits of the operand and the mask number, performing lower micro-operations and storing the result of the lower micro-operations in the lower bits of the physical register for the instruction; based on the obtained upper bits of the operand and the mask number, performing upper micro-operations and storing the result of the upper micro-operations in the upper bits of the physical register; According to the bit-width type of the physical register of the first instruction and the bit-width type of the source logical register of the second instruction, based on the lower bits and / or upper bits of the physical register of the first instruction, determining the lower bits of the source logical register for the lower micro-operations of the second instruction, and based on the lower bits of the operand of the second instruction and the lower bits of the source logical register for the lower micro-operations of the second instruction, performing the lower micro-operations of the second instruction, where the lower bits of the source logical register for the lower micro-operations of the second instruction are used to store the mask number of the lower micro-operations of the second instruction; According to the bit-width type of the physical register of the first instruction and the bit-width type of the source logical register of the second instruction, based on the lower bits or upper bits of the physical register of the first instruction, determining the upper bits of the source logical register for the upper micro-operations of the second instruction, and based on the upper bits of the operand of the second instruction and the upper bits of the source logical register for the upper micro-operations of the second instruction, performing the upper micro-operations of the second instruction, where the upper bits of the source logical register for the upper micro-operations of the second instruction are used to store the mask number of the upper micro-operations of the second instruction; Wherein, the bit-width type is any one of byte, word, double-word, and quadruple-word, and different bit-width types have different effective bit-widths.

9. An instruction processing apparatus, the apparatus includes: An operand acquisition module configured to respectively acquire the upper and lower bits of the operand of the instruction; A micro-operation processing module configured to perform lower micro-operations based on the obtained lower bits of the operand, allocate a physical register for the destination logical register of the lower micro-operations of the instruction, and store the result of the lower micro-operations in the lower bits of the physical register for the instruction; and perform upper micro-operations based on the obtained upper bits of the operand, map the destination logical register of the upper micro-operations of the instruction to the physical register, and store the result of the upper micro-operations in the upper bits of the physical register, Wherein, the micro-operation processing of the upper bits is independent of the micro-operation processing of the lower bits.

10. The instruction processing apparatus according to claim 9, wherein, performing the low-level micro-operation based on the low-order bits of the obtained operand further includes: allocating the physical register to the destination logic register of the low-level micro-operation of the instruction; performing the high-level micro-operation based on the high-order bits of the obtained operand further includes: mapping the destination logic register of the high-level micro-operation of the instruction to the physical register.

11. The instruction processing apparatus according to claim 10, further including: broadcasting the low-order bits stored in the physical register based on the low-order bits of the physical register; and broadcasting the high-order bits stored in the physical register based on the high-order bits of the physical register.

12. The instruction processing apparatus according to claim 11, wherein, An identifier for identifying that the low-order bits or high-order bits of the physical register are valid bits is assigned to the physical register, wherein, broadcasting the low-order bits stored in the physical register includes: broadcasting the low-order bits and high-order bits stored in the physical register, and broadcasting an identifier for identifying that the low-order bits of the physical register are valid bits; and broadcasting the high-order bits stored in the physical register includes: broadcasting the low-order bits and high-order bits stored in the physical register, and broadcasting an identifier for identifying that the high-order bits of the physical register are valid bits.

13. The instruction processing device according to claim 12, wherein, The instruction includes a first instruction and a second instruction, and the second instruction depends on the result of the first instruction; wherein, the instruction processing apparatus further includes: a correlation determination module, which determines the correlation relationship between the low-order bits and high-order bits of the source logic register of the second instruction and the low-order bits and high-order bits of the destination logic register of the first instruction according to the bit-width type of the physical register of the first instruction and the bit-width type of the source logic register of the second instruction, and the micro-operation processing module is further configured to: determine the low-order bits of the source logic register for the low-level micro-operation of the second instruction and the high-order bits of the source logic register for the high-level micro-operation of the second instruction according to the bit-width type of the physical register of the first instruction and the bit-width type of the source logic register of the second instruction; perform the low-level micro-operation of the second instruction based on the low-order bits of the operand of the second instruction and the low-order bits of the source logic register of the low-level micro-operation of the second instruction; and perform the high-level micro-operation of the second instruction based on the high-order bits of the operand of the second instruction and the high-order bits of the source logic register of the second instruction, wherein the low-order bits and high-order bits of the source logic register of the low-level micro-operation of the second instruction are respectively used to store the mask numbers of the low-level micro-operation and high-level micro-operation of the second instruction.

14. The instruction processing apparatus according to claim 13, wherein, The effective bit-width of the bit-width type of the physical register of the first instruction is equal to the effective bit-width of the bit-width type of the source logic register of the second instruction; the low-order bits of the source logic register for the low-level micro-operation of the second instruction depend on the low-order bits in the effective bit-width of the physical register of the first instruction; the high-order bits of the source logic register for the high-level micro-operation of the second instruction depend on the high-order bits in the effective bit-width of the physical register of the first instruction.

15. The instruction processing apparatus according to claim 13, wherein, The effective bit width of the bit width type of the physical register of the first instruction is less than the effective bit width of the bit width type of the source logical register of the second instruction; The lower bits of the source logical register for the lower micro-operation of the second instruction depend on the lower bits of the effective bit width of the physical register of the first instruction; The upper bits of the source logical register for the upper micro-operation of the second instruction depend on the lower bits of the effective bit width of the physical register of the first instruction.

16. The instruction processing device according to claim 13, wherein, The effective bit width of the bit width type of the physical register of the first instruction is greater than the effective bit width of the bit width type of the source logical register of the second instruction; The lower bits of the source logical register for the lower micro-operation of the second instruction depend on the lower bits and upper bits of the effective bit width of the physical register of the first instruction; The upper bits of the source logical register for the upper micro-operation of the second instruction do not depend on the effective bit width of the physical register of the first instruction.

17. An apparatus for instruction processing includes: One or more processors; And One or more memories storing computer-executable programs that, when executed by the processors, perform the method according to any one of claims 1-8.

18. A computer program product comprising computer software code that, when run by a processor, is used to implement the method according to any one of claims 1-8.

19. A computer-readable storage medium having stored thereon computer-executable instructions that, when executed by a processor, are used to implement the method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Single-precision floating-point data storing method and processor

    CN101539850A

  • Method and apparatus for implementing a dynamic out-of-order processor pipeline

    CN104951281A