Instruction processing method and device, equipment, readable storage medium and program product
By introducing the instant index register (IIR) into the thread controller, the index value is judged and read during decoding, which solves the problem of additional instruction cycle consumption in the thread controller, and improves the efficiency of instruction processing, especially the execution efficiency of loop instructions.
Patent Information
- Application Number
- CN202510444804.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-08-05
AI Technical Summary
In the prior art, thread controllers need to consume additional instruction cycles when decoding index values, resulting in reduced instruction execution efficiency, especially when facing cyclic instructions.
The memory space is allocated in the thread controller as the instant index register (IIR), and when decoding, it determines whether the index value contains the IIR index. If so, read the numerical value from the IIR and redecode the instruction, and transmit it directly to the logical operation unit.
By reducing multiple accesses to registers, it saves instruction cycles, improves index decoding efficiency, and improves thread execution efficiency, especially when processing loop instructions, it significantly improves decoding efficiency.
Smart Images

Figure CN120429013A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer processing technology, and in particular to a method, apparatus, device, readable storage medium, and program product for instruction processing. Background Art
[0002] The time it takes for the instruction computation unit to fetch and execute an instruction from main memory is called an instruction cycle. This cycle consists of five stages: instruction fetch, decode, issue, execute, and writeback. Each stage of instruction execution is performed by a different functional unit. To enhance the speed and efficiency of register communication and large amounts of parallel data processing, the Wave Controller (WVC) expands upon pipeline execution by providing a variety of data manipulation capabilities.
[0003] To accommodate the processor's current parallel computing needs, thread controllers must decode memory types that also include index registers. This means the address written to the index register field in an instruction is not a real address, but rather an index. The thread controller must decode the index in the first instruction cycle without sending the instruction to the logic unit. The thread controller then accesses the corresponding memory contents based on the decoding result of the first instruction cycle.
[0004] However, accessing index values in this way consumes additional instruction cycles, reducing instruction execution efficiency. This is especially true for looping instructions, where decoding each instruction code round consumes additional instruction cycles, multiplying the number of instruction cycles executed by the thread due to accessing the index value in the register. Summary of the Invention
[0005] Based on this, it is necessary to provide a method, device, equipment, readable storage medium and program product for instruction processing that can avoid multiple accesses to registers, save instruction cycles consumed by access, and improve index decoding efficiency to address the above technical problems.
[0006] In a first aspect, the present application provides a method for processing an instruction, comprising:
[0007] Allocate a portion of storage space in the thread controller as an immediate index register (IIR), and add an operation type encoded into the instruction word for the IIR, wherein the IIR is an immediate index register;
[0008] After the thread controller decodes the instruction containing the IIR, it determines whether the index value obtained by decoding contains the IIR index;
[0009] If the index value contains an IIR index, the stored value is read from the corresponding IIR according to the IIR index;
[0010] Generate a new instruction according to the read value, and re-decode the new instruction;
[0011] The re-decoded instruction is issued to the logic operation unit.
[0012] In one embodiment, allocating a portion of storage space in the thread controller as an IIR and adding an operation type to the instruction word for the IIR includes:
[0013] Allocate at least one register-sized space in the thread controller to store the immediate value in the IIR. Assuming that each IIR occupies m bits of the instruction content, each IIR can store 0 to 2 m-1 An immediate number is used as the index value, and m is a natural number greater than 0;
[0014] Read and write operands to the IIR are programmed into instruction words.
[0015] In one embodiment, after determining whether the decoded index value contains an IIR index, the method further includes:
[0016] If the index value does not contain an IIR index, the instruction is directly sent to the logic operation unit.
[0017] In one embodiment, the method further comprises:
[0018] When the IIR needs to be updated, the logic operation unit writes the updated value into the corresponding position in the IIR.
[0019] In one embodiment, the method further comprises:
[0020] When the IIR needs to be updated, the thread controller decodes the instruction for self-increment update in the current instruction cycle;
[0021] Get the address of the IIR that needs to be updated, and access the IIR to get the corresponding value;
[0022] After the accessed value is incremented by 1, it is rewritten into the IIR, completing the decoding of the instruction for self-increment update.
[0023] In one embodiment, the instructions for self-increment update are combined into the to-be-executed instructions of the thread controller;
[0024] When the thread controller decodes the instruction to be executed, the instruction for self-increment update is decoded within the instruction cycle of the instruction to be executed.
[0025] In a second aspect, the present application further provides an instruction processing device, comprising:
[0026] an allocation module, configured to allocate a portion of storage space in the thread controller as an IIR and add an operation type programmed into an instruction word for the IIR, wherein the IIR is an immediate index register;
[0027] a judgment module, configured to judge whether the index value obtained by decoding contains an IIR index after the thread controller decodes the instruction containing the IIR;
[0028] A reading module, configured to read a stored value from a corresponding IIR according to the IIR index when the index value contains an IIR index;
[0029] The sending module is used to send the re-decoded instruction to the logic operation unit.
[0030] In a third aspect, the present application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0031] Allocate a portion of storage space in the thread controller as an immediate index register (IIR), and add an operation type encoded into the instruction word for the IIR, wherein the IIR is an immediate index register;
[0032] After the thread controller decodes the instruction containing the IIR, it determines whether the index value obtained by decoding contains the IIR index;
[0033] If the index value contains an IIR index, the stored value is read from the corresponding IIR according to the IIR index;
[0034] Generate a new instruction according to the read value, and re-decode the new instruction;
[0035] The re-decoded instruction is issued to the logic operation unit.
[0036] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the following steps are implemented:
[0037] Allocate a portion of storage space in the thread controller as an immediate index register (IIR), and add an operation type encoded into the instruction word for the IIR, wherein the IIR is an immediate index register;
[0038] After the thread controller decodes the instruction containing the IIR, it determines whether the index value obtained by decoding contains the IIR index;
[0039] If the index value contains an IIR index, the stored value is read from the corresponding IIR according to the IIR index;
[0040] Generate a new instruction according to the read value, and re-decode the new instruction;
[0041] The re-decoded instruction is issued to the logic operation unit.
[0042] In a fifth aspect, the present application further provides a computer program product, comprising a computer program, which, when executed by a processor, implements the following steps:
[0043] Allocate a portion of storage space in the thread controller as an immediate index register (IIR), and add an operation type encoded into the instruction word for the IIR, wherein the IIR is an immediate index register;
[0044] After the thread controller decodes the instruction containing the IIR, it determines whether the index value obtained by decoding contains the IIR index;
[0045] If the index value contains an IIR index, the stored value is read from the corresponding IIR according to the IIR index;
[0046] Generate a new instruction according to the read value, and re-decode the new instruction;
[0047] The re-decoded instruction is issued to the logic operation unit.
[0048] The above-mentioned instruction processing method, apparatus, computer device, computer-readable storage medium, and computer program product allocate a portion of storage space within the thread controller as an IIR and add an operation type encoded into the instruction word for the IIR, wherein the IIR is an immediate index register. Since the IIR is internal to the thread controller, accessing the IIR does not require consuming additional instruction cycles. After the thread controller decodes an instruction containing the IIR, it determines whether the decoded index value contains the IIR index. Thus, the thread controller can decode the index value and determine whether the IIR needs to be accessed based on whether the index value contains the IIR index. If the index value contains the IIR index, the stored value is read from the corresponding IIR based on the IIR index. Thus, the value can be read from the IIR within the current instruction cycle, saving instruction cycles. A new instruction is generated based on the read value, and the new instruction is re-decoded. The re-decoded instruction is transmitted to the logic operation unit. This avoids multiple accesses to the register, saves instruction cycles consumed by access, and improves index decoding efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments of the present application or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying any creative work.
[0050] Figure 1 A schematic diagram of the working principle of the thread controller in the existing method;
[0051] Figure 2 A schematic diagram of the working principle of the decoding buffer unit in the existing method;
[0052] Figure 3 A flowchart of a method for processing instructions in one embodiment;
[0053] Figure 4 This is a schematic diagram of the working principle of the thread controller after adding IIR;
[0054] Figure 5 A schematic diagram of the working principle of a decoding buffer unit for reading and writing IIR instructions in one embodiment;
[0055] Figure 6 is a flowchart of a method for processing instructions in another embodiment;
[0056] Figure 7 is a flowchart of a method for processing instructions in yet another embodiment;
[0057] Figure 8 1 is a flow chart of a method for processing instructions in a fourth embodiment;
[0058] Figure 9 FIG1 is a schematic diagram of a decoding process of a UPIIA instruction in one embodiment;
[0059] Figure 10 A schematic diagram of the principle of the indexing method for loop instructions in the existing method;
[0060] Figure 11 A schematic diagram of the UPIIA instruction optimization principle for loop instructions in one embodiment;
[0061] Figure 12 A result block diagram of an apparatus for processing instructions according to an embodiment;
[0062] Figure 13 FIG. 1 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0063] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0064] Before describing the embodiments of the present application, a brief description of the professional terms appearing in the embodiments of the present application is now provided to facilitate understanding of the professional terms.
[0065] 1) Instruction Cycle: The time it takes for the instruction computation unit to fetch and execute an instruction from main memory is called an instruction cycle. The instruction cycle consists of five stages: instruction fetch, decode, issue, execute, and write back.
[0066] 2) Instruction pipeline: Instructions are executed on different functional components at each stage of the execution process.
[0067] 3) Functional expansion of the thread controller: To improve the ability and efficiency of rapid communication with registers and parallel processing of large amounts of data, the thread controller has expanded the ability to perform multiple data operations based on the execution of the instruction pipeline. It obtains data from the memory, places them in a large sequential register file, and then performs logical operations on the data in the register file, and finally places the results back into the memory.
[0068] 4) The basic structure of an instruction: It consists of an opcode and an operand. The opcode determines the operation to be performed, and the operand refers to the data involved in the calculation and the address of the unit in which it is located. In computers, the operation request and the operand address are both represented by binary digits, called the opcode and address code, respectively. The entire instruction is stored in memory in binary format. Read and write operands can be encoded into the instruction word. Operand types can include immediate values, vector registers, scalar registers, etc. Different operand types are decoded by the thread controller.
[0069] 5) The specific decoding operation of the thread controller is as follows: decoding in the current instruction cycle (cycle0) obtains the real address of the register, and the instruction is transmitted after decoding is completed in cycle0.
[0070] Figure 1 The working principle diagram of the thread controller in the existing method is as follows: Figure 1As shown, it includes labels 1 to 8, where label 1: the instruction fetch unit, which sends instruction fetch requests to the instruction cache; label 2: the instruction cache, which stores instructions. After receiving a read instruction request, it returns an instruction to the instruction fetch unit; label 3: the decode unit, which parses the instruction's source operand, destination operand, and preprocessing mask; label 4: the decode buffer unit, which operates on the decoded register type that requires further processing. For example, it locates the real register type storage space based on the index value of the index register and accesses it, waiting for the decode unit to re-decode other registers that need to be accessed within the instruction in the next instruction cycle; or it splits instructions with loops into multiple instructions for re-decoding; label 5: the instruction issue unit, which issues instructions to the logic operation unit; label 6: the write back unit, which writes the operation results back to the corresponding memory; label 7: the ALU (Arithmetic and Logic Unit), which performs the logical calculations of the instruction; label 8: the register unit, which stores data and requires at least 8 instruction cycles for the thread controller to access.
[0071] Figure 2 The schematic diagram of the working principle of the decoding buffer unit in the existing method is as follows: Figure 2 As shown in the figure, when the decoded instruction does not contain an index register, the WVC will directly transmit the decoded instruction information to the ALU without buffering. When the decoded instruction contains an index register, the WVC will decode the index value and wait for at least 8 instruction cycles until the index value stored in the register data is read back. The WVC then passes it to the decode unit for re-decoding along with other operands, and finally transmits the instruction.
[0072] Therefore, when WVC is decoded to the index register, it consumes additional instruction cycles, causing the number of instruction cycles executed by the thread to increase exponentially due to the access to the index value in the register.
[0073] The embodiments of the present application are intended to reduce the extra instruction cycles generated by index registers or instruction loops in a decoding buffer unit.
[0074] In an exemplary embodiment, Figure 3 As shown, a method for instruction processing is provided, which is applied to a thread controller and includes the following steps 301 to 305. In which:
[0075] Step 301: allocate a portion of storage space in the thread controller as an IIR, and add an operation type programmed into the instruction word for the IIR.
[0076] In this embodiment, an index register space for storing immediate values may be added inside the thread controller, and is named an Immediate Index Register (IIR).
[0077] For example, at least one register-sized space is allocated in the thread controller to store the immediate value in the IIR. Assuming that each IIR occupies m bits of the instruction content, each IIR can store 0 to 2 m-1 An immediate number is used as an index value, and m is a natural number greater than 0; the read and write operands for the IIR are encoded into the instruction word.
[0078] For example, two registers can be allocated in the thread controller to store the immediate value in the IIR for the thread controller to call. In order to increase the operand type of the instruction word required for IIR, the thread controller can arrange the possible IIRs into IIR0 and IIR1 according to the sequence number. Each IIR occupies m bits of the instruction content, and can represent 0~2 m-1 For example, if IIR is 6 bits, it can represent numbers 0 to 63 as index values.
[0079] For example, Figure 4 The working principle diagram of the thread controller after adding IIR is as follows: Figure 4 As shown, mark 9: IIR0, IIR1 (located in WVC, WVC can independently access and update the data stored in the module); mark 10: read and write instructions containing IIR.
[0080] Step 302: After the thread controller decodes the instruction containing the IIR, it is determined whether the index value obtained by decoding contains the IIR index.
[0081] For example, Figure 5 FIG. 1 is a schematic diagram showing the working principle of a decoding buffer unit for reading and writing IIR instructions in one embodiment. Figure 5 As shown, when WVC executes decoding containing IIR instructions, it first determines whether the index value contains an IIR index (ie, IIR_index).
[0082] In this embodiment, since the IIR is inside the WVC, accessing the IIR does not require consuming additional instruction cycles. This indexing method greatly improves the decoding efficiency of instructions containing indexes.
[0083] Step 303: If the index value contains an IIR index, the stored value is read from the corresponding IIR according to the IIR index.
[0084] For example, combined Figure 5If it contains an IIR index, the value stored in the corresponding IIR is accessed and obtained according to the needs, and then re-decoded; if it does not contain an IIR index, the instruction is directly issued.
[0085] Step 304: Generate a new instruction according to the read value, and re-decode the new instruction.
[0086] For example, assuming that the instruction to be decoded is IADD VR009, SR[IIR1], 0xf0f0, the thread controller decodes the second source operand to access the IIR index value, and the address is 0x1, that is, the WVC will immediately access the data stored in IIR1. If the data is 0x002, the generated instruction IADD VR009, SR002, 0xf0f0 is decoded again, that is, the instruction can be transmitted within the same instruction cycle.
[0087] In this embodiment, if the thread controller finds that the decoded result contains an IIR index during the first decoding process, indicating that the IIR needs to be accessed, it first accesses the corresponding address in the IIR, reads the data stored in the IIR, and then substitutes the read data into the instruction to form a new instruction. The thread controller then decodes the new instruction a second time within the same instruction cycle and, after decoding is complete, sends the instruction to the logic operation unit.
[0088] Step 305: Send the re-decoded instruction to the logic operation unit.
[0089] The above-mentioned instruction processing method allocates a portion of storage space as IIR in the thread controller and adds an operation type programmed into the instruction word for the IIR, wherein the IIR is an immediate index register; since the IIR is inside the thread controller, access to the IIR does not require consuming additional instruction cycles to complete. After the thread controller decodes the instruction containing the IIR, it determines whether the decoded index value contains the IIR index; thereby, the thread controller can decode the index value and determine whether the IIR needs to be accessed based on whether the index value contains the IIR index. If the index value contains the IIR index, the stored value is read from the corresponding IIR based on the IIR index; thereby, the value can be read from the IIR within the current instruction cycle, saving instruction cycles. A new instruction is generated based on the read value, and the new instruction is re-decoded; the re-decoded instruction is transmitted to the logic operation unit. Thus, multiple accesses to the register can be avoided, instruction cycles consumed by access can be saved, and index decoding efficiency can be improved.
[0090] In another exemplary embodiment, a method for instruction processing is provided, such as Figure 6 As shown, it may include: Step 601 to Step 606. Among them:
[0091] Step 601: Allocate a portion of storage space in the thread controller as an IIR, and add an operation type to the instruction word for the IIR.
[0092] Step 602 , after the thread controller decodes the instruction containing the IIR, it determines whether the decoded index value contains the IIR index. If so, step 603 is executed; if not, step 606 is executed.
[0093] Step 603: Read the stored value from the corresponding IIR according to the IIR index.
[0094] Step 604: Generate a new instruction based on the read value and re-decode the new instruction.
[0095] Step 605: Send the re-decoded instruction to the logic operation unit.
[0096] Step 606: Send the instruction directly to the logic operation unit.
[0097] In this embodiment, a judgment process is performed in the decoding buffer unit of the thread controller, that is, whether the result obtained by the first decoding contains an IIR index. If it contains an IIR index, it is necessary to access the IIR located inside the thread controller and read the value in the corresponding address. Then, a new instruction is generated based on the read value. At this time, the instruction generally no longer contains the IIR index, so it can be regarded as an ordinary instruction. After the thread controller decodes the new instruction for the second time, it sends the instruction to the logical operation unit. Correspondingly, if the result of the first decoding does not contain an IIR index, there is no need to access the IIR at this time, and the decoded instruction is directly sent to the logical operation unit.
[0098] In another exemplary embodiment, a method for instruction processing is provided, such as Figure 7 As shown, it may include: step 701 to step 702. Among them:
[0099] Step 701: Allocate a portion of storage space in the thread controller as an IIR, and add an operation type to the instruction word for the IIR.
[0100] For the specific implementation process and technical effects of step 701 in this embodiment, please refer to Figure 3 The description of step 301 is omitted here.
[0101] Step 702: When the IIR needs to be updated, the logic operation unit writes the updated value into the corresponding position in the IIR.
[0102] In this embodiment, since the IIR is located in the WVC, it can be read, written and updated autonomously by the WVC.
[0103] For example, since WVC does not require additional instruction cycles to read or write the IIR space, it can be incorporated into the instruction word as an operand and updated by the logical calculation unit as a normal instruction. For example, in IADD IIR0, SR001, 0xff, IIR0 is the destination operand. After WVC decodes and sends it to the ALU, the ALU calculates and writes the corresponding IIR register, completing the update of IIR0. This update method, written by the ALU, requires a normal instruction.
[0104] In this embodiment, when IIR is applied, the update operation for IIR can be pre-programmed into the instruction word, so that it can be treated as a common instruction as a type of operand and updated by the logic calculation unit.
[0105] It should be noted that step 702 in this embodiment can also be combined with Figure 3 、 Figure 6 In the method steps of the embodiment shown, there is no restriction on the execution order, that is, the IIR can be updated at any time.
[0106] In a fourth exemplary embodiment, a method for instruction processing is provided, such as Figure 8 As shown, it may include: step 801 to step 802. Among them:
[0107] Step 801: Allocate a portion of storage space in the thread controller as an IIR, and add an operation type programmed into the instruction word for the IIR.
[0108] For the specific implementation process and technical effects of step 801 in this embodiment, please refer to Figure 3 The description of step 301 is omitted here.
[0109] Step 802: When the IIR needs to be updated, the thread controller decodes the instruction for self-increment update in the current instruction cycle.
[0110] In this embodiment, the instruction for self-increment and update can be combined into the to-be-executed instruction of the thread controller; when the thread controller decodes the to-be-executed instruction, the instruction for self-increment and update is decoded within the instruction cycle of the to-be-executed instruction.
[0111] Exemplarily, the instruction used for auto-increment updates is a UPIIA instruction. When the same instruction requires accessing multiple registers at different addresses in a loop, the traditional approach is to encode the index value into multiple registers or lanes of a vector register. However, each access to the index value consumes a significant number of instruction cycles, significantly reducing operational efficiency. Using the UPIIA instruction for updates can be combined with other instructions without requiring additional instruction cycles. Therefore, combining the UPIIA instruction with the original instruction can be used for frequently updated index values, saving instruction cycles and improving decoding efficiency.
[0112] Step 803: Get the address of the IIR that needs to be updated, and access the IIR to obtain the corresponding value.
[0113] For example, Figure 9 FIG. 1 is a schematic diagram of a decoding process of a UPIIA instruction in one embodiment. Figure 9 As shown, when decoding the UPIIA instruction of self-increment update, WVC will first record the address of the IIR register that needs to be operated and access it to obtain the value.
[0114] Step 804 , increment the accessed value by 1 and rewrite it into the IIR, thus ending the decoding of the instruction for the self-increment update.
[0115] It should be noted that steps 802 to 804 in this embodiment can also be combined with Figure 3 、 Figure 6 The method steps in the embodiment shown are not limited in the execution order, that is, the UPIIA instruction can be combined with other instructions at any time to implement the update of the IIR.
[0116] For example, combined Figure 9 As shown, WVC will also increment the returned IIR value by 1 and rewrite it into the IIR space. After the update is completed, WVC ends decoding the UPIIA instruction.
[0117] For example, Figure 10 This is a schematic diagram of the principle of the indexing method for loop instructions in the existing method, such as Figure 10 As shown, the 8 instructions that need to be executed are:
[0118] IADD VR001, SR0, 0x001
[0119] IADD VR001, SR1, 0x001
[0120] …
[0121] IADD VR001, SR7, 0x001
[0122] In existing methods, the actual addresses of the eight instructions (SR0-7) might be encoded into two vector registers VR010 for indexing. The value of each register is distributed across each lane of the vector register. A new instruction, IADDVR001, SR[VR010], 0x001, replaces the eight instructions, and the initial values of each lane of VR010 need to be set as shown in the figure. While this method can reduce the amount of code and instruction memory, it expands the index instructions during WVC decoding. This expansion process consumes additional instruction cycles to frequently access register space. Other operands are kept in a waiting state during these additional instruction cycles, and each access requires a minimum of eight instruction cycles to write the index value back from the register. When executing the above instruction, an additional 56 instruction cycles are required, resulting in significant resource waste and reduced decoding efficiency.
[0123] For example, Figure 11 FIG. 1 is a schematic diagram of the UPIIA instruction optimization principle when facing a loop instruction in one embodiment. Figure 11 As shown, the instruction can be optimized to: IADD VR001,SR[IIR0],0x001 + UPIIA IIR0, incr, 0x1 to replace the original 8 instructions. Figure 11 As shown, the initial value of IIR is first set to 0. After WVC completes decoding the UPIIA instruction in the current instruction cycle of the loop, it updates IIR0, incrementing it by 1 to end the decoding. In the next instruction cycle, when IIR0 is accessed again, the updated value of IIR0 is obtained, and it is incremented by 1 again after decoding, until the loop ends. Using this indexing method, WVC only takes 8 instruction cycles to complete, without consuming additional instruction cycles. This improves decoding efficiency and significantly improves the execution efficiency of loop instructions.
[0124] This embodiment significantly improves the execution efficiency of instructions in parallel operations by adding an IIR to the thread controller. Compared to register indexing, it consumes fewer instruction cycles, effectively improving the decoding efficiency of WVC. Furthermore, WVC can autonomously update the IIR value for both regular instruction updates and combined UPIIA updates, significantly improving the execution efficiency of loop instructions.
[0125] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.
[0126] Based on the same inventive concept, the embodiments of the present application also provide an instruction processing device for implementing the instruction processing method involved above. The implementation solution provided by this device is similar to the implementation solution described in the above method. Therefore, the specific limitations of one or more instruction processing device embodiments provided below can be found in the above-mentioned limitations of the instruction processing method and will not be repeated here.
[0127] In an exemplary embodiment, Figure 12 As shown, a device for instruction processing is provided, including: an allocation module 1201, a judgment module 1202, a reading module 1203 and a sending module 1204, wherein:
[0128] an allocation module 1201 for allocating a portion of storage space in the thread controller as an IIR and adding an operation type programmed into an instruction word for the IIR, wherein the IIR is an immediate index register;
[0129] A judgment module 1202 is configured to judge whether an index value obtained by decoding the instruction containing the IIR contains an IIR index after the thread controller decodes the instruction containing the IIR;
[0130] A reading module 1203 is configured to read a stored value from a corresponding IIR according to the IIR index when the index value contains an IIR index;
[0131] The sending module 1204 is configured to send the re-decoded instruction to the logic operation unit.
[0132] Exemplarily, the allocation module 1201 is specifically configured to allocate at least one register-sized space in the thread controller for storing immediate values in the IIR, wherein, assuming that each IIR occupies m bits of the instruction content, each IIR can store 0 to 2 m-1An immediate number is used as an index value, and m is a natural number greater than 0; the read and write operands for the IIR are encoded into the instruction word.
[0133] Exemplarily, the sending module 1204 is further configured to: directly send an instruction to the logic operation unit when the index value does not contain an IIR index.
[0134] Exemplarily, the above apparatus may further include: a first updating module 1205, configured to: when the IIR needs to be updated, write the updated value into the corresponding position in the IIR through the logic operation unit.
[0135] Exemplarily, the above-mentioned device may also include: a second update module 1206, which is used to decode the instruction for self-increment update through the thread controller in the current instruction cycle when the IIR needs to be updated; obtain the address of the IIR that needs to be updated, and access the IIR to obtain the corresponding value; after incrementing the accessed value by 1, rewrite it into the IIR to end the decoding of the instruction for self-increment update.
[0136] Exemplarily, the instruction for self-increment update is combined in the to-be-executed instruction of the thread controller; when the thread controller decodes the to-be-executed instruction, the instruction for self-increment update is decoded within the instruction cycle of the to-be-executed instruction.
[0137] Each module in the above-mentioned instruction processing device can be implemented in whole or in part through software, hardware, or a combination thereof. Each of the above-mentioned modules can be embedded in or independent of the processor of the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the corresponding operations of each of the above modules.
[0138] In an exemplary embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as shown in FIG. Figure 13As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O) and a communication interface. The processor, memory and input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store instruction data. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a method for instruction processing is implemented.
[0139] Those skilled in the art will understand that Figure 13 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0140] In an exemplary embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.
[0141] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.
[0142] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.
[0143] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.
[0144] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the various embodiments provided herein may be, but are not limited to, general-purpose processors, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), programmable logic devices (PLDs), quantum computing-based data processing logic devices, artificial intelligence (AI) processors, and the like.
[0145] The technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0146] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.
Claims
1. A method for instruction processing, characterized in that: Applied to a thread controller, the method includes: Allocate a portion of storage space in the thread controller as an immediate index register (IIR), and add an operation type encoded into the instruction word for the IIR, wherein the IIR is an immediate index register; After the thread controller decodes the instruction containing the IIR, it determines whether the decoded index value contains the IIR index; If the index value contains an IIR index, the stored value is read from the corresponding IIR according to the IIR index; Generate a new instruction according to the read value, and re-decode the new instruction; The re-decoded instruction is issued to the logic operation unit.
2. The method according to claim 1, characterized in that The allocating a portion of storage space in the thread controller as an IIR and adding an operation type into the instruction word for the IIR includes: Allocate at least one register size space in the thread controller to store the immediate value in the IIR. Assuming that each IIR occupies m bits of the instruction content, each IIR can store 0~2 m-1 An immediate number is used as the index value, and m is a natural number greater than 0; Read and write operands to the IIR are programmed into instruction words.
3. The method according to claim 1, characterized in that After determining whether the decoded index value contains an IIR index, the method further includes: If the index value does not contain an IIR index, the instruction is directly sent to the logic operation unit.
4. The method according to any one of claims 1 to 3, characterized in that The method further comprises: When the IIR needs to be updated, the logic operation unit writes the updated value into the corresponding position in the IIR.
5. The method according to any one of claims 1 to 3, characterized in that The method further comprises: When the IIR needs to be updated, the thread controller decodes the instruction for self-increment update in the current instruction cycle; Get the address of the IIR that needs to be updated, and access the IIR to get the corresponding value; After the accessed value is incremented by 1, it is rewritten into the IIR, completing the decoding of the instruction for self-increment update.
6. The method according to claim 5, characterized in that The instructions for self-increment update are combined into the to-be-executed instructions of the thread controller; When the thread controller decodes the instruction to be executed, the instruction for self-increment update is decoded within the instruction cycle of the instruction to be executed.
7. A device for processing instructions, characterized in that: Applied to a thread controller, the device comprises: an allocation module, configured to allocate a portion of storage space in the thread controller as an IIR and add an operation type programmed into an instruction word for the IIR, wherein the IIR is an immediate index register; a judgment module, configured to judge whether the index value obtained by decoding contains an IIR index after the thread controller decodes the instruction containing the IIR; A reading module, configured to read a stored value from a corresponding IIR according to the IIR index when the index value contains an IIR index; The sending module is used to send the re-decoded instruction to the logic operation unit.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.