Processor, chip product, computer device and operand acquisition method
By introducing a pre-instruction unit into the processor, generating read requests and reading operands based on the free space of the memory, the problem of difficult to take into account both processing delay and hardware costs in the prior art is solved, and more efficient operand acquisition and storage are achieved.
Patent Information
- Application Number
- CN202410955145.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-16
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-07-16
AI Technical Summary
When obtaining operands, existing processors find it difficult to take into account the processing delay of instructions and the savings of hardware costs.
A pre-instruction processing unit is introduced in the processor, which generates a read request based on the free storage space in the memory, reads operands from the registers and stores them into the memory.
By effectively managing the storage space, the memory size requirement is reduced, the hardware cost is reduced, and the delay increase caused by the instruction execution unit due to waiting is avoided.
Smart Images

Figure CN118760472B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of chip technology, and in particular to a processor, a chip product, a computer device, and an operand acquisition method. Background Art
[0002] Currently, a processor generally includes an instruction issuing unit, an instruction executing unit and a register.
[0003] After the instruction issuing unit issues an instruction to be executed, it needs to obtain operands related to the instruction from the register, and then the instruction executing unit executes the instruction according to the operands related to the instruction.
[0004] However, how to obtain operands to balance the processing delay of instructions and the hardware cost of the processor is still a problem that needs further research. Summary of the invention
[0005] The embodiments of the present application provide a processor, a chip product, a computer device, and an operand acquisition method. The technical solutions provided by the embodiments of the present application include the following aspects.
[0006] According to one aspect of an embodiment of the present application, a processor is provided, the processor comprising: an instruction issuing unit, an instruction pre-processing unit, a data reading unit, a register and a memory;
[0007] The instruction transmitting unit is used to transmit a first instruction;
[0008] The instruction pre-processing unit is used to generate a read request corresponding to the first instruction according to the free storage space contained in the memory, and send the read request to the data reading unit, wherein the read request is used to request to read an operand related to the first instruction;
[0009] The data reading unit is used to read the operands related to the first instruction from the register according to the read request, and store the operands related to the first instruction in the memory.
[0010] According to one aspect of an embodiment of the present application, a chip product is provided, wherein the chip product includes the processor as described above.
[0011] According to one aspect of an embodiment of the present application, a computer device is provided, wherein the computer device includes the processor as described above.
[0012] According to one aspect of an embodiment of the present application, there is provided an operand acquisition method applied to a processor, the processor comprising: an instruction issuing unit, an instruction pre-processing unit, a data reading unit, a register and a memory; the method comprising:
[0013] The instruction transmitting unit transmits a first instruction;
[0014] The instruction pre-processing unit generates a read request corresponding to the first instruction according to the free storage space contained in the memory, and sends the read request to the data reading unit, wherein the read request is used to request to read an operand related to the first instruction;
[0015] The data reading unit reads operands related to the first instruction from the register according to the read request, and stores the operands related to the first instruction in the memory.
[0016] The technical solution provided in the embodiments of the present application can bring the following beneficial effects:
[0017] On the one hand, by setting an instruction pre-processing unit in the processor, it generates a read request corresponding to the instruction according to the free storage space contained in the memory, so that it can be ensured that when there is enough free storage space, the operands related to the instruction are read, and the operands related to the instruction can be effectively stored in the memory. This method can ensure the reliability of operand reading and writing without having to design the storage space of the memory to be very large, thereby saving the hardware cost of the processor; on the other hand, after the instruction is issued, the operands corresponding to the instruction are read first, and then the operands corresponding to the instruction are stored in the memory. The operands stored in the memory can be directly used by the instruction execution unit when executing the instruction, without the need for the instruction execution unit to request to read the operands in the register after starting to execute the instruction, which avoids the hole caused by the instruction execution unit being in a waiting state, thereby not increasing the delay of instruction execution. Combining the above two aspects, the technical solution provided by the embodiment of the present application can take into account the processing delay of the instruction and the hardware cost of the processor. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 is a simplified structural diagram of a GPU provided by an embodiment of the present application;
[0019] Figure 2 It is a schematic diagram of an embodiment of the present application in which an instruction issuing unit and an instruction executing unit are both deployed in a core;
[0020] Figure 3 It is a schematic diagram of an embodiment of the present application in which an instruction issuing unit is deployed inside a core and an instruction executing unit is deployed outside the core;
[0021] Figure 4 It is a schematic diagram of an operand acquisition method provided by the related art;
[0022] Figure 5 is a schematic diagram of another operand acquisition method provided by the related art;
[0023] Figure 6 is a structural block diagram of a processor provided in a possible implementation of the present application;
[0024] Figure 7 is a structural block diagram of a processor provided in another possible implementation of the present application;
[0025] Figure 8 is a structural block diagram of a processor provided in another possible implementation of the present application;
[0026] Fig. 9 is a schematic diagram of information included in a single instruction execution request provided in a possible implementation of the present application;
[0027] Fig.10 is a schematic diagram of a block in a memory provided in a possible implementation of the present application;
[0028] Fig.11 is a schematic diagram of an operand storage method provided in a possible implementation of the present application;
[0029] Fig.12 It is a schematic diagram of multiple stream processors sharing the same block counter provided in a possible implementation of the present application;
[0030] Fig.13 It is a flowchart of an operand acquisition method applied to a processor provided in a possible implementation of the present application. DETAILED DESCRIPTION
[0031] In order to make the objectives, technical solutions and advantages of the present application clearer, the implementation methods of the present application will be further described in detail below with reference to the accompanying drawings.
[0032] GPU (Graphics Processing Unit), also known as display core, visual processor, display chip, is a microprocessor that specializes in performing image and graphics related calculations on personal computers, workstations, game consoles and some mobile devices (such as tablets, smart phones, etc.).
[0033] A GPU can include one or more SMs (Streaming Multiprocessors). SM can be seen as the heart of the GPU, also called the GPU core. SM can be compared to the CPU core.
[0034] SP (Streaming Processor) is the most basic processing unit of GPU. Each SM contains dozens or hundreds of SPs, depending on the GPU architecture. The "hundreds" in the commonly mentioned GPU with hundreds of cores refers to the number of SPs. SP only runs one thread, and each SP has its own register file. The SP of GPU is only equivalent to the execution unit in the CPU (Central Processing Unit), which is responsible for executing instructions and performing calculations, and does not include a control unit.
[0035] Specific instructions and tasks are processed on SPs. The reason why GPU can perform parallel computing is that many SPs can perform processing at the same time in the GPU architecture.
[0036] Each SP has its own registers, which are a scarce resource and limit the number of programs that can run in parallel.
[0037] In theory, the computing power of the GPU determines the maximum number of thread blocks that can reside on the SM. It can be understood that the SM defines the maximum number of thread blocks that can run. The number of thread blocks may also be limited by the following two aspects, so in actual use it may be lower than the maximum value: (1) the hardware limit on the number of registers of the SP; (2) the hardware limit on the amount of shared memory consumed by the thread blocks. Registers and shared memory are scarce resources of the SM. These resources can be allocated to all blocks residing in the SM. Therefore, these resources are the main factors limiting the active thread warps in the SM, which in turn limits the parallel capability.
[0038] Take the processor as GPU as an example, Figure 1 The following is an example of a simplified structure diagram of a GPU. Figure 1 As shown, the GPU 10 may include: an instruction issuing unit 11, an instruction executing unit 12 and a register 13.
[0039] The instruction issuing unit 11 is mainly responsible for processing the instruction scheduling, thread management and execution control functions of the GPU. The design and implementation of the instruction issuing unit 11 has an important impact on the performance and parallel computing efficiency of the GPU.
[0040] The instruction execution unit 12 is mainly responsible for executing instructions, such as performing operations on operands related to the instruction according to the instruction. The instruction execution unit 12 can perform operations on multiple groups of operands related to the instruction in a pipeline manner, thereby improving parallelism.
[0041] Register 13 is used to store operands related to the instruction execution process, such as source operands required to execute the instruction and destination operands generated by executing the instruction.
[0042] In some embodiments, Figure 2 As shown, the instruction issuing unit 11 and the instruction executing unit 12 are both deployed in the core. The GPU 10 includes multiple SPs, each SP is equivalent to a computing core of the GPU, and each SP includes an instruction issuing unit 11, an instruction executing unit 12 and a register 13. Figure 2 In the illustrated case, for each SP, the instruction execution unit 12 in the SP can only execute instructions issued by the instruction issuing unit 11 in the SP.
[0043] In some embodiments, Figure 3 As shown, the instruction issuing unit 11 is deployed in the core, and the instruction executing unit 12 is deployed outside the core. The GPU 10 includes multiple SPs, each SP is equivalent to a computing core of the GPU, and each SP includes an instruction issuing unit 11 and a register 13. In addition, the GPU may also include one or more instruction executing units 12. Figure 3 In the illustrated case, for each instruction execution unit 12 , the instruction execution unit 12 can execute instructions issued by different instruction issuing units 11 .
[0044] Compared to Figure 2 The architecture shown deploys the instruction execution unit 12 in the core. Figure 3 The architecture shown in which the instruction execution unit 12 is deployed outside the core helps to reduce the number of instruction execution units 12, thereby saving the hardware cost and area overhead of the GPU 10.
[0045] The related art provides the following two solutions for obtaining operands.
[0046] Solution 1
[0047] like Figure 4 As shown, after the instruction emission unit 11 issues an instruction, the data reading unit 14 reads the operands related to the instruction from the register 13 according to the instruction, and then stores the instruction and the operands related to the instruction together in the memory 15 at the front end of the instruction execution unit 12. After the instruction and the operands related to the instruction are all stored in the memory 15, the instruction execution unit 12 executes the instruction.
[0048] For the first solution, the instruction issuing unit 11 and the instruction executing unit 12 can be both deployed in the core, or the instruction issuing unit 11 can be deployed in the core and the instruction executing unit 12 can be deployed outside the core. Exemplarily, the first solution is mainly used for the situation where the instruction issuing unit 11 and the instruction executing unit 12 are both deployed in the core, for example, it is mainly used for the preparation operation of the data related to the ALU (Arithmetic and Logic Unit) in the core. Before the instruction enters the ALU for calculation, the operands related to the instruction must be prepared first, and the operands related to the instruction must be taken out from the register, stored in the cache in front of the ALU, and then submitted to the ALU for execution.
[0049] This solution places high demands on the storage space of the memory, requiring the memory to have a large storage space. It can be imagined that if there are multiple SPs in the GPU, if it is necessary to store instructions and related operands passed down by multiple SPs, a lot of information will need to be stored, which has a large storage overhead. For this solution, if the storage space of the memory is designed to be too small, the newly stored information will overwrite the previously stored information. If the instructions that have not been executed or their related operands are overwritten, this will affect the normal execution of the instructions and cause errors. Therefore, in order to avoid errors, the storage space of the memory is usually designed to be very large, but this will increase the hardware cost of the GPU.
[0050] Solution 2
[0051] like Figure 5 As shown, the instruction issuing unit 11 sends the instruction to the instruction executing unit 12. When the instruction executing unit 12 executes the instruction, the instruction executing unit 12 reads the operands related to the instruction from the register 13. After all the operands related to the instruction have been read, the instruction executing unit 12 executes the instruction.
[0052] For the second solution, the instruction issuing unit 11 and the instruction executing unit 12 can be both deployed in the core, or the instruction issuing unit 11 can be deployed in the core and the instruction executing unit 12 can be deployed outside the core. Exemplarily, the second solution is mainly used for the situation where the instruction issuing unit 11 is deployed in the core and the instruction executing unit 12 is deployed outside the core, for example, it is mainly used for the instruction execution pipeline outside the core to perform instruction execution operations.
[0053] In this solution, when the instruction execution unit executes an instruction, it needs to first issue a read request, read the operands related to the instruction from the register, and then wait to receive the operands. After all the operands related to the instruction are read, it returns to the instruction execution unit to execute the instruction. In the above process, from the time the instruction execution unit issues a read request to the time it receives all the required operands, the instruction execution unit is in a waiting state, which adds a large gap to the execution process, resulting in a large delay in instruction execution.
[0054] Based on the shortcomings of the above-mentioned solutions 1 and 2, there is currently no solution that can take into account both the processing delay of instructions and the hardware cost of the processor.
[0055] Please refer to Figure 6 , which shows a structural block diagram of a processor provided in a possible implementation of the present application. The processor 10 includes: an instruction issuing unit 11, an instruction pre-processing unit 16, a data reading unit 14, a register 13 and a memory 15.
[0056] The instruction transmitting unit 11 is used to transmit a first instruction.
[0057] The instruction pre-processing unit 16 is used to generate a read request corresponding to the first instruction according to the free storage space contained in the memory 15, and send the read request to the data reading unit 14, where the read request is used to request to read the operand related to the first instruction.
[0058] The data reading unit 14 is used to read the operands related to the first instruction from the register 13 according to the read request, and store the operands related to the first instruction in the memory 15 .
[0059] In some embodiments, the processor 10 may be a GPU. Of course, in some other embodiments, the processor 10 may also be a CPU or other processors with instruction execution capability, which is not limited in the present application.
[0060] In some embodiments, the first instruction may be any unexecuted instruction. In some embodiments, the first instruction may be an instruction that has not been executed and satisfies the execution condition. Exemplarily, the instruction emission unit 11 may select the first instruction that satisfies the execution condition from the unexecuted instructions according to the current instruction execution status and emit it downward for execution. Among them, the execution condition may be pre-set according to the situation, such as the execution condition including that the previous instruction has been handed over to the instruction execution unit 12 for execution and is at the head of the instruction queue, etc., which is not limited in this application. Among them, the instruction queue stores unexecuted instructions. In addition, in the embodiment of the present application, the function of the first instruction is not limited. The first instruction may be a calculation instruction, such as a summation instruction, a multiplication instruction, etc., which is used to perform operations such as summation and multiplication. The first instruction may also be a non-calculation instruction, such as a storage instruction, a sampling instruction, etc., which is used to perform operations such as storage and sampling.
[0061] In some embodiments, the instruction transmitting unit 11 transmits a first instruction to the instruction pre-processing unit 16 , and correspondingly, the instruction pre-processing unit 16 receives the first instruction transmitted by the instruction transmitting unit 11 .
[0062] In some embodiments, the memory 15 is used as a cache to store instructions and operands required by the instruction execution unit 12. The free storage space contained in the memory 15 refers to the unused storage space in the memory 15, or the remaining available storage space in the memory 15. The size of the free storage space of the memory 15 is equal to the size of the total storage space of the memory 15 minus the size of the used storage space of the memory 15. For example, if the size of the total storage space of the memory 15 is 10 KB (kilobytes), the size of the used storage space of the memory 15 is 6 KB, and the size of the free storage space of the memory 15 is 4 KB.
[0063] It should be noted that some numerical examples given in this application are given to facilitate the reader's understanding of the content of this solution. For example, the size of the storage space mentioned here and the number of blocks mentioned below are only examples given to facilitate the reader's understanding. The actual storage space size and the number of blocks can be reasonably set based on actual conditions, and this application does not limit this.
[0064] In some embodiments, operands refer to data to be operated or obtained by operation. Operands can be divided into source operands and destination operands. Source operands refer to data to be operated. Destination operands refer to operands obtained by executing instruction operations.
[0065] In some embodiments, the operands related to the first instruction include source operands related to the first instruction. For example, assuming that the first instruction is to calculate A×B+C, the source operands related to the first instruction include at least one group of source operands, and each group of source operands includes a value of A, a value of B, and a value of C. Since the processor 10 supports parallel computing, the source operands related to the first instruction may include multiple groups of values of A, B, and C, such as 32 groups of values of A, B, and C. The specific number of groups of source operands included is related to the parallel capability of the processor 10 or the size of a single thread bundle, and this application does not limit this.
[0066] In some embodiments, the instruction pre-processing unit 16 obtains the free storage space contained in the memory 15, and generates a read request corresponding to the first instruction when the free storage space is sufficient to store the operands related to the first instruction. For example, assuming that the size of the free storage space of the memory 15 is 4KB, and the operands related to the first instruction need to occupy 2KB of storage space, the free storage space is sufficient to store the operands related to the first instruction, and the instruction pre-processing unit 16 generates a read request corresponding to the first instruction.
[0067] In some embodiments, the number of read requests corresponding to the first instruction may be one or more. When the number of read requests corresponding to the first instruction is one, the one read request is used to request to read all operands related to the first instruction. When the number of read requests corresponding to the first instruction is multiple, each read request is used to request to read part of the operands related to the first instruction, and the multiple read requests can read all operands related to the first instruction.
[0068] In some embodiments, the pre-instruction processing unit 16 may split or generate multiple read requests corresponding to the first instruction, each read request being used to request to read a portion of operands related to the first instruction.
[0069] In some embodiments, the instruction pre-processing unit 16 can split or generate multiple read requests corresponding to the first instruction based on at least one of the following factors: the type of operand, the register address of the operand, the number of operands, the free storage space contained in the memory 15, etc.
[0070] For example, in the case where the operands related to the first instruction include multiple different types of operands, the instruction pre-processing unit 16 can split or generate multiple read requests corresponding to the first instruction, each read request is used to request to read one type of operand related to the first instruction. Exemplarily, assuming that the first instruction is to calculate A×B+C, the operands related to the first instruction include three different types of operands, namely A, B, and C. The instruction pre-processing unit 16 can split or generate three read requests corresponding to the first instruction, and the three read requests are used to request to read an operand of type A, an operand of type B, and an operand of type C, respectively.
[0071] For example, in the case where the operands related to the first instruction are stored in multiple different registers 13, the instruction pre-processing unit 16 can split or generate multiple read requests corresponding to the first instruction, each read request is used to request to read part of the operands related to the first instruction from a register. Exemplarily, assuming that the first instruction is to calculate A×B+C, the operands related to the first instruction include three different types of operands, namely A, B, and C. Assuming that the operands of type A and type B are stored in the same register, and the operands of type C are stored in another register, the instruction pre-processing unit 16 can split or generate 2 read requests corresponding to the first instruction, one of which is used to request to read the operands of type A and type B from one register, and the other is used to request to read the operand of type C from another register. Alternatively, the instruction pre-processing unit 16 can also split or generate 3 read requests corresponding to the first instruction, wherein the first read request is used to request to read the operand of type A from a register, the second read request is used to request to read the operand of type B from the register, and the third read request is used to request to read the operand of type C from another register.
[0072] Of course, in some embodiments, for operands of the same type, multiple operands of the same type may be read separately through multiple read requests, and this application does not limit this.
[0073] In some embodiments, the instruction pre-processing unit 16 may split or generate multiple read requests corresponding to the first instruction according to the free storage space included in the memory 15 .
[0074] In some embodiments, the instruction pre-processing unit 16 is also used to split or generate multiple read requests corresponding to the first instruction according to the free storage space when there is insufficient free storage space, that is, when the free storage space is insufficient to store the operands related to the first instruction, and each read request is used to request to read part of the operands related to the first instruction. Before each splitting or generating of a read request, the instruction pre-processing unit 16 obtains the size of the free storage space contained in the memory 15, and then splits or generates a read request, and the amount of data of the part of the operands related to the first instruction read by the read request is less than or equal to the size of the free storage space contained in the memory 15 obtained by the instruction pre-processing unit 16, so as to ensure that the part of the operands read by the read request can be smoothly written into the memory 15.
[0075] For example, assuming that the size of the free storage space of the memory 15 is 4KB, and the operands related to the first instruction need to occupy 6KB of storage space, since the free storage space is not enough to store all the operands related to the first instruction, the instruction pre-processing unit 16 can split or generate at least two read requests corresponding to the first instruction, and each read request is used to request to read part of the operands related to the first instruction. For example, the first read request is used to request to read a part of the operands that occupy about 3KB of storage space, and then when the free storage space of the memory 15 can store the remaining operands related to the first instruction, the second read request is generated to request to read the remaining operands related to the first instruction.
[0076] In some embodiments, the instruction pre-processing unit 16 may split or generate multiple read requests corresponding to the first instruction according to at least two factors of the type of operand, the register address where the operand is located, the number of operands, the free storage space contained in the memory 15, and the like. Taking the free storage space contained in the memory 15, the type of operand, and the register address where the operand is located as an example, assuming that the size of the free storage space of the memory 15 is 4KB, and the first instruction is to calculate A×B+C, then the operands related to the first instruction include three different types of operands, namely A, B, and C. Assuming that the operand of type A and the operand of type B are stored in the same register, and the operand of type A occupies 3KB, the operand of type B also occupies 3KB, and the two require a total of 6KB of storage space, the operand of type C is stored in another register, and the operand of type C occupies 2KB. In this case, the instruction pre-processing unit 16 may split or generate three read requests corresponding to the first instruction, such as the first read request is used to request to read an operand of type A from a register. Afterwards, when the free storage space of the memory 15 can store an operand of type B, a second read request is generated, and the second read request is used to request to read an operand of type B from the above register. Afterwards, when the free storage space of the memory 15 can store an operand of type C, a third read request is generated, and the third read request is used to request to read an operand of type C from another register.
[0077] Through the above method, the instruction pre-processing unit 16 can split or generate multiple read requests corresponding to the first instruction, and read the operands related to the first instruction in multiple times, which fully improves the flexibility of operand reading. For example, when the operands are in different registers, the operands can be read from different registers respectively through multiple read requests. For example, when the free storage space contained in the memory 15 is not sufficient to store all the operands related to the first instruction, a part of the operands can be read first, and then the remaining operands can be read after the storage space is released. This not only improves flexibility, but also helps to speed up the efficiency and reliability of operand reading and storage.
[0078] In some embodiments, after receiving the read request, the data reading unit 14 reads the operand related to the first instruction from the register 13 according to the read request, and stores the operand related to the first instruction in the memory 15. Optionally, in the case where the data reading unit 14 receives multiple read requests, the data reading unit 14 can perform read register arbitration, that is, execute the multiple read requests in sequence in a certain order, and the specific method of determining the execution order is not limited in this application. For example, the data reading unit 14 executes each read request in sequence according to the order in which the read requests are received, and the earliest received read request is executed first, and the latest received read request is executed last.
[0079] In some embodiments, Figure 7 As shown, the processor 10 further includes a storage management unit 17. After receiving the read request, the data reading unit 14 reads the operands related to the first instruction from the register 13 according to the read request, and sends the operands related to the first instruction to the storage management unit 17.
[0080] In some embodiments, the storage management unit 17 is mainly responsible for managing the memory 15, such as allocating storage space of the memory 15, writing data into the memory 15, releasing storage space of the memory 15, etc. After receiving the operands related to the first instruction, the storage management unit 17 stores the operands related to the first instruction in the memory 15.
[0081] In some embodiments, Figure 7 As shown, the processor 10 further includes an instruction execution unit 12. The instruction execution unit 12 is used to execute the first instruction according to the operands related to the first instruction stored in the memory 10.
[0082] In some embodiments, the storage management unit 17 is further configured to send an instruction execution request to the instruction execution unit 12, the instruction execution request being used to request execution of the first instruction. The instruction execution unit 12 is configured to execute the first instruction according to the instruction execution request. Optionally, after all operands related to the first instruction are stored in the memory 15, the storage management unit 17 sends the instruction execution request to the instruction execution unit 12.
[0083] In some embodiments, the instruction execution request includes a first instruction and an operand related to the first instruction. After receiving the instruction execution request, the instruction execution unit 12 operates on the operand related to the first instruction according to the first instruction carried in the instruction execution request. In some embodiments, the instruction execution unit 12 may process the instruction execution request in a pipeline manner and operate on the operand related to the first instruction.
[0084] In summary, the technical solution provided by the embodiment of the present application, on the one hand, by setting an instruction pre-processing unit in the processor, it generates a read request corresponding to the instruction according to the free storage space contained in the memory, so that it can be ensured that when there is enough free storage space, the operands related to the instruction are read again, ensuring that the operands related to the instruction can be effectively stored in the memory. This method can ensure the reliability of operand reading and writing without having to design the storage space of the memory to be very large, thereby saving the hardware cost of the processor; on the other hand, after the instruction is issued, the operands corresponding to the instruction are read first, and then the operands corresponding to the instruction are stored in the memory. The operands stored in the memory can be directly used by the instruction execution unit when executing the instruction, without the need for the instruction execution unit to request to read the operands in the register after starting to execute the instruction, which avoids the hole caused by the instruction execution unit being in a waiting state, thereby not increasing the delay of instruction execution. Combining the above two aspects, the technical solution provided by the embodiment of the present application can take into account the processing delay of the instruction and the hardware cost of the processor.
[0085] Please refer to Figure 8 , which shows a structural block diagram of a processor provided in another possible implementation of the present application. The processor 10 includes: an instruction issuing unit 11, an instruction pre-processing unit 16, a block counter 18, a data reading unit 14, a storage management unit 17, an instruction execution unit 12, a register 13 and a memory 15.
[0086] The instruction transmitting unit 11 is used to transmit a first instruction.
[0087] The instruction pre-processing unit 16 is used to determine the number of blocks required to store operands related to the first instruction. When the number of free blocks contained in the memory 15 is greater than or equal to the above number of blocks, a read request corresponding to the first instruction is generated and sent to the data reading unit 14. The read request is used to request to read the operands related to the first instruction.
[0088] The pre-instruction processing unit 16 is further configured to update the value of the block counter 18 based on the number of blocks when the number of free blocks contained in the memory 15 is greater than or equal to the above number of blocks.
[0089] The instruction pre-processing unit 16 is further used to send a resource allocation request to the storage management unit 17, where the resource allocation request is used to request to allocate storage space for operands related to the first instruction.
[0090] The storage management unit 17 is used to allocate storage space for storing operands related to the first instruction according to the resource allocation request, and record the corresponding relationship between the number information of the first instruction and the number information of the allocated storage space.
[0091] The data reading unit 14 is used to read the operand related to the first instruction from the register 13 according to the read request, and send a data storage request to the storage management unit 17, wherein the data storage request includes the operand related to the first instruction and the number information of the first instruction.
[0092] The storage management unit 17 is used to determine the numbering information of the storage space corresponding to the numbering information of the first instruction after receiving the data storage request, store the operand related to the first instruction at the location of the storage space according to the numbering information of the storage space, and send an instruction execution request to the instruction execution unit 12, where the instruction execution request is used to request the execution of the first instruction.
[0093] The instruction execution unit 12 is configured to execute the first instruction according to the instruction execution request.
[0094] The storage management unit 17 is further used to recycle the blocks used to store operands related to the first instruction in the memory 15 as free blocks after the first instruction is executed, and to update the value of the block counter 18.
[0095] In some embodiments, the first instruction emitted by the instruction emission unit 11 includes the instruction content of the first instruction, and the instruction content is used to represent the calculation or operation performed by the first instruction. Optionally, the first instruction emitted by the instruction emission unit 11 also includes at least one of the following information: the thread warp number corresponding to the first instruction, and the PC (Program Counter) corresponding to the first instruction. The instruction counter is used to store the address of the next instruction of the first instruction.
[0096] In some embodiments, the storage space of the memory 15 may be divided into a plurality of blocks, and each block may be understood as a portion of storage resources of the memory 15. The storage space of the memory 15 may be allocated and used in blocks.
[0097] In some embodiments, in order to record the free storage space contained in the memory 15 , a block counter 18 is provided in the processor 10 , and the value of the block counter 18 is used to determine the number of free blocks contained in the memory 15 .
[0098] Exemplarily, the value of the block counter 18 is equal to the number of free blocks contained in the memory 15. The free blocks refer to unused blocks, that is, blocks that have not stored any data.
[0099] Exemplarily, the value of the block counter 18 is equal to the number of used blocks contained in the memory 15. The used blocks refer to used blocks, that is, blocks in which data has been stored. According to the total number of blocks contained in the memory 15 and the number of used blocks, the number of free blocks contained in the memory 15 can also be determined.
[0100] In some embodiments, after receiving the first instruction, the instruction pre-processing unit 16 first determines the number of blocks required to store the operands related to the first instruction, which refers to the number of blocks required to store the operands related to the first instruction, that is, the number of blocks required to be occupied by the operands related to the first instruction. When the number of free blocks contained in the memory 15 is greater than or equal to the above number of blocks, it means that there is enough storage space in the memory 15 to store the operands related to the first instruction. The instruction pre-processing unit 16 generates a read request corresponding to the first instruction and sends the read request to the data reading unit 14.
[0101] In the above manner, the instruction pre-processing unit 16 first determines the number of blocks required to store the operands related to the first instruction. When the number of free blocks contained in the memory 15 is greater than or equal to the number of blocks, that is, when there is sufficient storage space in the memory 15 to store the operands related to the first instruction, the instruction pre-processing unit 16 generates a read request corresponding to the first instruction, which helps to ensure that the operands subsequently read can be successfully written into the memory 15, thereby ensuring the reliability of operand reading and writing, and then ensuring the accuracy of instruction execution.
[0102] In addition, a block counter 18 is provided in the processor 10 , and the number of free blocks is recorded by the block counter 18 , so that the pre-instruction processing unit 16 can efficiently and accurately know the size of the free storage space in the memory 15 .
[0103] Furthermore, the pre-instruction processing unit 16 updates the value of the block counter 18 based on the number of blocks.
[0104] Exemplarily, when the value of the block counter 18 is equal to the number of free blocks contained in the memory 15, the instruction pre-processing unit 16 subtracts the number of blocks from the value of the block counter 18 and updates the value of the block counter 18. For example, the value of the block counter 18 is equal to 20, indicating that the number of free blocks contained in the memory 15 is 20. Assuming that the number of blocks required by the operands related to the first instruction is 4, at this time, the number of free blocks is greater than the number of blocks required by the operands related to the first instruction, a read request is sent, and the value of the block counter 18 is updated to 20-4=16.
[0105] Exemplarily, when the value of the block counter 18 is equal to the number of used blocks contained in the memory 15, the instruction pre-processing unit 16 adds the value of the block counter 18 to the above number of blocks, and updates the value of the block counter 18. For example, the value of the block counter 18 is equal to 40, indicating that the number of used blocks contained in the memory 15 is 40. Assuming that the total number of blocks contained in the memory 15 is 60, the number of blocks required by the operands related to the first instruction is 5, and the number of free blocks is 20 at this time, which is greater than the number of blocks required by the operands related to the first instruction, a read request is sent, and the value of the block counter 18 is updated to 40+5=45.
[0106] In the above manner, when the instruction pre-processing unit 16 determines that the first instruction needs to occupy part of the storage space in the memory 15 , it updates the value of the block counter 18 , thereby ensuring that the value of the block counter 18 can accurately reflect the number of free blocks in the memory 15 .
[0107] In some embodiments, the instruction pre-processing unit 16 is further configured to, when the free storage space is insufficient, wait for the free storage space to be sufficient to store the operands related to the first instruction, and then execute the step of generating a read request corresponding to the first instruction. In some embodiments, the free storage space is insufficient, that is, the number of free blocks included in the memory 15 is less than the number of blocks required to store the operands related to the first instruction.
[0108] In some embodiments, the pre-instruction processing unit 16 is further used to split or generate multiple read requests corresponding to the first instruction when the number of free blocks contained in the memory 15 is less than the number of blocks required to store the operands related to the first instruction, and each read request is used to request to read part of the operands related to the first instruction. Before each splitting or generating a read request, the pre-instruction processing unit 16 obtains the number of free blocks contained in the memory 15, and then splits or generates a read request, and the number of blocks required for the read request to read the part of the operands related to the first instruction is less than or equal to the number of free blocks contained in the memory 15 obtained by the pre-instruction processing unit 16, so as to ensure that the part of the operands read by the read request can be smoothly written into the memory 15. In this way, when the free storage space contained in the memory 15 is not enough to store all the operands related to the first instruction, a part of the operands can be read first, and the remaining operands can be read after the storage space is released, which helps to speed up the efficiency and reliability of operand reading and storage while improving flexibility.
[0109] In some embodiments, the instruction pre-processing unit 16 is further used to send feedback information to the instruction transmitting unit 11 when there is insufficient free storage space, and the feedback information is used to indicate that the first instruction cannot be executed temporarily. The instruction transmitting unit 11 is also used to transmit a second instruction, which is different from the first instruction.
[0110] The second instruction is another instruction different from the first instruction. Wherein, the second instruction is different from the first instruction, and the function of the second instruction and the first instruction may be different, such as the operation performed by the second instruction is different from that performed by the first instruction, the operands related to the second instruction are different from the operands related to the first instruction, etc. In some embodiments, the second instruction is an instruction that is not executed and does not depend on the execution result of the first instruction. In some embodiments, the data volume of the operands related to the second instruction is less than or equal to the size of the free storage space contained in the memory 15. Optionally, the number of blocks required to store the operands related to the second instruction is less than or equal to the number of free blocks contained in the memory 15. In order to save time, if the first instruction cannot be executed temporarily, if there is still a second instruction in the instruction queue that is not executed and does not depend on the execution result of the first instruction, the second instruction can be executed first, thereby avoiding waiting and speeding up the execution efficiency of the instruction. Of course, if there is an unexecuted instruction in the instruction queue, but the unexecuted instruction needs to depend on the execution result of the first instruction, it is impossible to skip the first instruction to execute the unexecuted instruction, and it is still necessary to wait for the first instruction to be executed before executing the unexecuted instruction.
[0111] In some embodiments, when the number of free blocks contained in the memory 15 is greater than or equal to the above-mentioned number of blocks, the pre-instruction processing unit 16 will also send a resource allocation request to the storage management unit 17 through the data path between the storage management unit 17, and the resource allocation request is used to request the allocation of storage space for the operands related to the first instruction.
[0112] In some embodiments, the resource allocation request includes the numbering information of the first instruction and the number of blocks required to store the operands related to the first instruction. After receiving the resource allocation request, the storage management unit 17 selects a corresponding number of free blocks from the free blocks contained in the memory 15 according to the above number of blocks and allocates them to the first instruction. After determining the blocks allocated to the first instruction, the storage management unit 17 records the correspondence between the numbering information of the first instruction and the numbering information of the allocated storage space (such as the numbering information of the allocated blocks). For example, the number of blocks required to store the operands related to the first instruction is 3, and the storage management unit 17 allocates the free blocks numbered 2, 5, and 11 to the first instruction. By recording the correspondence between the numbering information of the first instruction and the numbering information of the allocated storage space, when the storage management unit 17 receives a data storage request later, it can determine the blocks used to store the operands related to the first instruction according to the numbering information of the first instruction contained in the data storage request, thereby ensuring that the operands are stored in the correct position.
[0113] In some embodiments, the read request sent by the instruction pre-processing unit 16 to the data reading unit 14 includes the serial number information of the first instruction.
[0114] In some embodiments, the numbering information of the first instruction includes an instruction number. Optionally, the numbering information of the first instruction also includes at least one of the following: a stream processor number, a thread bundle number, an operand number, and a channel number. Among them, the instruction number refers to the number of the first instruction, which is used to distinguish different instructions, and different instructions have different numbers. The stream processor number refers to the number of the stream processor that issues the first instruction, which is used to distinguish different stream processors, and different stream processors have different numbers. The thread bundle number refers to the number of the thread bundle corresponding to the first instruction, which is used to distinguish different thread bundles, and different thread bundles have different numbers. The operand number refers to the number of the operand corresponding to the first instruction, which is used to distinguish different types of operands, and different types of operands have different numbers. The channel number refers to the number of the channel of the operand corresponding to the first instruction, which is used to distinguish operands of different channels, and operands of different channels have different numbers.
[0115] In some embodiments, after the data reading unit 14 reads the operands related to the first instruction from the register 13 according to the read request, it sends a data storage request to the storage management unit 17, and the data storage request includes the operands related to the first instruction and the serial number information of the first instruction. A direct data path can be established between the data reading unit 14 and the storage management unit 17, such as Figure 7As shown, the data reading unit 14 sends the data storage request to the storage management unit 17 through the direct data path. Alternatively, an indirect data path can also be established between the data reading unit 14 and the storage management unit 17, such as Figure 8 As shown in , the data reading unit 14 first sends the data storage request to the instruction pre-processing unit 16 , and then the instruction pre-processing unit 16 sends the data storage request to the storage management unit 17 .
[0116] In some embodiments, after receiving the data storage request, the storage management unit 17 determines the numbering information of the storage space corresponding to the numbering information of the first instruction according to the correspondence relationship recorded above, and stores the operands related to the first instruction at the location of the storage space according to the numbering information of the storage space. For example, the storage management unit 17 determines the numbering information of the block corresponding to the numbering information of the first instruction according to the correspondence relationship recorded above, that is, determines the block for storing the operands related to the first instruction, and then stores the operands related to the first instruction in the above-determined block. In the above manner, the storage management unit 17 can accurately store the received operands in the previously allocated location.
[0117] After all operands related to the first instruction are stored in the memory 15 , the storage management unit 17 may generate one or more instruction execution requests, and then send the instruction execution requests to one or more instruction execution units 12 , where the instruction execution requests are used to request execution of the first instruction.
[0118] In some embodiments, the storage management unit 17 may split the operands related to the first instruction into smaller granularities than a single thread bundle according to the processing capability of the instruction execution unit 12, and then read the required operands from the memory 15, and generate an instruction execution request together with the first instruction, and send it to the instruction execution unit 12. For example, assuming that the first instruction is a store instruction, the first instruction is used to write different image channel data of R (red), G (green), B (blue), and A (alpha) into a certain image, and a single thread bundle includes 32 groups of operands, each group of operands includes a value of R, a value of G, a value of B, and a value of A. Assuming that the instruction execution unit 12 can write 4 groups of operands into the image in parallel at a time, the storage management unit 17 may generate 8 instruction execution requests, each of which includes the instruction content of the first instruction and 4 groups of operands.
[0119] For example, the information included in a single instruction execution request may be as follows: Fig. 9 As shown, it includes instruction information and data information; wherein the instruction information includes the instruction content of the first instruction, and the data information includes at least one group of operands.
[0120] In some embodiments, after the first instruction is executed, the storage management unit 17 will promptly reclaim the block in the memory 15 used to store operands related to the first instruction, thereby releasing the storage space of the memory 15 in time and updating the value of the block counter 18.
[0121] Exemplarily, when the value of the block counter 18 is equal to the number of free blocks contained in the memory 15, when the blocks in the memory 15 used to store operands related to the first instruction are recovered as free blocks, so that the number of free blocks contained in the memory 15 increases by the first value, the value of the block counter 18 is added to the first value.
[0122] Exemplarily, when the value of the block counter 18 is equal to the number of used blocks contained in the memory 15, when a block in the memory 15 used to store operands related to the first instruction is recycled as a free block, so that the number of used blocks contained in the memory 15 is reduced by a first value, the value of the block counter 18 is subtracted by the first value.
[0123] Through the above method, the storage space of the memory 15 can be released in time for use by other instructions, thereby improving the utilization rate of the memory 15.
[0124] In some embodiments, the memory 15 includes a plurality of blocks, each of which includes a plurality of storage units distributed in an array. Fig.10 As shown, the memory 15 includes a plurality of blocks, each of which includes a plurality of storage units distributed in an array. Fig.10 In the diagram, each small square represents a storage unit, and each block includes 16 storage units distributed in a 4*4 array.
[0125] In some embodiments, different blocks may be used to store different data, such as different types of data. Fig.10 In the figure, the memory 15 includes odd-numbered blocks (blocks shown without filling in the figure) and even-numbered blocks (blocks shown with diagonal filling in the figure), and the odd-numbered blocks and the even-numbered blocks are used to store different types of data. Of course, the division of blocks in the memory 15 and the functions of each block can also be implemented in other ways, and this application does not limit this.
[0126] In some embodiments, for a block, storage cells in different rows are used to store operands of different types. Fig.10In the example, each block includes 4 rows, and each row of storage units is used to store one type of operand. Assuming that the first instruction is a store instruction, the first instruction is used to write R, G, B, and A into a certain image. A single thread warp includes 32 groups of operands, and each group of operands includes a value of R, a value of G, a value of B, and a value of A. The first row is used to store R data, the second row is used to store G data, the third row is used to store B data, and the fourth row is used to store A data.
[0127] In some embodiments, for a block, storage cells in different columns are used to store different operands of the same type. Fig.10 In the example, each block includes 4 columns, and each column of storage cells is used to store different operands of the same type.
[0128] It should be noted that the present application does not limit the number of blocks included in the storage unit, the number of rows, columns, and quantity of storage units included in each block, and the number of operands stored in each storage unit. The numerical examples given above are only examples given for ease of understanding and do not constitute a limitation on the technical solution of the present application.
[0129] In some embodiments, the storage management unit 17 is used to read multiple operands required for a single instruction execution request from different columns of a block; convert the multiple operands into n groups of operands, each group of operands includes at least one operand, n is a positive integer, and when n is greater than 1, operations on the n groups of operands are executed in parallel; and generate an instruction execution request, the instruction execution request includes the first instruction and n groups of operands.
[0130] Since the structural design of the memory 15 determines that it is impossible to read data of different rows from the same column of a block at one time, but it is possible to read data of different rows from different columns of a block at one time, therefore, when the storage management unit 17 reads operands from a block, it can read multiple operands required for a single instruction execution request from different columns of the block.
[0131] Afterwards, the storage management unit 17 may perform a swizzle process on the read operands, such as performing a vertical to horizontal conversion, and converting the multiple operands into n groups of operands, each group of operands including at least one operand required for one operation. The operations of the n groups of operands may be performed in parallel.
[0132] For example, Fig.11As shown, it is assumed that 4 operands of type A need to be read, which are recorded as operands A1, A2, A3 and A4; 4 operands of type B are recorded as operands B1, B2, B3 and B4; 4 operands of type C are recorded as operands C1, C2, C3 and C4; and 4 operands of type D are recorded as operands D1, D2, D3 and D4. When writing operands into the block of the memory 15, the operands A1, A2, A3 and A4 can be written into the first column of the first row, the operands B1, B2, B3 and B4 can be written into the second column of the second row, the operands C1, C2, C3 and C4 can be written into the third column of the third row, and the operands D1, D2, D3 and D4 can be written into the fourth column of the fourth row, so that all the above 16 operands can be directly read from different columns at one time.
[0133] Afterwards, the storage management unit 17 performs a swizzle process on the 16 operands read, such as performing a vertical to horizontal conversion, and converting the 16 operands into 4 groups of operands, namely: (A1, B1, C1, D1), (A2, B2, C2, D2), (A3, B3, C3, D3), (A4, B4, C4, D4). The operations of the above 4 groups of operands can be executed in parallel.
[0134] Through the above method, the storage management unit 17 can read multiple groups of operands at one time, thereby ensuring the efficiency of data reading and helping to improve the efficiency of instruction execution.
[0135] In some embodiments, Fig.12 As shown, the processor 10 includes a plurality of stream processors, each of which includes an instruction emission unit 11, an instruction pre-processing unit 16, a data reading unit 14, and a register 13. In some embodiments, each of the plurality of stream processors has an independent block counter 18, that is, different stream processors do not share the same block counter 18, each stream processor corresponds to a block counter 18, and different stream processors correspond to different block counters. In some embodiments, at least two of the plurality of stream processors share the same block counter 18. Fig.12 In the figure, stream processor 1 represents one stream processor, and the stacked boxes below it represent that processor 10 includes multiple stream processors. By at least two stream processors sharing the same block counter 18, the number of block counters 18 required to be included in processor 10 is reduced, thereby saving the hardware cost of processor 10.
[0136] In some embodiments, each instruction execution unit 12 uses an independent memory 15, that is, different instruction execution units 12 do not share the same memory 15, each instruction execution unit 12 corresponds to a memory 15, and different instruction execution units 12 correspond to different memories 15. In some embodiments, there are at least two instruction execution units 12 that share the same memory 15. In this way, it is not necessary to allocate a set of memory resources to each instruction execution unit 12, and the same memory 15 resource can be reused by multiple instruction execution units 12 and multiple instructions, thereby improving utilization, avoiding the need for each instruction execution unit 12 to be individually allocated a set of memory resources, wasting storage space, and saving the hardware cost of the processor 10.
[0137] It should be noted that the processor provided in this application can adopt an in-core deployment architecture or an out-of-core deployment architecture.
[0138] In some embodiments, if the in-core deployment architecture is adopted, the instruction issuing unit 11 and the instruction execution unit 12 are both deployed in the core. The processor 10 includes a plurality of stream processors, each of which includes an instruction issuing unit 11, an instruction pre-processing unit 16, a data reading unit 14, a storage management unit 17, an instruction execution unit 12, a register 13, and a memory 15. For each stream processor, the instruction execution unit 12 in the stream processor can only execute instructions issued by the instruction issuing unit 11 in the stream processor.
[0139] In some embodiments, if an out-of-core deployment architecture is adopted, the instruction emission unit 11 is deployed in the core, and the instruction execution unit 12 is deployed outside the core. The processor 10 includes a plurality of stream processors, each of which includes an instruction emission unit 11, an instruction pre-processing unit 16, a data reading unit 14, and a register 13. In addition, the processor 10 may also include one or more instruction execution units 12. Optionally, the storage management unit 17 and the memory 15 may also be deployed out of the core, so that the instruction execution unit 12 can efficiently read data from the memory 15 when executing instructions. For each instruction execution unit 12, the instruction execution unit 12 can execute instructions issued by the instruction emission units 11 in different stream processors.
[0140] Compared with the in-core deployment architecture, the out-of-core deployment architecture helps to reduce the number of instruction execution units 12, thereby saving the hardware cost and area overhead of the processor 10.
[0141] In addition, regardless of whether the processor provided in the present application adopts an in-core deployment architecture or an out-of-core deployment architecture, the functions of the various components in the processor are basically the same under these two architectures, and can be referred to the introduction and description in the above embodiments.
[0142] The following is an embodiment of the method of the present application. For details not disclosed in the embodiment of the method of the present application, please refer to the above embodiment.
[0143] Please refer to Fig.13 , which shows a flowchart of an operand acquisition method applied to a processor provided in a possible implementation of the present application. The components of the processor can be found in the description of the above embodiment. Fig.13 As shown, the method may include the following steps:
[0144] Step 1310: the instruction issuing unit issues a first instruction.
[0145] Step 1320: the instruction pre-processing unit generates a read request corresponding to the first instruction according to the free storage space contained in the memory, and sends the read request to the data reading unit, where the read request is used to request to read an operand related to the first instruction.
[0146] In some embodiments, the pre-instruction processing unit generates a read request corresponding to the first instruction based on the free storage space contained in the memory, including: the pre-instruction processing unit determines the number of blocks required to store operands related to the first instruction, and generates a read request corresponding to the first instruction when the number of free blocks contained in the memory is greater than or equal to the number of blocks.
[0147] In some embodiments, the method further includes: the pre-instruction processing unit reads the value of the block counter, and determines the number of free blocks contained in the memory according to the value of the block counter.
[0148] In some embodiments, the method further comprises: updating the value of the block counter based on the number of blocks if the number of free blocks is greater than or equal to the number of blocks.
[0149] In some embodiments, the method further includes: the instruction pre-processing unit sends a resource allocation request to the storage management unit, the resource allocation request is used to request to allocate storage space for operands related to the first instruction. The storage management unit allocates storage space for storing operands related to the first instruction according to the resource allocation request, and records the correspondence between the number information of the first instruction and the number information of the allocated storage space.
[0150] In some embodiments, the method further includes: when the free storage space is insufficient, the instruction pre-processing unit waits for the free storage space to be sufficient to store operands related to the first instruction, and then executes the step of generating a read request corresponding to the first instruction.
[0151] In some embodiments, the method further includes: when the free storage space is insufficient, the instruction pre-processing unit sends feedback information to the instruction emission unit, the feedback information is used to indicate that the first instruction cannot be executed temporarily. The instruction emission unit emits a second instruction, the second instruction is different from the first instruction. Optionally, the second instruction is an instruction that has not been executed and does not depend on the execution result of the first instruction.
[0152] Step 1330: The data reading unit reads operands related to the first instruction from the register according to the read request, and stores the operands related to the first instruction in the memory.
[0153] In some embodiments, the data reading unit sends operands related to the first instruction to the storage management unit.
[0154] In some embodiments, the data reading unit sends the operand related to the first instruction to the storage management unit, including: the data reading unit sends a data storage request to the storage management unit, and the data storage request includes the operand related to the first instruction and the number information of the first instruction.
[0155] In some embodiments, the storage management unit stores operands related to the first instruction in the memory, and sends an instruction execution request to the instruction execution unit, where the instruction execution request is used to request execution of the first instruction.
[0156] In some embodiments, the storage management unit stores the operands related to the first instruction in the memory, including: after receiving the data storage request, the storage management unit determines the numbering information of the storage space corresponding to the numbering information of the first instruction, and stores the operands related to the first instruction at the location of the storage space according to the numbering information of the storage space.
[0157] In some embodiments, the method also includes: a storage management unit reads multiple operands required for a single instruction execution request from different columns of a block; converts the multiple operands into n groups of operands, each group of operands includes at least one operand, n is a positive integer, and when n is greater than 1, operations on the n groups of operands are executed in parallel; and generates an instruction execution request, the instruction execution request including a first instruction and n groups of operands.
[0158] In some embodiments, the method further includes: an instruction execution unit executing the first instruction according to an operand related to the first instruction stored in a memory.
[0159] In some embodiments, the instruction execution unit receives an instruction execution request sent by the storage management unit, and executes the first instruction according to the instruction execution request.
[0160] In some embodiments, the method further includes: after the first instruction is executed, the storage management unit reclaims the block in the memory used to store the operand related to the first instruction as a free block, and updates the value of the block counter.
[0161] An exemplary embodiment of the present application further provides a chip product, the chip product includes the processor described above. Optionally, the chip product may be a GPU chip product, the processor described above is a GPU, and the GPU chip product includes the GPU described above. Optionally, the above chip product may be implemented as a graphics card, the graphics card includes the processor described above, such as a GPU.
[0162] An exemplary embodiment of the present application further provides a computer device, which includes the processor described above. Optionally, the computer device can be a personal computer, a workstation, a game console, and some mobile devices (such as a tablet computer, a smart phone, etc.), or a vehicle-mounted terminal device, a smart home device, a smart TV, an intelligent robot, etc., or a server, a server cluster, an artificial intelligence computing cluster, a cloud computing cluster, etc., wherein the artificial intelligence computing cluster can also be referred to as an intelligent computing cluster or a smart computing cluster, which is not limited in the present application.
[0163] It should be understood that the "multiple" mentioned in this article refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the objects associated before and after are in an "or" relationship. In addition, the step numbers described in this article only illustrate a possible execution sequence between the steps. In some other embodiments, the above steps may not be executed in the order of the numbers, such as two steps with different numbers are executed at the same time, or two steps with different numbers are executed in the opposite order to the diagram. The embodiments of the present application are not limited to this.
[0164] The above description is only an exemplary embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. A processor, characterized in that: The processor comprises: an instruction issuing unit, an instruction pre-processing unit, a data reading unit, a register and a memory; The instruction transmitting unit is used to transmit a first instruction; The instruction pre-processing unit is used to generate a plurality of read requests corresponding to the first instruction according to the free storage space contained in the memory, and send the read requests to the data reading unit, each of the read requests is used to request to read an operand related to the first instruction, and the operand is a part of the operand corresponding to the first instruction; The data reading unit is used to read the operands related to the first instruction from the register according to the read request, and store the operands related to the first instruction in the memory.
2. The processor according to claim 1, characterized in that The instruction pre-processing unit is used for: determining a number of blocks required to store operands associated with the first instruction; When the number of free blocks contained in the memory is greater than or equal to the number of blocks, a read request corresponding to the first instruction is generated.
3. The processor according to claim 2, characterized in that The processor further includes: a block counter, wherein a value of the block counter is used to determine the number of the free blocks; The instruction pre-processing unit is further configured to update the value of the block counter based on the number of blocks when the number of free blocks is greater than or equal to the number of blocks.
4. The processor according to claim 3, characterized in that The processor also includes a storage management unit; The storage management unit is used to recycle the block used to store the operand related to the first instruction in the memory as the free block and update the value of the block counter after the first instruction is executed.
5. The processor according to claim 3, characterized in that The processor comprises a plurality of stream processors, each of which comprises the instruction issuing unit, the instruction pre-processing unit, the data reading unit and the register; Each of the plurality of stream processors corresponds to one block counter, and different stream processors correspond to different block counters; or, At least two of the plurality of stream processors share the same block counter.
6. The processor according to claim 1, wherein: The processor also includes a storage management unit and an instruction execution unit; The data reading unit is used to send the operand related to the first instruction to the storage management unit; The storage management unit is used to store the operands related to the first instruction in the memory, and send an instruction execution request to the instruction execution unit, wherein the instruction execution request is used to request execution of the first instruction; The instruction execution unit is used to execute the first instruction according to the instruction execution request.
7. The processor according to claim 6, characterized in that The data reading unit is used to send a data storage request to the storage management unit, wherein the data storage request includes an operand related to the first instruction and number information of the first instruction; The storage management unit is used to determine the numbering information of the storage space corresponding to the numbering information of the first instruction after receiving the data storage request, and store the operand related to the first instruction at the location of the storage space according to the numbering information of the storage space.
8. The processor according to claim 6, characterized in that The memory includes a plurality of blocks, each block includes a plurality of storage units distributed in an array, storage units in different rows are used to store different types of operands, and storage units in different columns are used to store different operands of the same type; The storage management unit is used for: reading a plurality of operands required for a single instruction execution request from different columns of the block; Convert the plurality of operands into n groups of operands, each group of operands including at least one operand, n being a positive integer, and when n is greater than 1, operations of the n groups of operands are performed in parallel; The instruction execution request is generated, wherein the instruction execution request includes the first instruction and the n groups of operands.
9. The processor according to claim 6, characterized in that: Each of the instruction execution units corresponds to one of the memories, and different instruction execution units correspond to different memories; or, There are at least two of the instruction execution units that share the same memory.
10. The processor according to any one of claims 1 to 9, characterized in that: The instruction pre-processing unit is further configured to, when the free storage space is insufficient, wait for the free storage space to be sufficient to store operands related to the first instruction, and then execute the step of generating a read request corresponding to the first instruction; or, The instruction pre-processing unit is further used to send feedback information to the instruction transmitting unit when the free storage space is insufficient, wherein the feedback information is used to indicate that the first instruction cannot be executed temporarily; The instruction transmitting unit is further used to transmit a second instruction, where the second instruction is different from the first instruction.
11. A chip product, characterized in that: The chip product comprises the processor according to any one of claims 1 to 10.
12. A computer device, characterized in that: The computer device comprises a processor as claimed in any one of claims 1 to 10.
13. An operand acquisition method applied to a processor, characterized in that: The processor comprises: an instruction issuing unit, an instruction pre-processing unit, a data reading unit, a register and a memory; the method comprises: The instruction transmitting unit transmits a first instruction; The instruction pre-processing unit generates a plurality of read requests corresponding to the first instruction according to the free storage space contained in the memory, and sends the read requests to the data reading unit, each of the read requests being used to request reading an operand related to the first instruction, the operand being a part of the operand corresponding to the first instruction; The data reading unit reads operands related to the first instruction from the register according to the read request, and stores the operands related to the first instruction in the memory.
Citation Information
Patent Citations
First-in first-out data transfer control device having a plurality of banks
US20010029558A1