Apparatus and method for configuring a warp of cooperative threads in a vector processing system

Through the dynamic configuration of thread beam instruction scheduler and resource register, the problem of thread beam access conflict in vector computers is solved, pipeline efficiency and adaptability are improved, and it is suitable for big data and artificial intelligence computing.

CN114816529BActive Publication Date: 2025-07-18SHANGHAI BIREN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210479765.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-10-21
Publication Date
2025-07-18
Estimated Expiration
2040-10-21

AI Technical Summary

Technical Problem

Multiple thread bundles of access to general register files in vector computers may conflict, resulting in inefficient processing, especially in the problem of insufficient storage space or mismatch in big data and artificial intelligence computing.

Method used

The thread bundle instruction scheduler dynamically configures the general registers to ensure that the access between different thread bundles does not overlap, and the thread bundle resource register is used to store the basic position information of each thread bundle, and dynamically adjust the usage space of the general registers.

Benefits of technology

It improves the pipeline utilization rate of vector computing systems, can more widely adapt to the needs of big data and artificial intelligence computing, avoids mutual interference between thread bundles, and improves processing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114816529B_ABST
    Figure CN114816529B_ABST
Patent Text Reader

Abstract

The present invention relates to an apparatus and method for configuring a warp in a vector operation system. The apparatus includes: general-purpose registers; an arithmetic logic unit; a warp instruction scheduler; and a plurality of warp resource registers. The warp instruction scheduler allows each of a plurality of warps to include a part of relatively independent instructions in the program core according to the warp allocation instruction in the program core, and allows each of the plurality of warps to access all or specified local data in the general-purpose registers through the arithmetic logic unit according to the configuration during software execution, and completes the operations of each of the above warps through the arithmetic logic unit. The present invention can more widely adapt to different applications, such as big data, artificial intelligence operations, etc., through the above-described components that can allow software to dynamically adjust and configure general-purpose registers for different warps.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of a Chinese patent application with an application date of October 21, 2020, an application number of 202011131448.9, and an invention title of "Device and Method for Configuring Warps in a Vector Arithmetic System". Technical Field

[0002] The present invention relates to a vector arithmetic device, and particularly to a device and method for configuring warps in a vector arithmetic system. Background Art

[0003] A vector computer is a computer equipped with specialized vector instructions for improving vector processing speed. A vector computer can simultaneously process data calculations of multiple warps. Therefore, in terms of processing warp data, a vector computer is much faster than a scalar computer. However, conflicts may occur when multiple warps access the General-Purpose Register File (GPR File). Therefore, the present invention proposes a method for a device that configures warps in a vector arithmetic system. Summary of the Invention

[0004] In view of this, how to alleviate or eliminate the deficiencies in the above-related fields is indeed a problem to be solved.

[0005] An embodiment of the present invention relates to a device for configuring warps in a vector arithmetic system, including: general-purpose registers; an arithmetic logic unit; a warp instruction scheduler, and multiple warp resource registers. The warp instruction scheduler allows each of the multiple warps to include a part of relatively independent instructions in the program core according to the warp allocation instructions in the program core, and allows each of the multiple warps to access all or specified local data in the general-purpose registers through the arithmetic logic unit according to the configuration during software execution, and completes the operations of each of the above warps through the arithmetic logic unit. Each of the warp resource registers is associated with one of the warps, and is used to allow each of the warps to match the content of the corresponding warp resource register, and map data access to a specified local area in the general-purpose registers, where the specified local areas in the general-purpose registers mapped between different warps do not overlap.

[0006] Embodiments of the present invention also relate to a method for configuring warps in a vector operation system, including: according to the warp allocation instructions in the program core, each of multiple warps includes a part of relatively independent instructions in the program core; according to the configuration during software execution, each of the multiple warps accesses all or specified local data in the general register through the arithmetic logic unit; and completing the operations of each of the above warps through the arithmetic logic unit. The method further includes: according to the content of multiple warp resource registers, mapping the data access of each warp to the specified local in the general register, and the specified locals in the general register mapped among different warps do not overlap.

[0007] One of the advantages of the above embodiments is that through the component and operation that can dynamically adjust and configure the general register for different warps as described above, it can more widely adapt to different applications, such as the operations of big data, artificial intelligence, etc.

[0008] One of the advantages of the above embodiments is that by dynamically configuring multiple relatively independent instruction segments in the program core for different warps, it is possible to avoid interference between warps to improve the utilization rate of the pipeline.

[0009] Other advantages of the present invention will be explained in more detail in conjunction with the following description and drawings. Description of the Drawings

[0010] The drawings described herein are used to provide a further understanding of the present application, form a part of the present application, and the illustrative embodiments and descriptions of the present application are used to explain the present application and do not constitute an improper limitation to the present application.

[0011] Figure 1 It is a block diagram of a vector operation system according to an embodiment of the present invention.

[0012] Figure 2 It is a block diagram of a streaming multi-processor according to an embodiment of the present invention.

[0013] Figure 3 It is a schematic diagram of the segmentation of a general register in some embodiments.

[0014] Figure 4 It is a schematic diagram of the dynamic segmentation of a general register in combination with warp resource registers according to an embodiment of the present invention.

[0015] Figure 5 It is a flowchart according to an embodiment of the present invention, applied to warps that execute tasks in parallel.

[0016] Figure 6 It is a schematic diagram of warps of producers and consumers according to an embodiment of the present invention.

[0017] Figure 7 A flowchart according to an embodiment of the present invention, applied to a cooperative thread bundle for executing producer - consumer tasks.

[0018] Among them, a simple description of the symbols in the drawings is as follows:

[0019] 10: Electronic device; 100: Streaming multi - processor; 210: Arithmetic logic unit; 220: Thread bundle instruction scheduler; 230: General - purpose register; 240: Instruction cache; 250: Barrier register; 260, 260#0 to 260#7: Each thread bundle resource register; 300#0 to 300#7: Storage blocks of general - purpose registers; Base#0 to Base#7: Base positions; S510 to S540: Method steps; 610: Consumer thread bundle; 621, 663: Barrier instructions; 623, 661: A series of instructions; 650: Producer thread bundle; S710 to S770: Method steps. Detailed implementation manners

[0020] The embodiments of the present invention will be described below in conjunction with relevant drawings. In these drawings, the same reference numerals represent the same or similar components or method flows.

[0021] It must be understood that the words "comprising", "including", etc. used in this specification are used to indicate the existence of specific technical features, numerical values, method steps, operations, components, and / or components, but do not exclude the addition of more technical features, numerical values, method steps, operations, components, components, or any combination of the above.

[0022] The words such as "first", "second", "third", etc. used in the present invention are used to modify the components in the claims, and do not indicate a priority order, precedence relationship, or that one component precedes another component, or the time sequence when performing method steps, but are only used to distinguish components with the same name.

[0023] It must be understood that when a component is described as "connected" or "coupled" to another component, it can be directly connected or coupled to other components, and intermediate components may appear. On the contrary, when a component is described as "directly connected" or "directly coupled" to another component, there are no intermediate components. Other words used to describe the relationship between components can be interpreted in a similar manner, such as "between" versus "directly between", or "adjacent" versus "directly adjacent", etc.

[0024] Reference Figure 1。The electronic device 10 can be implemented in electronic products such as mainframes, workstations, personal computers, laptop computers, tablet computers, mobile phones, digital cameras, digital video cameras, etc. The electronic device 10 can set up a Streaming Multiprocessor Cluster (SMC) in a vector computing system, which includes multiple Streaming Multiprocessors (SMs) 100. The instruction executions between different Streaming Multiprocessors 100 can be synchronized with each other using signals. After being programmed, the Streaming Multiprocessor 100 can execute various application tasks, including but not limited to: linear and non-linear data conversion, database operations, big data operations, artificial intelligence calculations, encoding, decoding, modeling operations, image rendering operations of audio and video data, etc. Each Streaming Multiprocessor 100 can execute multiple warps simultaneously. Each warp is composed of a group of threads, and a thread is the smallest unit that runs using hardware and has its own life cycle. The warp can be associated with Single Instruction Multiple Data (SIMD) instructions, Single Instruction Multiple Thread (SIMT) technology, etc. The executions between different warps can be independent or sequential. A thread can represent a task associated with one or more instructions. For example, each Streaming Multiprocessor 100 can execute 8 warps simultaneously, and each warp contains 32 threads. Although Figure 1 4 Streaming Multiprocessors 100 are described, those skilled in the art can set more or fewer Streaming Multiprocessors in the vector computing system according to different needs, and the present invention is not limited thereby.

[0025] Reference Figure 2。Each streaming multi-processor 100 includes an Instruction Cache 240 for storing multiple instructions of a program kernel. Each streaming multi-processor 100 also includes a Warp Instruction Scheduler 220 for fetching a series of instructions for each warp and storing them in the Instruction Cache 240, and for fetching instructions to be executed from the Instruction Cache 240 for each warp according to the program counter. Each warp has an independent Program Counter (PC) register for recording the position of the currently executing instruction (i.e., the instruction address). Whenever an instruction is fetched from the Instruction Cache for a warp, the corresponding program counter is incremented by one. The Warp Instruction Scheduler 220 sends the instructions to the Arithmetic Logical Unit (ALU) 210 for execution at an appropriate time point, and these instructions are defined in the Instruction Set Architecture (ISA) of a specific computing system. The Arithmetic Logical Unit 210 can perform various operations, such as addition and multiplication calculations of integers and floating-point numbers, comparison operations, Boolean operations, bit shifts, and algebraic functions (such as planar interpolation, trigonometric functions, exponential functions, logarithmic functions), etc. During the execution process, the Arithmetic Logical Unit 210 can read data from a specified position (also known as the source address) of the General-Purpose Registers (GPRs) 230 and write back the execution result to a specified position (also known as the destination address) of the General-Purpose Registers 230. Each streaming multi-processor 100 also includes a Barriers Register 250, which can be used to enable software to synchronize the execution between different warps, and a Resource-per-warp Register 230, which can be used to enable software to dynamically configure the space range of the General-Purpose Registers 230 that each warp can use during execution. Although Figure 2 only components 210 to 260 are listed, but this is only for briefly illustrating the technical features of the present invention, and those skilled in the art understand that each streaming multi-processor 100 also includes more components.

[0026] In some embodiments, the General-Purpose Registers 230 can be physically or logically fixedly divided into multiple Blocks, and the storage space of each block is only allocated for access by one warp. The storage spaces between different blocks do not overlap, which is used to avoid access conflicts between different warps. Refer to Figure 3, for example, when a streaming multi-processor 100 can process data of eight warps and the general register 230 has a storage space of 256 kilobytes (KB), the storage space of the general register 230 can be divided into eight blocks 300#0 to 300#7, each block containing non-overlapping 32KB storage space and provided for the use of a designated warp. However, since vector operation systems are often applied to the calculations of big data and artificial intelligence nowadays, the amount of data to be processed is huge, resulting in the fixed partitioned space being possibly insufficient for a warp and unable to meet the operation needs of a large amount of data. For applications in the calculations of big data and artificial intelligence, these embodiments can be modified such that each streaming multi-processor 100 only processes data of one warp, and the entire storage space of the general register 230 is only used by this warp. However, when two consecutive instructions have data dependency, that is, the input data of the second instruction is the output result after the execution of the first instruction, it will cause the operation of the arithmetic logic unit 210 to be inefficient. Specifically, the second instruction must wait for the execution result of the first instruction to be ready in the general register 230 before it can start execution. For example, assume that each instruction takes 8 clock cycles from initialization to output result to the general register 230 in the pipeline: when the second instruction needs to wait for the execution result of the first instruction, the second instruction can only start execution from the 9th clock cycle. At this time, the instruction execution latency is 8 clock cycles, resulting in a very low utilization rate of the pipeline. In addition, since the streaming multi-processor 100 only processes one warp, the instructions that could originally be executed in parallel need to be arranged for sequential execution, which is inefficient.

[0027] To solve the above problems, on the one hand, the warp instruction scheduler 220 allows each of multiple warps to access all or specified partial data in the general register 230 through the arithmetic logic unit 210 according to the configuration during software execution, and completes the operations of each warp through the arithmetic logic unit 210. Through the above-mentioned components that can enable software to dynamically adjust and configure the general register for different warps, it can be more widely adapted to different applications, such as the operations of big data and artificial intelligence.

[0028] On the other hand, the embodiments of the present invention provide an environment that enables software to determine the instruction segments included in each warp. In some embodiments, a program core can divide instructions into multiple segments, each segment of instructions being independent of each other and executed by a warp. Table 1 is an example of the virtual code of the program core:

[0029] Table 1

[0030]

[0031] Assume that each streaming multi-processor 100 runs at most eight warps, and each warp has a unique identifier: when the warp instruction scheduler 220 fetches the instructions of the program core shown in Table 1, it can check the warp identifier for a specific warp, jump to the instruction segment associated with this warp, and store it in the instruction cache 240. Then, it fetches the instructions from the instruction cache 240 according to the value of the corresponding program counter and sends them to the arithmetic logic unit 210 to complete a specific calculation. In this case, each warp can execute tasks independently, and all warps can run simultaneously, keeping the pipeline in the arithmetic logic unit 210 as busy as possible to avoid bubbles. The instructions of each segment in the same program core can be called relatively independent instructions. Although the example in Table 1 adds conditional judgment instructions to the program core to achieve instruction segmentation, those skilled in the art can use other different instructions that can achieve the same or similar effects to complete the instruction segmentation in the program core. In the program core, these instructions used to segment instructions for multiple warps can also be called warp allocation instructions. Generally speaking, the warp instruction scheduler 240 can, according to the warp allocation instructions in the program core, make each of the multiple warps contain a part of the relatively independent instructions in the program core, so that the arithmetic logic unit 210 can independently and parallelly execute the multiple warps.

[0032] The barrier register 250 can store information for synchronizing the execution of different warps, including the number of other warps that need to wait for execution to complete, and the number of warps that are currently waiting to continue execution. To coordinate the execution processes of different warps, software can set the content of the barrier register 250 to record the number of warps that need to wait for execution to complete. Each instruction segment in the program core can add barrier instructions at appropriate places according to system requirements. When the warp instruction scheduler 220 fetches a barrier instruction for a warp, it increments by 1 the number of warps that are currently waiting to continue execution recorded in the barrier register 250, and puts this warp into a waiting state. Then, the warp instruction scheduler 220 checks the content of the barrier register 250 to determine whether the number of warps that are currently waiting to continue execution is equal to or greater than the number of other warps that need to wait for execution to complete. If so, the warp instruction scheduler 220 wakes up all the waiting warps so that these warps can continue execution. Otherwise, the warp instruction scheduler 220 fetches the instructions of the next warp from the instruction cache.

[0033] In addition, Figure 3The division of the described embodiments is a pre - configuration of the streaming multi - processor 100 and cannot be modified by software. However, the storage spaces required by different warps may not be consistent. For some warps, the pre - divided storage space may exceed the requirement, while for others, it may be insufficient.

[0034] On the other hand, although a streaming multi - processor 100 can execute multiple warps and all warps execute the same program core, the embodiments of the present invention do not pre - divide the general - purpose registers 230 for different warps. Specifically, in order to more widely adapt to different applications, the streaming multi - processor 100 does not fixedly divide the general - purpose registers 230 into multiple storage spaces for multiple warps, but provides an environment that allows software to dynamically adjust and configure the general - purpose registers 230 for different warps, so that software can make each warp use all or part of the general - purpose registers 230 according to application requirements.

[0035] In some other embodiments, each streaming multi - processor 100 may include warp resource registers 260 for storing information on the base position of each warp, and each base position points to a specific address in the general - purpose registers 230. In order to allow different warps to access non - overlapping storage spaces in the general - purpose registers 230, software can dynamically change the content of the warp resource registers 260 to set the base position of each warp. For example, referring to Figure 4 , software can divide the general - purpose registers 230 into eight blocks for eight warps. Block 0 is associated with the 0th warp, and its address range includes Base#0 to Base#1 - 1; Block 1 is associated with the 1st warp, and its address range includes Base#1 to Base#2 - 1; and so on. Software can set the content of the warp resource registers 260#0 to 260#7 before the execution of the program core or at the beginning of the execution to point to the base addresses in the general - purpose registers 230 associated with each warp. After the warp instruction scheduler 220 extracts an instruction from the instruction cache 240 for the ith warp, it can adjust the source address and destination address of the instruction according to the content of the warp resource register 260#i to map to the storage space dynamically configured for the ith warp in the general - purpose registers 230. For example, the original instruction is:

[0036] Dest_addr=Instr_i(Src_addr0,Src_addr1)

[0037] where Instr_i represents the operation code (OpCode) of the instruction assigned to the ith warp, Src_addr0 represents the 0th source address, Src_addr0 represents the 1st source address, and Dest_addr represents the destination address.

[0038] The warp instruction scheduler 220 modifies the instructions as described above to:

[0039] Base#i + Dest_addr = Instr_i(Base#i + Src_addr0, Base#i + Src_addr1)

[0040] where Base#i represents the base position recorded in the warp resource register 260#i. That is, the warp instruction scheduler 220 adjusts the source address and destination address of the instruction according to the content of each warp resource register 260, so that the specified locals in the general registers mapped between different warps do not overlap.

[0041] In some other embodiments, not only can a program core divide instructions into independent segments, and each segment of instructions is executed by a warp, but software can also set the content of each warp resource register 260#0 to 260#7 before or at the beginning of the execution of the program core, for pointing to the base address associated with each warp in the general register 230.

[0042] The present invention can be applied to cooperative warps that execute tasks in parallel. Refer to Figure 5 the example flowchart shown.

[0043] Step S510: The warp instruction scheduler 220 starts to extract the instructions of each warp and stores them in the instruction cache 240.

[0044] Steps S520 to S540 form a loop. The warp instruction scheduler 220 can use a scheduling method (for example, the round-robin scheduling algorithm) to sequentially obtain the specified instructions from the instruction cache 240 according to the corresponding program counter of each warp, and send them to the arithmetic logic unit 210 for execution. The warp instruction scheduler 220 can sequentially obtain the instructions pointed to by the program counter of the 0th warp, the instructions pointed to by the program counter of the 1st warp, the instructions pointed to by the program counter of the 2nd warp, and so on, from the instruction cache 240.

[0045] Step S520: The warp instruction scheduler 220 obtains the instructions of the 0th or the next warp from the instruction cache 240.

[0046] Step S530: The warp instruction scheduler 220 sends the obtained instructions to the arithmetic logic unit 210.

[0047] Step S540: The arithmetic logic unit 210 performs a specified operation according to the input instruction. In the instruction execution pipeline, the arithmetic logic unit 210 obtains data from the source address of the general register 230, performs the specified operation on the obtained data, and stores the result of the operation in the destination address in the general register 230.

[0048] In some embodiments of step S510, to prevent different warps from conflicting when accessing the general register 230, the warp instruction scheduler 220 may, according to the warp allocation instructions in the program core (such as the example shown in Table 1), assign each warp to process multiple instructions (also known as relatively independent instructions) in a specified segment of the same program core. These instruction segments are arranged to be independent of each other and can be executed in parallel.

[0049] In some other embodiments, the warp instruction scheduler 220 may adjust the source address and destination address in the instruction according to the content of the corresponding warp resource register 260 before the instruction is sent to the arithmetic logic unit 210 (before step S530), for mapping to the storage space dynamically configured for this warp in the general register 230.

[0050] In still some other embodiments, the warp instruction scheduler 220 may, in addition to assigning each warp to process multiple instructions in a specified segment of the same program core according to the relevant instructions in the program core (such as the example shown in Table 1) in step S510, also adjust the source address and destination address in the instruction according to the content of the corresponding warp resource register 260 before the instruction is sent to the arithmetic logic unit 210 (before step S530).

[0051] The present invention can be applied to cooperative warps that execute producer - consumer type tasks. Refer to Figure 6 , assuming that warp 610 is a consumer of data and warp 650 is a producer of data. In other words, the execution of a part of the instructions in warp 610 requires referring to the execution results of a part of the instructions in warp 650. In some embodiments, when the software is executing, it can configure the content of the corresponding warp resource registers for warps 610 and 650, so that the instructions of warps 610 and 650 can access an overlapping block in the general register 230 through the arithmetic logic unit 210.

[0052] Regarding the execution of producer - consumer type tasks, specifically, refer to Figure 7 the flowchart example shown in

[0053] Step S710: The warp instruction scheduler 220 starts to extract the instructions of each warp and store them in the instruction cache 240. The warp instruction scheduler 220 can, according to the relevant instructions in the program core (such as the example shown in Table 1), make each warp responsible for processing multiple instructions in a specified segment of the same program core. These instruction segments form a producer-consumer relationship under the arrangement.

[0054] Step S720: The warp instruction scheduler 220 obtains the barrier instruction 621 of the consumer warp 610 from the instruction cache 240 and makes the consumer warp 610 enter the waiting state accordingly.

[0055] Step S730: The warp instruction scheduler 220 obtains a series of instructions 661 of the producer warp 650 from the instruction cache 240 and sequentially sends the obtained instructions to the arithmetic logic unit 210.

[0056] Step S740: The arithmetic logic unit 210 performs the specified operation according to the input instruction 661. In the instruction execution pipeline, the arithmetic logic unit 210 obtains data from the source address of the general register 230, performs the specified operation on the obtained data, and stores the result of the operation in the destination address in the general register 230.

[0057] Step S750: The warp instruction scheduler 220 obtains the barrier instruction 663 of the producer warp 650 from the instruction cache 240 and wakes up the consumer warp 610 accordingly. In some embodiments, the warp instruction scheduler 220 can also make the producer warp 650 enter the waiting state.

[0058] Step S760: The warp instruction scheduler 220 obtains a series of instructions 623 of the consumer warp 610 from the instruction cache 240 and sequentially sends the obtained instructions to the arithmetic logic unit 210.

[0059] Step S770: The arithmetic logic unit 210 performs the specified operation according to the input instruction 623. In the instruction execution pipeline, the arithmetic logic unit 210 obtains data (including the data previously generated by the producer warp 650) from the source address of the general register 230, performs the specified operation on the obtained data, and stores the result of the operation in the destination address in the general register 230.

[0060] It should be noted here that the content regarding steps S730, S740, S760, and S770 is only a brief description for easy understanding. During the execution of steps S730 and S740, and S760 and S770, the warp instruction scheduler 220 can also obtain instructions of other warps (that is, any warp other than warp 610 and warp 610) from the instruction cache 240 and drive the arithmetic logic unit 210 to perform operations.

[0061] Although Figure 1 , Figure 2 contains the components described above, it does not exclude the use of more other additional components without violating the spirit of the invention to achieve better technical effects. In addition, although Figure 5 , Figure 7 the flowchart is executed in the specified order, those skilled in the art can modify the order between these steps on the premise of achieving the same effect without violating the spirit of the invention. Therefore, the present invention is not limited to only using the order described above. In addition, those skilled in the art can also integrate several steps into one step, or perform more steps sequentially or in parallel in addition to these steps, and the present invention should not be limited thereby.

[0062] The above is only a preferred embodiment of the present invention, but it is not used to limit the scope of the present invention. Any person familiar with this technology can make further improvements and changes on this basis without departing from the spirit and scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the scope defined by the claims of this application.

Claims

1. An apparatus for configuring a warp of threads in a vector operation system, characterized in that, Comprising: General registers; An arithmetic logic unit coupled to the general registers; A warp instruction scheduler coupled to the arithmetic logic unit, the warp instruction scheduler allocates instructions according to warps in a program core such that each of a plurality of warps contains a part of relatively independent instructions in the program core, and according to the configuration during software execution, each of the plurality of warps accesses specified local data in the general registers through the arithmetic logic unit, and executes operations of each of the above warps independently and in parallel through the arithmetic logic unit; And A plurality of warp resource registers, wherein each of the warp resource registers is associated with one of the warps, and is used to enable each of the warps to match the content of the corresponding warp resource register, map data access to the specified local in the general registers, wherein the content of each warp resource register is used to point to the base address in the general registers associated with each warp, and the specified locals in the general registers mapped between different warps do not overlap.

2. The apparatus for a cooperative warp in the configuration vector operation system according to claim 1, wherein The device does not pre-configure a specified local in the general registers for each of the warps.

3. The apparatus for a cooperative warp in the configuration vector operation system according to claim 1, wherein The warp includes a first warp and a second warp. When the warp instruction scheduler obtains a barrier instruction of the first warp from an instruction cache, the first warp is put into a waiting state, and when the warp instruction scheduler obtains a barrier instruction of the second warp from the instruction cache, the first warp is awakened, wherein the first warp and the second warp are configured to be associated with an overlapping block in the general registers.

4. The apparatus for a cooperative warp in the configuration vector operation system according to claim 3, wherein The first warp is a consumer warp, and the second warp is a producer warp.

5. The apparatus for a cooperative warp in the configuration vector operation system according to claim 1, wherein The warps are independent of each other, and each of the warps is configured to be associated with a non-overlapping block in the general registers.

6. The apparatus for a cooperative warp in the configuration vector operation system according to claim 1, wherein The warp instruction scheduler maintains an independent program counter for each of the warps.

7. A method for configuring a warp of threads in a vector operation system, which is executed in a streaming multi-processor, characterized in that, Comprising: Allocating instructions according to warps in a program core such that each of a plurality of warps contains a part of relatively independent instructions in the program core; According to the configuration during software execution, enabling each of the plurality of warps to access specified local data in general registers through an arithmetic logic unit; And Executing operations of each of the above warps independently and in parallel through the arithmetic logic unit, wherein the method further includes: According to the content of a plurality of warp resource registers, mapping data access of each of the warps to the specified local in the general registers, and wherein the content of each warp resource register is used to point to the base address in the general registers associated with each warp, and the specified locals in the general registers mapped between different warps do not overlap.

8. The method for a cooperative warp in the configuration vector operation system according to claim 7, wherein The streaming multi-processor does not pre-configure a specified local in the general registers for each of the warps.

9. The method for a collaborative warp in the configuration vector operation system according to claim 7, wherein The warp includes a first warp and a second warp, the first warp and the second warp are configured to be associated with an overlapping block in the general registers, and the method includes: When obtaining the barrier instruction of the first warp from the instruction cache, putting the first warp into a waiting state; and When obtaining the barrier instruction of the first warp from the instruction cache, waking up the second warp.

10. The method for a cooperative warp in the configuration vector operation system according to claim 9, wherein, The first warp is a consumer warp, and the second warp is a producer warp.

11. The method of a collaborative warp in the configuration vector operation system according to claim 7, wherein, Including: Maintaining an independent program counter for each warp.

Citation Information

Patent Citations

  • Multithreading processor and multithreading processing method

    CN101344842A

  • Instruction and logic for a vector format for processing computations

    CN106575219A