Processor, graphics card, computer equipment and register allocation method and device

By dividing the GPU's registers into configuration block sets and allocating them into thread groups, the problem of register fragmentation in the prior art is solved, and more efficient resource management and allocation is achieved.

CN120578422AActive Publication Date: 2025-09-02MOORE THREADS TECH CO LTD
View PDF 11 Cites 0 Cited by

Patent Information

Application Number
CN202511093566.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-06
Publication Date
2025-09-02
Estimated Expiration
2045-08-06

AI Technical Summary

Technical Problem

In the prior art, the GPU register allocation method is not conducive to management, resulting in frequent cutting of continuous registers, resulting in fragmentation, and affecting resource allocation and release efficiency.

Method used

Multiple registers are divided into a configuration block set, each configuration block set is used to allocate to a thread group, and resource management is carried out in the form of configuration blocks, avoiding frequent cutting of continuous registers, and improving allocation and release efficiency.

Benefits of technology

Register resource allocation is performed by configuring blocks, avoiding register fragmentation, improving resource allocation and release efficiency, and optimizing GPU management and performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120578422A_ABST
    Figure CN120578422A_ABST
Patent Text Reader

Abstract

The invention discloses a processor, a graphics card, computer equipment and a register allocation method and device, and belongs to the technical field of register management. The processor comprises a plurality of registers, a register management unit and a task assembly unit; the register management unit is used for dividing the plurality of registers into at least one configuration block set, each configuration block set in the at least one configuration block set comprises at least one configuration block, and the at least one configuration block comprises at least two registers in the plurality of registers; the task assembly unit is used for distributing a first configuration block set for a first thread group, and the first configuration block set is one of the at least two configuration block sets. The processor allocates register resources in the form of configuration blocks, and the allocation and release efficiency of the register resources can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of register management technology, and in particular to a processor, a graphics card, a computer device, and a register allocation method and apparatus. Background Art

[0002] Registers are crucial components of a graphics processing unit (GPU). They are the fastest-access memory in the GPU and provide fast data readout for the GPU's processing core.

[0003] In related technologies, when allocating resources to thread groups, the allocation is usually done at the register granularity. For example, multiple registers may be allocated to a thread group, but these registers may be continuous or non-contiguous, and the register resources allocated to different registers may be different each time.

[0004] However, this register allocation method is not conducive to GPU register management. Summary of the Invention

[0005] The present application provides a processor, graphics card, computer equipment, register allocation method and apparatus, and the technical solution is as follows.

[0006] According to one aspect of the present application, a processor is provided, comprising a plurality of registers, a register management unit, and a task assembly unit; The register management unit is configured to divide the plurality of registers into at least one configuration block set, each configuration block set in the at least one configuration block set includes at least one configuration block, and the at least one configuration block includes at least two registers in the plurality of registers; The task assembly unit is configured to allocate a first configuration block set to the first thread group, where the first configuration block set is one of the at least two configuration block sets.

[0007] According to one aspect of the present application, a graphics card is provided, comprising the above-mentioned processor.

[0008] According to one aspect of the present application, a computer device is provided, comprising the above-mentioned processor.

[0009] According to one aspect of the present application, a register allocation method is provided, the method being executed by the above-mentioned processor, the processor comprising a plurality of registers, a register management unit, and a task assembly unit; The method comprises: The register management unit divides the plurality of registers into at least one configuration block set, each configuration block set in the at least one configuration block set includes at least one configuration block, and the at least one configuration block includes at least two registers in the plurality of registers; The task assembly unit allocates a first configuration block set to the first thread group, where the first configuration block set is one of the at least two configuration block sets.

[0010] According to one aspect of the present application, a register allocation device is provided, the device comprising: The register usage management module divides the plurality of registers into at least one configuration block set, each configuration block set in the at least one configuration block set includes at least one configuration block, and the at least one configuration block includes at least two registers in the plurality of registers; The task assembly module allocates a first configuration block set to the first thread group, where the first configuration block set is one of the at least two configuration block sets.

[0011] The beneficial effects brought about by the technical solution provided in this application include at least the following:

[0012] By first dividing multiple registers into at least one configuration block set, each configuration block set is used to allocate to a thread group. By dividing register resources into multiple configuration blocks, register resources are allocated in the form of configuration blocks. Each thread group supports the allocation of at least one configuration block (i.e., a configuration block set). Compared with the traditional method of continuous allocation based on registers, this method of allocating register resources directly allocates configuration blocks instead of multiple continuous registers during allocation, which can avoid frequent cutting of continuous registers and cause register fragmentation. In addition, allocation through configuration blocks can achieve management through configuration blocks instead of management based on registers, which can improve the efficiency of register resource allocation and release. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0014] Figure 1 A schematic diagram of a processor provided by an exemplary embodiment of the present application is shown; Figure 2 A schematic diagram of a processor provided by yet another exemplary embodiment of the present application is shown; Figure 3 A schematic diagram of a processor provided by another exemplary embodiment of the present application is shown; Figure 4 A schematic diagram of an address conversion rule provided by an exemplary embodiment of the present application is shown; Figure 5 A schematic diagram showing a loading process of a processor provided by an exemplary embodiment of the present application is shown; Figure 6 A schematic diagram showing a loading process of a processor provided by yet another exemplary embodiment of the present application is shown; Figure 7 A flow chart of a register allocation method provided by an exemplary embodiment of the present application is shown; Figure 8 A structural block diagram of a register allocation device provided by an exemplary embodiment of the present application is shown. DETAILED DESCRIPTION

[0015] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.

[0016] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present application. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present application, as detailed in the appended claims.

[0017] The terms used in this disclosure are for the purpose of describing particular embodiments only and are not intended to limit the disclosure. As used in this disclosure and the appended claims, the singular forms "a," "an," "the," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise.

[0018] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, storage, and display, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the information such as the settings operations involved in this application is obtained with full authorization.

[0019] It should be understood that although the terms first, second, etc. may be used in this disclosure to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, a first parameter may also be referred to as a second parameter, and similarly, a second parameter may also be referred to as a first parameter without departing from the scope of this disclosure. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".

[0020] First, the relevant terms involved in this application are introduced.

[0021] Cache: A small, high-speed memory, part of a storage system, that stores instructions and data frequently used by programs. It can also be called cache memory. Processors typically include multiple levels of cache, each with a different capacity. Smaller caches typically have higher read and write efficiency because they need to process less data.

[0022] The cache is divided into multiple levels, such as the first level (L1), second level (L2) and third level (L3) cache. The cache capacity of each level gradually increases, while the speed gradually slows down. The last level cache (LLC) is the last level cache of the CPU core, such as the third level cache. When the processing core needs to access data, if the data does not hit in the first level cache or the second level cache, it will continue to search the last level cache. If the data still does not hit in the last level cache, the processing core needs to read the required data from the main memory.

[0023] Caches can be divided into instruction caches and data caches based on the different information stored therein. Instruction caches are caches used to store instructions, and data caches are caches used to store data. The present embodiment of the application mainly uses data caches as an example, but the present embodiment of the application can also be used to support other caches that store data, and the present embodiment of the application is not limited to this.

[0024] Network-on-Chip (NoC): A network-based communication subsystem within an integrated circuit. It is typically used for data transmission and communication between different modules in a system-on-chip (SoC). NoC technology connects processor cores, memory, and various peripherals through routers, forming a highly parallel communication architecture that effectively improves data transmission efficiency and communication bandwidth. Compared to traditional shared bus approaches, NoC technology offers greater scalability and performance, making it particularly suitable for multi-core systems.

[0025] An instruction is a command that instructs a computer to perform a certain operation and is the smallest functional unit of computer operation. An instruction is a statement in machine language, or a set of meaningful binary codes. The collection of all instructions for a computer constitutes its instruction set, also known as its instruction set.

[0026] Instruction format: A human-readable representation of an instruction. An instruction typically consists of an opcode and operands. The opcode describes the type of operation to be performed, while the operands provide the data or addresses required to execute the instruction. While the opcode is essential, an instruction may contain no operands, one operand, or two operands.

[0027] Dead registers: Registers no longer used by a thread group. In other words, dead registers are registers that are no longer used during the lifetime of a thread group. Alternatively, dead registers are registers that are no longer used during the lifetime of a thread in a thread group. A wave is a SIMD (Single Instruction Multiple Data) thread bundle structure in the OpenCL (Open Computing Language) processing framework. In some processors, a wave is also called a wavefront. A thread group typically consists of 32 or 64 threads. In CUDA (Compute Unified Device Architecture), a thread group is called a warp, and a warp typically consists of 32 threads. Of course, the aforementioned terms wave, warp, etc. can also be referred to more generally as a thread group.

[0028] Figure 1 FIG. 1 is a schematic diagram of a processor provided by an exemplary embodiment of the present application. The processor 100 includes a plurality of registers 110 , a register management unit 120 , and a task assembly unit 130 .

[0029] The register management unit 120 is configured to divide the plurality of registers 110 into at least one configuration block set, where each configuration block set includes at least one configuration block, and at least one configuration block includes at least two registers from the plurality of registers 110 .

[0030] Optionally, the number of configuration blocks included in each configuration block set in at least one configuration block set is the same or different. For example, each configuration block set includes 5 configuration blocks, or each configuration block set includes 6 configuration blocks, or each configuration block set includes 8 configuration blocks, or each configuration block set includes 10 configuration blocks, and so on. For another example, at least one configuration block set includes 4 configuration block sets, and each configuration block set includes a different number of configuration blocks, such as configuration block set 1 includes 5 configuration blocks, configuration block set 2 includes 6 configuration blocks, configuration block set 3 includes 8 configuration blocks, and configuration block set 4 includes 10 configuration blocks. For another example, at least one configuration block set includes four groups of configuration block sets, the number of configuration blocks included in each group of configuration block sets is different, and the number of configuration blocks included in the configuration block sets included in each group of configuration block sets is the same, such as the first group of configuration block sets includes 16 configuration block sets, each of which includes 5 configuration blocks; the second group of configuration block sets includes 13 configuration block sets, each of which includes 6 configuration blocks; the third group of configuration block sets includes 10 configuration block sets, each of which includes 8 configuration blocks; and the fourth group of configuration block sets includes 8 configuration block sets, each of which includes 10 configuration blocks. It should be understood that when at least one configuration block set includes multiple groups of configuration block sets, the number of configuration block sets included in each group of configuration block sets can be different as shown above, or can be the same, and this embodiment of the present application is not limited to this.

[0031] Optionally, the sizes of the configuration blocks included in each configuration block set in at least one configuration block set are the same or different. Exemplarily, the sizes of the configuration blocks included in a configuration block set are the same, that is, the number of registers included in each configuration block included in a configuration block set is the same. However, the sizes of the configuration blocks included in different configuration block sets may be different. For example, if configuration block set 1 includes 5 configuration blocks, each configuration block in configuration block set 1 includes 32 registers; if configuration block set 2 includes 10 configuration blocks, each configuration block in configuration block set 2 includes 16 registers.

[0032] Optionally, the register management unit 120 evenly divides the multiple registers 110 to obtain at least one configuration block set, wherein each configuration block set in the at least one configuration block set has the same number and size of configuration blocks. That is, even division means dividing the multiple registers 110 into at least one configuration block set, wherein each configuration block set in the at least one configuration block set has the same number and size of configuration blocks. Alternatively, the register usage unit unevenly divides the multiple registers 110 to obtain at least one configuration block set, wherein two configuration block sets in the at least one configuration block set differ in at least one of the number and size of configuration blocks. That is, uneven division means dividing the multiple registers 110 into at least one configuration block set, wherein two configuration block sets in the at least one configuration block set differ in at least one of the number and size of configuration blocks. For example, two configuration block sets may have different numbers of configuration blocks, or two configuration block sets may have different sizes of configuration blocks, or two configuration block sets may have different numbers and sizes of configuration blocks.

[0033] Optionally, the register usage unit evenly divides the multiple registers 110 to obtain at least one configuration block, wherein each configuration block in the at least one configuration block has the same size; evenly divides the at least one configuration block to obtain at least one configuration block set, wherein each configuration block set in the at least one configuration block set includes the same number of configuration blocks. That is, evenly dividing the multiple registers 110 means dividing the multiple registers 110 into at least one configuration block, wherein each configuration block in at least one configuration class has the same size; evenly dividing the at least one configuration block means dividing the at least one configuration block into at least one configuration block set, wherein each configuration block set in the at least one configuration block set includes the same number of configuration blocks. Alternatively, the register usage unit evenly divides the multiple registers 110 to obtain at least one configuration block; unevenly divides the at least one configuration block to obtain at least one configuration block set, wherein two configuration block sets in the at least one configuration block set have different numbers of configuration blocks. That is, unevenly dividing the at least one configuration block means dividing the at least one configuration block into at least one configuration block set, wherein two configuration block sets in the at least one configuration block set have different numbers of configuration classes. Alternatively, the register usage unit may perform a non-uniform division of the plurality of registers 110 to obtain at least one configuration block, wherein two configuration blocks in the at least one configuration block have different sizes; and then perform a uniform division of the at least one configuration block to obtain at least one configuration block set. In other words, the non-uniform division of the plurality of registers 110 refers to dividing the plurality of registers 110 into at least one configuration block, wherein two configuration blocks in the at least one configuration block have different sizes. Alternatively, the register usage unit may perform a non-uniform division of the plurality of registers 110 to obtain at least one configuration block; and then perform a non-uniform division of the at least one configuration block to obtain at least one configuration block set.

[0034] Optionally, the specific division method for the multiple registers can be determined based on multiple factors such as the specific architecture of the processor 100, the type of task currently being executed, and the user's subjective settings, and the embodiments of the present application are not limited to this.

[0035] The task assembly unit 130 is configured to allocate a first configuration block set to a first thread group (wave), where the first configuration block set is one of at least two configuration block sets.

[0036] The task assembling unit 130 selects a first configuration block set from at least two configuration block sets for the first thread group.

[0037] Optionally, the task assembly unit 130 allocates a first configuration block set to the first thread group according to the tasks executed by the first thread group, wherein the tasks executed by the first thread group are used to determine at least one of the number of configuration blocks included in the configuration block set and the size of the configuration blocks.

[0038] Optionally, the task assembly unit may also be called a thread assembly and resource usage control unit, or a program distribution scheduler, etc.

[0039] It should be noted that the names of the various units shown in the embodiments of the present application are for illustration only. In specific implementation, other names may also be used to refer to the corresponding units. In addition, the unit division method in the embodiments of the present application is also for illustration only. In specific implementation, the number of units, functional boundaries, etc. may be adjusted. Specific adjustments may include merging, splitting, renaming, etc. For example, the resource usage management unit and the task assembly unit may also be adjusted to a task assembly and resource usage control unit. The embodiments of the present application do not limit this, but the scope of protection of the embodiments of the present application is not limited thereto.

[0040] In summary, the processor provided by the embodiment of the present application divides a plurality of registers into at least one configuration block set, and each configuration block set is used to allocate to a thread group. By dividing register resources into multiple configuration blocks, register resources are allocated in the form of configuration blocks. Each thread group supports the allocation of at least one configuration block (i.e., a configuration block set). Compared with the traditional method of continuous allocation based on registers as the granularity, this method of allocating register resources directly allocates configuration blocks instead of multiple continuous registers during allocation, which can avoid frequent cutting of continuous registers and cause register fragmentation. In addition, allocation in the form of configuration blocks can achieve management through configuration blocks instead of management based on registers, which can improve the efficiency of allocation and release of register resources.

[0041] In some embodiments, the plurality of registers 110 are a plurality of scalar registers, or the plurality of registers 110 are a plurality of vector registers, or the plurality of registers 110 are a plurality of scalar registers and a plurality of vector registers, etc. It should be noted that the execution logic of the processor 100 shown in the embodiments of the present application is applicable to different types of registers. In addition to the scalar registers and vector registers shown above, other types of registers may also be supported, and the embodiments of the present application are not limited thereto.

[0042] Next, the division method of registers will be described in detail.

[0043] 1. Division of registers.

[0044] In some embodiments, the register management unit 120 is used to determine quantity information based on register usage configuration information, where the quantity information includes at least one of the number of configuration blocks, the number of active thread groups, a first data amount, and a second data amount; based on the quantity information, the multiple registers 110 are divided into at least one configuration block set; wherein the number of configuration blocks is used to indicate the number of configuration blocks in each configuration block set in at least one configuration block set, the number of active thread groups is used to indicate the maximum number of thread groups supported for parallel execution by the processor 100, the first data amount is used to indicate the maximum amount of data supported for storage by each configuration block set in at least one configuration block set, and the second data amount is used to indicate the maximum amount of data supported for storage by each configuration block set.

[0045] Optionally, the register usage configuration information is set by a user or developer. The register usage configuration information can be modified through an interface, method, or parameter opened by the designer of the processor 100, or modified through software designed by the designer.

[0046] Optionally, based on the register usage configuration information, quantity information corresponding to the plurality of registers 110 is determined; or based on the register usage configuration information, quantity information corresponding to some of the plurality of registers 110 is determined. In other words, the register usage configuration information can be used to perform configuration division for some of the plurality of registers 110.

[0047] Optionally, the register usage configuration information is used to indicate or identify at least one of the number of configuration blocks, the number of active thread groups, the first number, and the second data amount.

[0048] Optionally, the plurality of registers 110 may be divided into at least one configuration block set, and each set in the at least one configuration block set supports configuration of the configuration block by using a register usage configuration information.

[0049] Exemplarily, the processor 100 includes at least one sub-processor, and each processor 100 in the at least one sub-processor corresponds to a configuration block set. Each sub-processor can divide registers within the at least one configuration block set corresponding to the sub-processor into at least one configuration block set based on register usage configuration information.

[0050] For example, if the processor 100 is a graphics processing unit (GPU), the sub-processor may be a streaming multiprocessor (SM); or, if the processor 100 is a GPU, the sub-processor may be a stream processor (SP); or, if the processor 100 is a streaming multiprocessor 100, the sub-processor may be a stream processor 100.

[0051] Optionally, the multiple registers 110 included in the above-mentioned processor 100 refer to registers used to be allocated to thread groups to implement instruction execution, or the above-mentioned multiple registers 110 may be thread-private registers, that is, in addition to the multiple registers 110 shown above, the processor 100 may also include other types of registers. These different types of registers have different uses, such as for storing information shared by multiple thread groups, for storing instructions corresponding to thread groups, and so on.

[0052] Exemplarily, the register management unit 120 is configured to query the register configuration information table based on the register usage configuration information to determine the quantity information, where the register configuration information table is used to indicate a mapping relationship between the register usage configuration information and the quantity information.

[0053] Optionally, the register configuration information table may be pre-set or may be set by a user.

[0054] For example, a register configuration information table is shown in Table 1 below. The quantity information includes the number of configuration blocks, the number of active thread groups, and the first data quantity. Register usage configuration information can also be referred to as quantity information identification, or enumeration type information. An enumeration type is a data type that defines a set of named constants (i.e., a finite, predefined set of options). Enumeration constants typically have fixed, well-defined meanings, and their values ​​are typically integers (but can also be other types, such as strings).

[0055] Table 1 Register configuration information table

[0056] Optionally, the number of registers in the processor 100 is fixed, or the number of registers in a sub-processor is fixed.

[0057] Optionally, all configuration blocks in the processor 100 support allocation to thread groups; or, some configuration blocks in the processor 100 will be allocated to thread groups. Typically, each thread group executed in the processor 100 will evenly divide the configuration blocks of the processor 100, or each thread group executed in a sub-processor will evenly divide the configuration blocks of the sub-processor. That is, each thread group belonging to the same processor 100 or sub-processor corresponds to the same number of configuration blocks (or registers). Therefore, for some numbers of configuration blocks, there may be a situation where the last remaining configuration blocks cannot be allocated to a thread group. For example, in the cases of configuration block numbers 6 and 13 in Table 1 above, there are 2 configuration blocks in the processor 100 that are idle and will not be allocated to a thread group.

[0058] Optionally, different register configuration information tables can be set for different second data volumes. For example, Table 1 above is a register configuration information table for a sub-processor. If the number of registers included in the sub-processor is 2560, and each register supports storing 1DW of data, then it can be seen that in the scenario shown in Table 1 above, the second data volume is 32DW, that is, each configuration block includes 32 registers. At this time, a sub-processor with a second data volume of 16DW may exist in the processor 100. This sub-processor corresponds to another register configuration information table. It should be understood that the second data volume can also be used as a parameter in the register configuration information table, as shown in Table 2 below for example.

[0059] Table 2 Register configuration information table

[0060] Optionally, the processor 100 may use one or more second data amounts when dividing the configuration blocks, that is, to obtain configuration blocks of the same size or configuration blocks of different sizes.

[0061] Optionally, the setting of register usage configuration information is usually related to the tasks performed by the processor 100. For example, for some tasks with large register usage, such as AI computing tasks, model training tasks, and rendering and drawing tasks of complex scenes, more registers can be allocated to a thread group, that is, register usage configuration information with a larger number of configuration blocks is used. However, since the total number of registers in the processor 100 is unchanged, when more registers are allocated to a thread group, the number of executable thread groups in the processor 100 will also decrease accordingly, that is, the number of active thread groups will decrease accordingly. For some tasks with less register usage, such as rendering and drawing of simple graphics, simple mathematical function calculation tasks, etc., fewer registers can be allocated to a thread group, that is, register usage configuration information with a smaller number of configuration blocks is used, and the number of active thread groups in the corresponding processor 100 will also increase.

[0062] Optionally, when the processor 100 includes multiple sub-processors, the register usage configuration information can be set separately for the multiple sub-processors, that is, each sub-processor can determine the division method of the registers in the configuration block set corresponding to the sub-processor through a register usage configuration information.

[0063] Optionally, when the processor 100 includes multiple sub-processors, the number of registers included in each of the multiple sub-processors is the same or different.

[0064] For example, taking the case where multiple sub-processors include the same number of registers, for the first sub-processor, it corresponds to the first register usage configuration information, and the quantity information determined by the first register usage configuration information is: {number of configuration blocks: 5, number of active thread groups: 16, first data volume: 160, second data volume: 32}; for the second sub-processor, it corresponds to the second register usage configuration information, and the quantity information determined by the second register usage configuration information is: {number of configuration blocks: 10, number of active thread groups: 8, first data volume: 320, second data volume: 32}; for the third sub-processor, it corresponds to the third register usage configuration information, and the quantity information determined by the third register usage configuration information is: {number of configuration blocks: 16, number of active thread groups: 5, first data volume: 512, second data volume: 32}.

[0065] To summarize, the processor provided in the embodiment of the present application shows that the division method of multiple registers in the processor is determined based on the register usage configuration information, that is, it supports flexible division of registers according to specific needs by setting the register usage configuration information, so that the configuration blocks and configuration block sets obtained by division can meet the current task requirements and improve the utilization of registers.

[0066] Furthermore, register usage configuration information is used to indicate corresponding quantities in the register configuration information table, avoiding the fragmentation or performance issues that can arise from fully opening up user-defined configuration blocks (e.g., improperly set quantities). By predefining reasonable quantity options, users can select the closest option based on their needs while avoiding the uncontrollable nature of fully free configuration.

[0067] Furthermore, the processor illustrated above can employ different register usage configuration information for different types of tasks. For example, in scenarios where the usage of some first-type data (constants or common data, etc.) is relatively large, a larger partitioning granularity can be used for multiple registers, such as including more configuration blocks and, in turn, more registers in a configuration block set. However, the number of thread groups supporting parallel execution in the processor (i.e., the number of active thread groups) will also be reduced. However, this flexibly adjustable configuration method enables the processor's registers to have a high occupancy rate in a variety of scenarios.

[0068] Next, we will explain in detail how the thread group uses the configuration block set.

[0069] 2. Loading of registers.

[0070] In some embodiments, the processor 100 further includes a data loading unit 140, such as Figure 2 As shown; a data loading unit 140 is used to load at least one first data for a first part of registers in a first configuration block set, where the first part of registers is a specified number of registers or a specified proportion of registers in the first configuration block set, and the first data belongs to a first type, and the first type of data includes at least one of a constant and a common data.

[0071] Optionally, the first part of registers includes at least one register, and there is a one-to-one relationship between the at least one register and the at least one first data.

[0072] Optionally, the first portion of registers comprises a specified number or a specified ratio of registers in the first configuration block set. This specified number or ratio can be pre-set or user-configured. Pre-setting refers to a method in which the processor designer determines and fixes this in the hardware logic during processor design, taking effect after power-up or reset. User-configured means that the user configures this through an interface or method provided by the designer. For example, the user may design a specified number or ratio based on the task being performed. For example, graphics rendering tasks typically require some first-type data, such as model transformation matrices (world matrix, view matrix, projection matrix, etc.), lighting parameters (light source position, color, intensity), and material properties (reflectivity, transparency). In addition to this first-type data, other types of data are also required, such as second-type data, which is data that is updated during calculations, such as vertex data and texture data. Therefore, the user needs to determine the number of registers in the first portion of processor 100 based on the specific task being performed. This register allocation method effectively utilizes register resources, improves access efficiency for different types of data, and optimizes task performance.

[0073] Optionally, the specified number and the specified ratio may also be determined by the processor 100 according to the task currently to be executed.

[0074] Optionally, the first type of data includes at least one of constants and public data. Constants are data that cannot be modified during program execution, or data that cannot be modified during the execution of a thread group, or data that cannot be modified during the execution of a task. Public data refers to data shared by a thread group, i.e., data shared by all or some threads in a thread group. If all threads require the same common storage space (e.g., a memory area or buffer), the starting address of that common storage space can be used as common data, i.e., first type of data.

[0075] Optionally, in addition to storing the first type of data, the register may also be used to store a second type of data, such as variables. Generally speaking, the second type of data is private data of threads in a thread group.

[0076] For example, the first configuration block set includes 160 registers, each corresponding to a data size of 1 DW (32 bits). In this case, the specified number or ratio of registers in the first portion can be determined based on the actual amount of data of the first type and the amount of data of the second type required by the first thread group. For example, if the first portion of registers consists of 128 registers, the maximum amount of data allocated to the first type is 128 DW; and if the second portion of registers consists of the remaining 32 registers, the maximum amount of data allocated to the second type is 32 DW.

[0077] Among them, the processor shown above loads the first type of data (i.e., at least one first data) into the first part of the registers of the configuration block set. Compared with the method of setting dedicated shared registers or constant caches for constants in the related art, it ensures the universality of the registers, avoids the waste of register resources caused by some tasks not using dedicated registers or caches, and improves the overall utilization of registers. Moreover, without setting up dedicated registers or constant caches, the area of ​​the processor can be reduced as much as possible. In addition, the universal register supports the division of a part of the registers (such as the first part of the registers) for storing the first type of data, and can also achieve the advantage of dedicated registers to protect key data (such as constants and public data).

[0078] In the related art, the thread group scheduling process is usually executed first, and the data required for the instruction is loaded before or when a specific instruction is executed. However, in the embodiment of the present application, a large block of first type data is usually pre-loaded, and the data type of the first type pre-loaded in the embodiment of the present application is different from the data type loaded in the related art. Therefore, the traditional thread group scheduling process and data initialization loading process should be updated to adapt to the register allocation and usage method shown in the embodiment of the present application. The specific loading timing of loading at least one first data into the first part of the registers is as follows.

[0079] (1) Loading timing.

[0080] In some embodiments, the above-mentioned process of loading the first data into the first part of the registers can be called an initialization process. In order to improve the execution efficiency of the thread group in the processor 100, the above-mentioned initialization process can be executed in parallel with the scheduling process of the thread group. Compared with the related art of first completing the scheduling process of the thread group and then executing the initialization process (also called the loading process) for the registers of the thread group, by executing the scheduling process and the initialization process in parallel, it helps to shorten the execution time of the thread group, improve the execution efficiency of the thread group, and thus reduce the overall delay of task execution. Exemplarily, the above-mentioned processor 100 also includes a scheduler; the scheduler is used to send a scheduling request, the scheduling request is used to trigger the scheduling process of the first thread group, and the scheduling process refers to the process of scheduling execution of the first thread group; the task assembly unit 130 is used to send an initialization load request, the initialization load request is used to trigger the initialization process of the first configuration block set, and the initialization process refers to initializing the first part of the registers in the first configuration block set; wherein, the scheduling process and the initialization process are executed in parallel.

[0081] Optionally, the scheduling request sent by the scheduler is used to trigger the scheduling process of the first thread group. For example, the scheduler sends the scheduling request to the thread group management unit. The thread group management unit selects an instruction based on the current execution status of the thread group and the dependencies between instructions and transmits it to an execution unit (such as a processing core in processor 100). It also removes data dependencies and resource dependencies for the instruction to ensure correct execution of the instruction. During the execution of the instruction, the execution unit or other unit initiates a data load request to read data from the storage unit and load it into the second portion of registers of the processor 100. The second portion of registers is the registers in the first configuration block set other than the first portion of registers.

[0082] Optionally, before the initialization process is complete, the thread group management unit may prioritize instructions that do not use data of the first type, or instructions whose usage of the first type of data is below a threshold. That is, the thread group management unit typically prioritizes instructions that use only data of the second type during the initial thread group scheduling phase to avoid wasting register resources due to repeated loading of some data of the first type.

[0083] Optionally, the initialization load request sent by the task assembly unit 130 is used to trigger the initialization process of the first configuration block set. The initialization process is the process in which the data loading unit 140 loads at least one first data into the first part of registers in the first configuration block set, which will not be repeated here.

[0084] Optionally, before and during the initialization process performed by the data loading unit 140, a data dependency check is not performed on at least one first data. That is, before or during the loading of the first data, there is no need to perform a data dependency check on the first data based on the instructions included in the first thread group, nor is there a need to confirm which instructions executed by the first thread group need to use the first data. Since the first data is data of the first type, it is usually at least one of a constant and public data. Constants are data that do not support modification during the execution of the thread group, so constants usually do not need to check the above-mentioned data dependencies. As for public data, it is usually some information such as a public starting address, and this type of information does not need to be checked for data dependencies before execution.

[0085] Optionally, before executing a first instruction, the processor 100 adds a data dependency check instruction. The first instruction is an instruction whose operand contains data of the first type. The data dependency check instruction is used to detect whether the first type of data required by the first instruction has been loaded into the first configuration block set. That is, before executing the first instruction, the added data dependency check instruction is executed to ensure correct execution of the first instruction.

[0086] Among them, in the processor shown above, the scheduling process of the thread group and the initialization process of the first part of the registers are executed in parallel. This is because the data loaded by the first part of the registers is the first type of data, rather than all the registers required by the thread group. Therefore, it supports the scheduling process of the thread group and the initialization process of the first part of the registers to be executed in parallel, thereby hiding the delay of loading part of the data corresponding to the thread group and improving the efficiency of the thread group execution.

[0087] In some embodiments, the actual amount of data required by the thread group is large. For example, the data size of the configuration block set is 160DW, but the actual amount of data required by the thread group is greater than the data size of the configuration block set or the data size of the first part of registers. In this case, dynamic loading can be used to solve this problem, as shown below.

[0088] (2) Dynamic loading.

[0089] In some embodiments, the data loading unit 140 is used to load at least one second data into at least one register of the first configuration block set during the execution of the first thread group, the second data is of the first type, and there is a one-to-one relationship between the at least one register and the at least one second data.

[0090] Optionally, during execution of the first thread group, the data loading unit 140 determines data of the first type required by at least one instruction, ie, at least one second data, and loads the at least one second data into at least one register of the first configuration block set.

[0091] Optionally, in addition to the first type of data, there is also a second type of data to be loaded, but the second type of data is usually loaded into the second part of registers. The second part of registers refers to the registers in the first configuration block set except the first part of registers.

[0092] Optionally, the at least one register in the first configuration block set may include at least one of a register in the first portion of registers and a register in the second portion of registers. That is, the at least one second data may be loaded into the first portion of registers, may be loaded into the second portion of registers, or may be partially loaded into the first portion of registers and partially loaded into the second portion of registers.

[0093] Exemplarily, the data loading unit 140 is used to load at least one second data into at least one register of the first part of registers during the execution of the first thread group, the data in the at least one register is data that the first thread group has finished using, and the at least one second data is used to overwrite the first data in the at least one register.

[0094] Optionally, the data in the at least one register is data that the first thread group has finished using, or in other words, the first thread group has finished using the first type or second type of data in the at least one register. In this case, the at least one second data is used to overwrite the data in the at least one register.

[0095] Optionally, the data in the at least one register may be the first data loaded during the initialization process, or may be the second data loaded during the last dynamic loading process.

[0096] Exemplarily, the data loading unit 140 is used to load at least one second data into at least one register of the second part of registers during the execution of the first thread group, where the second part of registers is the registers in the first configuration block set except the first part of registers; wherein the registers in the at least one register meet at least one of the following characteristics: the data in the register is data that has been finished being used by the first thread group; the register is an empty register.

[0097] Optionally, the data in the register is data that the first thread group has finished using, that is, the first thread group has finished using the first type or second type of data in the register. In this case, the second data is loaded into the register to overwrite the data originally stored in the register.

[0098] Optionally, the register is an empty register. An empty register is a register that does not store valid data, i.e., data stored due to the execution of the first thread group; or an empty register is an uninitialized register. In this case, the second data is directly loaded into the register. Optionally, the empty register may store data loaded by a previous thread group, or may be reset to all zeros due to a clear operation.

[0099] In other embodiments, the data loading unit 140 is configured to load at least one second data into at least one register in the first portion of registers and the second portion of registers during execution of the first thread group. For example, the first configuration block set includes 160 registers, the first portion of registers includes 128 registers, the second portion of registers includes 32 registers, and the at least one second data is 32 second data. The data loading unit 140 determines the register for loading the second data based on the usage of each register in the first configuration block set. For example, if data in 24 registers in the first portion of registers has ended, and data in 16 registers in the second portion of registers has ended, the 24 registers in the first portion of registers are preferentially selected. Since the 24 registers in the first portion of registers are less than the number of second data, 8 registers in the second portion of registers are further selected for loading the second data.

[0100] In some embodiments, the above-mentioned dynamic loading process is triggered by a dynamic loading instruction, and the dynamic loading instruction is used to instruct that the first type of data be loaded into the first configuration block set during the execution of the first thread group. Optionally, the dynamic loading instruction is used to load the first type of data required by the first instruction into the first configuration block set before the first instruction is executed, and the first instruction is an instruction corresponding to the first thread group; or, the dynamic loading instruction is used to load the first type of data required by at least one instruction into the first configuration block set before the execution of at least one instruction, and the at least one instruction is an instruction corresponding to the first thread group. Optionally, the at least one instruction may be one or more instructions to be executed by the first thread group, and the number of the at least one instruction may be pre-set or configured by the user.

[0101] Among them, the processor shown above supports dynamic loading of the first type of data during the execution of the first thread group, avoiding the problem of too few thread groups that can be executed in parallel in the processor due to the excessive amount of the first type of data required by the first thread group, ensuring that the processor can have a certain degree of parallelism, especially for the GPU, avoiding too many computing cores in the processor being idle, resulting in the loss of the processor's advantages.

[0102] In addition, during dynamic loading, it supports loading the first type of data into the first part of registers, or loading the first type of data into the second part of data, or loading the first type of data into the first part of registers and the second part of registers. That is, it supports dynamic loading during the initialization process to load the first type of data into the second part of registers, thereby ensuring the parallel execution of the scheduling process and the initialization process of the thread group, and improving the execution efficiency of the thread group. On the other hand, it also supports loading the first type of data into the first part of registers according to the execution status of the thread group, overwriting the data in the first part of registers that the thread group has finished using, reusing limited registers, alleviating the lack of resources caused by the hardware scale setting of the processor, and indirectly reducing the demand for register resources in the processor, thereby reducing the area of ​​the processor.

[0103] Next, the above initialization process and dynamic loading process are specifically described by taking the example of a first configuration block set including n first registers (registers for storing first type data) and m second registers (registers for storing first type and second type data).

[0104] In some embodiments, when the number of first-type data required by the first thread group is less than or equal to the number of first registers, loading of the first-type data by the first thread group can be completed only through initial loading. That is, the data loading unit 140 is configured to load a first number of first data into a first number of first registers based on the initial load request when the first number is less than or equal to n, where the first number is the number of first-type data corresponding to the first thread group.

[0105] Optionally, if the amount of first type data required by the first thread group is less than or equal to the number of first registers, the remaining first registers may be left vacant or used to load second type data. The remaining first registers refer to registers among the n first registers that were not loaded with first data during the initial process.

[0106] In some embodiments, when the number of first type data required by the first thread group is greater than the number of first registers, initial loading and dynamic loading should be used to implement the execution of the first thread group. Specifically, the initial loading process is as follows. The first configuration block set includes n first registers, and the n first registers are used to store first type data, where n is a specified number or determined by a specified ratio; the data loading unit 140 is used to load the n first data into the n first registers based on the initialization load request when the first number is greater than n, where the first number is the number of first type data corresponding to the first thread group.

[0107] Optionally, the n first data are part of a first number of data. The first number of data refers to data of a first type corresponding to the first thread group.

[0108] Optionally, the initialization process can be executed in parallel with the scheduling process of the thread group.

[0109] Optionally, the n first registers are the above-mentioned first part of registers, that is, the first part of registers includes n registers.

[0110] Optionally, the n first data are loaded into the n first registers, that is, each of the n first data is loaded into one of the n first registers, or in other words, there is a one-to-one relationship between the n first data and the n first registers.

[0111] In some embodiments, before the initialization process is completed, the thread group management unit typically schedules execution of instructions in the first thread group that do not require the use of the first type of data. Alternatively, before the initialization process is completed, the thread group management unit may load the first type of data into the second register to enable execution of the first thread group.

[0112] In some embodiments, after the initialization process is completed, the remaining data in the first number of data corresponding to the first thread group can be loaded into the register using a dynamic loading method. Exemplarily, the processor 100 further includes an instruction control unit; the instruction control unit is configured to execute a dynamic loading instruction when the first number is greater than n, the dynamic loading instruction being configured to instruct that the first type of data be loaded into the first configuration block set during the execution of the first thread group; and a data loading unit 140 is configured to load k second data into k registers based on the dynamic loading instruction, where k is a positive integer and the k registers are free registers in the first configuration block set, where the free registers are empty registers or registers whose data have finished being used.

[0113] Optionally, k is less than or equal to the total number of registers in the first configuration block set. Optionally, k is less than or equal to the total number of free registers in the first configuration block set.

[0114] Optionally, k is related to an instruction to be executed by the first thread group. The k second data are data to be used by the first thread group.

[0115] Optionally, the dynamic load instruction is used to instruct the loading of a first type of data into the first configuration block set during the execution of the first thread group; or, the dynamic load instruction is used to instruct the loading of a first type of data required by at least one instruction into the first configuration block set, where the at least one instruction is an instruction corresponding to and not yet executed by the first thread group. Optionally, the at least one instruction is a second number of instructions to be executed by the first thread group, where the second number may be pre-set or user-configured.

[0116] Optionally, the dynamic load instruction may also be used to instruct that data required by at least one instruction be loaded into the first configuration block set, where the data required by the at least one instruction includes at least one of a first type of data and a second type of data.

[0117] Optionally, the k registers are k first registers, or the k registers are k second registers, or the k registers include i first registers and j second registers, where i and j are both positive integers, and i+j=k.

[0118] Exemplarily, the data loading unit 140 is configured to load k second data into k first registers based on a dynamic load instruction, where the data in the k first registers is data that has been finished using by the first thread group, and the k second data is used to overwrite the data in the k first registers. Alternatively, based on a dynamic load instruction, the k second data are loaded into k second registers, where each of the k second registers satisfies at least one of the following characteristics: the data in the register is data that has been finished using by the first thread group; the register is an empty register. Alternatively, based on a dynamic load instruction, the k second data are loaded into i first registers and j second registers, where i+j=k, where the screening conditions for the i first registers and the j second registers are as described above for the k first registers and the k second registers, and are not further described here.

[0119] Among them, the processor shown above, when the first number is less than or equal to the number of the first registers, only uses the initial loading method to load the first type of data into the first register, thereby realizing the loading of the first type of data required by the first thread group. When the first number is greater than the number of the first registers, the first type of data is loaded into at least one of the first register and the second register using the initial loading and dynamic loading methods, thereby avoiding the deadlock problem caused by loading too much first type of data at one time, which causes the thread group to be unable to load other types of data. It also supports loading the first type of data into the first part of the registers according to the execution status of the thread group, overwriting the data in the first part of the registers that the thread group has finished using, reusing limited registers, alleviating the lack of resources caused by the hardware scale setting of the processor, and indirectly reducing the demand for register resources in the processor, thereby reducing the area of ​​the processor.

[0120] In addition to the above-mentioned configuration method for dividing the registers of each thread group and loading the first type of data into the registers by combining initial loading and dynamic loading, in order to further reduce the area of ​​the processor 100, a method is also designed for multiple sub-processors to share a set of write-back paths, as shown below.

[0121] 3. Shared return data path.

[0122] In some embodiments, the processor 100 further includes at least one data loading unit 140, a first cache, and multiple sub-processors; each data loading unit 140 in the at least one data loading unit 140 corresponds to m sub-processors, the m sub-processors are all or part of the multiple sub-processors included in the processor 100, and each of the m sub-processors includes at least one configuration block set; the data loading unit 140 includes a data movement executor and a storage entry buffer; the data movement executor is used to read s third data from the first cache using a first bandwidth; and cache the s third data to a first first-in-first-out (FIFO) queue of the storage entry buffer, the first bandwidth is related to the cache line width of the first cache; the storage entry buffer is used to write the s third data from the first FIFO queue to at least two registers corresponding to the first sub-processor in the m sub-processors using a second bandwidth, and the second bandwidth is related to the size of the configuration block; wherein the first bandwidth is greater than the second bandwidth.

[0123] Optionally, the data loading unit 140 is configured to load the first type of data into the configuration block set. In other words, the data loading unit 140 is configured to load the first type of data into registers in the configuration block set.

[0124] Optionally, the data movement executor reads s third data from the first cache using a first bandwidth, where the first bandwidth is related to a cache line width of the first cache.

[0125] Optionally, each cache line in the first cache is used to cache s third data, or multiple cache lines in the first cache are used to cache s third data.

[0126] Optionally, the data mover uses the first bandwidth to read s third data from the first cache at one time or in one cycle. Alternatively, the data mover uses the first bandwidth to read s third data from the first cache in multiple cycles.

[0127] Optionally, the third data belongs to the first type; or, the third data belongs to the second type; or, the s third data include data belonging to the first type and data belonging to the second type.

[0128] Optionally, the FIFO queue is a data structure that processes data in chronological order. For example, if a cache line contains 256 bits and the cache line is stored in the FIFO queue in order from bit 0 to bit 255, then when writing data to a register through the FIFO queue, the data is also written starting from bit 0 and ending with bit 255.

[0129] Optionally, the storage entry buffer uses the second bandwidth to write s third data from the first FIFO queue to the first sub-processor among the m sub-processors. If the second bandwidth corresponds to t third data, the storage entry buffer uses the second bandwidth to read t third data from the first FIFO queue and write the t third data into t registers; then, the storage entry buffer uses the second bandwidth again to read t third data from the first FIFO queue and write the t third data into t registers.

[0130] For example, Figure 3 As shown, assuming the first cache is a level 1 cache, the data mover executor uses a first bandwidth to read data 30 corresponding to the first bandwidth from the first cache and caches it in a first FIFO queue 40 of the storage entry buffer. The storage entry buffer then uses a second bandwidth to read data 31 corresponding to the second bandwidth from the first FIFO queue and load it into the register corresponding to the first sub-processor. If the first bandwidth is 16 times the second bandwidth, that is, if the first bandwidth is 128 DW / cycle and the second bandwidth is 8 DW / cycle, it takes 16 cycles to load all the data in the first FIFO queue into the first sub-processor. If the storage entry buffer corresponds to only one sub-processor, the data mover executor must idle for at least 15 cycles while waiting for the storage entry buffer to load the data 30 corresponding to the first bandwidth into the register of the first sub-processor. This wastes the performance of the data mover executor. Moreover, if there are multiple sub-processors in the processor 100, a corresponding return data path (i.e., data mover executor-storage entry buffer, or data loading unit) must be set up for each sub-processor, which significantly increases the area of ​​the processor 100. Therefore, multiple FIFO queues can be set up in the storage entry buffer and connected to multiple sub-processors at the same time. When the first FIFO queue is writing data to the register of the first sub-processor, the data movement executor can continue to read the data required by the second sub-processor from the first cache and load it into the second FIFO queue 41. This can reduce the waiting period of the data movement executor and improve the utilization rate of the data movement executor. Similarly, a larger number of sub-processors can be set up for the storage entry buffer to ensure that the data movement executor is in a non-idle state as much as possible, such as setting more than 16 FIFO queues, that is, connecting more than 16 sub-processors.

[0131] Optionally, the storage entry buffer includes m FIFO queues, and each FIFO queue corresponds to a sub-processor.

[0132] Optionally, the second bandwidth is related to the size of the configuration block. For example, if the configuration block is arranged in i rows and j columns, the second bandwidth may be the data bandwidth corresponding to column j. In other words, the second bandwidth supports writing an entire row (i.e., j items) of data at once. Alternatively, the second bandwidth may be related to the size of the third data. For example, the second bandwidth may be a multiple of the number of bits corresponding to the third data.

[0133] Optionally, one data loading unit 140 corresponds to all or part of the sub-processors in the processor 100. The number of sub-processors corresponding to one data loading unit 140 can be flexibly set.

[0134] For example, if the processor 100 is a graphics processing unit (GPU), the sub-processor may be a streaming multiprocessor (SM); or, if the processor 100 is a GPU, the sub-processor may be a stream processor (SP); or, if the processor 100 is a streaming multiprocessor 100, the sub-processor may be a stream processor 100.

[0135] In summary, the processor provided in the embodiments of the present application illustrates a processor in which multiple sub-processors share a single return data path. This approach ensures the overall processor register write bandwidth while minimizing area overhead. This achieves the effect of a large bandwidth entering the data loading unit and multiple smaller bandwidths exiting the data loading unit, enabling fast writes without congesting the return data path pipeline.

[0136] Since the embodiment of the present application adopts the form of configuration blocks to configure register resources, corresponding changes should be made when the processor 100 accesses the registers, as shown below.

[0137] 4. Address access process.

[0138] In some embodiments, the processor 100 is used to determine, in response to an access request for a register to be accessed, a physical address of the register to be accessed based on an identifier of a thread group corresponding to the register to be accessed and a logical address of the register to be accessed, where the physical address is used to indicate that the register to be accessed is located in the bth row and cth column of the ath configuration block, where a, b, and c are all positive integers.

[0139] The registers in each configuration block of the at least one configuration block are arranged in i rows and j columns, where i and j are both positive integers. b is less than or equal to i, and c is less than or equal to j.

[0140] Optionally, the logical address of the register to be accessed refers to the relative address of the register to be accessed in the configuration block set corresponding to the thread group. That is, registers with the same logical address may exist in the configuration block sets corresponding to different thread groups, while the configuration block set corresponding to the same thread group generally does not have the same logical address. In other words, the logical address of the register to be accessed is used to identify the register in the configuration block set corresponding to the thread group.

[0141] Optionally, the physical address of the register to be accessed refers to the actual physical location of the register to be accessed in all registers of the processor 100 or the sub-processor, and is usually represented by a combination of a row address and a column address. Of course, in some special scenarios, a complete physical address can also be used to represent it. The complete physical address is usually related to the row address and the column address. For example, the complete physical address is obtained by splicing the row address and the column address. For example, if the row address is 0x0001 and the column address is 0x0101, the complete physical address can be expressed as 0x0001 0101. Of course, other splicing or mapping methods can also be used, and the embodiments of the present application are not limited to this.

[0142] Optionally, the access request for the register to be accessed may be a read request for the register to be accessed, a write request for the register to be accessed, or other operation request for the register to be accessed, which is not limited in the embodiment of the present application.

[0143] Optionally, the processor 100 determines the physical address of the register to be accessed based on the identifier of the thread group corresponding to the register to be accessed and the logical address of the register to be accessed. The identifier of the thread group is used to determine the starting position of the configuration block set. The logical address of the register to be accessed can be used to determine the position of the register to be accessed in the configuration block set, which can be broken down into the starting position of the configuration block in the configuration block set, the row position and the column position of the register in the multiple registers 110, wherein the starting position of the configuration block in the configuration block set can be said to be the position of the configuration block where the register to be accessed is located in the configuration block set.

[0144] Exemplarily, the processor 100 is configured to, in response to an access request for a register to be accessed, determine an identifier of a configuration block corresponding to the register to be accessed based on an identifier of a thread group corresponding to the register to be accessed and a logical address of the register to be accessed; determine a row start address corresponding to the configuration block based on the identifier of the configuration block; determine a row offset address and a column address corresponding to the register to be accessed based on the logical address of the register to be accessed; determine a row address of the register to be accessed based on the row start address and the row offset address; and determine a physical address of the register to be accessed based on the row address and the column address.

[0145] Exemplarily, the logical address of a register is used to indicate the relative address of the register in the configuration block set corresponding to the thread group. For example, if the logical address of the register is 0x0000 0000, it means that the register is R0 in the configuration block set corresponding to the thread group, that is, the first register; if the logical address of the register is 0x0010 0110, it means that the register is R70 in the configuration block set corresponding to the thread group, that is, the 71st register; if the logical address of the register is 0x0011 0100, it means that the register is R52 in the configuration block set corresponding to the thread group, that is, the 53rd register. Therefore, the meaning of each bit in the logical address of the register is related to factors such as the size of the configuration block set corresponding to the thread group and the size of the configuration block. For example, the logical address of the register represents the column index, row index, and configuration block index from low to high, and there is no invalid bit, such as Figure 4 As shown, the configuration blocks are arranged in 4 rows and 8 columns, with the lower logical address 10. (= =3) bit indicates column index 14, j is the number of columns in the configuration block; +1(= +1=4) to the + (= + =3+2=5) bits represent the row index 13, i is the number of rows in the configuration block; the remaining high bits represent the configuration block index 12. Based on the above configuration block index 12, row index 13 and column index 14, the physical position 11 of the register in the configuration block set can be determined. Since the registers are arranged as follows Figure 4 As shown, no matter how many configuration blocks there are, the column address 17 of each register is always its column index; and the row address of each register needs to determine the starting address of the configuration block set, that is, the starting address of the configuration block with index 0 in the configuration block set. Based on the starting address, the row starting address 15 of the configuration block corresponding to the logical address can be known, and finally the row address 16 is obtained based on the row starting address 15 and the row index 13. The row address 16 and the column address 17 together constitute the physical address 18.

[0146] It should be noted that the above-mentioned conversion method between logical address and physical location (i.e., the method of calculating by logarithm) is for illustration only. Since logarithm operation has a high cost, other methods with lower cost can be used in actual implementation, such as table lookup method, shift operation method, etc. For example, Figure 4 As shown in code block 19 in , a table lookup method is used to determine the configuration block index and a shift operation method is used to determine the row address and column address.

[0147] Exemplarily, taking the arrangement shown above as an example, the conversion process of logical address 10 and physical address 18 can refer to code block 19. Among them, ">>" is a right shift symbol. If the logical address is 8 bits, then shifting right by 5 bits means taking the upper 3 bits of the logical address. It can also be understood as dividing the logical address by 32 (that is, the configuration block size) to determine the configuration block index corresponding to the logical address, that is, "waveid+(logic_scalar_addr>>5) to locate chunk_idx" means determining the configuration block index based on the thread group identifier (waveid) and the logical address (logic_scalar_addr) to determine the configuration block identifier (chunkidx). If the configuration blocks are arranged in other forms, such as the above i rows and j columns, the number of bits that the above logical address determines the configuration block index should be shifted right will change, and the specific number of bits that should be shifted right will change. Bit. It should be understood that the above-mentioned configuration block index refers to the index of the configuration block corresponding to the logical address in the configuration block set, that is, if the number of configuration blocks is 5, the value range of the configuration block index is 0-4, which needs to be represented by 3 bits; and the configuration block identifier is used to indicate the identifier of the configuration block in all configuration blocks of the processor or sub-processor. For example, if the number of configuration blocks corresponding to the sub-processor is 5 and the number of active thread groups is 16, the value range of the identifier of the configuration block is 0-79. The identifier of the configuration block can be understood as the starting address of the row, or the starting address of the row with the configuration block index of 0 in the configuration block. After determining the row start address, the row address (line_addr) in the physical address can be determined based on the row start address and the 4th bit (b3) and the 5th bit (b4) of the logical address. The corresponding calculation formula is "line_addr=(chunk_idx<<2)+((logic_scalar_addr>>3)&0x3)", where "(logic_scalar_addr>>3)&0x3" is to extract the 4th bit (b3) and the 5th bit (b4) of the logical address; and "chunk_idx<<2" indicates that the row start address and the row offset address (or row index) are determined by splicing. Since the row index is represented by 2 bits in the current scenario, for the configuration block identifier (chunk_idx), it only needs to be shifted left by 2 bits. If the configuration blocks are arranged in other forms, such as the i row and j column mentioned above, the identifier of the configuration block here should be shifted left accordingly. The row index should be taken from the logical address. +1 to + Similarly, since multiple configuration blocks are stacked in columns (or horizontally), the column index is the column address (bank_addr), which means that you only need to take the lower 3 bits of the corresponding column index in the logical address, such as "bank_addr=logic_scalar_addr&0x7". It should be understood that if the configuration blocks are arranged in other forms, such as row i and column j as mentioned above, the column address here should be the lower 3 bits of the logical address. Bit.

[0148] For example, Figure 4 The logical address corresponding to register 20 can be expressed as 0x00110100, and its corresponding configuration block index is 1, row index is 2, and column index is 4. Figure 4 The row starting address of chunk0 in the register is 0x0000 0000 (usually, the row starting address is obtained by querying the thread group identifier), and the row starting address of chunk1 is 0x0000 0100; therefore, the row address corresponding to this register is 0x0000 0100 + 0x00000010 (i.e. 2) = 0x0000 0110 (i.e. the 6th row starting from 0); the corresponding column address is the column index, which is 0x00000100 (i.e. the 4th column starting from 0).

[0149] To sum up, the processor provided in the embodiment of the present application shows the conversion method of logical addresses and physical addresses based on the configuration block division method shown in the embodiment of the present application, ensuring that the conversion of logical addresses and physical addresses can be achieved based on this method, thereby ensuring the normal execution of the program.

[0150] In order to improve the efficiency of the above initial loading and dynamic loading, the data required by the thread group may be pre-fetched into the first cache, as shown below.

[0151] 5. Data pre-fetching and persistence.

[0152] In some embodiments, the processor 100 also includes a first cache; a task assembly unit 130, for sending an initialization load request to the first cache; the first cache, for loading the first type of data corresponding to the first thread group into the first cache based on the initialization load request; and setting the first type of data to a persistent storage state, the persistent storage state being used to indicate that the first type of data will remain cached in the first cache until the first thread group ends.

[0153] Optionally, both the initial loading and the dynamic loading load data of the first type from the first cache. That is, the data loading unit is configured to load at least one first data from a first portion of registers in the first configuration block set in the first cache. The data loading unit is configured to load at least one second data from the first cache to at least one register in the first configuration block set during execution of the first thread group.

[0154] Optionally, the first type of data is set to a persistent storage state, or in other words, the cache line storing the first type of data is set to a persistent storage state.

[0155] Optionally, the persistent storage state means that data of the first type will remain cached in the first cache until the first thread group terminates. Alternatively, the persistent storage state means that data of the first type in the first cache cannot be replaced with other data until the first thread group terminates. It should be understood that the first type of data is specific to the first thread group. That is, if other thread groups terminate, the cache line used by that thread group to cache data of the first type can be replaced.

[0156] Optionally, the first cache is a first-level cache in a processor or a sub-processor.

[0157] In some embodiments, the processor 100 also includes a computing data master control unit and a second cache, the second cache being a next-level cache of the first cache; the computing data master control unit is used to load the first type of data from the buffer to the second cache; the first cache is used to load the first type of data corresponding to the first thread group from the second cache to the first cache based on an initialization load request.

[0158] Optionally, the computing data main control unit triggers loading of the first type of data from the buffer zone to the second cache, where the buffer zone refers to a buffer zone with constants obtained after program compilation.

[0159] Optionally, the second cache is a shared cache of multiple sub-processors, such as a second-level cache or a last-level cache; and the first cache is a private cache of the sub-processor, such as a first-level cache.

[0160] To sum up, the processor shown in the embodiment of the present application, by pre-fetching the first type of data into the first cache, can achieve that during the initial loading and dynamic loading process, the processor can quickly read the corresponding first type of data from the first cache, thereby reducing the loading delay of the first type of data and improving the execution efficiency of the thread group.

[0161] Furthermore, the first type of data is loaded from the buffer into the second cache by the computational data master unit, and then loaded from the second cache into the first cache, reusing the cache hierarchy used in related technologies to improve the adaptability of the processor. Furthermore, the second cache acts as a shared cache, while the first cache acts as a private cache. This caching method facilitates the concurrency of multiple sub-processors and reduces access latency.

[0162] It should be noted that the above "1. Register Division," "2. Register Loading," "3. Shared Return Data Path," "4. Address Access Process," and "5. Data Prefetching and Persistence" can be implemented as independent embodiments or as combined embodiments. The present application does not limit the combination of the above methods, i.e., combinations of two, three, four, or five can be used.

[0163] In some embodiments, when the first thread group ends its use of the first configuration block in the first configuration block set, the first configuration block is released. That is, embodiments of the present application support the early release of the thread group's register resources before the end of the first thread group's lifecycle, so that the register resources can be allocated to other thread groups for use, thereby improving the overall register occupancy rate.

[0164] Optionally, when the first thread group finishes using some registers (such as one or more rows of registers) in the first configuration block set, some registers are released in advance.

[0165] In parallel computing tasks, a task typically requires some shared registers (also known as constant registers) and some slot registers. These registers are used to store constants and common data for use by the arithmetic logic unit (ALU). These registers have low latency and high usage frequency.

[0166] However, these constant registers and scalar registers are always limited in number. As a result, when some applications are running, constants and public data information cannot be stored in these registers. Reg spills must be used or additional storage space in global memory must be allocated. However, when loading from global memory frequently, there will be a relatively large delay, and the delay is difficult to hide. In addition, once the registers used to store constants in the public registers are initialized, they cannot be overwritten or rewritten during operation. Therefore, the register setting method in the related art can significantly affect the performance of the processor.

[0167] Related technologies typically add a constant cache to store these values, loading them from the constant cache into registers when needed. However, this requires an independent constant cache data path, which incurs significant area overhead.

[0168] In order to solve the capacity limitations of dedicated shared registers and slot registers, the area overhead caused by independently opening a constant cache, and loading delays, the embodiment of the present application designs a configurable scalar processing register that supports initial loading + dynamic runtime loading.

[0169] When resource usage is high, each thread group (such as a wave) can be configured to use more scalar processing registers. In high-volume scenarios, dynamic runtime loading is used to pre-load constant values ​​that cannot be stored, as the execution approaches the instruction location. To address the issue of long initialization delays, we designed a design that initializes constant data and eliminates corresponding data dependencies.

[0170] For example, the general scheme of the embodiment of the present application is as follows.

[0171] When tasks are dispatched, register usage configuration information, referred to as scalar register usage configuration information, is carried along the task dispatch path. During task assembly, the thread assembly and resource usage control unit deducts the scalar register usage, and simultaneously dispatches thread group information and the number of chunks corresponding to the scalar register usage to the scalar register management unit.

[0172] After obtaining the corresponding configuration management information, the scalar processing register management unit performs allocation management and records thread group information and chunk tag information. It also numbers the chunk tag information as a reference for subsequent access and search. The drawid or kernelid information of indirect calls is directly written into the scalar processing registers by the thread assembly and resource utilization control unit.

[0173] The thread assembly and resource usage control unit then issues a load initialization request to initialize the allocated scalar processing registers. To reduce latency, the thread group scheduling process and the initialization of the allocated scalar processing registers can be performed in parallel.

[0174] In order to reduce the delay during multiple loading processes, the constant data that needs to be loaded is written into the last-level cache in the form of pre-loading. When it is used, it is loaded into the first-level cache for persistent storage.

[0175] Since the capacity of scalar processing registers is limited and cannot meet the needs of large-scale usage scenarios, the scalar processing registers have been made configurable. By configuring the number of scalar processing registers, the number of scalar processing registers corresponding to a thread group can be increased, which will correspondingly reduce the number of active thread groups.

[0176] For scenarios that exceed the configured number limit, dynamic loading instructions for loading constant data are added to dynamically load them at runtime. In this case, the dependency is removed by relying on the instruction number bar.

[0177] Add the early release function of scalar processing registers, and use it together with the early release function of vector processing registers. Just select the corresponding type when releasing.

[0178] Next, the specific implementation details of the above solution are described.

[0179] First, data dependencies are checked post-hoc during allocation management.

[0180] Specifically, the compiler uses a constant buffer + scalar register format during compilation. After compilation, the driver places all constant data into a designated constant buffer. During initialization, a limited number of constant values ​​(such as Limited_Numb) are written to scalar registers, e.g., 128 DWs. To reduce initialization latency, dependency checking is not performed during the initial loading phase. Data dependency checking instructions are inserted before any instructions are used to ensure data correctness.

[0181] To reduce subsequent access latency, data is initially loaded into the first-level cache (such as the data cache) through prefetching. To prevent cache lines from being replaced, persistence information is added to keep constant data permanently resident in the first-level cache data cache.

[0182] The initial loading process is as follows Figure 5 shown.

[0183] First, the computational data master control unit 60 pre-fetches large blocks of constant data from the constant buffer into the L2 cache or final-level cache 70. The program dispatch scheduler 61 then assembles a thread group, subtracting the number of scalar chunks corresponding to that thread group. The thread group configuration information is then sent to the scalar processing register 65 for chunk allocation. For each assembled thread group, the program dispatch scheduler 61 sends an initialization load request to the L1 cache 67 to load the constant data. This request is also marked with persistence information, ensuring that the loaded constant data remains persistently in the L1 cache 67.

[0184] Optionally, the constant data in the first-level cache 67 is pre-fetched from the second-level cache or the last-level cache 70 through the bus interface 68 and the on-chip network 69. The pre-fetching process may be earlier than the thread group assembly process, or may occur together with the thread group assembly process, or may be pre-fetched after the thread group assembly.

[0185] At this time, the thread group is dispatched to the thread group management unit 63 through the scheduler 62. When the thread group management unit 63 detects that all related resources have been allocated, it releases the basic scheduling dependency and sends the instruction to the instruction control unit 64 for instruction fetching, decoding and issuing.

[0186] When constant data is needed, a data dependency check is performed by executing the newly added data dependency check instruction to check whether the data is ready. When this instruction is executed, the corresponding thread group is blocked (or invalidated) and prevented from continuing to issue instructions. Then, a data dependency check is performed in the control pipeline to check whether the required constant data has been initialized and loaded. Specifically, this can be determined by judging the initialization signal corresponding to each thread group maintained by the scalar register 65. If the initialization signal is received, it means that the initialization load has been completed. If the initialization signal is not received, it means that the initialization load has not been completed. If the initialization load of the required data is not completed, the process will continue to wait; if the initialization has been completed, the thread group will be activated to continue execution.

[0187] After the initial load, if there is constant data beyond the limited number, these constants can be dynamically loaded via runtime instructions. During runtime, constants are loaded from the constant buffer, the first-level cache, the second-level cache, or the final-level cache via newly added dynamic load instructions. Specifically, the instruction control unit 64 triggers the load storage unit 71 to retrieve the corresponding constant data from the first-level cache 67. Data dependencies during the dynamic loading process are resolved via write barriers, read barriers, and wait barriers. A write barrier is a memory barrier instruction. When executed, it ensures that all write operations preceding it are completed, and that subsequent write operations will not begin until the previous write operations have completed. A read barrier is also a memory barrier instruction. It ensures that all read operations preceding it are completed, and that subsequent read operations will not begin until the previous read operations have completed. Wait barriers are primarily used to handle dependencies between instructions. They can cause one instruction to wait for another instruction to complete a specific operation before continuing execution.

[0188] When the kernel function is executed, the cache line permanently stored by the kernel function in the first-level cache 67 can be set to invalid. At this time, when other data arrives at the first-level cache 67, the data in the corresponding cache line can be overwritten.

[0189] The request and data flow during initial loading and dynamic loading can be referred to Figure 6 During initial loading, the program dispatch scheduler 61 sends an initialization load request to the first-level cache 67, indicating that the target is a scalar register; and the program dispatch scheduler 61 sends an initialization write request to the data loading unit 66, indicating that the target is a scalar register. At this time, the data loading unit 66 receives the data sent by the first-level cache 67 and is responsible for writing it into the corresponding scalar register 65. During dynamic loading, the first-level cache 67 is triggered to send data to the data loading unit 66 through a dynamic load instruction, and the dynamic load instruction is used to indicate that the target of the dynamic load is a scalar register.

[0190] Specifically, when the data loading unit 66 is writing, it will determine the corresponding 6-bit waveid (thread group identifier), 9-bit register logical address, 32-bit data and mask, where the waveid and register logical address are used to determine the location of the register used to store data next, and the 32-bit data will be stored in the register.

[0191] To maintain sufficient bandwidth without incurring additional area overhead during the initial and dynamic load processes, multiple stream processors (e.g., two stream processors) can be designed to share a common write-back logic within the storage entry buffer. L1 cache writebacks utilize the cache line width, or throughput. However, the bandwidth associated with scalar processing registers is typically much smaller than the L1 cache line width. Therefore, a configuration block structure using eight banks is designed. When the L1 cache writes back 128 DWs via the data movement executor, eight entries are stored in the storage entry buffer's eight-level deep FIFO, with one entry per level, corresponding to 128 / 8 = 16 DWs. The storage entry buffer then writes half of the data (8 DWs) at a time to the corresponding eight banks, retrieving half the data from a level of the FIFO and storing it in a row of the configuration block. This completes the write of 128 DWs in 16 cycles.

[0192] However, when loading 128 DWs from another thread group into the scalar processing register of another stream processor core, it can be written into the 8-level deep FIFO of the storage entry buffer below. Therefore, when writing to the scalar processing register of another stream processor core, it is also 8 DWs. By using a large bandwidth input and two small bandwidth outputs, the writing effect is fast and the pipeline on the return path is not congested.

[0193] The specific contents of the above multiple stream processors sharing a set of storage entry buffers can be referred to above Figure 3 The corresponding embodiments will not be described in detail here.

[0194] To effectively allocate thread groups, the resource management unit adds a scalar register chunk counter to identify the number of available chunks within the stream processor core when grouping tasks. This counter is used to count the number of available chunks within the stream processor core. As shown in Table 3 below, thread groups can be allocated and managed based on the specific chunk counter count. This information effectively controls the number of active thread groups. Table 3 provides basic configuration information for configuring the number of scalar register chunks used and the corresponding number of active thread groups.

[0195] Table 3 Scalar processing register configuration information table

[0196] When accessing a specific physical space location, calculation is required based on the waveid and the logical address of the scalar processing register. The specific calculation method is as follows.

[0197] waveid+(logic_scalar_addr>>5) to locate chunkidx.

[0198] line_addr=(chunk_idx<<2)+((logic_scalar_addr>>3)&3).

[0199] bank_addr=logic_scalar_addr&0x7.

[0200] The processor illustrated in the embodiments of this application primarily addresses the issue of using a large number of shared registers and slot registers during program execution. It designs configurable scalar processing registers in the form of initial loading and dynamic runtime loading. When the driver knows how many shared registers and slot registers the compiler uses, it dynamically configures the number of active thread groups to achieve precise utilization of scalar processing registers.

[0201] When the number of shared registers and slot registers exceeds the limit, constant data load requests are issued in advance, prefetching and persistence support are added, and constant data exceeding the limit can be quickly loaded back during operation.

[0202] This allows the compiler to load and switch quickly with almost no delay when running normal kernel calls, especially for accessing constants in indirect kernel calls.

[0203] Specifically, the embodiments of the present application include the following key points: (1) scalar register discrete register allocation management strategy before thread group execution; (2) constant value initialization loading, post-checking data dependency, and data dependency checking process; (3) dynamic loading of constant data during program execution; (4) chunk configuration management of scalar processing registers; (5) access address calculation logic of scalar processing registers; and (6) register resource release management when the thread group exits.

[0204] Figure 7 A flowchart of a register allocation method provided by an exemplary embodiment of the present application is shown. The method is executed by a processor comprising a plurality of registers, a register management unit, and a task assembly unit. The method comprises at least one of the following steps.

[0205] Step 210: The register management unit divides the plurality of registers into at least one configuration block chunk set, each configuration block set in the at least one configuration block set includes at least one configuration block, and the at least one configuration block includes at least two registers in the plurality of registers.

[0206] Optionally, the number of configuration blocks included in each configuration block set in at least one configuration block set is the same or different. For example, each configuration block set includes 5 configuration blocks, or each configuration block set includes 6 configuration blocks, or each configuration block set includes 8 configuration blocks, or each configuration block set includes 10 configuration blocks, and so on. For another example, at least one configuration block set includes 4 configuration block sets, and each configuration block set includes a different number of configuration blocks, such as configuration block set 1 includes 5 configuration blocks, configuration block set 2 includes 6 configuration blocks, configuration block set 3 includes 8 configuration blocks, and configuration block set 4 includes 10 configuration blocks. For another example, at least one configuration block set includes four groups of configuration block sets, the number of configuration blocks included in each group of configuration block sets is different, and the number of configuration blocks included in the configuration block sets included in each group of configuration block sets is the same, such as the first group of configuration block sets includes 16 configuration block sets, each of which includes 5 configuration blocks; the second group of configuration block sets includes 13 configuration block sets, each of which includes 6 configuration blocks; the third group of configuration block sets includes 10 configuration block sets, each of which includes 8 configuration blocks; and the fourth group of configuration block sets includes 8 configuration block sets, each of which includes 10 configuration blocks. It should be understood that when at least one configuration block set includes multiple groups of configuration block sets, the number of configuration block sets included in each group of configuration block sets can be different as shown above, or can be the same, and this embodiment of the present application is not limited to this.

[0207] Optionally, the sizes of the configuration blocks included in each configuration block set in at least one configuration block set are the same or different. Exemplarily, the sizes of the configuration blocks included in a configuration block set are the same, that is, the number of registers included in each configuration block included in a configuration block set is the same. However, the sizes of the configuration blocks included in different configuration block sets may be different. For example, if configuration block set 1 includes 5 configuration blocks, each configuration block in configuration block set 1 includes 32 registers; if configuration block set 2 includes 10 configuration blocks, each configuration block in configuration block set 2 includes 16 registers.

[0208] Optionally, the register management unit divides the multiple registers evenly to obtain at least one configuration block set, wherein each configuration block set in the at least one configuration block set has the same number and size of configuration blocks. That is, even division refers to dividing the multiple registers into at least one configuration block set, wherein each configuration block set in the at least one configuration block set has the same number and size of configuration blocks. Alternatively, the register usage unit divides the multiple registers unevenly to obtain at least one configuration block set, wherein at least one of the number and size of configuration blocks included in two configuration block sets in the at least one configuration block set differs. That is, uneven division refers to dividing the multiple registers into at least one configuration block set, wherein at least one of the number and size of configuration blocks included in two configuration block sets in the at least one configuration block set differs. For example, two configuration block sets may have different numbers of configuration blocks, or two configuration block sets may have different sizes of configuration blocks, or two configuration block sets may have different numbers and sizes of configuration blocks.

[0209] Optionally, the register usage unit evenly divides the multiple registers to obtain at least one configuration block, wherein each configuration block in the at least one configuration block has the same size; evenly divides the at least one configuration block to obtain at least one configuration block set, wherein each configuration block set in the at least one configuration block set includes the same number of configuration blocks. That is, evenly dividing the multiple registers means dividing the multiple registers into at least one configuration block, wherein each configuration block in at least one configuration class has the same size; evenly dividing the at least one configuration block means dividing the at least one configuration block into at least one configuration block set, wherein each configuration block set in the at least one configuration block set includes the same number of configuration blocks. Alternatively, the register usage unit evenly divides the multiple registers to obtain at least one configuration block; unevenly divides the at least one configuration block to obtain at least one configuration block set, wherein two configuration block sets in the at least one configuration block set have different numbers of configuration blocks. That is, unevenly dividing the at least one configuration block means dividing the at least one configuration block into at least one configuration block set, wherein two configuration block sets in the at least one configuration block set have different numbers of configuration classes. Alternatively, the register usage unit may perform non-uniform division of the multiple registers to obtain at least one configuration block, wherein two configuration blocks in the at least one configuration block have different sizes; and then perform uniform division of the at least one configuration block to obtain at least one configuration block set. In other words, the non-uniform division of the multiple registers refers to dividing the multiple registers into at least one configuration block, wherein two configuration blocks in the at least one configuration block have different sizes. Alternatively, the register usage unit may perform non-uniform division of the multiple registers to obtain at least one configuration block; and then perform non-uniform division of the at least one configuration block to obtain at least one configuration block set.

[0210] Optionally, the specific division method for multiple registers can be determined based on multiple factors such as the specific architecture of the processor, the type of task currently being executed, and the user's subjective settings, and the embodiments of the present application are not limited to this.

[0211] Step 220: The task assembly unit allocates a first configuration block set to the first thread group, where the first configuration block set is one of at least two configuration block sets.

[0212] The task assembly unit selects a first configuration block set from at least two configuration block sets for the first thread group.

[0213] Optionally, the task assembly unit allocates a first configuration block set to the first thread group according to tasks executed by the first thread group, wherein the tasks executed by the first thread group are used to determine at least one of the number of configuration blocks included in the configuration block set and the size of the configuration blocks.

[0214] Optionally, the task assembly unit may also be called a thread assembly and resource usage control unit, or a program distribution scheduler, etc.

[0215] It should be noted that the names of the various units shown in the embodiments of the present application are for illustration only. In specific implementation, other names may also be used to refer to the corresponding units. In addition, the unit division method in the embodiments of the present application is also for illustration only. In specific implementation, the number of units, functional boundaries, etc. may be adjusted. Specific adjustments may include merging, splitting, renaming, etc. For example, the resource usage management unit and the task assembly unit may also be adjusted to a task assembly and resource usage control unit. The embodiments of the present application do not limit this, but the scope of protection of the embodiments of the present application is not limited thereto.

[0216] In summary, the processor provided by the embodiment of the present application divides a plurality of registers into at least one configuration block set, and each configuration block set is used to allocate to a thread group. By dividing register resources into multiple configuration blocks, register resources are allocated in the form of configuration blocks. Each thread group supports the allocation of at least one configuration block (i.e., a configuration block set). Compared with the traditional method of continuous allocation based on registers as the granularity, this method of allocating register resources directly allocates configuration blocks instead of multiple continuous registers during allocation, which can avoid frequent cutting of continuous registers and cause register fragmentation. In addition, allocation in the form of configuration blocks can achieve management through configuration blocks instead of management based on registers, which can improve the efficiency of allocation and release of register resources.

[0217] In some embodiments, the plurality of registers are a plurality of scalar registers, or the plurality of registers are a plurality of vector registers, or the plurality of registers are a plurality of scalar registers and a plurality of vector registers, etc. It should be noted that the execution logic of the processor shown in the embodiments of the present application is applicable to different types of registers. In addition to the scalar registers and vector registers shown above, other types of registers may also be supported, and the embodiments of the present application are not limited thereto.

[0218] Next, the division method of registers will be described in detail.

[0219] 1. Division of registers.

[0220] In some embodiments, the above-mentioned step 210 can be implemented as follows: the register management unit determines quantity information based on the register usage configuration information, where the quantity information includes at least one of the number of configuration blocks, the number of active thread groups, the first data amount, and the second data amount; based on the quantity information, the multiple registers are divided into at least one configuration block set; wherein the number of configuration blocks is used to indicate the number of configuration blocks in each configuration block set in at least one configuration block set, the number of active thread groups is used to indicate the maximum number of thread groups that the processor supports for parallel execution, the first data amount is used to indicate the maximum amount of data supported for storage by each configuration block set in at least one configuration block set, and the second data amount is used to indicate the maximum amount of data supported for storage by each configuration block set.

[0221] Exemplarily, the register management unit queries the register configuration information table based on the register usage configuration information to determine the quantity information, where the register configuration information table is used to indicate a mapping relationship between the register usage configuration information and the quantity information.

[0222] For details, please refer to "1. Register Division" corresponding to the above processor, which will not be repeated here.

[0223] Next, we will explain in detail how the thread group uses the configuration block set.

[0224] 2. Loading of registers.

[0225] In some embodiments, the processor further includes a data loading unit, such as Figure 2 As shown; the data loading unit loads at least one first data for the first part of the registers in the first configuration block set, the first part of the registers is a specified number of registers or a specified proportion of registers in the first configuration block set, the first data belongs to a first type, and the first type of data includes at least one of constants and common data.

[0226] For details, please refer to "2. Register Loading" corresponding to the above processor, which will not be repeated here.

[0227] In the related art, the thread group scheduling process is usually executed first, and the data required for the instruction is loaded before or when a specific instruction is executed. However, in the embodiment of the present application, a large block of first type data is usually pre-loaded, and the data type of the first type pre-loaded in the embodiment of the present application is different from the data type loaded in the related art. Therefore, the traditional thread group scheduling process and data initialization loading process should be updated to adapt to the register allocation and usage method shown in the embodiment of the present application. The specific loading timing of loading at least one first data into the first part of the registers is as follows.

[0228] (1) Loading timing.

[0229] In some embodiments, the above-mentioned process of loading the first data into the first part of the registers can be called an initialization process. In order to improve the execution efficiency of the thread group in the processor, the above-mentioned initialization process can be executed in parallel with the scheduling process of the thread group. Compared with the related art of first completing the scheduling process of the thread group and then executing the initialization process (also called the loading process) for the registers of the thread group, by executing the scheduling process and the initialization process in parallel, it helps to shorten the execution time of the thread group, improve the execution efficiency of the thread group, and thus reduce the overall delay of task execution. Exemplarily, the above-mentioned processor also includes a scheduler; the scheduler sends a scheduling request, and the scheduling request is used to trigger the scheduling process of the first thread group, and the scheduling process refers to the process of scheduling execution of the first thread group; the task assembly unit sends an initialization load request, and the initialization load request is used to trigger the initialization process of the first configuration block set, and the initialization process refers to initializing the first part of the registers in the first configuration block set; wherein, the scheduling process and the initialization process are executed in parallel.

[0230] In some embodiments, the actual amount of data required by the thread group is large. For example, the data size of the configuration block set is 160DW, but the actual amount of data required by the thread group is greater than the data size of the configuration block set or the data size of the first part of registers. In this case, dynamic loading can be used to solve this problem, as shown below.

[0231] (2) Dynamic loading.

[0232] In some embodiments, during the execution of the first thread group, the data loading unit loads at least one second data into at least one register of the first configuration block set, the second data is of the first type, and there is a one-to-one relationship between the at least one register and the at least one second data.

[0233] Optionally, during execution of the first thread group, the data loading unit determines data of the first type required by at least one instruction, ie, at least one second data, and loads the at least one second data into at least one register of the first configuration block set.

[0234] Exemplarily, during the execution of the first thread group, the data loading unit loads at least one second data into at least one register of the first part of registers, the data in the at least one register is data that the first thread group has finished using, and the at least one second data is used to overwrite the first data in the at least one register.

[0235] Exemplarily, during the execution of the first thread group, the data loading unit loads at least one second data into at least one register of the second part of registers, where the second part of registers is the registers in the first configuration block set except the first part of registers; wherein the registers in the at least one register meet at least one of the following characteristics: the data in the register is data that has been finished being used by the first thread group; the register is an empty register.

[0236] In other embodiments, the data loading unit loads at least one second data into at least one register in the first portion of registers and the second portion of registers during execution of the first thread group. For example, the first configuration block set includes 160 registers, the first portion of registers includes 128 registers, the second portion of registers includes 32 registers, and the at least one second data is 32 second data. The data loading unit determines the register for loading the second data based on the usage of each register in the first configuration block set. For example, if data in 24 registers in the first portion of registers has ended, and data in 16 registers in the second portion of registers has ended, the 24 registers in the first portion of registers are preferentially selected. Since the 24 registers in the first portion of registers are less than the number of second data, 8 registers in the second portion of registers are further selected for loading the second data.

[0237] Next, the above initialization process and dynamic loading process are specifically described by taking the example of a first configuration block set including n first registers (registers for storing first type data) and m second registers (registers for storing first type and second type data).

[0238] In some embodiments, when the number of first-type data required by the first thread group is less than or equal to the number of first registers, loading of the first-type data by the first thread group can be completed only through initial loading. That is, when the first number is less than or equal to n, the data loading unit loads a first number of first data into a first number of first registers based on the initial load request, where the first number is the number of first-type data corresponding to the first thread group.

[0239] Optionally, if the amount of first type data required by the first thread group is less than or equal to the number of first registers, the remaining first registers may be left vacant or used to load second type data. The remaining first registers refer to registers among the n first registers that were not loaded with first data during the initial process.

[0240] In some embodiments, when the amount of first type data required by the first thread group is greater than the number of first registers, initial loading and dynamic loading should be used to implement the execution of the first thread group. Specifically, the initial loading process is as follows. The first configuration block set includes n first registers, and the n first registers are used to store first type data, where n is a specified number or determined by a specified ratio; when the first number is greater than n, the data loading unit loads the n first data into the n first registers based on the initialization load request, where the first number is the number of first type data corresponding to the first thread group.

[0241] In some embodiments, after the initialization process is completed, the remaining data in the first number of data corresponding to the first thread group can be loaded into the register using a dynamic loading method. Exemplarily, the processor also includes an instruction control unit; when the first number is greater than n, the instruction control unit executes a dynamic loading instruction, the dynamic loading instruction being used to instruct that the first type of data be loaded into the first configuration block set during the execution of the first thread group; and a data loading unit being used to load k second data into k registers based on the dynamic loading instruction, where k is a positive integer and the second data is of the first type.

[0242] Exemplarily, the data loading unit, based on a dynamic load instruction, loads k second data into k first registers, where the data in the k first registers is data that has been finished using by the first thread group, and the k second data is used to overwrite the data in the k first registers. Alternatively, based on a dynamic load instruction, the k second data are loaded into k second registers, where each of the k second registers satisfies at least one of the following characteristics: the data in the register is data that has been finished using by the first thread group; the register is an empty register. Alternatively, based on a dynamic load instruction, the k second data are loaded into i first registers and j second registers, where i+j=k, where the screening conditions for the i first registers and the j second registers are as described above for the k first registers and the k second registers, and are not further described here.

[0243] In addition to the configuration method used to divide the registers of each thread group and the combination of initial loading and dynamic loading for the first type of data loaded into the registers as shown above, in order to further reduce the area of ​​the processor, a method is also designed in which multiple sub-processors share a set of write-back paths, as shown below.

[0244] 3. Shared return data path.

[0245] In some embodiments, the processor further includes at least one data loading unit, a first cache, and multiple sub-processors; each data loading unit in the at least one data loading unit corresponds to m sub-processors, the m sub-processors are all or part of the multiple sub-processors included in the processor, and each of the m sub-processors includes at least one configuration block set; the data loading unit includes a data movement executor and a storage entry buffer; the data movement executor uses a first bandwidth to read s third data from the first cache; and caches the s third data to a first first-in-first-out (FIFO) queue of the storage entry buffer, the first bandwidth is related to the cache line width of the first cache; the storage entry buffer uses a second bandwidth to write the s third data from the first FIFO queue to at least two registers corresponding to the first sub-processor in the m sub-processors, the second bandwidth is related to the size of the configuration block; wherein the first bandwidth is greater than the second bandwidth.

[0246] For details, please refer to "3. Shared return data path" corresponding to the above processor, which will not be repeated here.

[0247] Since the embodiment of the present application adopts the form of configuration blocks to configure register resources, corresponding changes should be made when the processor accesses the registers, as shown below.

[0248] 4. Address access process.

[0249] In some embodiments, the processor determines, in response to an access request for a register to be accessed, a physical address of the register to be accessed based on an identifier of a thread group corresponding to the register to be accessed and a logical address of the register to be accessed, where the physical address is used to indicate that the register to be accessed is located in the bth row and cth column of the ath configuration block, where a, b, and c are all positive integers.

[0250] For details, please refer to "4. Address Access Process" corresponding to the above processor, which will not be repeated here.

[0251] In order to improve the efficiency of the above initial loading and dynamic loading, the data required by the thread group may be pre-fetched into the first cache, as shown below.

[0252] 5. Data pre-fetching and persistence.

[0253] In some embodiments, the processor also includes a first cache; the task assembly unit sends an initialization load request to the first cache; the first cache loads the first type of data corresponding to the first thread group into the first cache based on the initialization load request; and sets the first type of data to a persistent storage state, the persistent storage state being used to indicate that the first type of data will remain cached in the first cache until the first thread group ends.

[0254] Optionally, the first cache is a first-level cache in a processor or a sub-processor.

[0255] In some embodiments, the processor also includes a computing data master control unit and a second cache, the second cache being a next-level cache of the first cache; the computing data master control unit loads the first type of data from the buffer to the second cache; the first cache loads the first type of data corresponding to the first thread group from the second cache to the first cache based on the initialization load request.

[0256] For details, please refer to "5. Data Prefetching and Persistence" corresponding to the above processors, which will not be repeated here.

[0257] In some embodiments, when the first thread group ends its use of the first configuration block in the first configuration block set, the first configuration block is released. That is, embodiments of the present application support the early release of the thread group's register resources before the end of the first thread group's lifecycle, so that the register resources can be allocated to other thread groups for use, thereby improving the overall register occupancy rate.

[0258] Optionally, when the first thread group finishes using some registers (such as one or more rows of registers) in the first configuration block set, some registers are released in advance.

[0259] Please refer to Figure 8 , which shows a block diagram of the register allocation device provided by an exemplary embodiment of the present application. This device has the functionality to implement the aforementioned register allocation method example. This functionality can be implemented by hardware, or by hardware executing corresponding software. This device can be the processor described above, or it can be provided within a processor. This device can include: a register usage management module 310 and a task assembly module 320.

[0260] a register usage management module 310, configured to divide the plurality of registers into at least one configuration block set, each configuration block set in the at least one configuration block set including at least one configuration block, and the at least one configuration block including at least two registers from the plurality of registers; The task assembly module 320 is configured to allocate a first configuration block set to the first thread group, where the first configuration block set is one of the at least two configuration block sets.

[0261] In some embodiments, the register usage management module 310 is configured to determine quantity information based on register usage configuration information, the quantity information including at least one of the number of configuration blocks, the number of active thread groups, the first data volume, and the second data volume; and divide the plurality of registers into the at least one configuration block set based on the quantity information; Among them, the number of configuration blocks is used to indicate the number of configuration blocks in each configuration block set in the at least one configuration block set, the number of active thread groups is used to indicate the maximum number of thread groups supported by the processor for parallel execution, the first data volume is used to indicate the maximum amount of data supported for storage by each configuration block set in the at least one configuration block set, and the second data volume is used to indicate the maximum amount of data supported for storage by each configuration block.

[0262] In some embodiments, the register usage management module 310 is used to query a register configuration information table based on the register usage configuration information to determine the quantity information, and the register configuration information table is used to indicate a mapping relationship between the register usage configuration information and the quantity information.

[0263] In some embodiments, the apparatus further comprises a data loading module; The data loading module is used to load at least one first data into a first part of registers in the first configuration block set, where the first part of registers is a specified number of registers or a specified proportion of registers in the first configuration block set, and the first data is of a first type, and the first type of data includes at least one of constants and common data.

[0264] In some embodiments, the apparatus further comprises a scheduling module; The scheduling module is configured to send a scheduling request, wherein the scheduling request is configured to trigger a scheduling process of the first thread group, wherein the scheduling process refers to a process of scheduling execution of the first thread group; The task assembly module 320 is configured to send an initialization load request, wherein the initialization load request is configured to trigger an initialization process of the first configuration block set, wherein the initialization process is to initialize the first portion of registers in the first configuration block set; The scheduling process and the initialization process are executed in parallel.

[0265] In some embodiments, the data loading module is used to load at least one second data into at least one register of the first configuration block set during the execution of the first thread group, the second data belongs to the first type, and there is a one-to-one relationship between the at least one register and the at least one second data.

[0266] In some embodiments, the data loading module is used to load the at least one second data into at least one register of the first part of registers during the execution of the first thread group, the first data in the at least one register is data that the first thread group has finished using, and the at least one second data is used to overwrite the first data in the at least one register.

[0267] In some embodiments, the data loading module is configured to load the at least one second data into at least one register of a second portion of registers during execution of the first thread group, where the second portion of registers is a register in the first configuration block set other than the first portion of registers; The register in the at least one register satisfies at least one of the following characteristics: The data in the register is data that has been finished being used by the first thread group; The register is an empty register.

[0268] In some embodiments, the first configuration block set includes n first registers, the n first registers are used to store the first type of data, and n is the specified number or is determined by the specified ratio; The data loading module is configured to load n first data into the n first registers based on the initialization load request when the first number is greater than n, where the first number is the number of first type data corresponding to the first thread group.

[0269] In some embodiments, the apparatus further comprises a command control module; the instruction control module is configured to execute a dynamic loading instruction when the first number is greater than n, the dynamic loading instruction being configured to instruct loading the first type of data into the first configuration block set during execution of the first thread group; The data loading module is configured to load k second data into k registers based on the dynamic loading instruction, where k is a positive integer and the second data belongs to the first type.

[0270] In some embodiments, the dynamic load instruction is used to load the first type of data required by the first instruction into the first configuration block set before the first instruction is executed.

[0271] In some embodiments, the first configuration block set includes n first registers, the n first registers are used to store the first type of data, and n is the specified number or is determined by the specified ratio; The data loading module is used to load the first number of first data into the first number of first registers based on the initialization load request when the first number is less than or equal to n, where the first number is the number of first type data corresponding to the first thread group.

[0272] In some embodiments, the apparatus further comprises at least one data loading module, a first cache, and a plurality of sub-processors; each of the at least one data loading module corresponds to m sub-processors, the m sub-processors being all or part of the plurality of sub-processors included in the processor, and each of the m sub-processors comprising at least one configuration block set; the data loading module comprises a data movement executor and a storage entry buffer; The data movement executor is configured to read s third data from the first cache using a first bandwidth; and cache the s third data in a first FIFO queue of the storage entry buffer, wherein the first bandwidth is related to a cache line width of the first cache; the storage entry buffer being configured to write the s third data from the first FIFO queue into at least two registers corresponding to a first sub-processor among the m sub-processors using a second bandwidth, wherein the first FIFO queue corresponds to the first sub-processor, and the second bandwidth is related to a size of the configuration block; The first bandwidth is greater than the second bandwidth.

[0273] In some embodiments, the apparatus further includes an instruction execution module. The instruction execution module is configured to, in response to an access request for a register to be accessed, determine a physical address of the register to be accessed based on an identifier of a thread group corresponding to the register to be accessed and a logical address of the register to be accessed, where the physical address indicates that the register to be accessed is located at the bth row and the cth column of the ath configuration block, where a, b, and c are all positive integers.

[0274] In some embodiments, the instruction execution module is used to determine, in response to an access request for the register to be accessed, an identifier of a configuration block corresponding to the register to be accessed based on an identifier of a thread group corresponding to the register to be accessed and a logical address of the register to be accessed; determine a row start address corresponding to the configuration block based on the identifier of the configuration block; determine a row offset address and a column address corresponding to the register to be accessed based on the logical address of the register to be accessed; determine a row address of the register to be accessed based on the row start address and the row offset address; and determine a physical address of the register to be accessed based on the row address and the column address.

[0275] In some embodiments, the task assembly module is used to send an initialization load request to the first cache; the first cache is used to load the first type of data corresponding to the first thread group into the first cache based on the initialization load request; and set the first type of data to a persistent storage state, and the persistent storage state is used to indicate that the first type of data will remain cached in the first cache until the first thread group ends.

[0276] In some embodiments, the apparatus further includes a computing data master control module and a second cache, wherein the second cache is a lower level cache than the first cache; The computing data main control module is configured to load the first type of data from the buffer into the second cache; The first cache is configured to load the first type of data corresponding to the first thread group from the second cache to the first cache based on the initialization load request.

[0277] In some embodiments, the plurality of registers is a plurality of scalar registers.

[0278] On the other hand, an embodiment of the present application provides a graphics card, which includes the atomic operation processing system described in the above embodiments.

[0279] In another aspect, embodiments of the present application provide a computer device comprising the atomic operation processing system described above. The computer device may be at least one of a portable computer, a desktop computer, a server, a server cluster, an artificial intelligence (AI) computing cluster, and a cloud computing cluster. The AI ​​computing cluster may also be referred to as an intelligent computing cluster or a smart computing cluster.

[0280] It should be understood that the "multiple" mentioned in this article refers to two or more. The character " / " generally indicates that the objects associated with each other are in an "or" relationship. In addition, the step numbers described in this article only illustrate a possible execution order between the steps. In some other embodiments, the above steps may also be executed in a non-numbered order, such as two steps with different numbers being executed simultaneously, or two steps with different numbers being executed in the opposite order to that shown in the figure. This embodiment of the application is not limited to this.

[0281] The above description is merely an optional embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.

Claims

1. A processor, characterized in that: The processor includes a plurality of registers, a register management unit and a task assembly unit; The register management unit is configured to divide the plurality of registers into at least one configuration block set, each configuration block set in the at least one configuration block set includes at least one configuration block, and the at least one configuration block includes at least two registers in the plurality of registers; The task assembly unit is configured to allocate a first configuration block set to the first thread group, where the first configuration block set is one of the at least two configuration block sets.

2. The processor according to claim 1, wherein: The register management unit is configured to determine quantity information based on register usage configuration information, the quantity information including at least one of a number of configuration blocks, a number of active thread groups, a first data volume, and a second data volume; and divide the plurality of registers into the at least one configuration block set based on the quantity information; Among them, the number of configuration blocks is used to indicate the number of configuration blocks in each configuration block set in the at least one configuration block set, the number of active thread groups is used to indicate the maximum number of thread groups supported by the processor for parallel execution, the first data volume is used to indicate the maximum data volume supported by each configuration block set in the at least one configuration block set for storage, and the second data volume is used to indicate the maximum data volume supported by each configuration block for storage.

3. The processor according to claim 2, wherein: The register management unit is configured to query a register configuration information table based on the register usage configuration information to determine the quantity information, wherein the register configuration information table is configured to indicate a mapping relationship between the register usage configuration information and the quantity information.

4. The processor according to any one of claims 1 to 3, characterized in that: The processor further includes a data loading unit; The data loading unit is used to load at least one first data for a first part of registers in the first configuration block set, where the first part of registers is a specified number of registers or a specified proportion of registers in the first configuration block set, and the first data is of a first type, and the first type of data includes at least one of constants and common data.

5. The processor according to claim 4, wherein: The processor also includes a scheduler; The scheduler is configured to send a scheduling request, wherein the scheduling request is configured to trigger a scheduling process of the first thread group, wherein the scheduling process refers to a process of scheduling execution of the first thread group; The task assembly unit is configured to send an initialization load request, where the initialization load request is configured to trigger an initialization process of the first configuration block set, where the initialization process is to initialize the first part of registers in the first configuration block set; The scheduling process and the initialization process are executed in parallel. The processor according to claim 4 , wherein: The data loading unit is configured to load at least one second data into at least one register of the first configuration block set during execution of the first thread group, where the second data belongs to the first type, and there is a one-to-one relationship between the at least one register and the at least one second data.

7. The processor according to claim 6, wherein: The data loading unit is configured to load the at least one second data into at least one register of the first part of registers during execution of the first thread group, wherein the at least one second data is configured to overwrite the first data in the at least one register.

8. The processor according to claim 6, wherein: The data loading unit is configured to load the at least one second data into at least one register of a second portion of registers during execution of the first thread group, where the second portion of registers is registers in the first configuration block set other than the first portion of registers; The register in the at least one register satisfies at least one of the following characteristics: The data in the register is data that has been finished being used by the first thread group; The register is an empty register.

9. The processor according to claim 6, wherein: The first configuration block set includes n first registers, the n first registers are used to store the first type of data, and n is the specified number or is determined by the specified ratio; The data loading unit is configured to load n first data into the n first registers based on an initialization load request when a first number is greater than n, where the first number is the number of first type data corresponding to the first thread group.

10. The processor according to claim 9, wherein: The processor further includes an instruction control unit; the instruction control unit being configured to execute a dynamic load instruction when the first number is greater than n, the dynamic load instruction being configured to instruct loading the first type of data into the first configuration block set during execution of the first thread group; The data loading unit is used to load k second data into k registers based on the dynamic load instruction, where k is a positive integer, and the k registers are free registers in the first configuration block set, and the free registers are empty registers or registers whose data have ended being used.

11. The processor according to claim 10, wherein: The dynamic load instruction is used to load the first type of data required by the first instruction into the first configuration block set before the first instruction is executed.

12. The processor according to claim 6, wherein: The first configuration block set includes n first registers, the n first registers are used to store the first type of data, and n is the specified number or is determined by the specified ratio; The data loading unit is used to load the first number of first data into the first number of first registers based on the initialization load request when the first number is less than or equal to n, where the first number is the number of first type data corresponding to the first thread group.

13. The processor according to claim 4, wherein: The processor further includes at least one data loading unit, a first cache, and a plurality of sub-processors; each of the at least one data loading unit corresponds to m sub-processors, the m sub-processors being all or part of the plurality of sub-processors included in the processor, and each of the m sub-processors including at least one configuration block set; the data loading unit includes a data movement executor and a storage entry buffer; The data movement executor is configured to read s third data from the first cache using a first bandwidth; and cache the s third data in a first FIFO queue of the storage entry buffer, wherein the first bandwidth is related to a cache line width of the first cache; the storage entry buffer being configured to write the s third data from the first FIFO queue into at least two registers corresponding to a first sub-processor among the m sub-processors using a second bandwidth, wherein the first FIFO queue corresponds to the first sub-processor, and the second bandwidth is related to a size of the configuration block; The first bandwidth is greater than the second bandwidth.

14. The processor according to any one of claims 1 to 3, characterized in that: The registers in each configuration block of the at least one configuration block are arranged in i rows and j columns; The processor is configured to, in response to an access request for a register to be accessed, determine a physical address of the register to be accessed based on an identifier of a thread group corresponding to the register to be accessed and a logical address of the register to be accessed, wherein the physical address is used to indicate that the register to be accessed is located in the bth row and cth column of the ath configuration block.

15. The processor according to claim 14, wherein: The processor is configured to, in response to an access request for the register to be accessed, determine an identifier of a configuration block corresponding to the register to be accessed based on an identifier of a thread group corresponding to the register to be accessed and a logical address of the register to be accessed; and determine a row start address corresponding to the configuration block based on the identifier of the configuration block; Based on the logical address of the register to be accessed, the row offset address and column address corresponding to the register to be accessed are determined; based on the row start address and the row offset address, the row address of the register to be accessed is determined; based on the row address and the column address, the physical address of the register to be accessed is determined.

16. The processor according to claim 4, wherein: The processor further includes a first cache; The task assembly unit is configured to send an initialization load request to the first cache; the first cache being configured to load data of the first type corresponding to the first thread group into the first cache based on the initialization load request; And setting the first type of data to a persistent storage state, where the persistent storage state is used to indicate that the first type of data will remain cached in the first cache before the first thread group ends.

17. The processor according to claim 16, wherein: The processor further includes a computing data main control unit and a second cache, wherein the second cache is a lower level cache than the first cache; The computing data main control unit is configured to load the data of the first type from the buffer into the second cache; The first cache is configured to load the first type of data corresponding to the first thread group from the second cache to the first cache based on the initialization load request.

18. The processor according to any one of claims 1 to 3, characterized in that: The plurality of registers are a plurality of scalar registers.

19. A graphics card, characterized in that: The graphics card includes the processor according to any one of claims 1 to 18.

20. A computer device, characterized in that: The computer device comprises the processor according to any one of claims 1 to 18.

21. A register allocation method, characterized in that: The method is executed by a processor according to any one of claims 1 to 18, wherein the processor comprises a plurality of registers, a register management unit and a task assembly unit; The method comprises: The register management unit divides the plurality of registers into at least one configuration block set, each configuration block set in the at least one configuration block set includes at least one configuration block, and the at least one configuration block includes at least two registers in the plurality of registers; The task assembly unit allocates a first configuration block set to the first thread group, where the first configuration block set is one of the at least two configuration block sets.

22. A register allocation device, characterized in that: The device comprises: The register usage management module divides the plurality of registers into at least one configuration block set, each configuration block set in the at least one configuration block set includes at least one configuration block, and the at least one configuration block includes at least two registers in the plurality of registers; The task assembly module allocates a first configuration block set to the first thread group, where the first configuration block set is one of the at least two configuration block sets.

Citation Information

Patent Citations

  • Methods and apparatus for source operand collector caching

    CN103197916A

  • Register allocation for clustered multi-level register files

    CN103870309A

  • GPGPU register file dynamic extension method

    CN108595258A

  • Method of managing register file units

    CN111459543A

  • Command distribution method, command distributor, chip and electronic device

    CN115129369A