Processor, graphics card, computer device, register allocation method and apparatus

By dividing registers into sets of configuration blocks and allocating them to thread groups, the problem of register fragmentation in existing technologies is solved, achieving more efficient resource management and release.

CN120578422BActive Publication Date: 2026-01-27MOORE THREADS TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511093566.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-06
Publication Date
2026-01-27
Estimated Expiration
2045-08-06

AI Technical Summary

Technical Problem

In existing technologies, register allocation methods are not conducive to GPU management, leading to frequent splitting of consecutive registers, resulting in register fragmentation and affecting resource allocation and release efficiency.

Method used

Multiple registers are divided into sets of configuration blocks, each set of configuration blocks is used to allocate to thread groups. Resource management is carried out in the form of configuration blocks, avoiding frequent cutting of consecutive registers and improving allocation and release efficiency.

Benefits of technology

By allocating register resources through configuration blocks, register fragmentation is avoided, resource allocation and release efficiency is improved, and GPU management and performance are optimized.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120578422B_ABST
    Figure CN120578422B_ABST
Patent Text Reader

Abstract

The application discloses a processor, a display card, a computer device, a register allocation method and device, and belongs to the technical field of register management. The processor comprises a plurality of registers, a register management unit and a task assembling unit. The register management unit is used for dividing the plurality of registers into at least one configuration block set, each configuration block set in the at least one configuration block set comprising at least one configuration block, and the at least one configuration block comprising at least two registers in the plurality of registers. The task assembling unit is used for allocating a first configuration block set to a first thread group, and the first configuration block set being one of the at least two configuration block sets. The processor allocates register resources in the form of configuration blocks, and can improve the allocation and release efficiency of register resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of register management technology, and in particular to a processor, graphics card, computer device, register allocation method and apparatus. Background Technology

[0002] Registers are crucial units in the Graphics Processing Unit (GPU). They are the fastest-accessible memory in the GPU, providing rapid data retrieval for the GPU's processing cores.

[0003] In related technologies, resource allocation for thread groups is typically performed at the register level. For example, multiple registers may be allocated to a thread group, but these registers may be contiguous or non-contiguous, and the register resources allocated for different registers may vary each time.

[0004] However, this register allocation method is not conducive to GPU register management. Summary of the Invention

[0005] This application provides a processor, a graphics card, a computer device, a register allocation method and apparatus, the technical solution of which is shown below.

[0006] According to one aspect of this application, a processor is provided, the processor including a plurality of registers, a register management unit and a task assembly unit;

[0007] The register management unit is configured to divide the plurality of registers into at least one set of configuration blocks, each set of configuration blocks including at least one configuration block, and each at least one configuration block including at least two registers from the plurality of registers;

[0008] The task assembly unit is used to allocate a first configuration block set to the first thread group, wherein the first configuration block set is one of the at least two configuration block sets.

[0009] According to one aspect of this application, a graphics card is provided, the graphics card including the processor described above.

[0010] According to one aspect of this application, a computer device is provided, the computer device including the processor described above.

[0011] According to one aspect of this application, a register allocation method is provided, the method being executed by the aforementioned processor, the processor including a plurality of registers, a register management unit, and a task assembly unit;

[0012] The method includes:

[0013] The register management unit divides the plurality of registers into at least one set of configuration blocks, each set of configuration blocks including at least one configuration block, and each at least one configuration block including at least two registers from the plurality of registers;

[0014] The task assembly unit allocates a first configuration block set to the first thread group, wherein the first configuration block set is one of the at least two configuration block sets.

[0015] According to one aspect of this application, a register allocation apparatus is provided, the apparatus comprising:

[0016] The register usage management module divides the plurality of registers into at least one set of configuration blocks. Each set of configuration blocks includes at least one configuration block, and each at least one configuration block includes at least two registers from the plurality of registers.

[0017] The task assembly module allocates a first configuration block set to the first thread group, and the first configuration block set is one of the at least two configuration block sets.

[0018] The beneficial effects of the technical solution provided in this application include at least the following:

[0019] By first dividing multiple registers into at least one set of configuration blocks, each set of configuration blocks is used to allocate register resources to thread groups. This allocation method, which allocates register resources directly to configuration blocks rather than multiple consecutive registers, avoids frequent fragmentation of consecutive registers. Furthermore, allocating through configuration blocks allows for management based on configuration blocks rather than individual registers, improving the efficiency of register resource allocation and release. Attached Figure Description

[0020] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 A schematic diagram of a processor provided in an exemplary embodiment of this application is shown;

[0022] Figure 2 A schematic diagram of a processor provided in yet another exemplary embodiment of this application is shown;

[0023] Figure 3 A schematic diagram of a processor provided in another exemplary embodiment of this application is shown;

[0024] Figure 4 A schematic diagram illustrating an address translation rule provided in an exemplary embodiment of this application is shown;

[0025] Figure 5 A schematic diagram illustrating the loading process of a processor provided in an exemplary embodiment of this application is shown;

[0026] Figure 6 This illustration shows a schematic diagram of the loading process of a processor provided in yet another exemplary embodiment of this application;

[0027] Figure 7 A flowchart illustrating a register allocation method provided in an exemplary embodiment of this application is shown;

[0028] Figure 8 A structural block diagram of a register allocation apparatus provided in an exemplary embodiment of this application is shown. Detailed Implementation

[0029] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0030] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0031] The terminology used in this disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of the disclosure. The singular forms “a,” “the,” and “the” as used in this disclosure and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.

[0032] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions. For example, the settings and operation information involved in this application were obtained with full authorization.

[0033] It should be understood that although the terms first, second, etc., may be used in this disclosure to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, a first parameter may also be referred to as a second parameter without departing from the scope of this disclosure, and similarly, a second parameter may also be referred to as a first parameter. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."

[0034] First, let me introduce the relevant terms used in this application.

[0035] Cache: A small-capacity, high-speed memory, part of the storage system, that stores frequently used instructions and data. It can also be called cache memory. Processors typically include multiple levels of cache, each with a different capacity. Smaller caches generally have higher read / write efficiency because they process less data.

[0036] Caches are divided into multiple levels, such as Level 1 (L1) cache, Level 2 (L2) cache, and Level 3 (L3) cache. The cache capacity of each level gradually increases, while the speed gradually decreases. The Last Level Cache (LLC) is the last level of cache for the CPU core, such as the Level 3 cache. When the processing core needs to access data, if the data is not found in the Level 1 or Level 2 cache, it will continue to search the Last Level Cache. If the data is still not found in the Last Level Cache, the processing core needs to read the required data from main memory.

[0037] Caches can be categorized into instruction caches and data caches based on the different information they store. Instruction caches store instructions, while data caches store data. This application primarily uses a data cache as an example, but it can also be used to support other caches for storing data; this application does not limit the scope of the examples.

[0038] Network-on-Chip (NoC): A network-based communication subsystem located within an integrated circuit. It is typically used for data transmission and communication between different modules in a System-on-Chip (SoC). NoC technology connects the processor core, memory, and various peripherals via a router, forming a highly parallel communication architecture that effectively improves data transmission efficiency and communication bandwidth. Compared to traditional shared bus methods, NoC technology offers higher scalability and better performance, making it particularly suitable for multi-core systems.

[0039] An instruction is a command that directs a computer to perform a specific operation; it is the smallest functional unit of computer operation. An instruction is a statement in machine language, or a set of meaningful binary code. The collection of all the instructions of a computer constitutes its instruction set, also known as its instruction system.

[0040] Instruction Format: A readable representation of an instruction. An instruction typically includes an opcode and operands. The opcode describes the type of operation the instruction will perform, while the operands provide the data or addresses of data required to execute the instruction. The opcode is indispensable in an instruction, but operands are optional, and there can be one or two operands.

[0041] Dead registers: Registers no longer used by a thread group. In other words, dead registers are registers that are no longer used during the lifetime of a thread group. Alternatively, dead registers are registers that are no longer used during the lifetime of the individual threads within a thread group. A wave is a SIMD (Single Instruction Multiple Data) thread bundle structure in the OpenCL (Open Computing Language) processing framework. In some processors, a wave can also be called a wavefront. A thread group typically includes 32 or 64 threads. In CUDA (Compute Unified Device Architecture), a thread group can be called a warp, and a warp typically includes 32 threads. Of course, the terms wave, warp, etc., can also be more generally used, namely, thread group.

[0042] Figure 1 A schematic diagram of a processor provided in an exemplary embodiment of this application is shown. The processor 100 includes a plurality of registers 110, a register management unit 120, and a task assembly unit 130.

[0043] Register management unit 120 is used to divide a plurality of registers 110 into at least one set of configuration blocks, each set of configuration blocks including at least one configuration block, and at least one configuration block including at least two registers from the plurality of registers 110.

[0044] Optionally, the number of configuration blocks included in each configuration block set within at least one configuration block set may be the same or different. For example, each configuration block set may include 5 configuration blocks, or each configuration block set may include 6 configuration blocks, or each configuration block set may include 8 configuration blocks, or each configuration block set may include 10 configuration blocks, and so on. As another example, at least one configuration block set may include 4 configuration block sets, each containing a different number of configuration blocks, such as configuration block set 1 containing 5 configuration blocks, configuration block set 2 containing 6 configuration blocks, configuration block set 3 containing 8 configuration blocks, and configuration block set 4 containing 10 configuration blocks. For example, at least one configuration block set may include four sets of configuration block sets. The number of configuration blocks included in each set may differ, but the number of configuration blocks included in each set may be the same. For instance, the first set may include 16 sets of configuration blocks, each containing 5 configuration blocks; the second set may include 13 sets of configuration blocks, each containing 6 configuration blocks; the third set may include 10 sets of configuration blocks, each containing 8 configuration blocks; and the fourth set may include 8 sets of configuration blocks, each containing 10 configuration blocks. It should be understood that when at least one configuration block set includes multiple sets of configuration block sets, the number of configuration blocks included in each set may be different or the same as shown above, and this embodiment does not limit this.

[0045] Optionally, the configuration blocks included in each configuration block set within at least one configuration block set may be of the same or different sizes. For example, the configuration blocks included in a configuration block set may be of the same size, meaning that the number of registers included in each configuration block within a configuration block set is the same. However, the configuration blocks included in different configuration block sets may be of different sizes. For instance, configuration block set 1 may include 5 configuration blocks, each of which includes 32 registers; configuration block set 2 may include 10 configuration blocks, each of which includes 16 registers.

[0046] Optionally, the register management unit 120 divides the multiple registers 110 evenly to obtain at least one set of configuration blocks, wherein each configuration block set in the at least one set contains the same number and size of configuration blocks. That is, even division means dividing the multiple registers 110 into at least one set of configuration blocks, where each configuration block set in the at least one set contains the same number and size of configuration blocks. Alternatively, the register usage unit divides the multiple registers 110 non-uniformly to obtain at least one set of configuration blocks, wherein at least two configuration block sets in the at least one set contain at least one different number and size of configuration blocks. That is, non-uniform division means dividing the multiple registers 110 into at least one set of configuration blocks, where at least two configuration block sets in the at least one set contain at least one different number and size of configuration blocks. For example, two configuration block sets may contain different numbers of configuration blocks, or different sizes of configuration blocks, or both different numbers and sizes of configuration blocks.

[0047] Optionally, the registers 110 can be evenly divided using a register variable unit to obtain at least one configuration block, where each configuration block in the at least one configuration block is of the same size; or the at least one configuration block can be evenly divided to obtain at least one set of configuration blocks, where each set of configuration blocks in the at least one set of configuration blocks includes the same number of configuration blocks. That is, even division of the multiple registers 110 means dividing the multiple registers 110 into at least one configuration block, where each configuration block in the at least one configuration class is of the same size; even division of the at least one configuration block means dividing the at least one configuration block into at least one set of configuration blocks, where each set of configuration blocks in the at least one set of configuration blocks includes the same number of configuration blocks. Alternatively, the registers 110 can be evenly divided using a register variable unit to obtain at least one configuration block; or the at least one configuration block can be non-uniformly divided to obtain at least one set of configuration blocks, where two sets of configuration blocks in the at least one set of configuration blocks include different numbers of configuration classes. That is, non-uniform division of the at least one configuration block means dividing the at least one configuration block into at least one set of configuration blocks, where two sets of configuration blocks in the at least one set of configuration blocks include different numbers of configuration classes. Alternatively, the registers 110 can be non-uniformly divided using variable units to obtain at least one configuration block, where two configuration blocks of different sizes exist within the at least one configuration block; then, the at least one configuration block can be uniformly divided to obtain at least one set of configuration blocks. In other words, non-uniform division of the multiple registers 110 means dividing the multiple registers 110 into at least one configuration block, where two configuration blocks of different sizes exist within the at least one configuration block. Alternatively, the registers 110 can be non-uniformly divided using variable units to obtain at least one configuration block; then, the at least one configuration block can be non-uniformly divided to obtain at least one set of configuration blocks.

[0048] Optionally, the specific partitioning method for multiple registers can be determined based on multiple factors such as the specific architecture of the processor 100, the type of task currently being executed, and the user's subjective settings. This application embodiment does not limit this.

[0049] Task assembly unit 130 is used to allocate a first configuration block set to a first thread group (wave), the first configuration block set being one of at least two configuration block sets.

[0050] Among them, the task assembly unit 130 selects the first configuration block set from at least two configuration block sets for the first thread group.

[0051] Optionally, the task assembly unit 130 allocates a first set of configuration blocks to the first thread group based on the tasks executed by the first thread group. The tasks executed by the first thread group are used to determine at least one of the number of configuration blocks included in the configuration block set and the size of the configuration blocks.

[0052] Alternatively, the task assembly unit may also be called a thread assembly and resource usage control unit, or a program dispatch scheduler, etc.

[0053] It should be noted that the names of the various units shown in the embodiments of this application are for illustrative purposes only. In actual implementation, other names may be used to refer to the corresponding units. Furthermore, the unit division method in the embodiments of this application is also for illustrative purposes only. In actual implementation, the number of units, functional boundaries, etc., can be adjusted. Specific adjustments include merging, splitting, renaming, etc. For example, the aforementioned resource usage management unit and task assembly unit can also be adjusted into a single task assembly and resource usage control unit. The embodiments of this application do not limit this, but the scope of protection of the embodiments of this application is not limited thereto.

[0054] In summary, the processor provided in this application first divides multiple registers into at least one set of configuration blocks, with each set of configuration blocks used to allocate to thread groups. By dividing register resources into multiple configuration blocks and allocating them in the form of configuration blocks, each thread group can be allocated at least one configuration block (i.e., one set of configuration blocks). Compared to the traditional method of allocating registers continuously at the register level, this method of allocating register resources directly to configuration blocks instead of multiple consecutive registers avoids frequent splitting of consecutive registers, thus preventing register fragmentation. Furthermore, allocating through configuration blocks allows for management based on configuration blocks rather than registers, improving the efficiency of register resource allocation and release.

[0055] In some embodiments, the plurality of registers 110 are a plurality of scalar registers, or a plurality of vector registers, or a plurality of scalar registers and a plurality of vector registers, etc. It should be noted that the execution logic of the processor 100 shown in the embodiments of this application is applicable to different types of registers. In addition to the scalar registers and vector registers shown above, other types of registers may also be supported, and the embodiments of this application do not limit this.

[0056] The following section will explain the specific partitioning method for registers.

[0057] 1. Register partitioning.

[0058] In some embodiments, the register management unit 120 is configured to determine quantity information based on register usage configuration information, the quantity information including at least one of the following: number of configuration blocks, number of active thread groups, first data volume, and second data volume; and to divide the plurality of registers 110 into at least one set of configuration blocks based on the quantity information; wherein the number of configuration blocks is used to indicate the number of configuration blocks in each of the at least one set of configuration blocks, the number of active thread groups is used to indicate the maximum number of thread groups that the processor 100 supports for parallel execution, the first data volume is used to indicate the maximum amount of data that each of the at least one set of configuration blocks can support for storage, and the second data volume is used to indicate the maximum amount of data that each configuration block can support for storage.

[0059] Optionally, the register usage configuration information is set by the user or developer. The register usage settings can be modified through interfaces, methods, or parameters exposed by the designer of the processor 100, or through software designed by the designer.

[0060] Optionally, the quantity information corresponding to multiple registers 110 is determined based on the register usage configuration information; or, the quantity information corresponding to a portion of the multiple registers 110 is determined based on the register usage configuration information. That is, it can be said that the register usage configuration information is used to configure and divide a portion of the multiple registers 110.

[0061] Optionally, register usage configuration information is used to indicate or identify at least one of the following: number of configuration blocks, number of active thread groups, first quantity, and second data quantity.

[0062] Optionally, the multiple registers 110 can be divided into at least one set of configuration blocks, and each set in the at least one set of configuration blocks supports the partitioning configuration of configuration blocks through a register usage configuration information.

[0063] For example, processor 100 includes at least one sub-processor, and each of the at least one sub-processor 100 corresponds to a set of configuration blocks. Each sub-processor can divide the registers within the at least one set of configuration blocks corresponding to the sub-processor into at least one set of configuration blocks based on register usage configuration information.

[0064] For example, if processor 100 is a graphics processing unit (GPU), then the subprocessor can be a streaming multiprocessor (SM); or, if processor 100 is a GPU, then the subprocessor can be a graphics processing unit (SP); or, if processor 100 is a streaming multiprocessor 100, then the subprocessor can be a streaming multiprocessor 100.

[0065] Optionally, the multiple registers 110 included in the processor 100 are registers used to allocate to thread groups to implement instruction execution, or the multiple registers 110 can be thread-private registers. That is, in addition to the multiple registers 110 shown above, the processor 100 may also include other types of registers. These different types of registers have different uses, such as storing information shared by multiple thread groups, storing instructions corresponding to thread groups, etc.

[0066] For example, the register management unit 120 is used to query the register configuration information table based on the register usage configuration information to determine the quantity information. The register configuration information table is used to indicate the mapping relationship between the register usage configuration information and the quantity information.

[0067] Optionally, the register configuration information table can be pre-set or set by the user.

[0068] For example, a register configuration information table is shown in Table 1 below. The quantity information includes the number of configuration blocks, the number of active thread groups, and the first data volume. Register usage configuration information can also be called quantity information identifiers, or enumeration type information. An enumeration type is a data type used to define a set of named constants (i.e., a finite, predefined set of options). Enumeration constants typically have fixed, explicit meanings, and their values ​​are usually integers (but can also be other types, such as strings).

[0069] Table 1 Register Configuration Information Table

[0070]

[0071] Optionally, the number of registers in processor 100 is fixed, or the number of registers in subprocessor is fixed.

[0072] Optionally, all configuration blocks in processor 100 can be allocated to thread groups; or, some configuration blocks in processor 100 may be allocated to thread groups. Typically, thread groups executing in processor 100 will equally share the configuration blocks of processor 100, or thread groups executing in a subprocessor will equally share the configuration blocks of the subprocessor. That is, thread groups belonging to the same processor 100 or subprocessor correspond to the same number of configuration blocks (or registers). Therefore, for some configuration block quantities, there may be situations where the remaining configuration blocks cannot be allocated to thread groups. For example, in the cases of configuration block quantities of 6 and 13 in Table 1 above, there are two idle configuration blocks in processor 100 that will not be allocated to thread groups.

[0073] Optionally, different register configuration information tables can be set for different amounts of second data. For example, Table 1 above is a register configuration information table for a subprocessor. If the subprocessor includes 2560 registers, and each register supports storing 1 DW of data, then in the scenario shown in Table 1, the amount of second data is 32 DW, that is, each configuration block includes 32 registers. In this case, a subprocessor with a second data amount of 16 DW can exist in the processor 100. This subprocessor corresponds to another register configuration information table. It should be understood that the amount of second data can also be used as a parameter in the register configuration information table, as exemplarily shown in Table 2 below.

[0074] Table 2 Register Configuration Information Table

[0075]

[0076] Optionally, the processor 100 may use one or more second data volumes when dividing the configuration blocks. That is, it may obtain configuration blocks of the same size or configuration blocks of different sizes.

[0077] Optionally, the register usage configuration information is typically related to the tasks performed by the processor 100. For example, for tasks with high register usage, such as AI computation, model training, and rendering of complex scenes, more registers can be allocated to a thread group, meaning a register usage configuration with more configuration blocks can be used. However, since the total number of registers in the processor 100 is constant, allocating more registers to a thread group will correspondingly reduce the number of executable thread groups in the processor 100, i.e., the number of active thread groups will decrease accordingly. For tasks with low register usage, such as simple graphics rendering and simple mathematical function calculation, fewer registers can be allocated to a thread group, i.e., a register usage configuration with fewer configuration blocks can be used, and the number of active thread groups in the processor 100 will increase accordingly.

[0078] Optionally, when the processor 100 includes multiple subprocessors, the register usage configuration information can be set separately for each subprocessor. That is, each subprocessor can determine the division of registers in the configuration block set corresponding to the subprocessor through a register usage configuration information.

[0079] Optionally, when the processor 100 includes multiple subprocessors, the number of registers included in each of the multiple subprocessors may be the same or different.

[0080] For example, taking multiple sub-processors that include the same number of registers as an example, for the first sub-processor, its corresponding first register usage configuration information determines the following quantities: {number of configuration blocks: 5, number of active thread groups: 16, first data volume: 160, second data volume: 32}; for the second sub-processor, its corresponding second register usage configuration information determines the following quantities: {number of configuration blocks: 10, number of active thread groups: 8, first data volume: 320, second data volume: 32}; for the third sub-processor, its corresponding third register usage configuration information determines the following quantities: {number of configuration blocks: 16, number of active thread groups: 5, first data volume: 512, second data volume: 32}.

[0081] In summary, the processor provided in this application embodiment shows that the division of multiple registers in the processor is determined based on register usage configuration information. That is, it supports flexible division of registers according to specific needs by setting register usage configuration information, so that the resulting configuration blocks and configuration block sets can meet the current task requirements and improve register utilization.

[0082] Furthermore, register usage configuration information is used to indicate the corresponding quantity information in the register configuration information table, avoiding fragmentation or performance issues that may result from completely open user-defined configuration blocks (such as unreasonable quantity information set by the user). By predefining reasonable quantity information options, users can choose the closest option according to their needs, while avoiding the uncontrollability of completely free configuration.

[0083] Furthermore, the processor described above allows for different register usage configurations to be applied to different types of tasks. For example, in scenarios with high usage of Type I data (constant or common data, etc.), a larger granularity of register partitioning can be used, such as including more configuration blocks and thus more registers in a configuration block set. However, this will correspondingly reduce the number of thread groups that can run in parallel (i.e., the number of active thread groups). But this flexible configuration method allows the processor's registers to maintain a high utilization rate in various scenarios.

[0084] The following section will explain in detail how thread groups use configuration block sets.

[0085] 2. Loading registers.

[0086] In some embodiments, the processor 100 further includes a data loading unit 140, such as Figure 2 As shown; the data loading unit 140 is used to load at least one first data into a first part register in the first configuration block set. The first part register is a specified number of registers or a specified proportion of registers in the first configuration block set. The first data belongs to a first type, and the data of the first type includes at least one of constants and common data.

[0087] Optionally, the first part of the register includes at least one register, and the at least one register has a one-to-one relationship with at least one first data.

[0088] Optionally, the first set of registers is a specified number or a specified proportion of registers in the first configuration block set. This specified number or proportion can be pre-set or user-defined. Pre-setting refers to a method where the processor designer determines and embeds it in the hardware logic during processor design, taking effect after power-on or reset. User-defined setting refers to the user setting it through interfaces or methods provided by the designer. For example, the user designs a specified number or proportion based on the currently executed task. For instance, for graphics rendering tasks, some first-type data are typically needed, such as model transformation matrices (world matrix, view matrix, projection matrix, etc.), lighting parameters (light source position, color, intensity), and material properties (reflectivity, transparency), etc. Besides these first-type data, there are other types of data, such as second-type data, which are data that are updated during calculation, such as vertex data and texture data. Therefore, the user needs to determine the number of first-type registers in the processor 100 based on the specific task being executed. This register allocation method is beneficial for maximizing the utilization of register resources and improving the access efficiency of different types of data, thus optimizing task performance.

[0089] Alternatively, the specified quantity and specified ratio can also be determined by the processor 100 itself based on the task currently required to be executed.

[0090] Optionally, the first type of data includes at least one of constants and public data. Constants are data that cannot be changed during program execution, or data that cannot be changed during the execution of a thread group, or data that cannot be changed during the execution of a task. Public data refers to data shared by the thread group, that is, data shared by all or some of the threads in the thread group. If all threads need the same common storage space (such as a region or buffer in memory), then the starting address of that common storage space can be considered as public data, i.e., data of the first type.

[0091] Optionally, in addition to storing the first type of data, registers can also be used to store the second type of data, such as variables. Typically, the second type of data is thread-private within the thread group.

[0092] For example, the first configuration block set includes 160 registers, each register corresponding to 1 DW (32 bits) of data. In this case, the specified number or proportion of registers in the first part can be determined based on the actual data volume of the first type and the second type required by the first thread group. For instance, if the first part consists of 128 registers, the maximum data volume allocated to the first type is 128 DW; and the second part consists of the remaining 32 registers, the maximum data volume allocated to the second type is 32 DW.

[0093] The processor described above loads first-type data (i.e., at least one piece of first data) into the first part of the registers in the configuration block set. Compared to related technologies that use dedicated shared registers or constant caches for constants, this ensures register generalization, avoids wasting register resources due to some tasks not using dedicated registers or caches, and improves the overall utilization of registers. Furthermore, without setting up dedicated registers or constant caches, the processor area can be minimized. In addition, the general-purpose registers support allocating a portion of the registers (such as the first part of the registers) to store first-type data, and can also achieve the advantage of dedicated registers protecting critical data (such as constants and common data).

[0094] In related technologies, the thread group scheduling process is usually executed first, and data is loaded before or during the execution of a specific instruction. However, in this embodiment, a large block of first-type data is usually preloaded, and the data type of the first-type data preloaded in this embodiment differs from the data type loaded in related technologies. Therefore, the traditional thread group scheduling process and data initialization loading process should be updated to adapt to the register allocation and usage method shown in this embodiment. Specifically, the loading timing of at least one piece of first data into the first part of the registers is as follows.

[0095] (1) Loading timing.

[0096] In some embodiments, the process of loading the first data into the first part of the registers can be referred to as the initialization process. To improve the execution efficiency of the thread group in the processor 100, the initialization process can be executed in parallel with the thread group's scheduling process. Compared to the related art where the thread group's scheduling process is completed first, and then the initialization process (also referred to as the loading process) is performed on the thread group's registers, executing the scheduling process and the initialization process in parallel helps to shorten the thread group's execution time, improve the thread group's execution efficiency, and thus reduce the overall latency of task execution. For example, the processor 100 further includes a scheduler; the scheduler is used to send a scheduling request, which triggers the scheduling process of the first thread group, the scheduling process referring to the process of scheduling the execution of the first thread group; the task assembly unit 130 is used to send an initialization loading request, which triggers the initialization process of the first configuration block set, the initialization process referring to the initialization of the first part of the registers in the first configuration block set; wherein the scheduling process and the initialization process are executed in parallel.

[0097] Optionally, the scheduling request sent by the scheduler is used to trigger the scheduling process of the first thread group. For example, the scheduler sends a scheduling request to the thread group management unit. The thread group management unit selects an instruction and issues it to the execution unit (such as the processing core in processor 100) based on the current execution status of the thread group and the dependencies between instructions, and removes data dependencies and resource dependencies for the instruction to ensure correct execution. During the execution of the instruction, the execution unit or other units initiate a data loading request to read data from the storage unit and load it into the second part of the registers of processor 100. The second part of the registers consists of the registers in the first configuration block set other than the first part of the registers.

[0098] Optionally, before the initialization process is complete, the thread group management unit will prioritize scheduling instructions that do not use the first type of data, or instructions whose usage of the first type of data is below a threshold. That is, the thread group management unit will usually prioritize scheduling instructions that only use the second type of data in the early stages of thread group scheduling to avoid wasting register resources due to repeated loading of some first type of data.

[0099] Optionally, the initialization loading request sent by the task assembly unit 130 is used to trigger the initialization process of the first configuration block set. The initialization process is the process by which the data loading unit 140 loads at least one piece of first data into the first part of the registers in the first configuration block set, which will not be described in detail here.

[0100] Optionally, before and during the initialization process executed by the data loading unit 140, no data dependency check is performed on at least one piece of first data. That is, before or during the loading of the first data, it is not necessary to perform a data dependency check on the first data based on the instructions included in the first thread group, nor is it necessary to determine which instructions executed by the first thread group need to use the first data. Since the first data is data of the first type, it is typically at least one of constants and public data. Constants are data that cannot be modified during the execution of the thread group, so constants typically do not require the aforementioned data dependency check. As for public data, it is typically information such as a public starting address; this type of information also does not require a data dependency check before execution.

[0101] Optionally, before executing the first instruction, the processor 100 adds a data dependency check instruction. The first instruction refers to an instruction whose operands contain data of the first type. The data dependency check instruction is used to detect whether the data of the first type required by the first instruction has been loaded into the first configuration block set. That is, the added data dependency check instruction is executed before the first instruction is executed to ensure the correct execution of the first instruction.

[0102] In the processor described above, the scheduling process of the thread group and the initialization process of the first part of the registers are executed in parallel. This is because the data loaded in the first part of the registers is of the first type, rather than all the registers required by the thread group. Therefore, it is possible to execute the scheduling process of the thread group and the initialization process of the first part of the registers in parallel, thereby hiding the delay in loading the part of the data corresponding to the thread group and improving the efficiency of the thread group execution.

[0103] In some embodiments, the actual amount of data required by the thread group is large, such as the configuration block set having a data size of 160 DW, but the actual amount of data required by the thread group is greater than the data size of the configuration block set or the data size of the first part register. In this case, dynamic loading can be used to solve the problem, as shown below.

[0104] (2) Dynamic loading.

[0105] In some embodiments, the data loading unit 140 is used to load at least one second data into at least one register of the first configuration block set during the execution of the first thread group, wherein the second data belongs to a first type and there is a one-to-one relationship between the at least one register and the at least one second data.

[0106] Optionally, during the execution of the first thread group, the data loading unit 140 determines the first type of data required by at least one instruction, namely at least one second data, and loads the at least one second data into at least one register of the first configuration block set.

[0107] Optionally, in addition to the first type of data mentioned above, a second type of data may also be loaded, but this second type of data is typically loaded into the second set of registers. The second set of registers refers to the registers in the first configuration block set other than the first set of registers.

[0108] Optionally, at least one register in the first configuration block set may include at least one of the registers in the first part register and the registers in the second part register. That is, at least one second data may be loaded into the first part register, or into the second part register, or partly loaded into the first part register and partly loaded into the second part register.

[0109] For example, the data loading unit 140 is used to load at least one second data into at least one register of the first part register during the execution of the first thread group, wherein the data in the at least one register is data that the first thread group has finished using, and the at least one second data is used to overwrite the first data in the at least one register.

[0110] Optionally, the data in at least one register is data that the first thread group has finished using, or in other words, the first thread group has finished using data of the first or second type in at least one register. In this case, at least one second type of data is used to overwrite the data in at least one register.

[0111] Optionally, the data in at least one register can be the first data loaded during initialization or the second data loaded during the previous dynamic loading process.

[0112] For example, the data loading unit 140 is used to load at least one second data into at least one register of a second part register during the execution of the first thread group, wherein the second part register is a register in the first configuration block set other than the first part register; wherein the register in the at least one register satisfies at least one of the following characteristics: the data in the register is data that has been used by the first thread group; the register is an empty register.

[0113] Optionally, the data in the register is data that the first thread group has finished using, meaning the first thread group has finished using the first or second type of data in the register. In this case, the second data is loaded into the register to overwrite the data originally stored in the register.

[0114] Optionally, the register is an empty register, which means a register that does not store valid data, i.e., data stored due to the execution of the first thread group; or, in other words, an empty register is an uninitialized register. In this case, the second data is directly loaded into this register. Optionally, the empty register may store data loaded by the previous thread group, or all-zero data reset due to a clear operation.

[0115] In other embodiments, the data loading unit 140 is configured to load at least one second data into at least one register in the first part register and the second part register during the execution of the first thread group. For example, the first configuration block set includes 160 registers, the first part register includes 128 registers, the second part register includes 32 registers, and the at least one second data is 32 second data. The data loading unit 140 determines the registers for loading the second data based on the usage of each register in the first configuration block set. For example, if the data in 24 registers in the first part register has been exhausted and the data in 16 registers in the second part register has been exhausted, the 24 registers in the first part register are preferentially selected. Since the 24 registers in the first part register are less than the number of second data, 8 registers in the second part register are also selected for loading the second data.

[0116] In some embodiments, the dynamic loading process described above is triggered by a dynamic loading instruction, which instructs the loading of data of the first type into the first configuration block set during the execution of the first thread group. Optionally, the dynamic loading instruction loads the data of the first type required by the first instruction into the first configuration block set before the first instruction is executed, where the first instruction is the instruction corresponding to the first thread group; or, the dynamic loading instruction loads the data of the first type required by at least one instruction into the first configuration block set before the execution of at least one instruction, where the at least one instruction is the instruction corresponding to the first thread group. Optionally, the at least one instruction may be one or more instructions that the first thread group is about to execute, and the number of at least one instructions may be preset or configured by the user.

[0117] The processor described above supports dynamically loading the first type of data during the execution of the first thread group. This avoids the problem that the number of threads that can be executed in parallel in the processor is too small due to the large amount of the first type of data required by the first thread group. This ensures that the processor has a certain degree of parallelism. In particular, for GPUs, this avoids too many computing cores in the processor being idle, which would cause the processor to lose its advantages.

[0118] Furthermore, during dynamic loading, it supports loading data of type 1 into the first part of the registers, or data of type 1 into the second part of the registers, or data of type 1 into both the first and second part of the registers. That is, it supports dynamic loading of data of type 1 into the second part of the registers during initialization, thereby ensuring the parallel execution of the thread group's scheduling and initialization processes and improving the execution efficiency of the thread group. On the other hand, it also supports loading data of type 1 into the first part of the registers based on the execution status of the thread group, overwriting data that the thread group has finished using in the first part of the registers, reusing limited registers, alleviating resource shortages caused by the processor's hardware size, and indirectly reducing the demand for register resources within the processor, thereby reducing the processor's area.

[0119] The following section will take the first configuration block set, which includes n first registers (registers for storing data of the first type) and m second registers (registers for storing data of the first type and the second type), as an example to explain the above initialization process and dynamic loading process in detail.

[0120] In some embodiments, when the amount of first type data required by the first thread group is less than or equal to the number of first registers, the loading of the first type data of the first thread group can be completed simply through initial loading. That is, the data loading unit 140 is used to load a first number of first data items into a first number of first registers based on an initial loading request when the first number is less than or equal to n, where the first number is the amount of first type data corresponding to the first thread group.

[0121] Optionally, if the amount of first type data required by the first thread group is less than or equal to the number of first registers, the remaining first registers can be left empty or used to load second type data. The remaining first registers refer to the registers among the n first registers that were not loaded with first data during the initial process.

[0122] In some embodiments, when the amount of data of the first type required by the first thread group is greater than the number of first registers, initial loading and dynamic loading should be used to implement the execution of the first thread group. Specifically, the initial loading process is as follows: The first configuration block set includes n first registers, which are used to store data of the first type, where n is a specified number or determined by a specified ratio; the data loading unit 140 is used to load n pieces of first data into the n first registers based on an initial loading request when the first number is greater than n, where the first number is the amount of data of the first type corresponding to the first thread group.

[0123] Optionally, the n first data points are a subset of the first number of data points. The first number of data points refers to the first type of data corresponding to the first thread group.

[0124] Optionally, this initialization process can be executed in parallel with the thread group's scheduling process.

[0125] Optionally, the n first registers are the aforementioned first part registers, that is, the first part registers include n registers.

[0126] Optionally, n first data items are loaded into n first registers; that is, each of the n first data items is loaded into one of the n first registers. In other words, there is a one-to-one relationship between the n first data items and the n first registers.

[0127] In some embodiments, before the initialization process is complete, the thread group management unit typically schedules and executes instructions in the first thread group that do not require data of the first type. Alternatively, before the initialization process is complete, the thread group management unit may load data of the first type into a second register to implement the execution of the first thread group.

[0128] In some embodiments, after the initialization process ends, the remaining data in the first number of data corresponding to the first thread group can be dynamically loaded into registers. For example, the processor 100 further includes an instruction control unit; the instruction control unit is configured to execute a dynamic loading instruction when the first number is greater than n, the dynamic loading instruction instructing that data of the first type be loaded into the first configuration block set during the execution of the first thread group; and a data loading unit 140 is configured to load k pieces of second data into k registers based on the dynamic loading instruction, where k is a positive integer, and the k registers are free registers in the first configuration block set, the free registers being empty registers or registers whose data has been no longer used.

[0129] Optionally, k is less than or equal to the total number of registers in the first configuration block set. Optionally, k is less than or equal to the total number of free registers in the first configuration block set.

[0130] Optionally, k is related to the instructions to be executed by the first thread group. The k second data are the data to be used by the first thread group.

[0131] Optionally, the dynamic loading instruction is used to instruct the loading of data of the first type into the first configuration block set during the execution of the first thread group; or, the dynamic loading instruction is used to instruct the loading of data of the first type required by at least one instruction into the first configuration block set, wherein the at least one instruction is an instruction corresponding to the first thread group that has not yet been executed. Optionally, the at least one instruction is a second number of instructions that the first thread group is about to execute, wherein the second number can be preset or configured by the user.

[0132] Optionally, the dynamic loading instruction can also be used to instruct the loading of data required by at least one instruction into the first configuration block set, wherein the data required by the at least one instruction includes at least one of a first type of data and a second type of data.

[0133] Optionally, the k registers can be k first registers, or k registers can be k second registers, or the k registers can include i first registers and j second registers, where i and j are both positive integers and i+j=k.

[0134] For example, the data loading unit 140 is used to load k second data items into k first registers based on a dynamic loading instruction. The data in the k first registers is data that has been no longer used by the first thread group, and the k second data items are used to overwrite the data in the k first registers. Alternatively, based on a dynamic loading instruction, the k second data items are loaded into k second registers, and each register in the k second registers satisfies at least one of the following characteristics: the data in the register is data that has been no longer used by the first thread group; or the register is an empty register. Alternatively, based on a dynamic loading instruction, the k second data items are loaded into i first registers and j second registers, where i + j = k, and the selection conditions for the i first registers and j second registers are as described above for the k first registers and k second registers, and will not be repeated here.

[0135] The processor described above, when the first quantity is less than or equal to the number of first registers, loads the first type of data into the first registers using only the initial load method, thus loading the first type of data required by the first thread group. When the first quantity is greater than the number of first registers, it uses both initial load and dynamic load methods to load the first type of data into at least one of the first and second registers. This avoids deadlock issues caused by loading too much first type of data at once, which would prevent the thread group from loading other types of data. Furthermore, it supports loading the first type of data into a first part of the registers based on the execution status of the thread group, overwriting data that the thread group has finished using in the first part of the registers. This reuses limited registers, alleviating resource shortages caused by the processor's hardware size, and indirectly reducing the demand for register resources within the processor, thereby reducing the processor's area.

[0136] In addition to the configuration-based allocation of registers for each thread group and the combination of initial loading and dynamic loading for the first type of data as shown above, a method for multiple sub-processors to share a single write-back path is also designed to further reduce the area of ​​the processor 100, as detailed below.

[0137] 3. Shared return data path.

[0138] In some embodiments, the processor 100 further includes at least one data loading unit 140, a first cache, and a plurality of subprocessors; each data loading unit 140 corresponds to m subprocessors, the m subprocessors being all or part of the plurality of subprocessors included in the processor 100, and each of the m subprocessors including at least one set of configuration blocks; the data loading unit 140 includes a data transfer executor and a storage entry buffer; the data transfer executor is used to read s third data from the first cache using a first bandwidth; and to cache the s third data to a first first-in-first-out (FIFO) queue of the storage entry buffer, the first bandwidth being related to the cache line width of the first cache; the storage entry buffer is used to write the s third data from the first FIFO queue to at least two registers corresponding to the first subprocessor among the m subprocessors using a second bandwidth, the second bandwidth being related to the size of the configuration block; wherein, the first bandwidth is greater than the second bandwidth.

[0139] Optionally, the data loading unit 140 is used to load the first type of data into the configuration block set. Alternatively, the data loading unit 140 is used to load the first type of data into registers within the configuration block set.

[0140] Optionally, the data transfer executor uses a first bandwidth to read s third data from the first cache, the first bandwidth being related to the cache line width of the first cache.

[0141] Optionally, each cache line in the first cache is used to cache s third data items, or multiple cache lines in the first cache are used to cache s third data items.

[0142] Optionally, the data transfer executor uses the first bandwidth to read *s* third data items from the first cache in one go, or within one cycle. Alternatively, the data transfer executor uses the first bandwidth to read *s* third data items from the first cache over multiple cycles.

[0143] Optionally, the third data belongs to the first type; or, the third data belongs to the second type; or, the s third data include data belonging to the first type and data belonging to the second type.

[0144] Optionally, a FIFO queue is a data structure that processes data according to the time sequence. For example, if a cache line consists of 256 bits, and the cache line is stored in the FIFO queue in the order of bit 0 to bit 255, then when writing data to the register through the FIFO queue, it will also be written sequentially starting from bit 0, and finally bit 255.

[0145] Optionally, the storage entry buffer uses the second bandwidth to write s third data items from the first FIFO queue to the first subprocessor among the m subprocessors. If the second bandwidth corresponds to t third data items, then the storage entry buffer uses the second bandwidth to read t third data items from the first FIFO queue and writes t third data items to t registers; then it again uses the second bandwidth to read t third data items from the first FIFO queue and writes t third data items to t registers.

[0146] For example, such as Figure 3 As shown, assuming the first cache is a level 1 cache, the data transfer executor uses the first bandwidth to read the data 30 corresponding to the first bandwidth from the first cache and caches it in the first FIFO queue 40 of the storage entry buffer. The storage entry buffer then uses the second bandwidth to read the data 31 corresponding to the second bandwidth from the first FIFO queue and loads it into the register corresponding to the first subprocessor. If the first bandwidth is 16 times the second bandwidth, that is, if the first bandwidth is 128 DW / cycle, then the second bandwidth is 8 DW / cycle. In this case, it takes 16 cycles to load all the data in the first FIFO queue into the first subprocessor. If the storage entry buffer corresponds to only one subprocessor, then the data transfer executor needs to be idle for at least 15 cycles, waiting for the storage entry buffer to load the data 30 corresponding to the first bandwidth into the register of the first subprocessor. This will waste the performance of the data transfer executor. Furthermore, when there are multiple subprocessors in the processor 100, a corresponding return data path (i.e., data transfer executor-storage entry buffer, or data loading unit) needs to be set up for each subprocessor, which will greatly increase the area of ​​the processor 100. Therefore, multiple FIFO queues can be set up in the storage entry buffer, connecting multiple subprocessors simultaneously. When the first FIFO queue is writing data to the registers of the first subprocessor, the data transfer executor can continue to read the data required by the second subprocessor from the first buffer and load it into the second FIFO queue 41. This reduces the waiting cycle of the data transfer executor and improves its utilization. Similarly, a larger number of subprocessors can be set up for the storage entry buffer to ensure that the data transfer executor is not idle as much as possible, such as setting up more than 16 FIFO queues, i.e., connecting more than 16 subprocessors.

[0147] Optionally, the storage entry buffer includes m FIFO queues, each FIFO queue corresponding to a subprocessor.

[0148] Optionally, the second bandwidth is related to the size of the configuration block. For example, if the configuration block is arranged in i rows and j columns, the second bandwidth can be the data bandwidth corresponding to the j columns, or in other words, the second bandwidth supports writing an entire row (i.e., j data) at once. Alternatively, the second bandwidth can be related to the size of the third data. For instance, the second bandwidth is a multiple of the number of bits corresponding to the third data.

[0149] Optionally, one data loading unit 140 may correspond to all or some of the subprocessors in the processor 100. The number of subprocessors corresponding to one data loading unit 140 is flexibly set.

[0150] For example, if processor 100 is a graphics processing unit (GPU), then the subprocessor can be a streaming multiprocessor (SM); or, if processor 100 is a GPU, then the subprocessor can be a graphics processing unit (SP); or, if processor 100 is a streaming multiprocessor 100, then the subprocessor can be a streaming multiprocessor 100.

[0151] In summary, the processor provided in this application embodiment demonstrates a processor in which multiple subprocessors share a single return data path. This approach can both guarantee the overall register write bandwidth of the processor and minimize area overhead. It achieves the effect of a large bandwidth entering the data loading unit and multiple smaller bandwidths outputting from the data loading unit, enabling fast writing without congesting the pipeline on the return data path.

[0152] Since this application embodiment uses configuration blocks to configure register resources, the processor 100 should make corresponding changes when accessing registers, as shown below.

[0153] 4. Address access process.

[0154] In some embodiments, the processor 100 is configured to, in response to an access request for a register to be accessed, determine the physical address of the register to be accessed based on the identifier of the thread group corresponding to the register to be accessed and the logical address of the register to be accessed, wherein the physical address is used to indicate that the register to be accessed is located in the b-th row and c-th column of the a-th configuration block, where a, b, and c are all positive integers.

[0155] In this configuration, the registers in each configuration block are arranged in rows i and columns j, where i and j are both positive integers. b is less than or equal to i, and c is less than or equal to j.

[0156] Optionally, the logical address of the register to be accessed refers to the relative address of the register within the configuration block set corresponding to the thread group. That is, different thread groups may have registers with the same logical address within their respective configuration block sets, while the same thread group typically does not have registers with the same logical address within its own configuration block set. In other words, the logical address of the register to be accessed is used to identify the register within the configuration block set corresponding to the thread group.

[0157] Optionally, the physical address of the register to be accessed refers to the actual physical location of the register in all registers of the processor 100 or the subprocessor, and is usually represented by a combination of row address and column address. Of course, in some special scenarios, a complete physical address can also be used. This complete physical address is usually related to the row address and column address. For example, the complete physical address is obtained by concatenating the row address and column address. If the row address is 0x0001 and the column address is 0x0101, then the complete physical address can be represented as 0x0001 0101. Of course, other concatenation or mapping methods can also be used, and this application embodiment does not limit this.

[0158] Optionally, the access request for the register to be accessed can be a read request for the register to be accessed, a write request for the register to be accessed, or other operation requests for the register to be accessed. This application embodiment does not limit this.

[0159] Optionally, the processor 100 determines the physical address of the register to be accessed based on the identifier of the thread group corresponding to the register and the logical address of the register. The thread group identifier is used to determine the starting position of the configuration block set. The logical address of the register to be accessed can be used to determine the position of the register in the configuration block set. This position can be broken down into the starting position of the configuration block in the configuration block set, the row position of the register in the multiple registers 110, and the column position. The starting position of the configuration block in the configuration block set can be considered as the position of the configuration block containing the register to be accessed within the configuration block set.

[0160] For example, processor 100 is configured to, in response to an access request for a register to be accessed, determine the identifier of the configuration block corresponding to the register to be accessed based on the identifier of the thread group corresponding to the register to be accessed and the logical address of the register to be accessed; determine the row start address corresponding to the configuration block based on the identifier of the configuration block; determine the row offset address and column address corresponding to the register to be accessed based on the logical address of the register to be accessed; determine the row address of the register to be accessed based on the row start address and the row offset address; and determine the physical address of the register to be accessed based on the row address and the column address.

[0161] For example, the logical address of a register is used to indicate the relative address of the register within the configuration block set corresponding to the thread group. For instance, a register with a logical address of 0x0000 0000 indicates that the register is R0 (the first register) in the configuration block set corresponding to the thread group; a register with a logical address of 0x0010 0110 indicates that the register is R70 (the 71st register) in the configuration block set corresponding to the thread group; and a register with a logical address of 0x0011 0100 indicates that the register is R52 (the 53rd register) in the configuration block set corresponding to the thread group. Therefore, the meaning of each bit in the register's logical address is related to factors such as the size of the configuration block set corresponding to the thread group and the size of the configuration block. Taking an example where the register's logical address represents the column index, row index, and configuration block index from low to high, and without any invalid bits, such as... Figure 4 As shown, the configuration block is arranged in a 4x8 format, with the lower logical address 10... (= =3) The bit represents column index 14, and j is the column number of the configuration block; the th +1 (= +1=4) to the 4th position + (= + =3+2=5) bits represent row index 13, where i is the row number of the configuration block; the remaining high bits represent configuration block index 12. Based on the configuration block index 12, row index 13, and column index 14, the physical location 11 of the register in the configuration block set can be determined. Because the registers are arranged as follows... Figure 4 As shown, regardless of the number of configuration blocks, the column address 17 of each register is always its column index; while the row address of each register needs to determine the starting address of the configuration block set, that is, the starting address of the configuration block with index 0 in the configuration block set. Based on this starting address, the row starting address 15 of the configuration block corresponding to the logical address can be obtained. Finally, the row address 16 is obtained based on the row starting address 15 and the row index 13. The row address 16 and the column address 17 together form the physical address 18.

[0162] It should be noted that the above-described method of converting logical addresses to physical locations (i.e., using logarithmic calculations) is for illustrative purposes only. Since logarithmic operations are expensive, other methods can be used in actual implementations, such as table lookup or bitwise shift operations, which have lower overhead. For example, ... Figure 4 As shown in code block 19, the configuration block index is determined by looking up a table and the row and column addresses are determined by shift operations.

[0163] For example, taking the arrangement shown above as an example, the conversion process between logical address 10 and physical address 18 can be referred to code block 19. Here, ">>" is the right shift symbol. If the logical address is 8 bits, then shifting right by 5 bits means taking the high 3 bits of the logical address. This can also be understood as dividing the logical address by 32 (i.e., the configuration block size) to determine the configuration block index corresponding to that logical address. That is, "waveid + (logic_scalar_addr >> 5) to locate chunk_idx" means determining the configuration block index based on the thread group identifier (waveid) and the logical address (logic_scalar_addr), thus determining the configuration block identifier (chunkidx). If the configuration block is arranged in other forms, such as row i and column j above, the number of bits that the configuration block index should be shifted right by to determine the logical address will change. Specifically, it should be shifted right... Bits. It should be understood that the configuration block index mentioned above refers to the index of the configuration block corresponding to the logical address within the configuration block set. That is, if the number of configuration blocks is 5, the value range of this configuration block index is 0-4, requiring 3 bits to represent. The configuration block identifier, on the other hand, indicates the identifier of the configuration block within all configuration blocks of the processor or subprocessor. For example, if the number of configuration blocks corresponding to the subprocessor is 5 and the number of active thread groups is 16, then the value range of the configuration block identifier is 0-79. This configuration block identifier can be understood as the row start address, or the starting address of the row with configuration block index 0 within the configuration block. After determining the row start address, the row address (line_addr) in the physical address can be determined based on the row start address and the 4th (b3) and 5th (b4) bits of the logical address. The corresponding calculation formula is "line_addr = (chunk_idx << 2) + ((logic_scalar_addr >> 3) & 0x3)", where "(logic_scalar_addr >> 3) & 0x3" is used to extract the 4th (b3) and 5th (b4) bits of the logical address; and "chunk_idx << 2" indicates that the row start address and the row offset address (or row index) are concatenated to determine the row address. Since the row index is represented by 2 bits in the current scenario, the identifier of the configuration block (chunk_idx) only needs to be shifted left by 2 bits. If the configuration block is arranged in other forms, such as row i and column j above, the identifier of the configuration block should be shifted left accordingly. The row index will also be retrieved in a different way; in this case, the row index should be retrieved from the logical address at the specified bit. +1 position to the + Similarly, since multiple configuration blocks are stacked column-wise (or horizontally), the column index is the column address (bank_addr). This means that only the lower 3 bits of the corresponding column index in the logical address need to be retrieved, such as "bank_addr = logic_scalar_addr & 0x7". It should be understood that if the configuration blocks are arranged in other ways, such as row i, column j as described above, then the column address here should be the lower 3 bits of the logical address. Bit.

[0164] For example, such as Figure 4 Register 20 in the configuration block has a logical address of 0x00110100, with a configuration block index of 1, a row index of 2, and a column index of 4. If... Figure 4 The starting address of the row corresponding to chunk0 is 0x0000 0000 (usually, this starting address is obtained by looking up the thread group identifier), so the starting address of the row corresponding to chunk1 is 0x0000 0100; therefore, the row address corresponding to this register is 0x0000 0100 + 0x00000010 (i.e., 2) = 0x0000 0110 (i.e., the 6th row starting from 0); the corresponding column address is the column index, which is 0x00000100 (i.e., the 4th column starting from 0).

[0165] In summary, the processor provided in this application embodiment demonstrates a method for converting logical addresses to physical addresses based on the configuration block partitioning method shown in this application embodiment. This ensures that the conversion between logical addresses and physical addresses can be achieved based on this method, thereby guaranteeing the normal execution of the program.

[0166] To improve the efficiency of the initial loading and dynamic loading mentioned above, the data required by the thread group can be prefetched into the first cache, as shown below.

[0167] 5. Data prefetching and persistence.

[0168] In some embodiments, the processor 100 further includes a first cache; a task assembly unit 130, configured to send an initialization load request to the first cache; a first cache, configured to load data of a first type corresponding to a first thread group into the first cache based on the initialization load request; and to set the data of the first type to a persistent storage state, the persistent storage state being used to indicate that the data of the first type will remain cached in the first cache until the first thread group ends.

[0169] Optionally, both the initial loading and dynamic loading described above load data of the first type from the first cache. That is, the data loading unit is used to load at least one piece of first data from a first portion of registers in the first configuration block set in the first cache. The data loading unit is also used to load at least one piece of second data from the first cache into at least one register of the first configuration block set during the execution of the first thread group.

[0170] Optionally, the data of the first type is set to persistent storage, or in other words, the cache line storing the data of the first type is set to persistent storage.

[0171] Optionally, persistent storage means that the data of the first type will remain cached in the first cache until the first thread group ends; or, persistent storage means that the data of the first type in the first cache cannot be replaced by other data until the first thread group ends. It should be understood that this data of the first type is specific to the first thread group. That is, if another thread group has ended, the cache line used by that thread group to cache the data of the first type can be replaced.

[0172] Optionally, the first cache is the L1 cache in the processor or subprocessor.

[0173] In some embodiments, the processor 100 further includes a computing data master control unit and a second cache, the second cache being the next level cache of the first cache; the computing data master control unit is used to load data of the first type from the buffer into the second cache; the first cache is used to load data of the first type corresponding to the first thread group from the second cache into the first cache based on an initialization loading request.

[0174] Optionally, the main control unit for computing data triggers the loading of the first type of data from the buffer to the second buffer, where the buffer is a constant buffer obtained after the program is compiled.

[0175] Optionally, the second cache is a shared cache for multiple subprocessors, such as a level 2 cache or a last-level cache; the first cache is a private cache for the subprocessors, such as a level 1 cache.

[0176] In summary, the processor shown in this embodiment of the application can quickly read the corresponding first type of data from the first cache during the initial loading and dynamic loading processes by prefetching the first type of data into the first cache, thereby reducing the loading latency of the first type of data and improving the execution efficiency of the thread group.

[0177] Furthermore, the first type of data is loaded from the buffer to the second cache by the main data control unit, and then from the second cache to the first cache. This reuse of cache hierarchy in related technologies helps improve the processor's adaptability. In addition, the second cache serves as a shared cache, while the first cache serves as a private cache. This caching method is beneficial for supporting concurrency across multiple subprocessors and reducing access latency.

[0178] It should be noted that the above-mentioned "1. Register partitioning", "2. Register loading", "3. Shared return data path", "4. Address access process", and "5. Data prefetching and persistence" can be implemented as independent embodiments or as combined embodiments. This application does not limit the combination of the above methods; they can be combined in pairs, in groups of three, in groups of four, in groups of five, etc.

[0179] In some embodiments, the first configuration block is released when the first thread group finishes using the first configuration block in the first configuration block set. That is, embodiments of this application support releasing the register resources of the thread group in advance before the end of the lifecycle of the first thread group, so that the register resources can be allocated to other thread groups for use, thereby improving the overall register utilization rate.

[0180] Optionally, some registers may be released early if the first thread group finishes using some registers (such as one or more rows of registers) in the first configuration block set.

[0181] In parallel computing tasks, a task typically needs to allocate some shared register (or constant register) resources and some slot register resources. These register resources are used to store constants and common data information for the use of the Arithmetic Logic Unit (ALU). At the same time, these registers have low latency and high usage frequency.

[0182] However, the number of these constant and scalar registers is always limited. This leads to situations where, during application runtime, constants and common data cannot fit within these registers, necessitating either register spillage or the allocation of additional space in global memory. However, frequent loading from global memory results in significant latency, which is difficult to conceal. Furthermore, for the common registers used to store constants, once initialized, they cannot be modified or rewritten during runtime. Therefore, the register configuration methods in these technologies can significantly impact processor performance.

[0183] In related technologies, a common approach is to add a constant cache to store these values, and then load them into registers from the constant cache when needed. However, this requires opening up a separate data path for the constant cache, which introduces a significant area overhead.

[0184] To address the issues of capacity limitations of dedicated shared registers and slot registers, area overhead caused by independently allocating constant caches, and loading latency, this application's embodiments design a configurable scalar processing register that supports initial loading plus dynamic runtime loading.

[0185] When resource usage is high, each thread group (such as wave) can be configured to use more scalar processing registers. Under heavy usage, a dynamic runtime loading approach is adopted, loading unstored constant values ​​and other data in advance as the execution approaches the instruction. To address the issue of significant initialization latency, we designed the system to initialize constant data and remove corresponding data dependencies.

[0186] As an example, the general scheme of this application embodiment is shown below.

[0187] When a task is issued, the register usage configuration information is carried through the task issuance path; this can be referred to as scalar processing register usage configuration information. During task assembly, the scalar processing register usage is deducted by the thread assembly and resource usage control unit, and at the same time, thread group-related information and the number of chunks corresponding to the scalar processing register usage are issued to the scalar processing register management unit.

[0188] After obtaining the corresponding configuration management information, the scalar processing register management unit performs allocation management and records thread group-related information and chunk block marking information. It also assigns a number to the chunk block marking information as a basis for subsequent access and lookup. For indirect calls, the drawid or kernelid information is directly written into the scalar processing register through the thread assembly and resource usage control unit.

[0189] Then, the thread assembly and resource usage control unit issues a load initialization request to initialize the allocated scalar processing registers. To reduce latency, the thread group's scheduling process and the initialization process of the allocated scalar processing registers can be performed in parallel.

[0190] To reduce latency during multiple loading processes, constant data that needs to be loaded is preloaded into the last-level cache, and then loaded into the first-level cache for persistent storage during actual use.

[0191] Since the capacity of scalar processing registers is limited and cannot meet the needs of high-usage scenarios, scalar processing registers have been made configurable. By configuring the number of scalar processing registers, the number of scalar processing registers corresponding to a thread group can be increased, which will correspondingly reduce the number of active thread groups.

[0192] For scenarios where the number of variables exceeds the configuration limit, a dynamic loading instruction for loading constant data is added for runtime dynamic loading. In this case, the dependency is removed via the instruction count dependency bar.

[0193] Add an early release function for scalar processing registers, which can be used in conjunction with the early release function for vector processing registers. Simply select the corresponding type when releasing.

[0194] The following section will explain the specific implementation details of the above solution.

[0195] First, data dependencies are checked after allocation management.

[0196] In other words, the compiler, during compilation, uses a combination of constant buffers and scalar registers. After compilation, the driver places all constant data together in the designated constant buffer. During initialization, a limited number of constant values ​​(e.g., Limited_Numb) are written to the scalar registers, such as 128 DWs. To reduce initialization latency, dependency checks are not performed during the initial loading phase. A data dependency check instruction is inserted before the usage instruction to ensure data correctness.

[0197] To reduce latency in subsequent accesses, data is initially prefetched into the first-level cache (such as the data cache). To prevent cache lines from being replaced, persistence information is added to keep constant data permanently residing in the first-level cache (data cache).

[0198] The initial loading process is as follows Figure 5 As shown.

[0199] First, the main data control unit 60 prefetches large blocks of constant data from the constant buffer into the second-level cache or the last-level cache 70. Then, the program dispatch scheduler 61 groups the thread groups, deducts the number of scalar chunks corresponding to the thread group, and sends the thread group configuration information to the scalar processing register 65 for chunk allocation. For the assembled thread group, the program dispatch scheduler 61 sends an initialization load request to the first-level cache 67 to load the constant data. It also includes persistence information to mark the loaded constant data, ensuring it can be persistently stored in the first-level cache 67.

[0200] Optionally, the constant data in the L1 cache 67 is prefetched from the L2 cache or the last cache 70 via the bus interface 68 or the on-chip network 69. This prefetching process can occur before the thread group assembly process, or it can occur together with the thread group assembly process, or it can be prefetched after the thread group is assembled.

[0201] At this point, the thread group is scheduled to the thread group management unit 63 via scheduler 62. Once the thread group management unit 63 detects that all relevant resources have been allocated, it releases the basic scheduling dependency and sends the instruction to the instruction control unit 64 for instruction fetching, decoding, and execution.

[0202] When execution reaches a point where constant data is required, a new data dependency check instruction is executed to check if the data is ready. Upon execution of this instruction, the corresponding thread group is blocked (or rendered inactive), preventing it from issuing further instructions. Then, the control pipeline performs another data dependency check to verify if the required constant data has been initialized and loaded. This is specifically determined by checking the initialization signal maintained by scalar register 65 for each thread group. If an initialization signal is received, it indicates that initialization loading is complete; otherwise, it indicates that initialization loading is incomplete. If the required data initialization loading is incomplete, the process continues waiting; if initialization is complete, the thread group is activated and execution continues.

[0203] After the initial loading, if there is a limit to the number of constant data, the excess constants can be dynamically loaded using runtime instructions. During runtime, newly added dynamic loading instructions load constants from the constant buffer, L1 cache, L2 cache, or last-level cache. Specifically, the instruction control unit 64 triggers the loading storage unit 71 to retrieve the corresponding constant data from the L1 cache 67. Data dependencies during dynamic loading are resolved using write barriers, read barriers, and wait barriers. A write barrier is a memory barrier instruction. When a write barrier is executed, it ensures that all write operations before it are completed, and that subsequent write operations will only begin after the previous write operations are completed. A read barrier is also a memory barrier instruction. It ensures that all read operations before it are completed, and that subsequent read operations will only begin after the previous read operations are completed. Wait barriers are mainly used to handle dependencies between instructions. They allow one instruction to wait for another instruction to complete a specific operation before continuing execution.

[0204] When the kernel function completes execution, the cache line persistently stored in L1 cache 67 can be invalidated. Then, when other data arrives in L1 cache 67, it can overwrite the data in the corresponding cache line.

[0205] For details on request and data flow during initialization loading and dynamic loading, please refer to [reference needed]. Figure 6 During initial loading, the program dispatch scheduler 61 sends an initialization load request to the L1 cache 67, indicating that the target is a scalar register; and the program dispatch scheduler 61 sends an initialization write request to the data loading unit 66, indicating that the target is a scalar register. At this time, the data loading unit 66 receives the data sent by the L1 cache 67 and is responsible for writing it into the corresponding scalar register 65. During dynamic loading, a dynamic load instruction is used to trigger the L1 cache 67 to send data to the data loading unit 66, and this dynamic load instruction is used to indicate that the target of the dynamic load is a scalar register.

[0206] Specifically, when the data loading unit 66 is writing, it determines the corresponding 6-bit waveid (thread group identifier), 9-bit register logical address, 32-bit data, and mask. The waveid and register logical address are used to determine the location of the register to store the data next, and the 32-bit data will be stored in that register.

[0207] During the initial loading and dynamic loading processes described above, to maintain sufficient bandwidth without increasing additional area overhead, multiple stream processors (e.g., two stream processors) can share a single write-back logic in the storage entry buffer. When writing back from the L1 cache, the bandwidth, or throughput, is used based on the cache line width. However, the bandwidth corresponding to scalar processing registers is typically much smaller than the L1 cache line width. Therefore, a configuration block structure using 8 banks is designed. When the L1 cache writes back 128 DWs via the data transfer executor, it is divided into 8 entries and stored in 8 deep FIFOs of the storage entry buffer. Each FIFO layer stores one entry, corresponding to 128 / 8 = 16 DWs. Then, the storage entry buffer writes half (8 DWs) of data at a time into the corresponding 8-bank structure, i.e., it retrieves half of the data from one FIFO layer and stores it in one line of the configuration block. This process takes 16 cycles to complete writing 128 DWs.

[0208] However, when loading 128DW from another thread group into the scalar processing register of another stream processor core, it can be written into the 8-level deep FIFO of the memory entry buffer below. Thus, when writing to the scalar processing register of another stream processor core, it is also 8DW. By using a large bandwidth input and two small bandwidth outputs, a fast write effect is achieved without congesting the pipeline on the return path.

[0209] For details regarding the shared storage entry buffer among the aforementioned multiple stream processors, please refer to the above. Figure 3 The corresponding implementation examples will not be described in detail here.

[0210] To effectively allocate thread groups, a scalar register chunk counter is added to the resource management unit during group task execution to indicate the number of available chunks within the current stream processor core. As shown in Table 3, thread groups can be allocated and managed based on the specific number of scalar register chunks. This information effectively controls the number of active thread groups. Table 3 contains basic configuration information that can be used to configure the number of scalar register chunks used and the corresponding number of active thread groups.

[0211] Table 3 Scalar Processing Register Configuration Information Table

[0212]

[0213] When accessing a specific physical location, it is necessary to perform calculations based on the waveid and the logical address of the scalar processing register, as shown below.

[0214] waveid+(logic_scalar_addr>>5) to locate chunkidx.

[0215] line_addr=(chunk_idx<<2)+((logic_scalar_addr>>3)&3).

[0216] bank_addr=logic_scalar_addr&0x7.

[0217] The processor illustrated in this application primarily addresses the problem of using a large number of shared registers and slot registers during program execution. It designs a configurable scalar processing register system with an initial load followed by dynamic runtime loading. When the driver knows how many shared registers and slot registers the compiler is using, it can dynamically configure and control the number of active thread groups, achieving precise utilization of the scalar processing registers.

[0218] When running beyond the shared registers and slot registers, by issuing constant data load requests in advance, and by adding prefetching and persistence support, the excess constant data can be quickly loaded back during runtime.

[0219] This allows the compiler to achieve fast loading and switching with almost no delay when running normal kernel calls, especially addressing the issue of accessing constants in indirect kernel calls.

[0220] Specifically, the following key points exist in the embodiments of this application: (1) scalar register allocation management strategy before the thread group is executed; (2) constant value initialization loading, post-check data dependency, and data dependency check process; (3) dynamic loading of constant data during program execution; (4) chunk configuration management of scalar processing registers; (5) access address calculation logic of scalar processing registers; (6) register resource release management when the thread group exits.

[0221] Figure 7 A flowchart illustrating a register allocation method provided in an exemplary embodiment of this application is shown. The method is executed by a processor including a plurality of registers, a register management unit, and a task assembly unit. The method includes at least one of the following steps.

[0222] Step 210: The register management unit divides multiple registers into at least one set of configuration chunks, each set of configuration chunks including at least one configuration block, and at least one configuration block including at least two registers from multiple registers.

[0223] Optionally, the number of configuration blocks included in each configuration block set within at least one configuration block set may be the same or different. For example, each configuration block set may include 5 configuration blocks, or each configuration block set may include 6 configuration blocks, or each configuration block set may include 8 configuration blocks, or each configuration block set may include 10 configuration blocks, and so on. As another example, at least one configuration block set may include 4 configuration block sets, each containing a different number of configuration blocks, such as configuration block set 1 containing 5 configuration blocks, configuration block set 2 containing 6 configuration blocks, configuration block set 3 containing 8 configuration blocks, and configuration block set 4 containing 10 configuration blocks. For example, at least one configuration block set may include four sets of configuration block sets. The number of configuration blocks included in each set may differ, but the number of configuration blocks included in each set may be the same. For instance, the first set may include 16 sets of configuration blocks, each containing 5 configuration blocks; the second set may include 13 sets of configuration blocks, each containing 6 configuration blocks; the third set may include 10 sets of configuration blocks, each containing 8 configuration blocks; and the fourth set may include 8 sets of configuration blocks, each containing 10 configuration blocks. It should be understood that when at least one configuration block set includes multiple sets of configuration block sets, the number of configuration blocks included in each set may be different or the same as shown above, and this embodiment does not limit this.

[0224] Optionally, the configuration blocks included in each configuration block set within at least one configuration block set may be of the same or different sizes. For example, the configuration blocks included in a configuration block set may be of the same size, meaning that the number of registers included in each configuration block within a configuration block set is the same. However, the configuration blocks included in different configuration block sets may be of different sizes. For instance, configuration block set 1 may include 5 configuration blocks, each of which includes 32 registers; configuration block set 2 may include 10 configuration blocks, each of which includes 16 registers.

[0225] Optionally, the register management unit divides multiple registers evenly to obtain at least one set of configuration blocks, where each configuration block set in the at least one set contains the same number and size of configuration blocks. That is, even division means dividing multiple registers into at least one set of configuration blocks, where each configuration block set in the at least one set contains the same number and size of configuration blocks. Alternatively, the register usage unit divides multiple registers non-uniformly to obtain at least one set of configuration blocks, where at least two configuration block sets contain configuration blocks of different numbers and sizes. That is, non-uniform division means dividing multiple registers into at least one set of configuration blocks, where at least two configuration block sets contain configuration blocks of different numbers and sizes. For example, two configuration block sets may contain different numbers of configuration blocks, or different sizes of configuration blocks, or both different numbers and sizes of configuration blocks.

[0226] Optionally, registers can be uniformly divided using variable units to obtain at least one configuration block, where each configuration block in the at least one configuration block is of the same size; or at least one configuration block can be uniformly divided to obtain at least one set of configuration blocks, where each set of configuration blocks in the at least one set of configuration blocks includes the same number of configuration blocks. That is, uniform division of multiple registers means dividing multiple registers into at least one configuration block, where each configuration block in the at least one configuration class is of the same size; uniform division of at least one configuration block means dividing at least one configuration block into at least one set of configuration blocks, where each set of configuration blocks in the at least one set of configuration blocks includes the same number of configuration blocks. Alternatively, registers can be uniformly divided using variable units to obtain at least one configuration block; or at least one configuration block can be non-uniformly divided to obtain at least one set of configuration blocks, where two sets of configuration blocks in the at least one set of configuration blocks include different numbers of configuration classes. That is, non-uniform division of at least one configuration block means dividing at least one configuration block into at least one set of configuration blocks, where two sets of configuration blocks in the at least one set of configuration blocks include different numbers of configuration classes. Alternatively, registers can be non-uniformly divided using variable units to obtain at least one configuration block, where at least two configuration blocks are of different sizes; or at least one configuration block can be uniformly divided to obtain at least one set of configuration blocks. In other words, non-uniform division of multiple registers means dividing multiple registers into at least one configuration block, where at least two configuration blocks are of different sizes. Alternatively, registers can be non-uniformly divided using variable units to obtain at least one configuration block; or at least one configuration block can be non-uniformly divided to obtain at least one set of configuration blocks.

[0227] Optionally, the specific partitioning method for multiple registers can be determined based on multiple factors such as the processor's specific architecture, the type of task currently being executed, and the user's subjective settings. This application embodiment does not limit this in any way.

[0228] Step 220: The task assembly unit allocates a first configuration block set to the first thread group. The first configuration block set is one of at least two configuration block sets.

[0229] The task assembly unit is the first thread group selecting the first configuration block set from at least two configuration block sets.

[0230] Optionally, the task assembly unit allocates a first set of configuration blocks to the first thread group based on the tasks executed by the first thread group. The tasks executed by the first thread group are used to determine at least one of the number of configuration blocks included in the configuration block set and the size of the configuration blocks.

[0231] Alternatively, the task assembly unit may also be called a thread assembly and resource usage control unit, or a program dispatch scheduler, etc.

[0232] It should be noted that the names of the various units shown in the embodiments of this application are for illustrative purposes only. In actual implementation, other names may be used to refer to the corresponding units. Furthermore, the unit division method in the embodiments of this application is also for illustrative purposes only. In actual implementation, the number of units, functional boundaries, etc., can be adjusted. Specific adjustments include merging, splitting, renaming, etc. For example, the aforementioned resource usage management unit and task assembly unit can also be adjusted into a single task assembly and resource usage control unit. The embodiments of this application do not limit this, but the scope of protection of the embodiments of this application is not limited thereto.

[0233] In summary, the processor provided in this application first divides multiple registers into at least one set of configuration blocks, with each set of configuration blocks used to allocate to thread groups. By dividing register resources into multiple configuration blocks and allocating them in the form of configuration blocks, each thread group can be allocated at least one configuration block (i.e., one set of configuration blocks). Compared to the traditional method of allocating registers continuously at the register level, this method of allocating register resources directly to configuration blocks instead of multiple consecutive registers avoids frequent splitting of consecutive registers, thus preventing register fragmentation. Furthermore, allocating through configuration blocks allows for management based on configuration blocks rather than registers, improving the efficiency of register resource allocation and release.

[0234] In some embodiments, the aforementioned plurality of registers are plurality of scalar registers, or plurality of registers are plurality of vector registers, or plurality of registers are a plurality of scalar registers and a plurality of vector registers, etc. It should be noted that the processor execution logic shown in the embodiments of this application is applicable to different types of registers. In addition to the scalar registers and vector registers shown above, other types of registers may also be supported, and the embodiments of this application do not limit this.

[0235] The following section will explain the specific partitioning method for registers.

[0236] 1. Register partitioning.

[0237] In some embodiments, step 210 above can be implemented as follows: the register management unit determines quantity information based on register usage configuration information, the quantity information including at least one of the number of configuration blocks, the number of active thread groups, a first data volume, and a second data volume; based on the quantity information, the multiple registers are divided into at least one set of configuration blocks; wherein, the number of configuration blocks is used to indicate the number of configuration blocks in each of the at least one set of configuration blocks, the number of active thread groups is used to indicate the maximum number of thread groups that the processor supports to execute in parallel, the first data volume is used to indicate the maximum amount of data that each of the at least one set of configuration blocks can support to store, and the second data volume is used to indicate the maximum amount of data that each configuration block can support to store.

[0238] For example, the register management unit queries the register configuration information table based on the register usage configuration information to determine the quantity information. The register configuration information table is used to indicate the mapping relationship between the register usage configuration information and the quantity information.

[0239] For details, please refer to "1. Register Division" for the processor mentioned above, which will not be repeated here.

[0240] The following section will explain in detail how thread groups use configuration block sets.

[0241] 2. Loading registers.

[0242] In some embodiments, the processor further includes a data loading unit, such as Figure 2 As shown; the data loading unit loads at least one first data into the first part registers in the first configuration block set. The first part registers are a specified number of registers or a specified proportion of registers in the first configuration block set. The first data belongs to a first type, and the data of the first type includes at least one of constants and common data.

[0243] For details, please refer to "2. Loading Registers" for the processor mentioned above, which will not be repeated here.

[0244] In related technologies, the thread group scheduling process is usually executed first, and data is loaded before or during the execution of a specific instruction. However, in this embodiment, a large block of first-type data is usually preloaded, and the data type of the first-type data preloaded in this embodiment differs from the data type loaded in related technologies. Therefore, the traditional thread group scheduling process and data initialization loading process should be updated to adapt to the register allocation and usage method shown in this embodiment. Specifically, the loading timing of at least one piece of first data into the first part of the registers is as follows.

[0245] (1) Loading timing.

[0246] In some embodiments, the process of loading the first data into the first part of the registers described above can be referred to as the initialization process. To improve the execution efficiency of the thread group in the processor, the initialization process can be executed in parallel with the thread group's scheduling process. Compared to the related art where the thread group's scheduling process is completed first, and then the initialization process (also referred to as the loading process) is performed on the thread group's registers, executing the scheduling process and the initialization process in parallel helps to shorten the thread group's execution time, improve the thread group's execution efficiency, and thus reduce the overall latency of task execution. For example, the processor also includes a scheduler; the scheduler sends a scheduling request to trigger the scheduling process of the first thread group, the scheduling process referring to the process of scheduling the execution of the first thread group; the task assembly unit sends an initialization loading request to trigger the initialization process of the first configuration block set, the initialization process referring to the initialization of the first part of the registers in the first configuration block set; wherein, the scheduling process and the initialization process are executed in parallel.

[0247] In some embodiments, the actual amount of data required by the thread group is large, such as the configuration block set having a data size of 160 DW, but the actual amount of data required by the thread group is greater than the data size of the configuration block set or the data size of the first part register. In this case, dynamic loading can be used to solve the problem, as shown below.

[0248] (2) Dynamic loading.

[0249] In some embodiments, during the execution of the first thread group, the data loading unit loads at least one second data into at least one register of the first configuration block set, wherein the second data belongs to a first type, and there is a one-to-one relationship between the at least one register and the at least one second data.

[0250] Optionally, during the execution of the first thread group, the data loading unit determines at least one type of data required by at least one instruction, namely at least one second data, and loads at least one second data into at least one register of the first configuration block set.

[0251] For example, during the execution of the first thread group, the data loading unit loads at least one second data into at least one register of the first part register, wherein the data in the at least one register is data that the first thread group has finished using, and the at least one second data is used to overwrite the first data in the at least one register.

[0252] For example, during the execution of the first thread group, the data loading unit loads at least one second data into at least one register of the second part register, wherein the second part register is a register in the first configuration block set other than the first part register; wherein the register in the at least one register satisfies at least one of the following characteristics: the data in the register is data that has been used by the first thread group; the register is an empty register.

[0253] In other embodiments, during the execution of the first thread group, the data loading unit loads at least one second data into at least one register in the first part register and the second part register. For example, the first configuration block set includes 160 registers, the first part register includes 128 registers, the second part register includes 32 registers, and the at least one second data is 32 second data. The data loading unit determines the registers to load the second data based on the usage of each register in the first configuration block set. For example, if the data in 24 registers in the first part register has been exhausted and the data in 16 registers in the second part register has been exhausted, the 24 registers in the first part register are preferentially selected. Since the 24 registers in the first part register are less than the number of second data, 8 registers in the second part register are also selected for loading the second data.

[0254] The following section will take the first configuration block set, which includes n first registers (registers for storing data of the first type) and m second registers (registers for storing data of the first type and the second type), as an example to explain the above initialization process and dynamic loading process in detail.

[0255] In some embodiments, when the amount of first type data required by the first thread group is less than or equal to the number of first registers, the loading of the first type data of the first thread group can be completed simply through initial loading. That is, when the first number is less than or equal to n, the data loading unit loads a first number of first data items into a first number of first registers based on the initial loading request, where the first number is the amount of first type data corresponding to the first thread group.

[0256] Optionally, if the amount of first type data required by the first thread group is less than or equal to the number of first registers, the remaining first registers can be left empty or used to load second type data. The remaining first registers refer to the registers among the n first registers that were not loaded with first data during the initial process.

[0257] In some embodiments, when the amount of data of the first type required by the first thread group is greater than the number of first registers, initial loading and dynamic loading should be used to implement the execution of the first thread group. Specifically, the initial loading process is as follows: The first configuration block set includes n first registers, which are used to store data of the first type, where n is a specified number or determined by a specified ratio; when the first number is greater than n, the data loading unit loads n pieces of first data into the n first registers based on the initial loading request, where the first number is the amount of data of the first type corresponding to the first thread group.

[0258] In some embodiments, after the initialization process is completed, the remaining data in the first number of data corresponding to the first thread group can be dynamically loaded into registers. For example, the processor further includes an instruction control unit; when the first number is greater than n, the instruction control unit executes a dynamic loading instruction, which instructs that data of the first type be loaded into the first configuration block set during the execution of the first thread group; a data loading unit is used to load k pieces of second data into k registers based on the dynamic loading instruction, where k is a positive integer and the second data belongs to the first type.

[0259] For example, the data loading unit loads k second data items into k first registers based on dynamic loading instructions. The data in the k first registers is data that the first thread group has finished using, and the k second data items are used to overwrite the data in the k first registers. Alternatively, based on dynamic loading instructions, k second data items are loaded into k second registers, where each register in the k second registers satisfies at least one of the following characteristics: the data in the register is data that the first thread group has finished using; or the register is an empty register. Alternatively, based on dynamic loading instructions, k second data items are loaded into i first registers and j second registers, where i + j = k. The selection criteria for the i first registers and j second registers are as described above for the k first registers and k second registers, and will not be repeated here.

[0260] In addition to the configuration-based allocation of registers for each thread group and the combination of initial loading and dynamic loading for the first type of data as shown above, a method was designed for multiple subprocessors to share a single write-back path in order to further reduce the processor area, as detailed below.

[0261] 3. Shared return data path.

[0262] In some embodiments, the processor further includes at least one data loading unit, a first cache, and a plurality of subprocessors; each data loading unit corresponds to m subprocessors, the m subprocessors being all or part of the plurality of subprocessors included in the processor, and each of the m subprocessors including at least one set of configuration blocks; the data loading unit includes a data transfer executor and a storage entry buffer; the data transfer executor reads s third data from the first cache using a first bandwidth; and caches the s third data to a first first-in-first-out (FIFO) queue of the storage entry buffer, the first bandwidth being related to the cache line width of the first cache; the storage entry buffer writes the s third data from the first FIFO queue to at least two registers corresponding to the first subprocessor among the m subprocessors using a second bandwidth, the second bandwidth being related to the size of the configuration block; wherein, the first bandwidth is greater than the second bandwidth.

[0263] For details, please refer to "3. Shared Return Data Path" for the processor mentioned above, which will not be repeated here.

[0264] Since this application embodiment uses configuration blocks to configure register resources, the processor should make corresponding changes when accessing registers, as shown below.

[0265] 4. Address access process.

[0266] In some embodiments, in response to an access request for a register to be accessed, the processor determines the physical address of the register to be accessed based on the identifier of the thread group corresponding to the register and the logical address of the register to be accessed. The physical address is used to indicate that the register to be accessed is located in the b-th row and c-th column of the a-th configuration block, where a, b, and c are all positive integers.

[0267] For details, please refer to "4. Address Access Process" for the processor mentioned above, which will not be repeated here.

[0268] To improve the efficiency of the initial loading and dynamic loading mentioned above, the data required by the thread group can be prefetched into the first cache, as shown below.

[0269] 5. Data prefetching and persistence.

[0270] In some embodiments, the processor further includes a first cache; the task assembly unit sends an initialization load request to the first cache; the first cache loads data of the first type corresponding to the first thread group into the first cache based on the initialization load request; and the data of the first type is set to a persistent storage state, the persistent storage state being used to indicate that the data of the first type will remain cached in the first cache until the first thread group ends.

[0271] Optionally, the first cache is the L1 cache in the processor or subprocessor.

[0272] In some embodiments, the processor further includes a computing data master control unit and a second cache, the second cache being the next level cache of the first cache; the computing data master control unit loads data of the first type from the buffer into the second cache; the first cache loads data of the first type corresponding to the first thread group from the second cache into the first cache based on an initialization load request.

[0273] For details, please refer to "5. Data Prefetching and Persistence" for the processor mentioned above, which will not be repeated here.

[0274] In some embodiments, the first configuration block is released when the first thread group finishes using the first configuration block in the first configuration block set. That is, embodiments of this application support releasing the register resources of the thread group in advance before the end of the lifecycle of the first thread group, so that the register resources can be allocated to other thread groups for use, thereby improving the overall register utilization rate.

[0275] Optionally, some registers may be released early if the first thread group finishes using some registers (such as one or more rows of registers) in the first configuration block set.

[0276] Please refer to Figure 8 This diagram illustrates a structural block diagram of a register allocation apparatus provided in an exemplary embodiment of this application. The apparatus has the functionality to implement the above-described register allocation method example; this functionality can be implemented in hardware or by hardware executing corresponding software implementation. The apparatus can be the processor described above, or it can be located within a processor. The apparatus may include: a register usage management module 310 and a task assembly module 320.

[0277] Register usage management module 310 is used to divide the plurality of registers into at least one configuration block set, each configuration block set in the at least one configuration block set includes at least one configuration block, and the at least one configuration block includes at least two registers from the plurality of registers;

[0278] The task assembly module 320 is used to allocate a first configuration block set to the first thread group, wherein the first configuration block set is one of the at least two configuration block sets.

[0279] In some embodiments, the register usage management module 310 is used to determine quantity information based on register usage configuration information, wherein the quantity information includes at least one of the number of configuration blocks, the number of active thread groups, a first data volume, and a second data volume; and based on the quantity information, divide the plurality of registers into the at least one set of configuration blocks.

[0280] Wherein, the number of configuration blocks is used to indicate the number of configuration blocks in each configuration block set in the at least one configuration block set, the number of active thread groups is used to indicate the maximum number of thread groups that the processor supports for parallel execution, the first data amount is used to indicate the maximum amount of data that each configuration block set in the at least one configuration block set can support for storage, and the second data amount is used to indicate the maximum amount of data that each configuration block can support for storage.

[0281] In some embodiments, the register usage management module 310 is used to query a register configuration information table based on the register usage configuration information to determine the quantity information, wherein the register configuration information table is used to indicate the mapping relationship between the register usage configuration information and the quantity information.

[0282] In some embodiments, the apparatus further includes a data loading module;

[0283] The data loading module is used to load at least one first data into a first part of the registers in the first configuration block set. The first part of the registers is a specified number of registers or a specified proportion of registers in the first configuration block set. The first data belongs to a first type, and the data of the first type includes at least one of constants and common data.

[0284] In some embodiments, the apparatus further includes a scheduling module;

[0285] The scheduling module is used to send a scheduling request, which is used to trigger the scheduling process of the first thread group, and the scheduling process refers to the process of scheduling and executing the first thread group.

[0286] The task assembly module 320 is used to send an initialization loading request, which is used to trigger the initialization process of the first configuration block set. The initialization process refers to initializing the first part of the registers in the first configuration block set.

[0287] The scheduling process and the initialization process are executed in parallel.

[0288] In some embodiments, the data loading module is configured to load at least one second data into at least one register of the first configuration block set during the execution of the first thread group, wherein the second data belongs to the first type, and the at least one register has a one-to-one relationship with the at least one second data.

[0289] In some embodiments, the data loading module is configured to load at least one second data into at least one register of the first partial register during the execution of the first thread group, wherein the first data in the at least one register is data that the first thread group has finished using, and the at least one second data is used to overwrite the first data in the at least one register.

[0290] In some embodiments, the data loading module is configured to load the at least one second data into at least one register of a second part register during the execution of the first thread group, wherein the second part register is a register in the first configuration block set other than the first part register;

[0291] Wherein, the register in the at least one register satisfies at least one of the following characteristics:

[0292] The data in the register is data that the first thread group has finished using;

[0293] The register in question is an empty register.

[0294] In some embodiments, the first configuration block set includes n first registers, the n first registers being used to store data of the first type, wherein n is the specified number or is determined by the specified ratio;

[0295] The data loading module is used to load n first data items into the n first registers based on the initialization loading request when the first quantity is greater than n, where the first quantity is the number of first type data items corresponding to the first thread group.

[0296] In some embodiments, the apparatus further includes an instruction control module;

[0297] The instruction control module is used to execute a dynamic loading instruction when the first quantity is greater than n. The dynamic loading instruction is used to instruct that the first type of data be loaded into the first configuration block set during the execution of the first thread group.

[0298] The data loading module is used to load k second data into k registers based on the dynamic loading instruction, where k is a positive integer and the second data belongs to the first type.

[0299] In some embodiments, the dynamic loading instruction is used to load data of a first type required by the first instruction into the first configuration block set before the first instruction is executed.

[0300] In some embodiments, the first configuration block set includes n first registers, the n first registers being used to store data of the first type, wherein n is the specified number or is determined by the specified ratio;

[0301] The data loading module is used to load the first number of first data into the first number of first registers based on the initialization loading request when the first number is less than or equal to n. The first number is the number of first type data corresponding to the first thread group.

[0302] In some embodiments, the apparatus further includes at least one data loading module, a first cache, and a plurality of subprocessors; each data loading module in the at least one data loading module corresponds to m subprocessors, the m subprocessors being all or part of the plurality of subprocessors included in the processor, and each of the m subprocessors including at least one set of configuration blocks; the data loading module includes a data transfer executor and a storage entry buffer;

[0303] The data transfer executor is configured to read s third data items from a first cache using a first bandwidth; and to cache the s third data items into a first first-in-first-out (FIFO) queue of the storage entry buffer, wherein the first bandwidth is related to the cache line width of the first cache.

[0304] The storage entry buffer is used to write the s third data from the first FIFO queue to at least two registers corresponding to the first subprocessor among the m subprocessors using the second bandwidth. The first FIFO queue has a corresponding relationship with the first subprocessor. The second bandwidth is related to the size of the configuration block.

[0305] Wherein, the first bandwidth is greater than the second bandwidth.

[0306] In some embodiments, the apparatus further includes an instruction execution module. The instruction execution module is configured to, in response to an access request for a register to be accessed, determine the physical address of the register to be accessed based on the identifier of the thread group corresponding to the register and the logical address of the register to be accessed, wherein the physical address indicates that the register to be accessed is located in the b-th row and c-th column of the a-th configuration block, where a, b, and c are all positive integers.

[0307] In some embodiments, the instruction execution module is configured to, in response to an access request for the register to be accessed, determine the identifier of the configuration block corresponding to the register to be accessed based on the identifier of the thread group corresponding to the register to be accessed and the logical address of the register to be accessed; determine the row start address corresponding to the configuration block based on the identifier of the configuration block; determine the row offset address and column address corresponding to the register to be accessed based on the logical address of the register to be accessed; determine the row address of the register to be accessed based on the row start address and the row offset address; and determine the physical address of the register to be accessed based on the row address and the column address.

[0308] In some embodiments, the task assembly module is configured to send an initialization loading request to the first cache; the first cache is configured to load data of a first type corresponding to the first thread group into the first cache based on the initialization loading request; and to set the data of the first type to a persistent storage state, the persistent storage state being used to indicate that the data of the first type will remain cached in the first cache until the first thread group ends.

[0309] In some embodiments, the apparatus further includes a computing data master control module and a second cache, wherein the second cache is the next level cache of the first cache;

[0310] The main control module for computing data is used to load the first type of data from the buffer into the second cache;

[0311] The first cache is used to load data of the first type corresponding to the first thread group from the second cache into the first cache based on the initialization loading request.

[0312] In some embodiments, the plurality of registers are a plurality of scalar registers.

[0313] On the other hand, embodiments of this application provide a graphics card that includes the atomic operation processing system described in the above embodiments.

[0314] On the other hand, embodiments of this application provide a computer device, which includes the atomic operation processing system described above. This computer device can be at least one of a portable computer, desktop computer, server, server cluster, artificial intelligence (AI) computing cluster, and cloud computing cluster. The AI ​​computing cluster can also be simply referred to as an intelligent computing cluster or smart computing cluster.

[0315] It should be understood that "multiple" as used herein refers to two or more. The character " / " generally indicates that the preceding and following objects are in an "or" relationship. Furthermore, the step numbers described herein are merely illustrative of one possible execution order between steps. In some other embodiments, the steps may not be executed in numerical order, such as two steps with different numbers being executed simultaneously, or two steps with different numbers being executed in the reverse order of the illustration. This application does not limit this approach.

[0316] The above description is merely an optional embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A processor, characterized in that, The processor includes multiple registers, a register management unit, and a task assembly unit; The register management unit is configured to determine quantity information based on register usage configuration information, wherein the quantity information includes at least one of the following: number of configuration blocks, number of active thread groups, first data volume, and second data volume; and based on the quantity information, divide the plurality of registers into at least one set of configuration blocks, wherein the register usage configuration information is related to the task executed by the processor, and each set of configuration blocks includes at least one configuration block, wherein the at least one configuration block includes at least two registers from the plurality of registers. The task assembly unit is used to allocate a first configuration block set to the first thread group, wherein the first configuration block set is one of the at least two configuration block sets; Wherein, the number of configuration blocks is used to indicate the number of configuration blocks in each configuration block set in the at least one configuration block set, the number of active thread groups is used to indicate the maximum number of thread groups that the processor supports for parallel execution, the first data amount is used to indicate the maximum amount of data that each configuration block set in the at least one configuration block set can support for storage, and the second data amount is used to indicate the maximum amount of data that each configuration block can support for storage.

2. The processor according to claim 1, characterized in that, The register management unit is used to query the register configuration information table based on the register usage configuration information to determine the quantity information. The register configuration information table is used to indicate the mapping relationship between the register usage configuration information and the quantity information.

3. The processor according to claim 1 or 2, characterized in that, The processor also includes a data loading unit; The data loading unit is used to load at least one first data into a first part of the registers in the first configuration block set. The first part of the registers is a specified number of registers or a specified proportion of registers in the first configuration block set. The first data belongs to a first type, and the data of the first type includes at least one of constants and common data.

4. The processor according to claim 3, characterized in that, The processor also includes a scheduler; The scheduler is used to send a scheduling request, which is used to trigger the scheduling process of the first thread group, and the scheduling process refers to the process of scheduling and executing the first thread group. The task assembly unit is used to send an initialization loading request, which is used to trigger the initialization process of the first configuration block set, and the initialization process refers to initializing the first part of the registers in the first configuration block set. The scheduling process and the initialization process are executed in parallel.

5. The processor according to claim 3, characterized in that, The data loading unit is used to load at least one second data into at least one register of the first configuration block set during the execution of the first thread group. The second data belongs to the first type, and there is a one-to-one relationship between the at least one register and the at least one second data.

6. The processor according to claim 5, characterized in that, The data loading unit is used to load at least one second data into at least one register of the first partial register during the execution of the first thread group, wherein the at least one second data is used to overwrite the first data in the at least one register.

7. The processor according to claim 5, characterized in that, The data loading unit is used to load the at least one second data into at least one register of the second part register during the execution of the first thread group, wherein the second part register is a register in the first configuration block set other than the first part register; Wherein, the register in the at least one register satisfies at least one of the following characteristics: The data in the register is data that the first thread group has finished using; The register in question is an empty register.

8. The processor according to claim 5, characterized in that, The first configuration block set includes n first registers, which are used to store data of the first type, where n is the specified number or is determined by the specified ratio; The data loading unit is used to load n first data items into the n first registers based on an initialization loading request when the first quantity is greater than n, wherein the first quantity is the number of first type data items corresponding to the first thread group.

9. The processor according to claim 8, characterized in that, The processor also includes an instruction control unit; The instruction control unit is configured to execute a dynamic loading instruction when the first quantity is greater than n. The dynamic loading instruction is configured to instruct the first type of data to be loaded into the first configuration block set during the execution of the first thread group. The data loading unit is used to load k second data into k registers based on the dynamic loading instruction, where k is a positive integer, and the k registers are free registers in the first configuration block set. The free registers are empty registers or registers whose data has been no longer in use.

10. The processor according to claim 9, characterized in that, The dynamic loading instruction is used to load the first type of data required by the first instruction into the first configuration block set before the first instruction is executed.

11. The processor according to claim 5, characterized in that, The first configuration block set includes n first registers, which are used to store data of the first type, where n is the specified number or is determined by the specified ratio; The data loading unit is configured to load the first number of first data items into the first number of first registers based on an initialization loading request when the first number is less than or equal to n. The first number is the number of first type data items corresponding to the first thread group.

12. The processor according to claim 3, characterized in that, The processor further includes at least one data loading unit, a first cache, and multiple sub-processors; each data loading unit corresponds to m sub-processors, the m sub-processors being all or part of the multiple sub-processors included in the processor, and each of the m sub-processors including at least one set of configuration blocks; the data loading unit includes a data transfer executor and a storage entry buffer; The data transfer executor is configured to read s third data items from a first cache using a first bandwidth; and to cache the s third data items into a first first-in-first-out queue of the storage entry buffer, wherein the first bandwidth is related to the cache line width of the first cache; The storage entry buffer is used to write the s third data from the first first-in-first-out queue to at least two registers corresponding to the first subprocessor among the m subprocessors using the second bandwidth. The first first-in-first-out queue has a corresponding relationship with the first subprocessor. The second bandwidth is related to the size of the configuration block. Wherein, the first bandwidth is greater than the second bandwidth.

13. The processor according to claim 1 or 2, characterized in that, The registers in each of the at least one configuration block are arranged in i rows and j columns; The processor is configured to, in response to an access request for a register to be accessed, determine the physical address of the register to be accessed based on the identifier of the thread group corresponding to the register to be accessed and the logical address of the register to be accessed, wherein the physical address is used to indicate that the register to be accessed is located in the b-th row and c-th column of the a-th configuration block.

14. The processor according to claim 13, characterized in that, The processor is configured to, in response to an access request for the register to be accessed, determine the identifier of the configuration block corresponding to the register to be accessed based on the identifier of the thread group corresponding to the register to be accessed and the logical address of the register to be accessed; and determine the row start address corresponding to the configuration block based on the identifier of the configuration block. Based on the logical address of the register to be accessed, determine the row offset address and column address corresponding to the register to be accessed; based on the row start address and the row offset address, determine the row address of the register to be accessed; based on the row address and the column address, determine the physical address of the register to be accessed.

15. The processor according to claim 3, characterized in that, The processor also includes a first cache; The task assembly unit is used to send an initialization loading request to the first cache; The first cache is used to load data of the first type corresponding to the first thread group into the first cache based on the initialization loading request; And set the data of the first type to a persistent storage state, the persistent storage state being used to indicate that the data of the first type will remain cached in the first cache until the first thread group ends.

16. The processor according to claim 15, characterized in that, The processor further includes a computing data master control unit and a second cache, wherein the second cache is the next level cache of the first cache; The main control unit for computing data is used to load the first type of data from the buffer into the second cache; The first cache is used to load data of the first type corresponding to the first thread group from the second cache into the first cache based on the initialization loading request.

17. The processor according to claim 1 or 2, characterized in that, The multiple registers are multiple scalar registers.

18. A graphics card, characterized in that, The graphics card includes the processor as described in any one of claims 1 to 17.

19. A computer device, characterized in that, The computer device includes the processor as described in any one of claims 1 to 17.

20. A register allocation method, characterized in that, The method is executed by a processor according to any one of claims 1 to 17, the processor comprising a plurality of registers, a register management unit, and a task assembly unit; The method includes: The register management unit determines quantity information based on register usage configuration information, the quantity information including at least one of the following: number of configuration blocks, number of active thread groups, first data volume, and second data volume; based on the quantity information, the plurality of registers are divided into at least one set of configuration blocks, the register usage configuration information being related to the task executed by the processor, each set of configuration blocks including at least one configuration block, and the at least one configuration block including at least two registers from the plurality of registers; The task assembly unit allocates a first configuration block set to the first thread group, wherein the first configuration block set is one of the at least two configuration block sets; Wherein, the number of configuration blocks is used to indicate the number of configuration blocks in each configuration block set in the at least one configuration block set, the number of active thread groups is used to indicate the maximum number of thread groups that the processor supports for parallel execution, the first data amount is used to indicate the maximum amount of data that each configuration block set in the at least one configuration block set can support for storage, and the second data amount is used to indicate the maximum amount of data that each configuration block can support for storage.

21. A register allocation device, characterized in that, The device includes: The register usage management module determines quantity information based on register usage configuration information. The quantity information includes at least one of the following: the number of configuration blocks, the number of active thread groups, the first data volume, and the second data volume. Based on the quantity information, the multiple registers are divided into at least one set of configuration blocks. The register usage configuration information is related to the task executed by the processor. Each set of configuration blocks includes at least one configuration block, and the at least one configuration block includes at least two registers from the multiple registers. The task assembly module allocates a first configuration block set to the first thread group, wherein the first configuration block set is one of the at least two configuration block sets; Wherein, the number of configuration blocks is used to indicate the number of configuration blocks in each configuration block set in the at least one configuration block set, the number of active thread groups is used to indicate the maximum number of thread groups that the processor supports for parallel execution, the first data amount is used to indicate the maximum amount of data that each configuration block set in the at least one configuration block set can support for storage, and the second data amount is used to indicate the maximum amount of data that each configuration block can support for storage.

Citation Information

Patent Citations

  • SIMT-based register allocation method and device, equipment and storage medium

    CN117573205A