A method, apparatus, and medium for data processing
Patent Information
- Application Number
- CN202610823017.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-09
- Publication Date
- 2026-08-28
- Estimated Expiration
- 2046-06-09
AI Technical Summary
[0006] In this way, by adjusting multiple shaders to the first mode and pausing the instruction scheduling of at least one of the second shaders, and using the first shader to uniformly initiate the first and second instructions to broadcast the second data to multiple shaders, it is possible to achieve synchronous reuse of the loaded second data by multiple shaders, reducing repeated data loading and on-chip bandwidth occupation. At the same time, with the first data being stored in a partially distributed manner in multiple shaders, multiple shaders can perform matrix multiplication operations in parallel on the corresponding parts of the stored first data and the broadcast second data, thereby improving the throughput of matrix multiplication operations and data reuse efficiency without introducing complex synchronization control overhead.
Smart Images

Figure CN122346452B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this disclosure generally relate to data processing techniques, and more specifically, to methods, apparatus and media for data processing. Background Technology
[0002] With the development of large model technology, matrix multiplication has become one of the most core and frequent computational operations in the training and inference process of large models, and its computational efficiency directly affects the overall performance of large models. Matrix multiplication is characterized by high data reuse rate and high parallelism requirements, and usually relies on high-throughput computing devices such as general-purpose graphics processing units (GPGPUs) for parallel processing. Summary of the Invention
[0003] In a first aspect, a data processing method is proposed. The method includes: using multiple shaders to obtain multiple portions of first data; storing the multiple portions of the first data into multiple on-chip memory units associated with the multiple shaders; in response to the completion of storing the first data, adjusting the multiple shaders to a first mode, the first mode indicating that the multiple shaders are divided into a first shader and at least one second shader, and the instruction scheduling of at least one second shader is suspended; in response to a first instruction initiated by the first shader, loading second data into a common memory unit; in response to a second instruction initiated by the first shader, broadcasting the second data in the common memory unit to the multiple shaders using an on-chip network; and using the multiple shaders to perform matrix multiplication operations based on the corresponding portions of the second data and the first data stored in the multiple on-chip memory units.
[0004] In a second aspect, an apparatus for data processing is proposed. The apparatus includes: an on-chip network; multiple memory units communicatively connected via the on-chip network; multiple shaders communicatively connected to the multiple memory units via the on-chip network; and a central scheduling unit communicatively connected to the multiple shaders via the on-chip network. The multiple shaders are configured to perform matrix multiplication based on first and second data stored in the multiple memory units. The central scheduling unit is configured to adjust the multiple shaders to a first mode, dividing them into a first shader and at least one second shader, and suspending instruction scheduling for at least one second shader. The first shader is configured to initiate a first instruction and a second instruction, the first instruction supporting the loading of second data into the multiple memory units, and the second instruction instructing the broadcast of the second data to the multiple shaders using the on-chip network.
[0005] In a third aspect, a non-transitory computer-readable storage medium is proposed. This non-transitory computer-readable storage medium stores instructions that cause a processor to perform the method according to the first aspect of this disclosure.
[0006] In this way, by adjusting multiple shaders to the first mode and pausing the instruction scheduling of at least one of the second shaders, and using the first shader to uniformly initiate the first and second instructions to broadcast the second data to multiple shaders, it is possible to achieve synchronous reuse of the loaded second data by multiple shaders, reducing repeated data loading and on-chip bandwidth occupation. At the same time, with the first data being stored in a partially distributed manner in multiple shaders, multiple shaders can perform matrix multiplication operations in parallel on the corresponding parts of the stored first data and the broadcast second data, thereby improving the throughput of matrix multiplication operations and data reuse efficiency without introducing complex synchronization control overhead.
[0007] The present invention is provided to present, in a simplified form, the selection of concepts further described below in the detailed description. The present invention is not intended to identify key or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. Attached Figure Description
[0008] The above and other objects, features, and advantages of exemplary embodiments of the present disclosure will become more apparent from the following detailed description with reference to the accompanying drawings. In the exemplary embodiments of the present disclosure, the same reference numerals generally refer to the same components.
[0009] Figure 1 Schematic diagrams of example environments for some embodiments are shown; Figure 2 A schematic structural block diagram of an example apparatus for data processing according to some embodiments is shown; Figure 3 A schematic diagram of an example data loading process for matrix multiplication operations in some embodiments is shown; Figure 4 A schematic structural block diagram of a first instruction scheduling unit in a first shader of some embodiments is shown; Figure 5 A schematic structural block diagram of a second instruction scheduling unit in a second shader of some embodiments is shown; Figure 6 Flowcharts of example processes for data processing in some embodiments are shown; and Figure 7 A block diagram of an electronic device capable of implementing multiple illustrative scenarios is shown. Detailed Implementation
[0010] The principles of this disclosure will now be described with reference to some embodiments. It should be understood that these embodiments are described for illustrative purposes only and to help those skilled in the art understand and implement this disclosure, and do not imply any limitation on the scope of this disclosure. In addition to the methods described below, the disclosure described herein can be implemented in various other ways.
[0011] In the following description and claims, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.
[0012] The terms "an embodiment," "embodiment," "example embodiment," etc., used in this disclosure refer to embodiments that may include specific features, structures, or characteristics, but not every embodiment is required to include that specific feature, structure, or characteristic. Furthermore, these phrases do not necessarily refer to the same embodiment. Moreover, when a specific feature, structure, or characteristic is described in conjunction with an example embodiment, it is claimed that, whether explicitly described or not, such a feature, structure, or characteristic affecting its relation to other embodiments is within the knowledge of those skilled in the art.
[0013] It should be understood that although the terms “first” and “second”, etc., may be used herein to describe various elements, these elements should not be limited to these terms. These terms are used only to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element, without departing from the scope of the exemplary embodiments. As used herein, the term “and / or” includes any and all combinations of one or more of the listed terms.
[0014] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments. As used herein, the singular forms “a,” “an,” and “the” are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the terms “comprising,” “including,” “having,” “containing,” and / or “comprising” as used herein indicate the presence of the said features, elements, and / or components, but do not exclude the presence or addition of one or more other features, elements, components, and / or combinations thereof.
[0015] The examples in this article may involve user data, data acquisition, and / or use. All of these aspects comply with relevant laws, regulations, and rules. In the examples, all data collection, acquisition, processing, manipulation, forwarding, and use are conducted with the user's knowledge and confirmation. Accordingly, in implementing each example, the type, scope of use, and usage scenarios of any data or information involved should be communicated to the user and their authorization obtained through appropriate means, in accordance with relevant laws and regulations. The specific methods of notification and / or authorization can vary depending on the actual situation and application scenario; the scope of the solution is not limited in this regard.
[0016] In this manual and the sample solutions, any processing of personal information will be conducted only under legal grounds (such as obtaining the consent of the data subject or being necessary for the performance of a contract) and will only be carried out within the scope stipulated or agreed upon. A user's refusal to process personal information beyond what is necessary for basic functions will not affect the user's use of basic functions.
[0017] The term "shader" as used in this article refers to the processing unit in a general-purpose graphics processing unit that performs operations and instruction scheduling. It typically includes, but is not limited to, instruction scheduling logic, arithmetic units, temporary caches, and memory access paths with on-chip memory units. A shader can independently schedule and execute its own instructions, or it can be configured to receive instructions from other shaders and perform operations accordingly.
[0018] As used in this document, “Network on Chip” (NoC) refers to an on-chip interconnect structure used to connect multiple functional units within a general-purpose graphics processor. This typically includes, but is not limited to, two-dimensional mesh (2DMesh) structures, three-dimensional mesh (3D Mesh) structures, ring structures, crossbar structures, or combinations thereof. Network on chip can be used to transfer data and instructions between shaders, on-chip memory units, central scheduling units, and other functional units.
[0019] As used in this document, the term "on-chip memory unit" refers to an on-chip storage component in a general-purpose graphics processing unit (GPU) used to store computational data. This typically includes, but is not limited to, static random-access memory (SRAM) units, but may also include other types of high-speed on-chip memory. Multiple on-chip memory units can be associated with multiple shaders to support fast access to the stored data by the shaders.
[0020] As used herein, the term "common storage unit" refers to an on-chip storage unit that can be accessed by multiple shaders. It may include all or part of the multiple on-chip storage units and may include at least two independent storage units to support alternating or parallel loading and reading of different data.
[0021] The term "central scheduling unit" as used in this article refers to the functional unit in a general-purpose graphics processor responsible for cross-shader coordination. It typically includes, but is not limited to, the Command Engine (CE) and its related control logic. It can be configured to perform control operations such as selecting the first shader, pausing and resuming instruction scheduling for other shaders, and synchronizing the states between shaders.
[0022] The term “matrix multiplication” as used in this article refers to a multiplication operation performed on at least two input matrices, which typically includes, but is not limited to, the General Matrix Multiplication (GEMM) operation, and may further include additional operations such as accumulation and scaling associated with the multiplication result.
[0023] The term "first data" as used in this paper refers to data that is preloaded and partially distributed across multiple shaders before the matrix multiplication operation is performed. It typically corresponds to one of the input matrices in the matrix multiplication operation (e.g., it can be denoted as matrix B) and can be reused by multiple shaders during the matrix multiplication operation. The term "second data" as used in this paper refers to data that is loaded by the first shader and broadcast to multiple shaders via an on-chip network during the matrix multiplication operation. It typically corresponds to another input matrix in the matrix multiplication operation (e.g., it can be denoted as matrix A) and may include multiple parts divided along one dimension of the matrix (e.g., it can be denoted as A[0], A[1], ..., A[n]).
[0024] The term "first mode" as used in this paper refers to an execution mode in which one of a group of shaders is selected as the first shader and uniformly distributes instructions to the other shaders. In this mode, the instruction scheduling of the other shaders is suspended. Therefore, the first mode can also be called synchronous execution mode or stream mode. The term "second mode" as used in this paper refers to an execution mode in which multiple shaders independently schedule and execute their own instructions. Therefore, the second mode can also be called asynchronous execution mode.
[0025] The term "first shader" as used in this paper refers to the shader responsible for unified scheduling and initiating instructions in the first mode, which may also be called the master shader; the term "second shader" as used in this paper refers to the shader whose instruction scheduling is suspended in the first mode and which receives instructions from the first shader, which may also be called the slave shader; in multiple shaders, there is usually one first shader and at least one second shader.
[0026] As used in this document, the term "first instruction" refers to an instruction initiated by the first shader to trigger the loading of second data from external storage to a common storage unit. This instruction typically includes, but is not limited to, Tensor Memory Access (TMA) instructions and their associated Memory Barrier (mbarrier) instructions to indicate the initiation and completion determination of the loading of the second data. As used in this document, the term "second instruction" refers to an instruction initiated by the first shader to instruct the broadcast of the second data to multiple shaders using an on-chip network and to trigger the multiple shaders to collaboratively perform matrix multiplication operations. This instruction typically includes, but is not limited to, Matrix Multiply-Accumulate (MMA) instructions.
[0027] The term "Program Counter" (PC) as used in this document refers to the register or equivalent logic within a shader that indicates the address of the instruction to be executed. The term "Wrap" as used in this document refers to a scheduling unit in a general-purpose graphics processor consisting of a group of threads executing in parallel. The term "Instruction Scheduling Unit" as used in this document generally includes, but is not limited to, the Wrap scheduler and its related dependency checks, dispatchers, and instruction buffers, which determine when a given shader schedules and which Wrap instruction to issue.
[0028] As used in this document, the term "temporary cache" refers to a cache deployed inside a shader to store data received by the shader from the on-chip network, which typically includes, but is not limited to, a source / destination cache; the term "data queue" refers to a queue structure deployed inside a shader to store operands to be sent to the arithmetic unit, which typically includes, but is not limited to, a reservation station.
[0029] As mentioned above, Generalized Matrix Multiplication (GEMM) is the core algorithm of the third stage in the basic linear algebra subroutine library. Its essence is a combination of matrix multiplication and accumulation. Matrix multiplication is highly versatile and can cover various matrix operation scenarios. It is one of the underlying supporting algorithms in modern scientific computing, graphics processing, and artificial intelligence.
[0030] With the rapid iteration and widespread application of large-scale modeling techniques, the core role of matrix multiplication has become increasingly prominent. In some examples, the core functional modules of large-scale models, such as querying, key-value (QKV) projection, and multi-head fusion in attention mechanisms, dimensionality scaling operations in feedforward networks (FFNs), and linear transformations between embedding and unembedding layers, are essentially different forms of matrix multiplication. The computational efficiency of matrix multiplication directly affects the training speed, inference latency, and overall performance of large-scale models.
[0031] From a computational perspective, matrix multiplication in large models is characterized by high arithmetic strength, high data reuse rate, and high parallelism requirements, making it suitable for parallel processing using high-throughput computing devices such as general-purpose graphics processing units (GPGPUs). During matrix multiplication, the same set of matrix elements may be accessed multiple times by multiple computational units within a computation cycle, placing high demands on the efficiency and consistency of data retrieval. Furthermore, the weight and activation matrices are typically large, requiring a highly parallel hardware architecture to fully realize the computational potential of matrix multiplication.
[0032] In traditional solutions, matrix multiplication operations on large models are primarily implemented using conventional general-purpose graphics processor architectures, employing a hierarchical cache structure, register files (RF), and a combination of synchronous and asynchronous execution modes. However, this approach still has several areas for improvement in meeting the high-performance computing demands of large-model matrix multiplication. For example, traditional hierarchical cache structures may experience data thrashing when dealing with the high data reuse characteristics of matrix multiplication, and multi-level caches consume significant on-chip storage resources, impacting data retrieval efficiency. Furthermore, the register files within shaders require pre-allocation and rearrangement of computational data before instruction issuance, increasing the complexity of data transmission and scheduling, and making further reduction of the shader's overall area difficult. Additionally, in asynchronous execution mode, each shader independently schedules its own instructions; even with shared memory (e.g., shared memory or tensor memory), the same data may still be read by multiple shaders, resulting in additional on-chip bandwidth consumption.
[0033] This paper proposes a scheme for data processing. According to this scheme, firstly, multiple shaders are used to obtain multiple parts of a first data set; these parts are then stored in multiple on-chip memory units associated with the shaders; after the first data storage is complete, the shaders are adjusted to a first mode, in which they are divided into one first shader and at least one second shader, and instruction scheduling for at least one second shader is suspended, allowing the first shader to uniformly distribute instructions to the other shaders; subsequently, the first shader initiates a first instruction to load the second data into a common memory unit and initiates a second instruction, which is then broadcast by the on-chip network to the multiple shaders; finally, the multiple shaders collaboratively perform matrix multiplication operations based on the corresponding parts of the locally stored first data and the broadcast second data.
[0034] By adjusting multiple shaders to the first mode and pausing the instruction scheduling of at least one of the second shaders, and using the first shader to uniformly initiate the first and second instructions to broadcast the second data to multiple shaders, the synchronous reuse of the loaded second data by multiple shaders can be achieved, reducing the repeated loading of data and the on-chip bandwidth occupation. At the same time, with the first data being stored in a partially distributed manner in multiple shaders, multiple shaders can perform matrix multiplication operations in parallel on the corresponding parts of the stored first data and the broadcast second data, thereby improving the throughput of matrix multiplication operations and data reuse efficiency without introducing complex synchronization control overhead.
[0035] The following section provides a detailed description of various example implementations of this scheme, with reference to the accompanying drawings.
[0036] Example Environment Figure 1 A schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented is shown. For example... Figure 1 As shown, example environment 100 includes electronic device 110. Electronic device 110 can be any electronic device capable of supporting graphics processing and parallel computing, and a graphics processing unit 120 is deployed in electronic device 110.
[0037] The graphics processing unit 120 may include multiple shaders (e.g. Figure 1 Shaders 122-1 to 122-3, etc., shown in the figure), and multiple on-chip memory units (e.g. Figure 1 The on-chip storage units 124-1 to 124-4, bus interface 130, and command engine 132 are shown in the figure. The graphics processing unit 120 may further include an on-chip network (e.g., for interconnecting the above-mentioned functional units) for interconnecting the above-mentioned functional units. Figure 1The on-chip networks 126-1 and 126-2 shown in the figure) and multiple routing nodes (e.g. Figure 1 The routing nodes shown are 128-1 to 128-5.
[0038] In this system, multiple shaders serve as basic processing units within the graphics processing unit 120, responsible for performing specific graphics rendering or general computing tasks. In the example environment 100, the multiple shaders are organized into an array. In embodiments of this disclosure, the multiple shaders can be configured to perform matrix multiplication operations.
[0039] Multiple on-chip memory units include public memory units and private memory units corresponding to the respective shaders (e.g., on-chip memory units 124-5 and 124-6). On-chip memory units can serve as local caches for shaders. They typically refer to shared memory or local memory, used for fast access to data being processed by the core.
[0040] The on-chip network includes horizontal channels (e.g., on-chip network 126-1) and vertical channels (e.g., on-chip network 126-2), which are responsible for data transmission within the graphics processing unit 120. On-chip network 126-1 can be used for data transmission between rows. On-chip network 126-2 can be used for data propagation between columns. The on-chip network includes multiple routing nodes (routing nodes 128-1 to 128-4). These routing nodes control the flow direction of data packets (instructions or data) or determine whether they are received by the current node.
[0041] Bus interface 130 is responsible for connecting the internal shader array of graphics processing unit 120 to an external system bus (such as CPU or main memory) to receive commands and transmit large amounts of data. Command engine 132 is responsible for parsing external instructions, configuring the working modes of multiple shaders, and assigning tasks to specific shaders.
[0042] Electronic device 110 can be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, handheld computers, portable gaming terminals, VR / AR devices, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some embodiments, electronic device 110 can also support any type of user-facing interface (such as "wearable" circuitry).
[0043] Some exemplary embodiments of this disclosure will be described in detail below. It should be noted that section headings are used in this document for ease of understanding and not to limit the embodiments disclosed in a section to that section. Furthermore, although some embodiments are described with reference to example scenario 100, the disclosed techniques are also applicable to other data processing techniques.
[0044] It should be understood that the structure and function of the various elements in environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of this disclosure.
[0045] The following description will continue with reference to the accompanying drawings, which will provide some exemplary embodiments of this disclosure.
[0046] Example execution method Figure 2 A schematic structural block diagram of an example apparatus 200 for data processing according to some embodiments is shown. Apparatus 200 may be implemented as or included in the above references. Figure 1 The described electronic device 110 is used to support the electronic device 110 in performing various data processing tasks, including matrix multiplication. For example, the device 200 can be implemented as a graphics processing unit 120.
[0047] like Figure 2As shown, the device 200 includes on-chip networks 126-1 and 126-2, multiple on-chip memory units, multiple shaders 122-1 to 122-3 and 202, a bus interface 130, a command engine 132, and a state information storage area associated with the command engine 132. The multiple on-chip memory units are communicatively connected via on-chip networks (e.g., on-chip networks 126-1 and 126-2), the multiple shaders 122-1 to 122-3 and 202 are communicatively connected to the multiple on-chip memory units via on-chip networks (e.g., on-chip networks 126-1 and 126-2), and the command engine 132 is communicatively connected to the multiple shaders 122-1 to 122-3 and 202 via on-chip networks (e.g., on-chip networks 126-1 and 126-2). The multiple on-chip memory units... Figure 2 The image is shown as a collection of multiple static random access memory (SRAM) cells, wherein some on-chip memory cells (e.g., on-chip memory cells 124-5 to 124-6) can be associated with each shader to store a corresponding portion of the first data required by the shader to perform matrix multiplication operations, and other on-chip memory cells (e.g., on-chip memory cells 124-1 to 124-4, etc.) can serve as a common memory cell to store second data loaded from the outside via bus interface 130.
[0048] In some examples, on-chip networks (e.g., on-chip networks 126-1 and 126-2) may include two-dimensional or three-dimensional on-chip networks, such as 2D mesh structures, 3D mesh structures, or combinations thereof, to provide high-bandwidth, low-latency data transmission paths between the functional units of device 200. In this way, device 200 can more efficiently complete data exchange between multiple shaders 122-1 to 122-3 and 202 and multiple on-chip memory units, reducing latency and bandwidth consumption introduced by data transmission.
[0049] In some examples, each of the multiple shaders 122-1 to 122-3 and 202 can be linked via a bypass path to at least two on-chip memory units. For example, shader 122-1 can be linked via a bypass path to at least two on-chip memory units (e.g., on-chip memory units 124-5 and 124-6), which are configured to alternately store data required for performing matrix multiplication operations. In this way, shader 122-1 can directly read the required data from adjacent on-chip memory units with lower latency, and the at least two on-chip memory units can operate in a ping-pong manner, with one on-chip memory unit being read while the other is being written, thereby achieving parallel data loading and computation and further reducing data latency.
[0050] In some examples, multiple shaders 122-1 to 122-3 and 202 can also be configured to retrieve data required for performing matrix multiplication operations from multiple on-chip memory units based on address information. The address information can be direct or indirect addresses. The shader can directly initiate memory access requests to the on-chip memory units based on the address information and use the read data for the matrix multiplication operation. In this way, the shader no longer needs to pre-allocate registers for the operation data before performing the operation; instead, it can read the data on demand based on the address information, reducing the pre-constraints on data layout on the software side and improving the flexibility of data usage.
[0051] Further reference Figure 2 Each of the multiple shaders 122-1 to 122-3 and 202 (e.g., shader 122-1) may include: an instruction scheduling unit 212, a temporary cache 214, a crossbar 216, data queues 218-1 to 218-8, and arithmetic units 220-1 to 220-3. The temporary cache 214 is configured to store data received by the corresponding shader from the on-chip network (e.g., on-chip network 126-1 or 126-2). In some examples, the temporary cache 214 may be implemented as a source / destination cache to cache data received from other shaders or on-chip storage units. Data queues 218-1 to 218-8 are configured to store operands to be sent to the arithmetic units, the operands being associated with first data and second data. In some examples, data queues 218-1 to 218-8 may be implemented as multiple reservation stations. Operation units 220-1 to 220-3 are configured to perform matrix multiplication in response to a data queue containing a first operand corresponding to first data and a second operand corresponding to second data. Figure 2 In this structure, the arithmetic units 220-1 to 220-3 may specifically include one or more of the following: Special Function Units (SFUs), Load / Store Units (LDSTs), integer and floating-point arithmetic units (e.g., FP32 / INT32 units), and dedicated tensor arithmetic units (e.g., Tensor FP16 units). Through this internal structure, the shader 122-1 can, without relying on a traditional register file, collaboratively complete the reception, temporary storage, and transmission of arithmetic data via temporary buffer 214 and data queues 218-1 to 218-8, alleviating waiting problems caused by inconsistent data arrival times and supporting efficient transmission of the arithmetic units.
[0052] In some examples, each of the multiple shaders 122-1 to 122-3 and 202 may further include an instruction scheduling unit 212, which may be configured to: determine the current execution mode of the corresponding shader based on a first flag; determine whether the corresponding shader unit is currently able to schedule instructions for its own thread bundle based on a second flag; and in the first mode, in response to the corresponding shader being assigned to a second shader, suspend the sending of instructions for the current thread bundle of the second shader and schedule the execution of instructions from the first shader in the second shader. In some examples, the first flag may correspond to the following in conjunction with... Figure 4 and Figure 5 The described flow mode valid flag field indicates whether the shader is in first mode (i.e., flow mode); the second flag may correspond to what will be combined below. Figure 4 and Figure 5 The described scheduling ready flag field indicates whether the shader is currently able to schedule its own thread bundle instructions. In this way, flexible switching between first and second modes for multiple shaders and fine-grained control over thread bundle scheduling can be achieved with fewer hardware status bits, without significantly altering the overall scheduling logic of the shaders.
[0053] In some examples, the command engine 132 in device 200 can act as a central scheduling unit, communicating with multiple shaders 122-1 to 122-3 and 202 via an on-chip network (e.g., on-chip networks 126-1 to 126-2). The command engine 132 can maintain relevant state information for indicating a first mode, such as those described below. Figure 3 Further description of the stream shader status information, which may record the master shader identifier (e.g., Master_shader_ID), workgroup identifier (WG_ID), and the current program counter (PC) value of the first shader, to support centralized management of instruction scheduling in the first mode.
[0054] In some examples, the first shader in device 200 can also be configured to issue multiple second instructions without waiting for all second data to be loaded; and the central scheduling unit (e.g., command engine 132) can also be configured to update the first value of the program counter of at least one second shader to the second value of the program counter of the first shader in response to the completion of the matrix multiplication operation. In this way, the first shader can continue to issue second instructions to drive multiple shaders to perform subsequent matrix multiplication operations even when the previous set of second data has been loaded and the next set of second data has not yet been loaded, realizing the parallel advancement of data loading and matrix multiplication operations, effectively hiding the latency introduced by data transmission; and after the matrix multiplication operation is completed, by uniformly aligning the program counters of at least one second shader to the program counter value of the first shader, it can be ensured that multiple shaders can then smoothly recover from the first mode to the second mode and continue to independently schedule subsequent instructions.
[0055] The modules included in device 200 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more modules can be implemented using software and / or firmware, such as machine-executable instructions stored on a storage medium. In addition to or as an alternative to machine-executable instructions, some or all of the units in device 200 can be implemented at least partially by one or more hardware logic components. By way of example and not limitation, exemplary types of hardware logic components that can be used include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard parts (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), and so on.
[0056] To facilitate understanding of the execution process of the above scheme, the following will be combined with... Figure 3 A sample data loading process for matrix multiplication is further described. Figure 3 The diagram illustrates example data loading procedures for matrix multiplication operations in some embodiments. For example... Figure 3As shown, in the example environment, the graphics processing unit 120 includes nine shaders, which may include shaders 0 to 8, each shader being associated with an on-chip memory unit for storing a corresponding portion B[0] to B[8] of matrix B. The graphics processing unit 120 also includes multiple on-chip memory units (including on-chip memory units 124-1 to 124-4) for storing multiple portions A[0] to A[8] of matrix A, a bus interface 130, a command engine 132, and a stream shader status recording area 134. It should be understood that... Figure 3 The number of shaders, the number of on-chip storage units, and the granularity of data partitioning shown are merely examples. In other examples, the number of shaders, the number of on-chip storage units, and the granularity of matrix partitioning can be adjusted according to actual needs.
[0057] In the example data loading process, the first data can correspond to matrix B, and the second data can correspond to matrix A. Specifically, matrix B can be divided into multiple parts B[0], B[1], ..., B[8], and loaded and stored in the on-chip storage units associated with shader 0 to shader 8 respectively; matrix A can be divided into multiple parts A[0], A[1], ..., A[n], and loaded sequentially into a common storage unit (e.g., the corresponding on-chip storage unit), and then broadcast to all shaders via the on-chip network. In some examples, the first data (matrix B) corresponds to the first input matrix in the matrix multiplication operation, and the second data (matrix A) corresponds to the second input matrix in the matrix multiplication operation. The first input matrix is reused by multiple shaders in the matrix multiplication operation, that is, B[0] can be repeatedly used in the operation by multiple parts of A[0], A[1], ..., A[n], thereby realizing data reuse.
[0058] The following is combined with Figure 3 Describe the instruction scheduling process in detail.
[0059] During the loading phase, shaders 0 through 8 each issue a `load B[n]` instruction to bus interface 130 and store their respective loaded `B[n]` (different parts of the first data) into the on-chip memory associated with their respective shader. Subsequently, shaders 0 through 8 jointly execute a barrier operation to wait for all parts of `B[n]` (different parts of the first data) to be successfully returned to their respective on-chip memory. In this way, multiple shaders can load the first data in parallel, improving data loading efficiency.
[0060] After the loading phase is complete, shaders 0 through 8 jointly execute the `stream_start` instruction to notify command engine 132 to enter the first mode (i.e., streaming mode). In response to the `stream_start` instruction, command engine 132 selects one shader from multiple shaders as the first shader (e.g., shader 202, also known as shader 7), and records the identifier of the selected first shader and its workgroup identifier in the status field of status record area 134. Simultaneously, command engine 132 sends halt instructions to the other shaders (i.e., second shaders) to suspend the instruction sending of their respective current thread bundles. Halting the instruction sending of a thread bundle can manifest as suspending the program counter of the second shader and / or suspending the instruction sending of the current thread bundle of at least one second shader. The second shaders may include shaders 302-1 and 302-2. During this period, the thread bundle scheduler of the second shader can select other non-streaming mode thread bundles to continue execution, ensuring that the overall computing resources of graphics processing unit 120 can still be utilized.
[0061] Subsequently, the first shader (i.e., shader 7) issues a first instruction to load A[0] and attaches an mbarrier instruction to wait for the data of A[0] to be returned to the corresponding common memory unit (in Figure 3 In the example, i.e., on-chip storage unit 124-1), the first instruction can be a Tensor Memory Accelerator (TMA) instruction; at the same time as or after A[0] begins loading, the first shader can also continue to issue a first instruction (e.g., a Tensor Memory Accelerator instruction) to load A[1] and attach an mbarrier instruction to wait for the data of A[1] to return to the corresponding common storage unit (in Figure 3 In the example, on-chip storage unit 124-2). After the data of A[0] is returned to on-chip storage unit 124-1, the mbarrier corresponding to A[0] is released, and the first shader can issue multiple second instructions (such as matrix multiplication and addition instructions) to perform matrix multiplication based on A[0] and B[n] stored in each shader; at this time, the first shader does not need to wait for A[1] to be loaded before it can start the matrix multiplication operation based on A[0], thereby realizing the parallel advancement of matrix multiplication operation and second data loading.
[0062] During the execution of the second instruction (e.g., matrix multiplication-addition instruction) by the first shader, the first shader initiates a read request to a common memory unit (e.g., on-chip memory unit 124-1) and initiates the second instruction (e.g., matrix multiplication-addition instruction). The A[0] data read from the common memory unit, along with the second instruction (e.g., matrix multiplication-addition instruction), is broadcast to all shaders via the on-chip network according to a predetermined broadcast route. After receiving the broadcast A[0] data, each shader stores the data in its own temporary cache (e.g., a reservation station or source / destination cache); and after receiving the second instruction (e.g., matrix multiplication-addition instruction), each shader initiates a read request to the local adjacent on-chip memory unit based on the corresponding part of its stored B[n], and stores the read B[n] data in the data queue (e.g., a reservation station) based on the address information. When both operands (i.e., the part B[n] corresponding to the first data and the part A[0] corresponding to the second data) are collected, each shader sends the set of operands to the arithmetic unit (e.g., ALU or tensor arithmetic unit) to perform matrix multiplication. In this way, each shader can directly read data from the on-chip memory based on the address information and issue it to the arithmetic unit via a temporary cache and a data queue, without having to load the data into the register file first and then read it out, thus reducing data transfer overhead. Furthermore, through the coordinated work of the temporary cache and the data queue, the problem of inconsistent arrival times of multiple operands due to different transmission paths can be solved, ensuring the stability of the arithmetic unit's issue timing.
[0063] In the above operation, only the program counter of the first shader increments with the execution of instructions, while the program counters of the other shaders remain unchanged, receiving instructions only from the first shader and executing the corresponding operations. For example, in Figure 3 In the process, the program counter of shader 7 increments from value 5 (corresponding to pc5) to value 50 (corresponding to pc50) after the instruction sequence following stream_start, while the program counters of other shaders (such as shader 0) remain at the value of a certain pause moment (e.g., 5) during this period, and are uniformly updated to the program counter value corresponding to the first shader (e.g., 50) by the command engine 132 when exiting the stream mode.
[0064] After the data of A[1] is returned to the on-chip storage unit 124-2, the mbarrier corresponding to A[1] is released, and the first shader can continue to issue the first instruction to load A[2] and attach the mbarrier instruction to wait for the data of A[2] to be returned to the corresponding common storage unit (e.g., on-chip storage unit 124-3). Similarly, while A[2] starts loading, the first shader does not need to wait for A[2] to finish loading before it can continue to issue the second instruction for A[1], thereby maintaining the parallel progress of matrix multiplication and second data loading.
[0065] In some examples, the second data may include a first part (e.g., A[0]) and a second part (e.g., A[1]). In this case, the process of loading the second data into the common storage unit includes: loading the first part of the second data into the first storage unit (e.g., on-chip storage unit 124-1) in the common storage unit; in response to the completion of loading the first part, initiating the second instruction using the first shader; and loading the second part of the second data into the second storage unit (e.g., on-chip storage unit 124-2) in the common storage unit during the matrix multiplication operation performed by multiple shaders. In this way, different parts of the second data can be stored alternately in multiple storage units in the common storage unit, and the loading process of different parts and the matrix multiplication operation process overlap in time, further improving the parallelism of data loading and matrix multiplication operation and shortening the overall operation time.
[0066] Once all parts of matrix A have been loaded and their corresponding matrix multiplication operations completed, the first shader (e.g., shader 7) sends a `stream_exit` instruction to notify the command engine 132 that the calculation is complete. In response to the `stream_exit` instruction, the command engine 132 updates the current value (e.g., 5) of the program counters of all shaders except the first shader to the value of the first shader's program counter (e.g., 50), and returns the updated program counter value to the thread bundle scheduler of each shader, allowing each shader to return to the second mode. At this point, each shader can continue to independently schedule its own thread bundle instructions and write the result D[n] of the matrix multiplication operation to the corresponding on-chip storage unit. In some examples, the result D[n] can be stored in the on-chip storage unit that stores matrix B[n] (e.g., if B[n] does not fill the storage unit), or it can be stored in another on-chip storage unit associated with that shader. In this way, the result of the matrix multiplication operation can be written back to the on-chip storage unit in a timely manner for subsequent storage or subsequent calculations.
[0067] In some examples, throughout the matrix multiplication operation, the initial data is consistently stored in multiple on-chip memory units associated with each shader, and the result of the matrix multiplication operation can also be stored in multiple on-chip memory units. This approach avoids repeated loading of the initial data and frequent access to multi-level caches, reducing on-chip storage resource consumption and improving data access efficiency.
[0068] Through the above combination Figure 3In the exemplary data loading process, the electronic device 110 can pre-distribute first data (e.g., matrix B) in multiple on-chip storage units of multiple shaders via the graphics processing unit 120. The first shader then loads the second data (e.g., matrix A) and broadcasts it to all shaders simultaneously via the on-chip network. This allows multiple shaders to perform matrix multiplication operations in parallel based on the shared second data and the corresponding portions of the first data stored in the multiple on-chip storage units. Compared to the traditional approach where each shader independently reads shared data, this method significantly reduces the number of times the second data is repeatedly read on-chip, reduces on-chip bandwidth usage, and further improves overall computational throughput through the parallel execution of data loading and matrix multiplication operations.
[0069] Further reference below Figure 4 and Figure 5 Describe the specific working process of the instruction scheduling unit inside the shader in the first mode. Figure 4 A schematic structural block diagram of a first instruction scheduling unit 400 in a first shader of some embodiments is shown. Figure 5 A schematic block diagram of a second instruction scheduling unit 500 in a second shader of some embodiments is shown.
[0070] like Figure 4 As shown, the first shader includes a first instruction scheduling unit 400, which includes a first instruction buffer unit 402, a first dependency checking unit 406, a first dispatch unit 410, and a first tensor memory accelerator 412.
[0071] The first instruction scheduling unit 400 can select a suitable thread bundle for instruction scheduling based on the current state information of multiple thread bundles (e.g., Wrap0, Wrap1, Wrap2, ..., Wrap31); the first instruction buffer unit 402 can be used to cache the instructions to be executed by each thread bundle; the first dependency checking unit 406 can maintain a dependency count vector (Dep cnt vec), a wait count check (Wait cnt chk), and a memory counter (Memcnt) for different thread bundles to determine whether the corresponding thread bundle currently meets the execution conditions; the first dispatch unit 410 is used to dispatch instructions after the dependency relationship is satisfied to the operation unit; and the first tensor memory accelerator 412 is used to record relevant information of the TMA request currently being processed (In-flight), such as the in-process memory counter (Inflight mem cnt).
[0072] In the first shader, the first dispatch unit 410 can maintain state information entries for each thread bundle. Each entry can include fields such as thread bundle identifier (Wrap_ID), workgroup identifier (WG_ID), program counter (PC), stream mode valid flag (stream_vld), scheduling ready flag (ready), and reserved station identifier (station_ID). The stream mode valid flag (stream_vld) indicates whether the corresponding thread bundle is in the first mode, and the scheduling ready flag (ready) indicates whether the corresponding thread bundle can currently schedule its own instructions. For the shader selected as the first shader, upon entering the first mode, the stream mode valid flag (stream_vld) of its corresponding thread bundle can be set to 1, and the scheduling ready flag (ready) can remain at 1, allowing the thread bundle to continue autonomously issuing its own instructions (e.g., TMA instructions, MMA instructions, and stream_exit instructions). For example, in... Figure 4 In the first mode, the program counter value of a certain thread bundle is 0, 1, 2, ..., 50, 51 in sequence, that is, the instruction sequence is executed in the order of program counter increment.
[0073] like Figure 5 As shown, the second shader includes... Figure 4 Similar components include, for example, the second instruction scheduling unit 500, the second instruction buffer unit 502, the second dependency checking unit 506, the second dispatch unit 510, and the second tensor memory accelerator 512. Figure 4 The difference lies in the fact that, in the second shader, since instruction scheduling is suspended in the first mode, the state information of the corresponding thread bundle changes accordingly. For example, in the second dispatch unit 510, the program counter value of a thread bundle can go through 0, 1, 2 (corresponding to instructions before stream_start), remain at 2 during the first mode, jump to 50 after being triggered by the stream_exit instruction, and then continue to increment to 51; the scheduling ready flag can go through 1 (before stream_start), 0 (no longer scheduling its own instructions during the first mode), and 1 (restored after the first mode ends); the stream mode valid flag (stream_vld) is set to 1 during the first mode and set to 0 after the first mode ends.
[0074] In some examples, adjusting multiple shaders to first mode may include one or more of the following: pausing the program counter of at least one second shader; and / or pausing instruction sending for the current thread bundle of at least one second shader. Combined Figure 5For example, in the first mode, the program counter value of the current thread bundle of the second shader is paused (e.g., held at a value of 2), and the instruction sending of the current thread bundle of the second shader is also paused (e.g., the scheduling ready flag is set to 0). In this way, the second shader can avoid conflicting with the instructions of the first shader in the first mode by neither issuing its own instructions nor updating its own program counter, and reduce the control complexity of the device's internal state machine.
[0075] In some examples, in the first mode, in response to the second shader receiving an instruction (e.g., an MMA instruction transmitted via broadcast) and its corresponding second data from the first shader, the second shader can read the portion corresponding to the first data from its adjacent on-chip memory based on the instruction from the first shader, and perform a matrix multiplication operation based on the corresponding portion of the first data and the broadcast second data. During this process, the second shader does not need to update its own program counter, and the instructions executed are entirely from the first shader, thus achieving synchronous multiplexing of the second data by multiple shaders. In other examples, even if the instruction sending of the second shader's current thread bundle is paused, the second instruction scheduling unit 500 can still select other non-streaming thread bundles to continue execution, ensuring that the internal computing resources of the shader are still fully utilized.
[0076] In some examples, after matrix multiplication is completed, multiple shaders can be switched from a first mode to a second mode, which indicates the resumption of instruction scheduling for at least one second shader. In some examples, in response to multiple shaders being switched to the second mode, the first value (e.g., value 2) of the program counter of at least one second shader can be updated to a second value (e.g., value 50) corresponding to the program counter of the first shader. In this way, the program counters of each second shader can be uniformly aligned to the program counter value of the first shader after the first mode ends, so that multiple shaders can then smoothly resume to the second mode, which independently schedules their own instructions, and continue to execute subsequent program flows (e.g., the store D[n] instruction corresponding to pc51), ensuring the overall consistency of program flows for multiple shaders and reducing the software's awareness and programming burden of mode switching.
[0077] Example process Figure 6 A flowchart of an example process 600 for data processing according to some embodiments is shown. Process 600 can be implemented as described above. Figure 1 The described electronic device 110, for example, can be implemented at the graphics processing unit 120 deployed in the electronic device 110. See below for reference. Figure 6 To describe process 600.
[0078] like Figure 6 As shown, in step 610, the electronic device 110 uses multiple shaders to obtain multiple portions of the first data.
[0079] In step 620, the electronic device 110 stores multiple portions of the first data into multiple on-chip memory units associated with multiple shaders.
[0080] In step 630, in response to the completion of the storage of the first data, the electronic device 110 adjusts the plurality of shaders to a first mode, the first mode indicating that the plurality of shaders are divided into a first shader and at least one second shader, and the instruction scheduling of at least one second shader is suspended.
[0081] In step 640, the electronic device 110, in response to a first instruction initiated by the first shader, loads the second data into the common storage unit.
[0082] In step 650, in response to a second instruction initiated by the first shader, the electronic device 110 broadcasts the second data in the common memory unit to multiple shaders using the on-chip network.
[0083] In step 660, the electronic device 110 uses multiple shaders to perform matrix multiplication based on the corresponding portions of the second data and the first data stored in multiple on-chip memory units.
[0084] In some embodiments, process 600 further includes: obtaining the result of matrix multiplication; and storing the result in a plurality of on-chip storage units.
[0085] In this way, the results of matrix multiplication can be directly stored in multiple on-chip storage units associated with each shader, so that subsequent operations can directly obtain the results from the on-chip storage units without going through multiple levels of cache or writing back to off-chip storage. This further reduces the overhead of transporting the results in the storage path and improves the reusability of the results in subsequent operations.
[0086] In some embodiments, each of the plurality of shaders may be configured with at least two on-chip storage units, which are configured to alternately store first data and second data.
[0087] In this way, at least two on-chip storage units can work in a ping-pong manner: while one on-chip storage unit is used to support the current operation, the other on-chip storage unit can preload the data required for the next round of operation, thereby further overlapping data loading and matrix multiplication operations in time and improving the parallelism of data loading and operation.
[0088] In some embodiments, adjusting multiple shaders to a first mode includes: pausing the program counter of at least one second shader; and / or pausing instruction sending for the current thread bundle of at least one second shader.
[0089] In this way, fine-grained control over the instruction scheduling of the second shader in the first mode can be achieved with fewer hardware state changes. This ensures that the second shader neither issues its own instructions nor updates its own program counter in the first mode, thereby avoiding conflicts with the instructions of the first shader in the first mode and reducing the control complexity of the internal state machine of the device.
[0090] In some embodiments, process 600 further includes: in response to the completion of matrix multiplication, adjusting a plurality of shaders from a first mode to a second mode, the second mode indicating the resumption of instruction scheduling for at least one second shader.
[0091] In this way, after the matrix multiplication operation is completed, each shader can independently reschedule its own instructions and continue to execute the subsequent program flow. This makes the device highly flexible in switching between different modes and can select the appropriate execution mode according to the characteristics of the subsequent operations.
[0092] In some embodiments, process 600 further includes: in response to a plurality of shaders being adjusted to a second mode, updating a first value of the program counter of at least one second shader to a second value corresponding to the program counter of the first shader.
[0093] In this way, the program counters of each second shader can be aligned to the program counter value of the first shader after the first mode ends. This ensures that multiple shaders can continue to execute subsequent instructions based on a consistent program flow position after exiting the first mode, reducing the synchronization overhead caused by inconsistencies in the program flow of each shader, and simplifying the software's perception and processing of mode switching.
[0094] In some embodiments, the common storage unit includes a plurality of storage units, and loading the second data into the common storage unit includes: loading a first portion of the second data into a first storage unit among the plurality of storage units; in response to the completion of loading the first portion, initiating a second instruction using a first shader; and during the matrix multiplication operation performed by the plurality of shaders, loading a second portion of the second data into a second storage unit among the plurality of storage units.
[0095] In this way, multiple storage units in the common storage unit can be used to alternately handle the loading and reading of different parts of the second data, and the loading process of the subsequent part of the second data and the matrix multiplication operation based on the preceding part of the second data can overlap in time, further improving the parallelism of data loading and matrix multiplication operations, and effectively hiding the latency introduced by the loading of the second data.
[0096] In some embodiments, multiple shaders are deployed with temporary caches and data queues, and each shader is configured to: in response to receiving second data, store the second data in a temporary cache; obtain the corresponding portion of the first data from an on-chip storage unit based on address information; store the corresponding portion of the first data and the second data in a data queue; and perform a matrix multiplication operation based on the corresponding portion of the second data and the first data.
[0097] In this way, each shader can use temporary buffers and data queues to work together to receive, temporarily store and transmit operands, which alleviates the problem of inconsistent operand arrival times caused by different transmission paths. It also eliminates the need for pre-allocating operation data in traditional register files, reducing the overhead of data handling within the shader and the complexity of hardware implementation.
[0098] In some embodiments, the first data may correspond to the first input matrix in a matrix multiplication operation, and the second data may correspond to the second input matrix in a matrix multiplication operation. The first input matrix may be reused by multiple shaders in a matrix multiplication operation.
[0099] In this way, the data reuse characteristics of the first input matrix in matrix multiplication can be fully utilized. The first input matrix is pre-distributed and stored in multiple shaders for repeated access, and the second input matrix is loaded by the first shader and broadcast to multiple shaders, thereby maximizing data reuse efficiency and significantly reducing the number of times the first input matrix is repeatedly read and the corresponding on-chip bandwidth usage.
[0100] Example devices and equipment Figure 7 A block diagram of a computing device 700 in which various embodiments of the present disclosure may be implemented is shown. The computing device 700 may be implemented as an electronic device 110, or may be included in an electronic device 110.
[0101] It should be understood that, Figure 7 The computing device 700 shown is for illustrative purposes only and is not intended to imply any limitation on the functionality and scope of the embodiments of this disclosure.
[0102] like Figure 7As shown, the computing device 700 includes a general-purpose computing device 700. The computing device 700 may include at least one or more processors or processing units 710, memory 720, storage units 730, one or more communication units 740, one or more input devices 750, and one or more output devices 760.
[0103] In some embodiments, the computing device 700 can be implemented as any user terminal or server terminal with computing capabilities. The server terminal can be a server, large computing device, etc., provided by a service provider. The user terminal can be, for example, any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, stations, units, devices, multimedia computers, multimedia tablet computers, internet nodes, communicators, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices, or any combination thereof. It is conceivable that the computing device 700 can support any type of interface to the user (such as "wearable" circuitry devices, etc.).
[0104] The processing unit 710 can be a physical processor or a virtual processor, and can perform various processes based on programs stored in the memory 720. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of the computing device 700. The processing unit 710 may also be referred to as a central processing unit (CPU), a microprocessor, a controller, or a microcontroller.
[0105] Computing device 700 typically includes various computer storage media. Such media can be any media accessible by computing device 700, including but not limited to volatile and non-volatile media, or removable and non-removable media. Memory 720 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), or flash memory) or any combination thereof. Storage cell 730 can be any removable or non-removable media and may include machine-readable media, such as memory, flash drives, disks, or other media that can be used to store information and / or data and can be accessed within computing device 700.
[0106] The computing device 700 may also include additional removable / non-removable storage media, volatile / non-volatile storage media. Although in Figure 7Not shown, but a disk drive for reading from and / or writing to a removable non-volatile disk, and an optical disc drive for reading from and / or writing to a removable non-volatile optical disc may be provided. In this case, each drive may be connected to a bus (not shown) via one or more data media interfaces.
[0107] The communication unit 740 communicates with another computing device via a communication medium. Furthermore, the functionality of the components in the computing device 700 can be implemented by a single computing cluster or multiple computing machines that can communicate via communication connections. Therefore, the computing device 700 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or other general-purpose network nodes.
[0108] Input device 750 can be one or more of a variety of input devices, such as a mouse, keyboard, trackball, voice input device, etc. Output device 760 can be one or more of a variety of output devices, such as a monitor, speaker, printer, etc. With the aid of communication unit 740, computing device 700 can also communicate with one or more external devices (not shown), such as storage devices and display devices. Computing device 700 can also communicate with one or more devices that enable a user to interact with computing device 700, or, if needed, with any device (e.g., network card, modem, etc.) that enables computing device 700 to communicate with one or more other computing devices. This communication can be performed via an input / output (I / O) interface (not shown).
[0109] In some embodiments, some or all of the components of computing device 700 may be arranged in a cloud computing architecture, rather than integrated into a single device. In a cloud computing architecture, components may be remotely provided and work together to achieve the functionality described herein. In some embodiments, cloud computing provides computing, software, data access, and storage services without requiring end users to know the physical location or configuration of the systems or hardware providing these services. In various embodiments, cloud computing provides services via a wide area network (WAN), such as the Internet, using suitable protocols. For example, a cloud computing provider provides applications via a WAN that can be accessed through a web browser or any other computing component. The software or components of the cloud computing architecture, along with the corresponding data, may be stored on servers at remote locations. Computing resources in a cloud computing environment may be consolidated or distributed across remote data center locations. Cloud computing infrastructure may provide services through shared data centers, although to users they appear as a single access point. Therefore, a cloud computing architecture can be used to provide the components and functionality described herein from service providers at remote locations. Alternatively, the components and functionality described herein may be provided by conventional servers or installed directly or otherwise on client devices.
[0110] In embodiments of this disclosure, computing device 700 may be used to perform data processing. Memory 720 may include one or more data processing modules 725 having one or more program instructions. These modules are accessible and executable by processing unit 710 to perform the functions of the various embodiments described herein.
[0111] In an example embodiment where data processing is performed, the virtual scene may be processed, for example, by the data processing module 725 to generate a picture of the virtual scene or a bitstream corresponding to the picture of the virtual scene. The picture or bitstream of the virtual scene may be provided as output 780 via output device 760.
[0112] While this disclosure has been specifically shown and described with reference to preferred embodiments, those skilled in the art will understand that various changes in form and detail may be made without departing from the spirit and scope of this application as defined by the appended claims. These variations are intended to be covered by the scope of this application. Therefore, the foregoing description of embodiments of this application is not intended to be limiting.
Claims
1. A data processing method, characterized in that, The method includes: Multiple shaders are used to obtain multiple parts of the first data respectively; The plurality of portions of the first data are respectively stored in a plurality of on-chip memory units associated with the plurality of shaders; In response to the completion of the storage of the first data, the plurality of shaders are adjusted to a first mode, the first mode indicating that the plurality of shaders are divided into a first shader and at least one second shader, and the instruction scheduling of the at least one second shader is suspended; In response to the first instruction initiated by the first shader, the second data is loaded into the common storage unit; In response to a second instruction initiated by the first shader, the second data in the common memory unit is broadcast to the plurality of shaders via the on-chip network; and Using the plurality of shaders, matrix multiplication is performed based on the corresponding portions of the second data and the first data stored in the plurality of on-chip memory units.
2. The method according to claim 1, characterized in that, The method further includes: Obtain the result of the matrix multiplication operation; and The calculation results are stored in the plurality of on-chip storage units.
3. The method according to claim 1, characterized in that, Each of the plurality of shaders is configured with at least two on-chip storage units, which are configured to alternately store the first data and the second data.
4. The method according to claim 1, characterized in that, Adjusting the plurality of shaders to the first mode includes: Pause the program counter of at least one second shader; and / or Pause the sending of instructions for the current thread bundle of the at least one second shader.
5. The method according to claim 1, characterized in that, The method further includes: In response to the completion of the matrix multiplication operation, the plurality of shaders are switched from the first mode to a second mode, the second mode indicating the resumption of instruction scheduling for the at least one second shader.
6. The method according to claim 5, characterized in that, The method further includes: In response to the plurality of shaders being adjusted to the second mode, the first value of the program counter of the at least one second shader is updated to a second value corresponding to the program counter of the first shader.
7. The method according to claim 1, characterized in that, The common storage unit includes multiple storage units, and loading the second data into the common storage unit includes: The first portion of the second data is loaded into the first storage unit among the plurality of storage units; In response to the completion of the first part loading, the second instruction is initiated using the first shader; and During the matrix multiplication operation performed by the plurality of shaders, the second portion of the second data is loaded into the second storage unit among the plurality of storage units.
8. The method according to claim 1, characterized in that, Each of the plurality of shaders is deployed with a temporary cache and a data queue, and each of the plurality of shaders is configured as follows: In response to receiving the second data, the second data is stored in the temporary cache; Based on the address information, the corresponding part of the first data is obtained from the on-chip storage unit; Store the corresponding portion of the first data and the second data into the data queue; as well as The matrix multiplication operation is performed based on the corresponding portions of the second data and the first data.
9. The method according to claim 1, characterized in that, The first data corresponds to the first input matrix in the matrix multiplication operation, and the second data corresponds to the second input matrix in the matrix multiplication operation. The corresponding parts of the first input matrix are reused by the plurality of shaders in the matrix multiplication operation.
10. A data processing apparatus, characterized in that, The device includes: On-screen network; Multiple storage units, which are connected via the on-chip network; Multiple shaders, which are communicatively connected to the multiple memory units via the on-chip network; and A central scheduling unit, which communicates with the plurality of shaders through the on-chip network; The plurality of shaders are configured to perform matrix multiplication based on the corresponding portions of the second data and the first data stored in the plurality of storage units; The central scheduling unit is configured to: adjust the plurality of shaders to a first mode to divide the plurality of shaders into a first shader and at least one second shader, and suspend instruction scheduling of the at least one second shader; The first shader is configured to: issue a first instruction and a second instruction, the first instruction instructing the loading of second data into the plurality of storage units, and the second instruction instructing the broadcasting of the second data to the plurality of shaders using the on-chip network.
11. The apparatus according to claim 10, characterized in that, Each of the plurality of shaders includes: A temporary cache, configured to store data received by the corresponding shader from the on-chip network; A data queue, configured to store operands to be sent to the processing unit, the operands being associated with the first data and the second data; and The arithmetic unit is configured to perform matrix multiplication in response to the data queue storing a first operand corresponding to the first data and a second operand corresponding to the second data.
12. The apparatus according to claim 10, characterized in that, The plurality of memory units include a plurality of on-chip memory units, and the plurality of shaders are further configured to: Based on the address information, the data required to perform the matrix multiplication operation is obtained from the plurality of on-chip storage units.
13. The apparatus according to claim 10, characterized in that, Each of the plurality of shaders is linked via a bypass path to at least two of the plurality of storage units, which are configured to alternately store the data required to perform the matrix multiplication operation.
14. The apparatus according to claim 10, characterized in that, The on-chip network includes a two-dimensional on-chip network or a three-dimensional on-chip network.
15. The apparatus according to claim 10, characterized in that, Each of the plurality of shaders includes an instruction scheduling unit, the instruction scheduling unit being configured to: Based on the first flag, the current execution mode of the corresponding shader is determined; Based on the second flag, determine whether the corresponding shader is currently able to schedule instructions for its own thread bundle; as well as In the first mode, in response to the corresponding shader being divided into a second shader, the instruction sending of the current thread bundle of the second shader is suspended, and the second shader is scheduled to execute instructions from the first shader.
16. The apparatus according to claim 10, characterized in that, The first shader is also configured to: Initiate multiple second instructions without waiting for all second data to finish loading; and The central scheduling unit is further configured to update the first value corresponding to the program counter of the at least one second shader to the second value corresponding to the program counter of the first shader in response to the end of the matrix multiplication operation.
17. A non-transitory computer-readable storage medium for storing instructions, characterized in that, The instructions cause the processor to execute the instructions of the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Dynamic integrity verification for shaders processing workloads
CN120632852A
Memory shader
CN121599823A