An adaptive scheduling reconfigurable matrix operation architecture
Patent Information
- Application Number
- CN202610909371.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-23
- Publication Date
- 2026-09-11
AI Technical Summary
随着应用场景多样化,矩阵运算呈现多精度、多尺寸、低时延、高吞吐的混合需求,传统固定架构加速器面临以下技术瓶颈:1、仅支持定点 INT8/INT16 或浮点 FP16/FP32 单一精度,无法动态切换定点与浮点运算,难以适配混合精度神经网络与多场景算法需求
Smart Images

Figure CN122734239A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of integrated circuit design and matrix operation acceleration technology, and in particular to an adaptive scheduling reconfigurable matrix operation architecture. Background Technology
[0002] Matrix multiplication and accumulation are core operations in artificial intelligence neural network inference / training, digital signal processing, image processing, and scientific computing. With the diversification of application scenarios, matrix operations present mixed requirements for multi-precision, multi-size, low latency, and high throughput. Traditional fixed-architecture accelerators face the following technical bottlenecks: 1. They only support single-precision fixed-point INT8 / INT16 or floating-point FP16 / FP32 operations, unable to dynamically switch between fixed-point and floating-point operations, making it difficult to adapt to the needs of mixed-precision neural networks and multi-scenario algorithms. 2. Traditional architectures use serial computation or coarse-grained parallelism; matrix partitioning, data transport, and computation execution cannot be pipelined and overlapped, resulting in significant memory access latency and severe latency degradation in small matrix operations. 3. Fixed computation array size, number of parallel paths, and precision mode prevent dynamic adjustment of parallelism and partitioning strategies based on matrix dimensions, leading to wasted computing power in small matrix scenarios and poor scalability in large matrix scenarios. 4. Traditional configurable architectures rely on external main cores to issue fixed configuration instructions. They cannot autonomously make adaptive decisions on the operation mode based on the characteristics of matrix operation tasks and the real-time status of the system. They require manual pre-configuration of array enablement schemes and block strategies, which cannot adapt to dynamically changing multi-task mixed scenarios. The computing power utilization and power consumption optimization effects are limited, and external configuration interaction introduces additional scheduling latency.
[0003] To address the aforementioned issues, there is an urgent need for a flexible matrix operation architecture that can achieve flexible configuration of computing power, high reuse of hardware circuits, extreme compression of computation latency, and fine-grained control of power consumption. Summary of the Invention
[0004] The purpose of this invention is to provide an adaptive scheduling reconfigurable matrix operation architecture to solve the problems in the background art.
[0005] To address the aforementioned technical problems, this invention provides an adaptive scheduling reconfigurable matrix operation architecture, including an operation scheduling module, a reusable matrix operation unit, and a general-purpose register module; The core control module of the computation scheduling module realizes adaptive decision-making and array control of the computation mode. It does not require the external main core to issue fixed computation mode configuration instructions. It can autonomously complete task feature acquisition, intelligent scheduling reasoning, mode judgment and array control, and also supports closed-loop scheduling optimization. The reusable matrix operation unit includes four sets of fixed-point operation units with identical structures and four sets of floating-point operation units with identical structures, used to perform matrix operations of corresponding precision; the hardware circuits of the four sets of fixed-point operation units are completely reused, and the hardware circuits of the four sets of floating-point operation units are completely reused, without the need to redesign the hardware for parallel operation mode or single operation mode. The general-purpose register module is used for storing and interacting with matrix operation data. It adopts a ping-pong caching mechanism to enable parallel execution of matrix data reading, writing and operation, and supports multi-step continuous matrix multiplication and accumulation operations.
[0006] In one implementation, the fixed-point arithmetic unit supports completing 8*8 INT8 and INT16 matrix multiplication and addition, matrix addition, and matrix multiplication operations within 10 cycles. Each fixed-point arithmetic unit adopts an 8×8 outer product multiplication and addition array, which consists of 64 multiplication and addition processing units forming an 8×8 two-dimensional array. Each multiplication and addition processing unit includes one 16-bit multiplier and four 32-bit adders, supporting one INT16 multiplication and addition operation or four INT8 multiplication and addition operations, realizing hardware circuit reuse for INT8 and INT16 precision.
[0007] In one embodiment, the fixed-point arithmetic unit adopts a 3-stage pipelined arithmetic design, including data fetching and register cycles, multiplication cycles, and addition and accumulation cycles, with an operating frequency of up to 1 GHz; the input design of the multiply-accumulate processing unit is such that multiply-accumulate processing units in the same row share the same row of input data of the matrix, and multiply-accumulate processing units in the same column share the same column of input data of the matrix, with an array utilization rate of 100%.
[0008] In one implementation, the floating-point arithmetic unit supports completing matrix multiplication and addition, matrix addition, and matrix multiplication operations of FP16 (M8N8K4 specification) and FP32 (M4N4K4 specification) within 10 cycles. Each floating-point arithmetic unit adopts an inner product arithmetic array, which consists of 16 inner product arithmetic units forming a 4×4 two-dimensional array. Each inner product arithmetic unit includes one multiply-add fused FMA unit, three floating-point multipliers, and three floating-point adders. It adopts a three-level tree computing topology and supports one FP32 inner product operation or four FP16 inner product operations, realizing hardware circuit reuse for FP16 and FP32 precision.
[0009] In one embodiment, the floating-point arithmetic unit adopts a 4-stage pipelined arithmetic design, including a data fetching and exponent alignment cycle, a multiplication and multiply-add fusion cycle, a reduction-addition and leading zero prediction cycle, and a normalization and anomaly detection cycle, with an operating frequency of up to 1 GHz.
[0010] In one embodiment, the computation scheduling module includes four sub-units: a task feature extraction unit, a lightweight neural network scheduling and inference unit, a pattern decision and encoding unit, and an array control signal generation unit. The task feature extraction unit is responsible for collecting relevant feature parameters of matrix operation tasks in real time, extracting and generating standardized task feature vectors and feature parameters, and providing input basis for subsequent intelligent scheduling. The lightweight neural network scheduling and inference unit adopts a hardware-friendly shallow multilayer perceptron network structure, with only 2-3 fully connected hidden layers, and is equipped with ReLU activation function and INT8 fixed-point quantization design. The lightweight neural network scheduling and inference unit receives the task feature vector and performs preprocessing, performs fast inference through pre-trained network weights, and outputs confidence data corresponding to the parallel working mode of 1-4 operation units. The mode decision and encoding unit performs optimal value filtering on the confidence data, determines the optimal number of parallel arrays for the current computing task, generates corresponding computing mode configuration instructions, and encodes the configuration instructions into hardware-recognizable control codes to achieve precise conversion of scheduling results into hardware control signals. The array control signal generation unit generates enable signals, start / stop signals, synchronization signals, and data path selection signals for the corresponding arithmetic units based on the control codes output by the mode decision and encoding unit. This precisely controls the start or stop of the corresponding number of arithmetic units, ensuring timing synchronization and conflict-free execution during parallel multi-array operations. Simultaneously, it feeds back the actual operation status to the lightweight neural network scheduling and inference unit, forming a closed-loop scheduling logic.
[0011] In one implementation, the general-purpose register module is divided into register area A, register area B, and register area C. Register area A consists of eight 1024-bit registers for storing left operand matrix data of matrix operations. Register area B consists of eight 1024-bit registers for storing right operand matrix data of matrix operations. Register area C consists of eight 2048-bit registers for storing accumulated operand matrix data and operation result matrix data of matrix operations, supporting intermediate result caching for multi-step continuous matrix multiplication and accumulation operations.
[0012] In one implementation, the ping-pong caching mechanism of the general-purpose register module divides register area A, register area B, and register area C into two cache partitions, Buffer0 and Buffer1, with completely identical capacity and structure. When the arithmetic unit reads data from Buffer0 to perform an operation, it simultaneously writes the data required for the next operation into Buffer1. After the operation is completed, the read / write state is immediately switched, the arithmetic unit reads data from Buffer1 to start the operation, and at the same time writes the operation result in Buffer0 back to external storage, thereby achieving complete parallelism between data reading / writing and arithmetic operations.
[0013] In one implementation, four sets of fixed-point arithmetic units correspond one-to-one with four sets of floating-point arithmetic units, forming four independent arithmetic paths. Each set of arithmetic paths can work independently or synchronously in parallel. When the four sets work in parallel, they can complete four times the matrix operation task of a single set of operations within the same operation cycle, thus achieving a fourfold increase in computing power.
[0014] In one embodiment, the computation scheduling module further includes an exception handling submodule and a clock gating submodule; the exception handling submodule receives computation exception signals reported by each group of computation units, encodes the exceptions and reports them to the external main core, and triggers pause and reset operations of the computation units according to the exception type. The clock gating submodule controls the input clock of each set of arithmetic units according to the enable signal: For arithmetic units with an active enable signal, the clock path is opened, providing a 1GHz operating clock. For arithmetic units with active shutdown signals, the clock path is cut off, causing the arithmetic unit to enter a low-power sleep state without clock flipping, thus completely eliminating dynamic power consumption; when fixed-point arithmetic units are working, the global clock of all floating-point arithmetic units is turned off; when floating-point arithmetic units are working, the global clock of all fixed-point arithmetic units is turned off, achieving intelligent optimization of global power consumption.
[0015] The present invention provides an adaptive scheduling and reconfigurable matrix operation architecture, which realizes flexible configuration of computing power, high reuse of hardware circuits, extreme compression of operation latency and fine control of power consumption. Attached Figure Description
[0016] Figure 1 This is an overall structural block diagram of an adaptive scheduling reconfigurable matrix operation architecture provided by the present invention.
[0017] Figure 2 This is a schematic diagram of the computing scheduling module architecture of the present invention.
[0018] Figure 3 This is a schematic diagram of the INT16 integer array operation scheme architecture of the present invention.
[0019] Figure 4 This is a schematic diagram of the architecture of the INT8 integer array operation scheme of the present invention.
[0020] Figure 5 This is a schematic diagram of the FP32 floating-point array operation scheme architecture of the present invention.
[0021] Figure 6 This is a schematic diagram of the FP16 floating-point array operation scheme architecture of the present invention. Detailed Implementation
[0022] The adaptive scheduling reconfigurable matrix operation architecture proposed in this invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. The advantages and features of this invention will become clearer from the following description. It should be noted that the accompanying drawings are all in a very simplified form and use non-precise scales, and are only used to facilitate and clarify the illustration of the embodiments of this invention.
[0023] This invention provides an adaptive scheduling and reconfigurable matrix operation architecture. The core of this architecture lies in achieving flexible computing power configuration, low latency, high hardware reuse, and low power consumption for matrix operations through configurable operation scheduling, fully reused multiple operation arrays, and a ping-pong cache register architecture.
[0024] like Figure 1 As shown, the matrix operation architecture of this invention consists of three core parts: an operation scheduling module, reusable matrix operation units, and a general-purpose register module. The operation scheduling module, as the control core of the architecture, collects task and system status characteristics in real time, performs adaptive scheduling inference, and outputs enable signals and operation control parameters to the reusable matrix operation units. The reusable matrix operation units include four sets of fixed-point operation units (denoted as fixed-point operation units 0-3) and four sets of floating-point operation units (denoted as floating-point operation units 0-3). The four sets of fixed-point operation units correspond one-to-one with the four sets of floating-point operation units (fixed-point operation unit 0 corresponds to floating-point operation unit 0, fixed-point operation unit 1 corresponds to floating-point operation unit 1, fixed-point operation unit 2 corresponds to floating-point operation unit 2, and fixed-point operation unit 3 corresponds to floating-point operation unit 3), forming four independent operation paths. Each operation path can work independently or in parallel synchronously according to the control signals of the operation scheduling module. The general-purpose register module, as the on-chip data cache core of the architecture, connects to external storage and also connects to each set of fixed-point and floating-point operation units, providing input data and caching operation results for the operation units.
[0025] The computation scheduling module is the core module of this invention for realizing adaptive decision-making of computation mode, dynamic reconfigurability of computing power, and intelligent optimization of global power consumption. Its specific structure includes four sub-units: task feature extraction unit, lightweight neural network scheduling and inference unit, mode decision and encoding unit, and array control signal generation unit. Each sub-unit works in a pipelined manner and can complete a complete scheduling decision within a single clock cycle, adapting to a system operating frequency of 1GHz.
[0026] The input terminals of the task feature extraction unit are connected to the external main core, general-purpose register module, reusable matrix operation unit, and system bus, respectively. It collects relevant feature parameters of the matrix operation task in real time and in parallel, including: input matrix dimension, data bit width type, operation task type (multiplication / addition / multiplication), remaining storage space of the general-purpose register module, working status of each path of the reusable matrix operation unit, and real-time load information of the system bus. After acquisition, the task feature extraction unit standardizes the multi-source feature parameters, converting feature parameters of different dimensions and bit widths into fixed-length standardized task feature vectors, which are then output to the lightweight neural network scheduling and inference unit, providing complete and accurate input for subsequent intelligent scheduling.
[0027] The lightweight neural network scheduling and inference unit adopts a hardware-friendly shallow multilayer perceptron (MLP) network structure, with only 2-3 fully connected hidden layers and a ReLU activation function. This eliminates the hardware overhead associated with complex network structures, achieving extreme lightweighting. After receiving standardized task feature vectors, the unit performs data preprocessing and normalization. Using pre-trained and fixed network weights in on-chip non-volatile memory, it completes fast inference in a single cycle, outputting confidence data corresponding to the parallel operation mode of 1-4 computation units. This offers advantages such as fast inference speed, low hardware resource consumption, and adaptability to real-time edge scheduling, preventing the computation scheduling module itself from becoming a system bottleneck. Simultaneously, the unit receives actual computational state data from the array control signal generation unit, performing online fine-tuning and optimization of network inference to achieve closed-loop scheduling.
[0028] The mode decision and encoding unit receives the confidence data of four parallel operating modes output by the lightweight neural network scheduling and inference unit. It then uses a maximum value filtering algorithm to select the optimal value, determining the best number of parallel arrays (1-4 groups) for the current computation task and generating corresponding computation mode configuration instructions. Simultaneously, the mode decision and encoding unit encodes the configuration instructions into hardware-recognizable fixed-format control codes, specifically covering array enable coding, computation block strategy coding, synchronization control coding, and precision mode control coding. This achieves precise conversion of scheduling results into hardware control signals, which are then output to the array control signal generation unit.
[0029] The array control signal generation unit generates enable signals, start / stop signals, synchronization signals, precision mode control signals, and data path selection signals for the corresponding arithmetic units based on the control codes output by the mode decision and encoding unit, thereby precisely controlling the start or stop of the corresponding number of arithmetic units. When the decision is a single-group operation mode, only one set of enable signals for the corresponding fixed-point or floating-point operation unit is generated, and the other three sets of operation units output disable signals. When the decision is a parallel operation mode of 2-3 groups, a synchronization enable signal is generated for the corresponding number of operation units, and the remaining units output a shutdown signal; When the decision is a 4-group fully parallel operation mode, synchronous enable signals for four groups of operation units (0-3) are generated simultaneously to ensure that the four groups of units start operation synchronously. The array control signal generation unit can guarantee the timing synchronization and conflict-free execution during multi-array parallel operation, and at the same time, feed back the actual operation status and array working status to the lightweight neural network scheduling and inference unit in real time to form a complete closed-loop scheduling logic.
[0030] The computation scheduling module incorporates an exception handling submodule, which receives computational exception signals (including data overflow, underflow, input NaN, illegal operands, etc.) reported by each group of computational units. After encoding the exceptions, it reports them to the external main core. Simultaneously, it can trigger pause and reset operations for the computational units based on the exception type. The computation scheduling module also incorporates a clock gating submodule, connected to the array control signal generation unit. Based on the enable signal, it gates the input clock of each group of computational units: for computational units with a valid enable signal, the clock path is opened, providing a 1GHz operating clock; for computational units with a valid disable signal, the clock path is cut off, causing the computational unit to enter a low-power sleep state without clock flipping, completely eliminating dynamic power consumption; when fixed-point computational units are operating, the global clock of all floating-point computational units is turned off; when floating-point computational units are operating, the global clock of all fixed-point computational units is turned off, achieving intelligent optimization of global power consumption.
[0031] The specific steps of the operation scheduling module's workflow are as follows: S1: The task feature extraction unit collects multi-source feature parameters of matrix operation tasks and system status in real time and generates standardized task feature vectors. S2: The lightweight neural network scheduling inference unit receives feature vectors, completes fast inference through pre-trained weights, and outputs 1-4 sets of confidence data in parallel working modes; S3: The mode decision and encoding unit completes the optimal mode selection, generates and encodes the operation mode configuration instruction; S4: The array control signal generation unit generates control signals and clock gating signals for the corresponding arithmetic units according to the control code, controls the start or stop of the corresponding number of arithmetic units, synchronously sends out arithmetic parameters, and starts matrix operations. S5: The exception handling submodule monitors the operation status and exception signals of the operation unit in real time, completes the exception reporting, and at the same time, the array control signal generation unit feeds back the operation status to the lightweight neural network scheduling and inference unit to realize closed-loop scheduling optimization.
[0032] The reusable matrix operation unit of this invention is divided into two categories: fixed-point operation unit and floating-point operation unit, with four groups in each category. Each group has a completely identical structure and fully reuses the hardware circuits, supporting independent operation of a single group and parallel operation of four groups.
[0033] The core of each fixed-point arithmetic unit is an 8×8 multiply-accumulate array, such as Figure 3 As shown, an outer product architecture is adopted. Each fixed-point arithmetic unit supports matrix operations of both INT8 and INT16 precision, and can complete matrix multiplication-addition (C=A×B+C), matrix multiplication (C=A×B), and matrix addition (C=A+B) operations of the M8N8K8 specification within 10 clock cycles. At INT8 precision, four columns of matrix A and four rows of matrix B can be input in a single cycle. Through four parallel multiplication channels, an 8×8 matrix multiplication operation can be completed in five cycles. At INT16 precision, a single-cycle input of one column of matrix A and one row of matrix B is used to complete an 8×8 matrix multiplication operation in 10 cycles via a multiply-accumulate pipeline.
[0034] The fixed-point arithmetic unit consists of 64 multiply-accumulate processing units (PEs) arranged in an 8×8 two-dimensional array. Each PE unit contains one 16-bit multiplier and four 32-bit adders, supporting one INT16 multiply-accumulate operation or four INT8 multiply-accumulate operations, realizing hardware multiplexing of INT8 / INT16 precision.
[0035] The input design of each PE cell is as follows: PE cells in the same row share the same row of input data in matrix A, and PE cells in the same column share the same column of input data in matrix B. This conforms to the outer product operation rules, eliminates the need for data start-up and emptying cycles of the pulsating array, and achieves 100% array utilization.
[0036] Each fixed-point arithmetic unit employs a 3-stage pipeline design, breaking down computation steps to shorten the critical path per cycle, and can operate at frequencies up to 1 GHz. Level 1: Data fetching and register cycle: Read matrix data A, B, and C from the general-purpose register module and latch it into the register submodule of the input fixed-point arithmetic unit; Second level: Multiplication cycle, the fixed-point arithmetic unit completes the multiplication operation of the corresponding elements and outputs the partial product; Level 3: Addition and accumulation cycle. The partial product is reduced and added by the adder in PE, and then accumulated with the C matrix data. The result is latched to the output register submodule and written back to the general-purpose register module.
[0037] The pipeline design enables overlapping execution of operations. In continuous matrix operation scenarios, a new set of operations can be started in each cycle, maximizing the throughput of operations.
[0038] The hardware circuits of the four fixed-point arithmetic units are completely identical. In single-operation mode, only one group is activated, while in parallel operation mode, all four groups work synchronously. The hardware circuits do not require any modification and are fully reused. At the same time, the INT8 / INT16 operations in each group share the same multiplication array and addition tree. The data bit width and operation path are switched only by the operation type configuration signal, achieving full hardware reuse between precision levels.
[0039] The core of each floating-point arithmetic unit is a fused multiply-accumulate (FMA) arithmetic array, designed using an inner product architecture, such as... Figure 4 As shown, it specifically includes an input alignment submodule, an FMA array submodule, a reduction addition submodule, a normalization submodule, an anomaly detection submodule, and a pipeline control submodule.
[0040] Each floating-point arithmetic unit supports matrix operations of both FP16 and FP32 precision, and can complete FP16 matrix operations (M8N8K4 specification) and FP32 matrix multiplication and addition (M4N4K4 specification) within 10 clock cycles. At FP16 precision, a single-cycle input of an 8×4 matrix A and a 4×8 matrix B is processed in parallel by 64 IPPE units, completing the 8×8 matrix multiplication operation in 10 cycles. Figure 6 As shown; At FP32 precision, a 4×4 matrix A and a 4×4 matrix B are input in a single cycle, and the 4×4 matrix multiplication operation is completed in 10 cycles through parallel processing using 16 IPPE units. Figure 5 As shown.
[0041] It consists of a 4×4 two-dimensional array of 16 inner product operation units (IPPE). Each IPPE unit contains one FMA multiply-accumulate fusion unit, three floating-point multipliers and three floating-point adders. It adopts a three-level tree computing topology and can complete one FP32 inner product operation or four FP16 inner product operations, realizing hardware reuse of FP16 / FP32 precision.
[0042] Each IPPE unit corresponds to one element of the result matrix. The multiplication and addition operations of the corresponding row and column are directly completed by the inner product method, without the need for data pulse transmission, the operation delay is fixed, and the control logic is simple.
[0043] Each floating-point unit employs a 4-stage pipeline design to meet the timing requirements of a 1GHz operating frequency.
[0044] Level 1: Number fetching and exponent alignment cycle, reads matrix data from the general-purpose register module, and completes the exponent alignment and mantissa preprocessing of floating-point numbers; Level 2: Multiplication and multiply-accumulate fusion cycle, the FMA multiply-accumulate fusion unit completes the multiplication and accumulation operations of the mantissa; Level 3: Reduced addition and leading zero prediction cycle, completes the reduced addition of intermediate results, and performs leading zero prediction in parallel to shorten the normalization delay; Level 4: Normalization and anomaly detection cycle. This stage normalizes the computation results, detects overflow, underflow, NaN, and other anomalies, and writes the latched results back to the general-purpose register module.
[0045] The hardware circuits of the four floating-point arithmetic units are completely identical, and the single-operation mode and parallel operation mode are fully reused. At the same time, the FP16 / FP32 operations in each unit share the same FMA array and reduction addition circuit. The hardware reuse between precisions is achieved by switching the operation bit width and pipeline parameters only by configuring signals.
[0046] The parallel operation implementation of the four sets of arithmetic units is as follows: when the parallel operation instruction of the four sets of arithmetic units is received, the four sets of fixed-point / floating-point arithmetic units synchronously receive the operation parameters and input data, start the operation synchronously, and independently complete a set of matrix operation tasks. The operation results are synchronously written back to the general-purpose register module.
[0047] Taking INT8 precision as an example, a single unit can complete one 8×8 matrix multiplication in 10 cycles, and four parallel units can complete four 8×8 matrix multiplications in 10 cycles, achieving a fourfold increase in computing power. At the same time, the four sets of operation units can work together to complete larger-scale matrix block operations, splitting a large 32×32 matrix into four 8×8 block matrices. Through parallel operation of the four sets of operation units, the overall operation cycle of large matrices is significantly shortened.
[0048] The general-purpose register module is the core of this invention for achieving parallel data reading and writing and operation, eliminating data waiting latency, and also undertakes the functions of on-chip data caching and reducing external memory access. The specific implementation structure includes register area A, register area B, register area C, ping-pong cache control submodule, and memory access interface submodule.
[0049] The A register area consists of eight 1024-bit registers, used only to store the left operand A matrix data for matrix operations. The bit width matches the input bit width of the fixed-point / floating-point arithmetic unit, and it supports single-cycle reading of 8×8 INT8 / 8×4 FP16 matrix data. The B register area consists of eight 1024-bit registers and is used only to store the right operand B matrix data for matrix operations. It has the same bit width and structure as the A register area. The C register area consists of eight 2048-bit registers, which can be used to store the accumulated operand C matrix data of matrix operations, and also to cache the data of the operation result matrix D. It supports caching of intermediate results of multi-step continuous multiplication and accumulation operations without interacting with external storage.
[0050] The general-purpose register module uses a double-buffered ping-pong buffer mechanism to divide the A, B, and C register areas into two independent buffer partitions, Buffer0 and Buffer1. The two partitions have the same capacity and structure, and the ping-pong buffer control submodule controls the switching between read and write states.
[0051] The specific workflow of the ping-pong caching mechanism is as follows: T1 cycle: The arithmetic unit reads data from Buffer0 to perform matrix operations, while the memory access interface submodule writes the matrix data required for the next operation from the external L2 SRAM to Buffer1; T2 cycle: The arithmetic unit completes the operation on the data corresponding to Buffer0, and the ping-pong caching control submodule immediately switches the read / write state. The arithmetic unit reads data from Buffer1, which has already been loaded with data, to start the next operation without any data waiting delay; at the same time, the memory access interface submodule writes the new operation data to Buffer0 and writes the operation result in Buffer0 back to the external L2 SRAM. Through the aforementioned ping-pong switching, the data reading and writing operations and matrix operation operations are executed in complete parallel, completely eliminating the idle period of the computing unit waiting for data loading in the traditional architecture, and significantly reducing the overall computing latency.
[0052] The C register area of the general-purpose register module supports intermediate result caching for multi-step continuous matrix multiplication and accumulation operations: In continuous multiplication and accumulation operations, the result of the first step can be directly cached in the C register area and directly used as the accumulation operand of the second step operation to be input to the operation unit. There is no need to write the intermediate result back to external storage and then read it again, which greatly reduces the number of data interactions with external storage and reduces memory access bandwidth usage and memory access latency.
[0053] The general-purpose register module connects to the external L2 SRAM through the memory access interface submodule. The interface has a bit width of 1024 bits, which supports reading and writing 1024 bits of data in a single cycle. This perfectly matches the operation bandwidth of the matrix operation unit, avoiding memory access bandwidth bottlenecks.
[0054] The following uses examples of different precision and different operation modes to illustrate the complete workflow of the architecture of this invention.
[0055] Single arithmetic unit operation mode (INT8 precision M8N8K8 matrix multiplication and addition) This embodiment demonstrates an INT8 precision 8×8 matrix multiplication and addition operation C=A×B+C. The operation scheduling module autonomously decides to adopt a single operation unit operation mode. The specific steps are as follows: Step 1: The external main core submits an INT8 precision 8×8 matrix multiplication and addition operation task to this architecture, and at the same time writes the data of matrices A, B, and C into the external L2 SRAM; Step 2: The operation scheduling module starts the adaptive scheduling process: (1) The task feature extraction unit collects feature parameters in real time: the input matrix dimension is 8×8 (small size), the data bit width is INT8, the operation task type is multiplication and addition, the remaining storage space of the general register module is 40%, there are no other operation units working at present, the system bus load is 75% (high), and generates a standardized task feature vector; (2) The lightweight neural network scheduling inference unit receives feature vectors, completes single-cycle inference through pre-trained weights, and outputs confidence data: 92% confidence for single-group operation mode, 5% confidence for two-group parallel operation mode, and <3% confidence for three-group / four-group parallel operation mode. (3) The maximum value of the mode decision and coding unit is selected to determine the single operation unit operation mode, and the corresponding control code is generated and encoded. (4) The array control signal generation unit enables only fixed-point arithmetic unit 0 according to the control code, and shuts down fixed-point arithmetic units 1~3 to put them into a low-power sleep state, while sending arithmetic parameters to fixed-point arithmetic unit 0. Step 3: The memory access interface submodule of the general-purpose register module writes the A, B, and C matrix data from the external L2 SRAM into Buffer0 of the ping-pong cache, completing the data preloading; Step 4: Fixed-point arithmetic unit 0 reads matrix data A, B, and C from Buffer0, starts pipelined operation, completes 8×8 matrix multiplication in 5 cycles at INT8 precision, and then completes the accumulation with matrix C in 1 cycle, for a total of 6 cycles to complete all operations, and writes the operation result back to the C register area of the general-purpose register module; Step 5: After the operation is completed, the general-purpose register module switches through the ping-pong buffer to write the operation result back to the external L2 SRAM, and loads the next set of operation data at the same time; the array control signal generation unit feeds back the operation status to the lightweight neural network scheduling and inference unit to complete the closed-loop scheduling optimization.
[0056] Parallel operation mode with four arithmetic units (INT8 precision, 4 sets of M8N8K8 matrix multiplication and addition) This embodiment involves four independent INT8 precision 8×8 matrix multiplication and addition operations. The operation scheduling module autonomously decides to adopt a four-operation-unit parallel operation mode. The specific steps are as follows: Step 1: The external main core submits four independent INT8 precision 8×8 matrix multiplication and addition tasks to this architecture in batches, and writes the corresponding A, B, and C matrix data of the four sets to the external L2 SRAM at the same time. Step 2: The operation scheduling module starts the adaptive scheduling process: (1) The task feature extraction unit collects feature parameters in real time: the input matrix is 4 independent 8×8 small matrices, the data bit width is INT8, the operation task type is batch multiplication and addition, the remaining storage space of the general register module is 90% (sufficient), there are no other matrix operation units working at present, the system bus load is 20% (low), and generates standardized task feature vectors; (2) The lightweight neural network scheduling inference unit receives the feature vectors, completes single-cycle inference through pre-trained weights, and outputs confidence data: the confidence of four parallel operation modes is 95%, the confidence of three parallel modes is 3%, and the confidence of single / two parallel modes is <2%; (3) The mode decision and encoding unit selects the maximum value, determines the use of the four operation units parallel operation mode, generates and encodes the corresponding control code (including synchronization control code); (4) The array control signal generation unit activates four fixed-point operation units (0~3) according to the control code, and sends operation parameters and synchronization signals to the four operation units synchronously to ensure that the four operation units enter the working state synchronously; Step 3: The general-purpose register module preloads the A, B, and C matrix data corresponding to the four sets of operations into the register partitions corresponding to the four sets of operation units, respectively, to complete the data preparation; Step 4: Fixed-point arithmetic units 0-3 start operation synchronously. Each unit completes one set of 8×8 matrix multiplication and addition operations within 6 cycles. The four sets of arithmetic units execute in parallel, completing 4 sets of matrix operations in a total of 6 cycles. The computing power is 4 times that of the single-set mode. Step 5: The operation results of the four sets of arithmetic units are synchronously written back to the general-purpose register module and written back to the external L2 SRAM in batches through the memory access interface; the array control signal generation unit feeds back the status of this batch operation to the lightweight neural network scheduling and inference unit to complete the closed-loop scheduling optimization.
[0057] Single arithmetic unit operation mode (FP16 precision M8N8K4 matrix multiplication) This embodiment demonstrates an FP16 precision 8×8 matrix multiplication operation. The operation scheduling module autonomously decides to adopt a single operation unit operation mode. The specific steps are as follows: Step 1: The external main core submits an FP16 precision 8×8 matrix multiplication task to this architecture, and simultaneously writes the data of matrix A (8×4 FP16) and matrix B (4×8 FP16) into the external L2 SRAM; Step 2: The operation scheduling module starts the adaptive scheduling process: (1) The task feature extraction unit collects feature parameters in real time: the input matrix dimension is 8×8 (FP16 small size), the data bit width is FP16, the operation task type is multiplication, the remaining storage space of the general register module is 50%, there are no other matrix operation units working at present, the system bus load is 65%, and generates a standardized task feature vector; (2) The lightweight neural network scheduling inference unit receives the feature vector, completes single-cycle inference through pre-trained weights, and outputs confidence data: the confidence of a single operation mode is 90%, the confidence of two parallel modes is 7%, and the confidence of three / four parallel modes is <3%; (3) The mode decision and encoding unit filters the maximum value, determines the single operation unit operation mode, generates and encodes the corresponding control code (including FP16 precision mode control code); (4) The array control signal generation unit enables floating point operation unit 0 according to the control code, shuts down the other three floating point operation units to put them into a low-power sleep state, and sends operation parameters and precision mode control signals to floating point operation unit 0 at the same time; Step 3: The general-purpose register module completes the data preloading of matrix A (8×4 FP16) and matrix B (4×8 FP16); Step 4: The floating-point unit 0 starts pipelined operation, completes the 8×8 matrix multiplication operation within 10 cycles, and writes the result back to the general-purpose register module; Step 5: The calculation result is written back to the external L2 SRAM through the general-purpose register module; the array control signal generation unit feeds back the current calculation status to the lightweight neural network scheduling and inference unit to complete the closed-loop scheduling optimization.
[0058] By using the clock gating submodule of the computing configuration module, the clock of idle computing units is cut off, causing them to enter a low-power sleep state. In single computing mode, only one set of computing units is in working state, while the other three sets of units do not have clock flipping, reducing dynamic power consumption by 75%. The arithmetic unit adopts a pipelined design to decompose the critical path. Low threshold voltage and large-size transistors are used for the circuits on the critical path, while standard threshold voltage and standard-size transistors are used for non-critical paths. While meeting the timing requirements of the 1GHz operating frequency, the proportion of low threshold transistors is controlled to within 20%. Compared with the full-domain low threshold design, leakage power consumption is reduced by more than 25%. When the fixed-point arithmetic unit is working, the global clock of the floating-point arithmetic unit is turned off; when the floating-point arithmetic unit is working, the global clock of the fixed-point arithmetic unit is turned off, further reducing global dynamic power consumption.
[0059] The above description is merely a description of preferred embodiments of the present invention and is not intended to limit the scope of the present invention in any way. Any changes or modifications made by those skilled in the art based on the above disclosure shall fall within the protection scope of the claims.
Claims
1. An adaptive scheduling reconfigurable matrix operation architecture, characterized by, It includes an operation scheduling module, a reusable matrix operation unit, and a general-purpose register module; The core control module of the computation scheduling module realizes adaptive decision-making and array control of the computation mode. It does not require the external main core to issue fixed computation mode configuration instructions. It can autonomously complete task feature acquisition, intelligent scheduling reasoning, mode judgment and array control, and also supports closed-loop scheduling optimization. The reusable matrix operation unit includes four sets of fixed-point operation units with identical structures and four sets of floating-point operation units with identical structures, used to perform matrix operations of corresponding precision; the hardware circuits of the four sets of fixed-point operation units are completely reused, and the hardware circuits of the four sets of floating-point operation units are completely reused, without the need to redesign the hardware for parallel operation mode or single operation mode. The general-purpose register module is used for storing and interacting with matrix operation data. It adopts a ping-pong caching mechanism to enable parallel execution of matrix data reading, writing and operation, and supports multi-step continuous matrix multiplication and accumulation operations.
2. The self-adapting dispatch reconfigurable matrix operation architecture of claim 1, wherein, The fixed-point arithmetic unit supports 8*8 matrix multiplication and addition, matrix addition and matrix multiplication operations of INT8 and INT16 within 10 cycles; each fixed-point arithmetic unit adopts an 8×8 external product multiplication and addition array, which consists of 64 multiplication and addition processing units forming an 8×8 two-dimensional array. Each multiplication and addition processing unit includes one 16-bit multiplier and four 32-bit adders, supporting one INT16 multiplication and addition operation or four INT8 multiplication and addition operations, realizing the multiplexing of hardware circuits for INT8 and INT16 precision.
3. The self-adapting dispatch reconfigurable matrix operation architecture of claim 2, wherein, The fixed-point arithmetic unit adopts a 3-stage pipelined arithmetic design, including data fetching and register cycles, multiplication cycles, and addition and accumulation cycles, with an operating frequency of 1GHz; the input design of the multiply-accumulate processing unit is as follows: multiply-accumulate processing units in the same row share the same row input data of the matrix, and multiply-accumulate processing units in the same column share the same column input data of the matrix, with an array utilization rate of 100%.
4. The self-adapting scheduling reconfigurable matrix operation architecture of claim 1, wherein, The floating-point arithmetic unit supports matrix multiplication and addition, matrix addition, and matrix multiplication operations of FP16 (M8N8K4 specification) and FP32 (M4N4K4 specification) within 10 cycles. Each floating-point arithmetic unit adopts an inner product arithmetic array, which consists of 16 inner product arithmetic units forming a 4×4 two-dimensional array. Each inner product arithmetic unit includes one multiply-add fused FMA unit, three floating-point multipliers, and three floating-point adders. It adopts a three-level tree computing topology and supports one FP32 inner product operation or four FP16 inner product operations, realizing the hardware circuit reuse of FP16 and FP32 precision.
5. The self-adapting dispatch reconfigurable matrix operation architecture of claim 4, wherein, The floating-point arithmetic unit adopts a 4-stage pipelined arithmetic design, including a data fetching and exponent alignment cycle, a multiplication and multiply-add fusion cycle, a reduction-addition and leading zero prediction cycle, and a normalization and anomaly detection cycle, with an operating frequency of 1GHz.
6. The self-adapting scheduling reconfigurable matrix operation architecture of claim 1, wherein, The computation scheduling module includes four sub-units: a task feature extraction unit, a lightweight neural network scheduling and inference unit, a pattern decision and encoding unit, and an array control signal generation unit. The task feature extraction unit is responsible for collecting relevant feature parameters of matrix operation tasks in real time, extracting and generating standardized task feature vectors and feature parameters, and providing input basis for subsequent intelligent scheduling. The lightweight neural network scheduling and inference unit adopts a hardware-friendly shallow multilayer perceptron network structure, with only 2-3 fully connected hidden layers, and is equipped with ReLU activation function and INT8 fixed-point quantization design. The lightweight neural network scheduling and inference unit receives the task feature vector and performs preprocessing, performs fast inference through pre-trained network weights, and outputs confidence data corresponding to the parallel working mode of 1-4 operation units. The mode decision and encoding unit performs optimal value filtering on the confidence data, determines the optimal number of parallel arrays for the current computing task, generates corresponding computing mode configuration instructions, and encodes the configuration instructions into hardware-recognizable control codes to achieve precise conversion of scheduling results into hardware control signals. The array control signal generation unit generates enable signals, start / stop signals, synchronization signals, and data path selection signals for the corresponding arithmetic units based on the control codes output by the mode decision and encoding unit. This precisely controls the start or stop of the corresponding number of arithmetic units, ensuring timing synchronization and conflict-free execution during parallel multi-array operations. Simultaneously, it feeds back the actual operation status to the lightweight neural network scheduling and inference unit, forming a closed-loop scheduling logic.
7. The self-adapting scheduling reconfigurable matrix operation architecture of claim 1, wherein, The general-purpose register module is divided into register area A, register area B, and register area C. Register area A consists of eight 1024-bit registers and is used to store the left operand matrix data of matrix operations. Register area B consists of eight 1024-bit registers and is used to store the right operand matrix data of matrix operations. Register area C consists of eight 2048-bit registers and is used to store the accumulated operand matrix data and the operation result matrix data of matrix operations, supporting intermediate result caching for multi-step continuous matrix multiplication and accumulation operations.
8. The self-adapting dispatch reconfigurable matrix operation architecture of claim 7, wherein, The ping-pong caching mechanism of the general-purpose register module divides register area A, register area B, and register area C into two cache partitions, Buffer0 and Buffer1, which have completely identical capacity and structure. When the arithmetic unit reads data from Buffer0 to perform an operation, it simultaneously writes the data required for the next operation into Buffer1. After the operation is completed, it immediately switches the read / write state, and the arithmetic unit reads data from Buffer1 to start the operation, while writing the operation result in Buffer0 back to external storage, thus achieving complete parallelism between data reading / writing and arithmetic operations.
9. The self-adapting scheduling reconfigurable matrix operation architecture of claim 1, wherein, The four fixed-point arithmetic units correspond one-to-one with the four floating-point arithmetic units, forming four independent arithmetic paths. Each arithmetic path can work independently or synchronously in parallel. When the four arithmetic paths work in parallel, they can complete four times the matrix operation tasks of a single arithmetic operation within the same operation cycle, thus achieving a fourfold increase in computing power.
10. The self-adapting scheduling reconfigurable matrix operation architecture of claim 1, wherein, The computation scheduling module also includes an exception handling submodule and a clock gating submodule; the exception handling submodule receives the computation exception signals reported by each group of computation units, encodes the exceptions and reports them to the external main core, and triggers the pause and reset operations of the computation units according to the exception type. The clock gating submodule controls the input clock of each set of arithmetic units according to the enable signal: For arithmetic units with an active enable signal, the clock path is opened, providing a 1GHz operating clock. For arithmetic units with active shutdown signals, the clock path is cut off, causing the arithmetic unit to enter a low-power sleep state without clock flipping, thus completely eliminating dynamic power consumption; when fixed-point arithmetic units are working, the global clock of all floating-point arithmetic units is turned off; when floating-point arithmetic units are working, the global clock of all fixed-point arithmetic units is turned off, achieving intelligent optimization of global power consumption.