Realization method of in-memory calculation standard extension instruction set based on RISC-V instruction set

Through the hierarchically decoupled programming model and hardware to implement parameterized definition, the problems of computing completeness and software ecology in in-memory computing hardware design are solved, efficient development and expansion of hardware are realized, and the universality and portability of hardware are improved.

CN120338058APending Publication Date: 2025-07-18HANG ZHOU NANO CORE CHIP ELECTRONIC TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510421110.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing in-memory computing hardware design lacks computing completeness and software ecosystem, resulting in insufficient hardware flexibility, high development costs, long development cycles, and lack of portability and scalability.

Method used

Using a hierarchical decoupling programming model, the in-memory computing hardware is abstracted into a software-visible architectural register, and it is converted into hardware-executable commands through specific instructions and opcodes, and hardware implementation parameters are defined to achieve the decoupling of hardware and software.

Benefits of technology

It reduces the cost and difficulty of software stack development, improves hardware development efficiency, enhances the scalability and portability of hardware, supports unified compilation tool chains and abstract operator libraries for different hardware, and improves hardware utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120338058A_ABST
    Figure CN120338058A_ABST
Patent Text Reader

Abstract

The invention relates to a method for realizing an in-memory computing standard extension instruction set based on an RISC-V instruction set, which comprises the following steps: a programming model realization step: abstracting in-memory computing hardware into a software-visible architecture register by adopting a hierarchical decoupling mode, the architecture register comprises a storage register for mapping internal storage into a two-dimensional data layout and a calculation register for mapping an in-storage calculation cluster into a three-dimensional data layout; and an instruction set implementation step: converting hardware abstraction and interface specifications in the programming model into commands which can be executed by hardware through specific instructions and operation codes. According to the hierarchical decoupling programming model, upper-layer compiling and lower-layer micro-architecture implementation can be independently developed, the software stack development cost and difficulty are reduced, and the hardware development and iteration process is accelerated. Due to the unified programming model, different in-memory computing hardware designs can adopt the same abstract operator library and compiling tool chain, and the portability and reusability of codes are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of chip design, and more particularly to a method for implementing a standard extended instruction set for in-memory computing based on the RISC-V instruction set. Background Art

[0002] Based on the research at the circuit level and macrocell level, integrating CIM into the system to accelerate neural network algorithms or specific applications has become a research hotspot. The research on CIM processor architectures mainly falls into four categories: single-core / pipelined CIM, homogeneous CIM, heterogeneous CIM, and CIM die expansion.

[0003] Single-core or pipelined CIM mainly focuses on the load mapping and integration of multiple CIM arrays, simply connecting through arithmetic circuits and buffers, assuming that most computational loads are processed in CIM, and a small amount of buffers are used for data transfer. It supports pipelined execution of multiple tasks, including training and inference. Homogeneous CIM integrates components such as input buffers, preprocessing units, CIM arrays, postprocessing units, and control circuits, aiming to support more complex neural network operations, and usually equipped with larger buffers to handle multiple iterations and intermediate results. Heterogeneous CIM architectures integrate multiple functional cores outside the CIM core to achieve end-to-end deployment of algorithms. The core modules include CPU cores, CIM cores, and digital co-processor cores. These cores may have independent local storage and share global on-chip storage. With the rise of 2.5D / 3D integration technology, chip-level CIM architectures have been explored and can be flexibly extended according to requirements. Large neural networks are divided into multiple CIM blocks, each storing weights and intermediate data. Advanced packaging technology enables lower data transfer power consumption between small chips.

[0004] Currently, there are many custom accelerators based on in-memory computing hardware implementations. Generally speaking, in-memory computing can significantly improve computing power density and energy efficiency. However, in further industrial application promotions, it faces two major problems: computational completeness and software ecosystem. In-memory computing technology is suitable for tensor computing, but its high computing power density sacrifices a certain degree of hardware flexibility, so it lacks computational completeness and has serious hardware utilization problems during scalar and vector calculations. At the same time, due to the lack of standard instruction sets and standard toolchains, even with chips with advanced performance, to achieve task deployment and efficient operation, support from the upper software stack and compilation toolchain is required to control the chip to execute corresponding functions. Since current accelerator designs usually adopt a bottom-up design approach and design instructions according to the designed hardware characteristics, the designed instructions are fragmented, resulting in the problem of reinventing the wheel for computational graph optimization, operator library design, and compilation optimization, greatly increasing the development cycle duration, leading to an increase in software development costs, reducing development efficiency, and lacking portability and scalability in hardware design.

[0005] Therefore, it is necessary to improve the existing in-memory computing instruction design. Summary of the Invention

[0006] In view of the above problems, the present invention provides a method for implementing a standard extended instruction set for in-memory computing based on the RISC-V instruction set, including the steps of:

[0007] Programming model implementation step: adopting a hierarchical decoupling method, abstracting the in-memory computing hardware into software-visible architecture registers, where the architecture registers include storage registers that map internal storage to a two-dimensional data layout and computing registers that map the in-memory computing cluster to a three-dimensional data layout; and

[0008] Instruction set implementation step: converting the hardware abstraction and interface specifications in the programming model into commands that can be executed by the hardware through specific instructions and operation codes.

[0009] Optionally, it further includes: Hardware implementation parameter definition step: defining the hardware implementation in a parameterized form to facilitate the application of the programming model to different types of hardware.

[0010] Optionally, the "defining the hardware implementation in a parameterized form" includes:

[0011] CIMC_NUM: The number of in-memory computing clusters inside the neural network processor core;

[0012] CIMC_CIME_NUM: The number of in-memory computing engines included in each in-memory computing cluster;

[0013] CIME_ROW_NUM: The number of rows of each of the in-memory computing engines;

[0014] CIME_COL_NUM: The number of columns of each of the in-memory computing engines;

[0015] CIMM_PG_NUM: In-memory computing macro storage-computation ratio;

[0016] SPAD_NUM: The number of scratchpad memories inside the neural network processor core;

[0017] SPAD_WD: The bit width of the scratchpad memory; and

[0018] SPAD_DP: The depth of the scratchpad memory.

[0019] Optionally, the storage hierarchy defined by the programming model includes:

[0020] External storage: It interacts with internal storage in a direct memory access manner through load / store instructions and is not directly mapped to architectural registers; and

[0021] Internal storage includes:

[0022] Scratchpad (SPAD) scratchpad memory: Storage registers mapped to a two-dimensional data layout, including scratchpad memory for storing scalar and tensor data;

[0023] In-memory computing cluster (CIM Cluster, CIMC): Computing registers mapped to a three-dimensional data layout in computing mode, used to store tensor data and support in-memory computing.

[0024] Optionally, in storage mode, the width of the in-memory computing cluster is: CIME_ROW_NUM; the depth is:

[0025] CIMC_CIME_NUM * CIME_COL_NUM * CIMM_PG_NUM.

[0026] Optionally, in computing mode, the storage space of the in-memory computing cluster is jointly determined by the in-memory computing macro storage-computing ratio, its effective number of rows, and effective number of columns.

[0027] Optionally, the data types of the neural network processor core include scalar, vector data, and tensor data.

[0028] Optionally, it further includes: register configuration steps: data type and precision configuration, data arrangement information configuration, and operation configuration.

[0029] Optionally, it further includes: instruction division steps: dividing instructions into control instructions, memory access instructions, data transfer and operation instructions, matrix calculation instructions, and element-level calculation instructions. Description of the Drawings

[0030] Figure 1 is an implementation method of the in-memory computing standard extension instruction set based on the RISC-V instruction set provided by the embodiments of the present invention.

[0031] Figure 2 is the architecture of the programming model provided by the embodiments of the present invention.

[0032] Figure 3 is a schematic diagram of the storage structure of the programming model provided by the embodiments of the present invention in computing mode.

[0033] Figure 4It is a schematic diagram of data interaction between the RISC-V processor core and the RoCC Accelerator provided in the embodiments of the present invention. Detailed implementation manners

[0034] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0035] As Figure 1 shown, this embodiment provides a method for implementing a standard extended instruction set for in-memory computing based on the RISC-V instruction set, including the steps of:

[0036] Programming model implementation step: Adopting a hierarchical decoupling method, abstracting the in-memory computing hardware into software-visible architecture registers, where the architecture registers include storage registers that map internal storage to a two-dimensional data layout and computing registers that map the in-memory computing cluster to a three-dimensional data layout; and

[0037] Instruction set implementation step: Through specific instructions and operation codes, transforming the hardware abstraction and interface specifications in the programming model into commands that the hardware can execute.

[0038] To complete the decoupling between the instruction set design, upper-layer compilation, and lower-layer microarchitecture implementation, at the beginning of the design, this embodiment constructs a programming model that abstracts the hardware into architecture registers. The storage parts of SPAD and CIM are respectively mapped to storage and computing registers. By abstracting the storage registers into a 2D data arrangement and the computing registers into a 3D data arrangement, a unified programming model is generated. Based on this programming model, all hardware designs adopting the in-memory computing paradigm can use the same set of abstract operator libraries and corresponding compilation toolchains, reducing the software stack development cost and difficulty and accelerating the hardware development and iteration process.

[0039] Optionally, the method for implementing the standard extended instruction set for in-memory computing based on the RISC-V instruction set further includes: Hardware implementation parameter definition step: Defining the hardware implementation in a parameterized form to facilitate the programming model for different types of hardware.

[0040] Such settings can improve the scalability and generality of the instruction set architecture for hardware implementation. In this embodiment, the hardware scale information is described in a parameterized form, so as to ensure that the hardware accelerators adopting the in-memory computing mechanism can all be adapted to the extended instruction set architecture. We define the following fixed parameters defined by the hardware implementation to provide the host with the hardware resource scale of the co-processor implemented based on CIM, so as to realize the control of the co-processor and describe the hardware resource situation of the computing components implemented based on CIM and the storage components implemented based on SRAM.

[0041] Specifically, "defining the hardware implementation in a parameterized form" includes:

[0042] CIMC_NUM: The number of in-memory computing clusters inside the neural network processor core;

[0043] CIMC_CIME_NUM: The number of in-memory computing engines included in each in-memory computing cluster;

[0044] CIME_ROW_NUM: The number of rows of each in-memory computing engine;

[0045] CIME_COL_NUM: The number of columns of each in-memory computing engine; CIMM_PG_NUM: The storage-computation ratio of the in-memory computing macro storage;

[0046] SPAD_NUM: The number of scratch memories inside the neural network processor core;

[0047] SPAD_WD: The bit width of the scratch memory; and

[0048] SPAD_DP: The depth of the scratch memory.

[0049] Through the above hardware parameter definitions, co-processors implemented based on CIM with different scales and different implementation methods can form a unified hardware information description method, so that the CIM standard extended instruction set can be widely applied to the control of various different hardware accelerators implemented based on the in-memory computing method. The decoupling between the accelerator microarchitecture design and the instruction set design is completed, avoiding the problems of insufficient scalability and portability in the bottom-up design project.

[0050] Compared with the hardware definition parameters required in the traditional computing architecture, the hardware parameters defined in this embodiment are more adapted to the hardware characteristics of CIM.

[0051] Specifically, inside the in-memory computing engine, since multiple storage arrays usually share the same computing resources and adjust the number of storage arrays sharing the computing resources according to the requirements of the algorithm for storage and computing resources. The ratio of storage resources to computing resources is defined as the storage-computation ratio, which is defined by CIMM_PG_NUM in the hardware parameters.

[0052] The in-memory computing cluster supports both storage and computing functions simultaneously. Therefore, its addressing mode is determined by its corresponding configuration scheme, and its addressing mode is configured through the corresponding configuration register. For the in-memory computing engine containing CIMC_CIME_NUM in-memory computing engines, its corresponding addressing configuration is represented by a one-hot code of log2(CIMC_CIME_NUM)+1 bits. For example, when CIM_ADDR_CFG = 0, the CIM enters the storage mode and is regarded as a memory. When CIM_ADDR_CFG = 1, the CIM is in the computing mode and is regarded as being spliced in the way of output channel fusion, that is, each CIM engine has independent input computing data. After completing the matrix-vector multiplication calculation responsible for by the CIM engine, the calculation results output by each engine are added up, and the final output result after fusion is output. Therefore, the equivalent input channel number row_eff = CIMC_CIME_NUM * CIME_ROW_NUM, and the equivalent output channel is col_eff =

[0053] CIME_COL_NUM. When CIM_ADDR_CFG = 2, every two CIM engines share the input computing data, and the calculation results of different input data engines are output-fused and the final result is output. Therefore, the equivalent input channel number row_eff =

[0054] CIMC_CIME_NUM * CIME_ROW_NUM / 2, and the equivalent output channel is col_eff = CIME_COL_NUM * 2. The configuration of other CIM_ADDR_CFGs can be deduced by analogy.

[0055] Reference Figure 2 , in the storage mode, the in-memory computing cluster is regarded as a two-dimensional storage structure, and its addressing method is similar to that of the scratchpad memory. At this time, the memory width is determined by the CIME_ROW_NUM hardware-defined parameter.

[0056] The depth is CIMC_CIME_NUM * CIME_COL_NUM * CIMM_PG_NUM.

[0057] In the computing mode, the in-memory computing cluster is regarded as a three-dimensional storage structure, and the data arrangement method is closely related to the corresponding computing expansion method. It consists of CIM_PG_NUM pages, and the effective number of rows (#row_eff) and the effective number of columns (#col_eff) of each page are determined according to the addressing configuration.

[0058] According to the corresponding addressing configuration, the addressing space of the in-memory computing cluster can be represented as a cube of #row_eff × #col_eff × CIMM_PG_NUM, where each row is divided into several consecutive address segments in units of 256 bits.

[0059] After that, referring to Figure 3 , these consecutive address segments will be organized into a continuous address space in the order of column (①) direction first, then row (②) direction, and finally page (③) direction.

[0060] For RISC-V, the internal memory and the control and status register file are organized in a unified internal storage space, and can be read / written through the RD_LMEM / WR_LMEM Primitive. The data width is the same as the RISC-V XLEN (32 for RV32 and 64 for RV64). The corresponding storage space addresses are organized as follows:

[0061] 1) The address width of the internal storage space is the same as the RISC-V XLEN;

[0062] 2) The internal storage space address is divided into four fields:

[0063] a. Zero bits: Since the internal memory takes 256 bits (one implementation) as the minimum unit, the 5 bits at the LSB of the storage space address are zero values;

[0064] b. LMEM offset: The width is the larger value of SPAD_ADDR_WD and CIMC_ADDR_WD, and is used for addressing within each internal memory hardware;

[0065] c. LMEM index: The width is log2(SPAD_NUM + CIMC_NUM), and is used for addressing between different internal memory hardwares;

[0066] d. The remaining fields are reserved temporarily;

[0067] The scratchpad memories SPADs are numbered as SPAD(0), SPAD(1),..., SPAD(NUM_SPAD - 1) in sequence; the CIMClusters are numbered as CIMC(0), CIMC(1),..., CIMC(CIMC_NUM - 1) in sequence;

[0068] There is an address mapping relationship between the internal memories of RISC-V and the NPU. For example, in a hardware design implementation, 4 Scratchpads and 4 CIMClusters are adopted:

[0069] The CPU passes the base addresses of 4 SPs:

[0070] Scratchpad0: 0x9000_0000; Scratchpad1: 0x9001_0000;

[0071] Scratchpad2: 0x9002_0000; Scratchpad3: 0x9003_0000;

[0072] CIMC base address:

[0073] Page0: 0x9000_4000; Page1: Page0 + 2^13;

[0074] Page2: Page1 + 2^13; Page3: Page2 + 2^13

[0075] Through the above configuration scheme, efficient mapping for different computing requirements can be achieved, and the hardware utilization efficiency can be improved according to task requirements in the end-to-end AI application scenario.

[0076] Optionally, continue to refer to Figure 2 , the storage hierarchy defined by the programming model includes:

[0077] External storage: Data interaction with internal storage is carried out in a direct memory access manner through load / store instructions and is not directly mapped to architecture registers; and

[0078] Internal storage includes:

[0079] Scratchpad memory: Mapped to storage registers with a two-dimensional data layout, including scratchpad memory for storing scalar and two-dimensional tensor data;

[0080] In-memory computing cluster: Mapped to computing registers with a three-dimensional data layout in the computing mode, used to store three-dimensional tensor data and support in-memory computing.

[0081] Optionally, in the storage mode, the width of the in-memory computing cluster is: CIME_ROW_NUM; the depth is:

[0082] CIMC_CIME_NUM * CIME_COL_NUM * CIMM_PG_NUM.

[0083] Optionally, in the computing mode, the storage space of the in-memory computing cluster is jointly determined by the in-memory computing macro storage-computation ratio, its effective number of rows, and effective number of columns.

[0084] Optionally, the data types of the neural network processor core include scalar data and three-dimensional tensor data.

[0085] Scalar data:

[0086] There is no explicit dedicated memory for storing scalar data inside the neural network processor core; the input scalar value needs to be passed by RISC-V in the form of commands; the output scalar value will be received by RISC-V in the form of responses.

[0087] Three-dimensional tensor data:

[0088] Stored in the global / temporary memory. The following assumes that the corresponding tensor dimensions are dim2×dim1×dim0. For computer vision, the three usually correspond to the height, width, and number of channels of the image, expressed as height×width×channels. When dim2 = 1, it can be called a matrix, and when dim2 = dim1 = 1, it can be called a vector.

[0089] Data type and precision:

[0090] From the perspective of data type and precision, the neural network processor core supports integer, floating-point, BF, and BBF data types, and 4, 8, 16, 32-bit data precisions.

[0091] Data arrangement:

[0092] For tensor data of any precision, it is stored in the global memory in a unified storage format. Assuming the tensor data dimensions are dim2×dim1×dim0, the following unified constraints are imposed on its storage format to ensure unified functional support.

[0093] Logical storage constraints:

[0094] For any precision, round to the scratchpad memory bit width in the dim0 dimension (pad with zeros for the insufficient part);

[0095] In dim0, split it into dim0b×dim0a with the scratchpad memory bit width as the minimum unit, where the number of bits in dim0a is fixed to the scratchpad memory bit width;

[0096] The tensor data dimensions are re-expressed as dim2×dim1×dim0b×dim0a.

[0097] Physical storage constraints:

[0098] The tensor is arranged in the order of dim0a-->dim0b-->dim1-->dim2;

[0099] The arrangement of the tensor in dim0a must be continuous;

[0100] The arrangement of the tensor in dim0b has a fixed stride, and it is always equal to 256 bits;

[0101] The arrangement of the tensor in dim1 has a fixed stride and must be an integer multiple of 256 bits;

[0102] The arrangement of the tensor in dim2 has a fixed stride and must be an integer multiple of 256 bits;

[0103] Representation method:

[0104] Based on the above constraints, for a tensor with a shape of size_dim2×size_dim1×size_dim0, where size_dim2, size_dim1, and size_dim0 are arbitrary integers not greater than 4096, its physical storage format in the global memory can be represented by three sets of data:

[0105] REM = rem_dim0, representing the remainder of size_dim0 modulo size_dim0a.

[0106] SIZE = (size_dim2, size_dim1, size_dim0b, size_dim0a), representing how many elements there are in each dimension;

[0107] STRIDE = (stride_dim2, stride_dim1, stride_dim0b), representing the stride of the arrangement in each dimension, with the unit of 256 bits;

[0108] And there are the following constraints:

[0109] For a specified precision, size_dim0a is a fixed value = scratch memory bit width / bit width, that is, it is taken as follows: for INT4, size_dim0a = 64; for INT8, size_dim0a = 32; for INT16 / FP16 / BF16, size_dim0a = 16; for INT32 / FP32, size_dim0a = 8; subsequent potential precision expansions follow the same pattern;

[0110] size_dim0b = ceil(size_dim0 / size_dim0a); rem_dim0 =

[0111] mod(size_dim0, size_dim0a);

[0112] stride_dim0b is always equal to 1;

[0113] stride_dim1 is not less than size_dim0b;

[0114] The stride_dim2 is not less than stride_dim1 * size_dim1 and must be an integer multiple of stride_dim1.

[0115] With this representation method, flexible slicing operations can be performed on tensor data, as shown in the following figure:

[0116] On this basis, the addressing of data in a certain tensor can be represented by three indices:

[0117] IDX = (idx_dim2, idx_dim1, idx_dim0b)

[0118] The address can be expressed as:

[0119] ADDR = idx_dim2 * stride_dim2 + idx_dim1 * stride_dim1 +

[0120] idx_dim0b

[0121] The base address (byte address) of tensor data storage in the scratchpad memory must be an integer multiple of 32; the rest is the same as in the global memory. In-memory computing cluster: The base address (byte address) of tensor data storage in storage mode addressing must be an integer multiple of 32; the rest is the same as in the global memory. In-memory computing cluster: In computing mode addressing, the in-memory computing cluster can only be used to store two-dimensional matrices (dim2 = 1); the row direction can only be stored along dim1; the column direction can only be stored along dim0; the page direction can be stored along dim0 or dim1; the base address (byte address) of storage must be an integer multiple of # valid rows * # valid columns / 256, that is, it can only start storing from the starting position of a certain page and occupy consecutive pages.

[0122] Optionally, an implementation method of a standard extension instruction set for in-memory computing based on the RISC-V instruction set further includes:

[0123] Register configuration steps: data type and precision configuration, data layout information configuration, and operation configuration.

[0124] Data type and precision configuration:

[0125] To support parallel operations of different hardware units in a neural network processor (NPU), corresponding data types and precisions need to be configured for various operations:

[0126] Memory access instruction:

[0127] cfg_type_ldst: Memory access data type configuration register

[0128] cfg_wd_ldst: Memory access data width configuration register

[0129] Data transfer operation:

[0130] cfg_type_dm: Data transfer data type configuration register

[0131] cfg_wd_dm: Data transfer data width configuration register

[0132] Matrix calculation instruction:

[0133] cfg_type_out / orig / in / wt_matrix: Matrix calculation output, original, input, weight data type configuration register

[0134] cfg_wd_out / orig / in / wt_matrix: Matrix calculation output, original, input, weight data width configuration register

[0135] Element-level calculation instruction:

[0136] cfg_type_out / in1 / in2_vector: Element-level calculation output, input1, input2 vector data type configuration register

[0137] cfg_wd_out / in1 / in2_vector: Element-level calculation output, input1, input2 vector data width configuration register

[0138] Data layout information configuration:

[0139] To ensure the correctness and efficiency of data access, it is necessary to configure the data layout information of various operations:

[0140] Memory access instruction:

[0141] cfg_size_dim0 / dim1 / dim2_ldst: Memory access data dimension size configuration register

[0142] cfg_stride_dim1 / dim2_gmem_ldst: Memory access global memory stride configuration register cfg_stride_dim1 / dim2_lmem_ldst: Memory access scratchpad memory stride configuration register

[0143] Data transfer instruction:

[0144] cfg_size_dim0 / dim1 / dim2_dm: Data transfer data dimension size configuration register

[0145] cfg_stride_dim1 / dim2_out_dm: Data transfer output stride configuration register

[0146] cfg_stride_dim1 / dim2_in_dm: Data transfer input stride configuration register

[0147] Matrix calculation instruction:

[0148] cfg_size_dim0b / dim1 / dim2_out / in_matrix: Matrix calculation output and input data dimension size configuration register

[0149] cfg_stride_dim1 / dim2_out / in_matrix: Matrix calculation output and input stride configuration register

[0150] Element-level calculation instruction:

[0151] cfg_size_dim0_ref_vector: Element-level calculation reference vector data dimension size configuration register

[0152] Operation configuration information:

[0153] To achieve flexible control of various computing operations, corresponding operation parameters need to be configured:

[0154] cfg_op_matrix: Matrix calculation accumulation, activation function, shift configuration register;

[0155] cfg_op_vv_v: Vector-vector single instruction multiple data streams (SIMD) operation configuration register;

[0156] cfg_op_vs_v: Vector-scalar single instruction multiple data streams (SIMD) operation configuration register;

[0157] cfg_op_v_s: Vector reduction operation configuration register;

[0158] Kernel: Convolution kernel configuration register;

[0159] Padding: Padding configuration register.

[0160] Optionally, the implementation method of the in-memory computing standard extension instruction set based on the RISC-V instruction set further includes:

[0161] Instruction division step: Divide instructions into control instructions, memory access instructions, data transfer and operation instructions, matrix calculation instructions, and element-level calculation instructions.

[0162] Instruction Type Classification: Instructions are generally divided into basic instructions and extended instructions. Among them, basic instructions are the minimum instruction set to implement support for the CIM extension coprocessor, while extended instructions provide a broad space for hardware design to provide efficient instruction support according to the design method of the microarchitecture. For the basic instruction set, our design covers control instructions, memory access instructions, data transfer and operation instructions, matrix calculation instructions, and element-level calculation instructions. A Primitive is composed of a fixed combination of multiple Instructions, and this level is docked with the AI compiler.

[0163] Control Instructions: Used to complete system control and register configuration functions.

[0164] SYNC Synchronization Instruction: Similar to the fence instruction in RISC-V, it is used for synchronization and ensuring data consistency. After all the instructions before this instruction are executed, the subsequent instructions will be executed.

[0165] SW_CIMC Instruction: Switches the addressing mode of the CIMCluster.

[0166] CONFIG Instruction: Writes the configuration value into the corresponding configuration register.

[0167] Memory Access Instructions: Used to complete data interaction between internal storage and external storage.

[0168] TLD Instruction: Transfers tensor data from external storage to internal storage. The TLD instruction supports flexible data segmentation: by separately specifying the corresponding information of the tensor in external storage and internal storage, as well as the size of the data after segmentation, the segmented tensor data can be extracted from the tensor address stored in external storage and stored in the corresponding position of the tensor in internal storage.

[0169] Indexed Tensor Load Instruction: Selects the data at the corresponding position according to the provided address index information for transfer, and needs to additionally provide the base address information of the storage space for storing the index information. The size of the storage for the index information is the same as the size of the transferred tensor, and the data type is fixed as integer. Since the required index widths for tensors of different sizes are different, a 6-bit configurable register cfg_index_wd_ldst is used for regulation, and the corresponding index bit width is (2^cfg_index_wd_ldst). At the same time, the stored index information is arranged in a tightly packed manner.

[0170] Transposed Tensor Load Instruction: Directly transfers the transposed data. The required base address, destination address, and registers to be configured are the same as those of the traditional fixed-stride Load / Store instructions, and the corresponding distinction is made in the operation code.

[0171] The Alignment Tensor Load instruction. Since CIM can only efficiently process integer or floating-point mantissa multiply-add calculations, it is necessary to align the mantissa of the floating-point number in advance and transfer the aligned floating-point data. The base address and destination address to be provided, as well as the registers to be configured, are the same as those of the traditional fixed-stride Load / Store instruction, and the corresponding distinction is made in the opcode.

[0172] The TST instruction: It is an instruction for moving tensor data from internal storage to external storage and is a mirror operation of the TLD instruction. It also supports flexible data slicing operations.

[0173] The Indexed Tensor Store instruction requires an additional base address information of the storage space for storing index information. The size of the stored index information is the same as the size of the transmitted tensor. At the same time, the stored index information is arranged in a tightly packed manner.

[0174] The Transposed Tensor Store instruction directly transfers the transposed data. The base address and destination address to be provided, as well as the registers to be configured, are the same as those of the traditional fixed-stride Load / Store instruction, and the corresponding distinction is made in the opcode.

[0175] The Alignment Tensor Store instruction. Since CIM can only efficiently process integer or floating-point mantissa multiply-add calculations, it is necessary to unalign the floating-point data and then re-store the floating-point data to external storage. The base address and destination address to be provided, as well as the registers to be configured, are the same as those of the traditional fixed-stride Load / Store instruction, and the corresponding distinction is made in the opcode.

[0176] Data transfer and operation instructions: Used to complete data interaction between internal storages.

[0177] The BC instruction: Broadcasts scalar data as a tensor and stores it in internal storage

[0178] Step 1: Address the three-dimensional Tensor O stored in internal storage according to the base address of the output tensor; jump to step 2;

[0179] Step 2: Broadcast Scalar I to the corresponding position of Tensor O according to the corresponding configuration signal; end the execution.

[0180] The MOV instruction: Moves a tensor within internal storage

[0181] Step 1: Address Tensor I stored in the internal memory according to the base address of the input tensor; Address Tensor O stored in the internal memory according to the base address of the output tensor;

[0182] Step 2: Read out Tensor I and then write it to the corresponding position of Tensor O; End the execution.

[0183] TRANS instruction: Transpose the tensor

[0184] Step 1: Address Tensor I stored in the internal memory according to the base address of the input tensor, read it out; Jump to Step 2;

[0185] Step 2: Transpose Tensor I (swap dim0 and dim1), and store it in the internal memory according to the base address of the output tensor; End the execution.

[0186] ALIGN instruction: Perform floating-point pre-alignment operation on the tensor

[0187] Step 1: Address Tensor I stored in the internal memory according to the base address of the input tensor, read it out;

[0188] Step 2: Align the mantissa of Tensor I according to the exponent situation

[0189] Step 3: Store it in the internal memory according to the base address of the output tensor; End the execution.

[0190] RE-ALIGN instruction: Re-convert the pre-aligned floating-point mantissa into floating-point form

[0191] Step 1: Address Tensor I stored in the internal memory according to the base address of the input tensor, read it out;

[0192] Step 2: Perform anti-alignment operation on Tensor I according to the exponent situation

[0193] Step 3: Store it in the internal memory according to the base address of the output tensor; End the execution.

[0194] Matrix multiplication instruction: Used to complete high-dimensional tensor calculation

[0195] MV_V instruction: Complete the general matrix-vector multiplication operation

[0196] Step 1: Select the input data according to the base address of the input tensor

[0197] Step 2: Select the weight data according to the base address of the weight tensor

[0198] Step3: Perform matrix-vector multiplication according to the operation method provided by the cfg_op_matrix field

[0199] Step4: Store the calculation result in the memory corresponding to the output tensor base address to complete the operation.

[0200] MM_M instruction: Complete the general matrix-matrix multiplication operation

[0201] Step1: Select input data according to the input tensor base address

[0202] Step2: Select weight data according to the weight tensor base address

[0203] Step3: Perform matrix-vector multiplication according to the operation method provided by the cfg_op_matrix field

[0204] Step4: Store the calculation result in the memory corresponding to the output tensor base address to complete the operation.

[0205] Element-level operation instruction: Used to complete SIMD element-level operations

[0206] Including SIMD calculations (VV_V Primitive, VS_V Primitive, V_V Primitive) and Reduce (V_SPrimitive) types; in terms of function only, it can be regarded as a partial cut, addition and restriction based on the RISC-V "V" Extension, which better meets the current scenario requirements.

[0207] VV_V instruction: Used to complete SIMD operations between vectors

[0208] step1: Read out two one-dimensional Vectors I1 and I2 with the same length from the internal storage according to the input vector base address 1 and the input vector base address 2 respectively;

[0209] step2: After performing SIMD calculations on each element of the one-dimensional Vectors I1 and I2, output a one-dimensional Vector O with the same length;

[0210] step3: Write Vector O to the internal storage according to the output vector base address, and then end the execution.

[0211] VS_V instruction: Used to complete SIMD operations between a vector and a scalar

[0212] step1: Read out the one-dimensional Vector I1 from the internal storage according to the input vector base address 1;

[0213] Step 2: Broadcast Scalar I2 to be the same length as I1. After performing element-by-element SIMD calculation with the one-dimensional Vector I1, output a one-dimensional Vector O of the same length.

[0214] Step 3: Write Vector O into the internal storage according to the output vector base address.

[0215] V_S instruction: Used to complete the vector reduction operation and finally produce a scalar result

[0216] Step 1: Read out the one-dimensional Vector I1 from the internal storage according to the input vector base address.

[0217] Step 2: After converting Vector I into Scalar O through the reduction operation;

[0218] Step 3: Return Scalar O to end the execution.

[0219] V_V instruction: Used to complete the data type conversion of vectors

[0220] Step 1: Read out the one-dimensional Vector I1 from the internal storage according to the input vector base address.

[0221] Step 2: Process each element of Vector I1 and output a one-dimensional Vector O;

[0222] Step 3: Write Vector O into the internal storage according to the output vector base address.

[0223] In summary, the core idea and innovation of the implementation method of the in-memory computing standard extension instruction set based on the RISC-V instruction set provided in this embodiment are mainly reflected in:

[0224] 1. Hierarchical decoupled programming model:

[0225] Core idea:

[0226] By abstracting the in-memory computing hardware into software-visible architecture registers, the decoupling of software and hardware is realized.

[0227] Map the storage parts of SPAD (Scratchpad Memory, also known as temporary storage memory) and CIM (In-Memory Computing Cluster) to storage registers with a two-dimensional data layout and computing registers with a three-dimensional data layout respectively to form a unified programming model.

[0228] Innovation:

[0229] This hierarchical decoupled programming model enables the independent development of upper-layer compilation and lower-layer microarchitecture implementation, reduces the cost and difficulty of software stack development, and accelerates the hardware development and iteration process.

[0230] The unified programming model allows different in-memory computing hardware designs to adopt the same abstract operator library and compilation toolchain, improving the portability and reusability of code.

[0231] 2. Parameterized definition of hardware implementation:

[0232] Core idea:

[0233] Define the hardware implementation in a parameterized form to facilitate the application of different types of hardware to the programming model.

[0234] Define a series of fixed parameters for hardware implementation to provide the host with the hardware resource scale of the co-processor implemented based on CIM.

[0235] Innovation:

[0236] This parameterized hardware description information enables co-processors implemented based on CIM with different scales and implementation methods to form a unified hardware description, allowing the CIM standard extended instruction set to be widely applied to various hardware accelerators.

[0237] Avoids the problems of insufficient scalability and portability in bottom-up design projects.

[0238] The defined hardware parameters are more adapted to the hardware characteristics of CIM. For example, the memory-computation ratio is defined by CIMM_PG_NUM, which better reflects the shared characteristics of internal storage and computing resources in the in-memory computing engine.

[0239] 3. Flexible addressing mode of in-memory computing cluster:

[0240] Core idea:

[0241] The in-memory computing cluster supports both storage and computing functions simultaneously, and its addressing mode is determined by the corresponding configuration scheme.

[0242] In the storage mode, the in-memory computing cluster is regarded as a two-dimensional storage structure, and the addressing method is similar to that of the scratchpad memory.

[0243] In the computing mode, the in-memory computing cluster is regarded as a three-dimensional storage structure, and the data arrangement method is closely related to the corresponding computing expansion method.

[0244] Innovation:

[0245] By configuring the addressing mode of the in-memory computing cluster through configuration registers, efficient mapping for different computing requirements is achieved, improving the hardware utilization efficiency.

[0246] According to the addressing configuration, the addressing space of the in-memory computing cluster can be represented as a three-dimensional cube and organized into a continuous address space in a specific order, facilitating data access and management.

[0247] 4. Address mapping scheme of RISC-V Local Memory:

[0248] Core idea:

[0249] Organize the internal storage of RISC-V and the control and status register file in a unified storage space and define the corresponding byte address organization method.

[0250] Define the address mapping relationship between RISC-V and the Scratchpad of NPU, facilitating data transfer and access.

[0251] Innovation:

[0252] This address mapping scheme realizes efficient mapping for different computing requirements and improves the hardware utilization efficiency according to task requirements in end-to-end AI application scenarios.

[0253] 5. Instruction set partitioning and design:

[0254] Core idea:

[0255] Divide instructions into control instructions, memory access instructions, data movement and operation instructions, matrix calculation instructions, and element-level calculation instructions.

[0256] Define specific instructions and operation codes, transforming the hardware abstraction and interface specifications in the programming model into commands that can be executed by the hardware.

[0257] Innovation:

[0258] This instruction set partitioning and design cover the basic functions and extended requirements of the in-memory computing coprocessor, providing a broad space for hardware design.

[0259] Form Primitive through fixed combinations of multiple instructions. This level interfaces with the AI compiler, facilitating software development and optimization.

[0260] In summary, through a hierarchical decoupled programming model, parameterized definition of hardware implementation, flexible addressing mode, address mapping scheme, and division and design of the instruction set, this technical solution realizes a standard extended instruction set for in-memory computing based on the RISC-V instruction set, improving the generality, scalability, and development efficiency of in-memory computing hardware.

[0261] This embodiment also provides an implementation example of RISC-V cooperating with a CIM-based NPU core. Specifically, as Figure 4 shown, it respectively shows the process of the RISC-V Core sending instructions and data to the RoCC Accelerator and the process of the RoCC Accelerator returning the calculation result to the RISC-V Core. RISC-V and RoCC extension:

[0262] This solution is based on the RISC-V instruction set architecture and uses the RoCC (Rocket Custom Coprocessor) extension mechanism of RISC-V to achieve integration with the CIM-based NPU Core.

[0263] RoCC allows users to customize coprocessors and add custom instructions to expand the functions of RISC-V.

[0264] Instruction encoding:

[0265] This solution selects the extended instruction operation codes (opcodes) reserved by the RISC-V ISA: custom-0 / 1 for the extended instructions of the coprocessor.

[0266] This selection ensures compatibility with future RISC-V standards.

[0267] Data interaction:

[0268] The RISC-V Core initiates a calculation request by passing the instruction field and the source register values (rs1val, rs2val) to the NPU Core.

[0269] After the NPU Core executes the instruction, it returns the destination register value (rdval) to the RISC-V Core.

[0270] Instruction classification:

[0271] Due to the limited information of a single instruction, this solution defines multiple hardware instructions and divides them into four categories to complete complex coarse-grained operations:

[0272] CFG Type (configuration type):

[0273] Used to configure the functional unit registers inside the NPU Core.

[0274] Carries configuration information unrelated to address mapping, such as dimension size, data format, computing unit functions, etc.

[0275] rs1 and rs2 correspond to the configuration register address and the written value respectively.

[0276] PRE Type (Preliminary Type):

[0277] Used to carry local / global memory addresses or scalar values that need to be passed.

[0278] Does not drive the actual execution unit.

[0279] DRV Type (Drive Type):

[0280] Used to carry local / global memory addresses or scalar values that need to be passed.

[0281] Activates the execution unit and starts to execute operations according to the information of the previous CFG instruction and PRE instruction.

[0282] According to the xd field, it may return a scalar calculation result to RISC-V through rd.

[0283] POST Type (Post-Processing Type):

[0284] When the bit width of the scalar value to be returned is greater than the RISC-V XLEN, it is used to complete the collection of the remaining part.

[0285] Function implementation:

[0286] Through the cooperation of these instructions, flexible control of the NPU Core can be achieved to complete various complex computing tasks.

[0287] For example, first configure the computing unit through the CFG instruction, then prepare the data through the PRE instruction, and finally start the calculation through the DRV instruction.

[0288] This solution realizes the efficient integration with the CIM-based NPU Core through the RoCC extension mechanism of RISC-V.

[0289] By defining different types of instructions, flexible control of the NPU Core and the execution of complex computing tasks are realized.

[0290] This design makes full use of the scalability of RISC-V and provides an efficient solution for the application of the NPU Core.

[0291] In summary, through the RoCC extension of RISC-V, this solution realizes the tight integration of RISC-V and the NPU Core, and designs a flexible instruction set to support various complex computing tasks.

[0292] So far, the technical solution of the present invention has been described in conjunction with the accompanying drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to the above specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the protection scope of the present invention.

Claims

1. An implementation method of an in-memory computing standard extended instruction set based on the RISC-V instruction set, characterized in that Including the steps: Steps for programming model implementation, which abstracts the in-memory computing hardware as software-visible architecture registers in a hierarchical decoupling manner. The architecture registers include a storage register that maps internal storage to a two-dimensional data layout and a computing register that maps the in-memory computing cluster to a three-dimensional data layout; and Steps for instruction set implementation: transforming the hardware abstraction and interface specifications in the programming model into commands executable by the hardware through specific instructions and operation codes.

2. The implementation method of the in-memory computing standard extended instruction set based on the RISC-V instruction set according to claim 1, characterized in that, It also includes: Steps for defining hardware implementation parameters: defining the hardware implementation in a parameterized form to facilitate the application of the programming model to different types of hardware.

3. The implementation method of the in-memory computing standard extended instruction set based on the RISC-V instruction set according to claim 2, characterized in that, The "defining the hardware implementation in a parameterized form" includes: CIMC_NUM: The number of in-memory computing clusters inside the neural network processor core; CIMC_CIME_NUM: The number of in-memory computing engines included in each in-memory computing cluster; CIME_ROW_NUM: The number of rows of each in-memory computing engine; CIME_COL_NUM: The number of columns of each in-memory computing engine; CIMM_PG_NUM: In-memory computing macro storage-computation ratio; SPAD_NUM: The number of scratch memories inside the neural network processor core; SPAD_WD: The bit width of the scratch memory; and SPAD_DP: The depth of the scratch memory.

4. The implementation method of the in-memory computing standard extension instruction set based on the RISC-V instruction set according to any one of claims 1-3, characterized in that, The storage hierarchy defined by the programming model includes: External storage: Interacts with internal storage in a direct memory access manner through load / store instructions and is not directly mapped to architecture registers; and Internal storage includes: Scratch memory: A storage register mapped to a two-dimensional data layout, including a scratch memory for storing scalar and tensor data; In-memory computing cluster: A computing register mapped to a three-dimensional data layout in computing mode, for storing tensor data and supporting in-memory computing.

5. The implementation method of the in-memory computing standard extended instruction set based on the RISC-V instruction set according to claim 4, wherein In storage mode, the width of the in-memory computing cluster is: CIME_ROW_NUM; the depth is: CIMC_CIME_NUM * CIME_COL_NUM * CIMM_PG_NUM.

6. The implementation method of the in-memory computing standard extended instruction set based on the RISC-V instruction set according to claim 4, characterized in that, In computing mode, the storage space of the in-memory computing cluster is jointly determined by the in-memory computing macro storage-computation ratio, its effective number of rows, and effective number of columns.

7. The implementation method of the in-memory computing standard extended instruction set based on the RISC-V instruction set according to claim 4, characterized in that, The data types of the neural network processor core include scalar, vector data, and tensor data.

8. The implementation method of the in-memory computing standard extended instruction set based on the RISC-V instruction set according to claim 4, characterized in that, It also includes: Steps for register configuration: Data type and precision configuration, data arrangement information configuration, and operation configuration.

9. The implementation method of the in-memory computing standard extended instruction set based on the RISC-V instruction set according to claim 4, wherein It also includes: Steps for instruction partitioning: Partitioning instructions into control instructions, memory access instructions, data movement and operation instructions, matrix calculation instructions, and element-level calculation instructions.