Processor, processing method of processor and chip
By designing a processor based on the RISC-V instruction set in edge devices and integrating execution units and register stacks, the problems of high power consumption, high latency and insufficient versatility in the execution of artificial intelligence algorithms on edge devices are solved, efficient and low-power computing capabilities are achieved, and the real-time processing capabilities of AI tasks are improved.
Patent Information
- Application Number
- CN202510479184.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-09-12
AI Technical Summary
In the existing technology, when edge devices execute artificial intelligence algorithms, they have problems such as limited computing resources, high power consumption, large latency and insufficient versatility, making it difficult to meet the requirements of efficient computing and low power consumption.
Design a processor based on the RISC-V instruction set that integrates dedicated execution units and register stacks to support the efficient execution of artificial intelligence operator instructions. Through collaborative processing, it achieves high-speed access and rapid feedback of data, reduces power consumption and latency, and improves computing efficiency.
It significantly reduces power consumption and latency during algorithm execution, improves the real-time processing capability of edge AI tasks, supports multiple AI operator operations, and provides a good infrastructure suitable for efficient computing on edge devices.
Smart Images

Figure CN120631435A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer technology, and in particular to a processor, a processing method of the processor, and a chip. Background Art
[0002] With the widespread deployment of artificial intelligence (AI) algorithms in scenarios such as smart terminals, IoT devices, and edge computing platforms, higher requirements are placed on the real-time performance, power consumption, and instruction efficiency of computing resources.
[0003] In related technologies, the edge side usually uses a graphics processing unit (GPU) or a central processing unit (CPU) to execute artificial intelligence algorithms. Among them, although the GPU has high computing power, it has high cost and power consumption, and is not suitable for large-scale deployment; while the CPU has strong versatility but low processing efficiency and high latency.
[0004] Therefore, there is an urgent need for a solution that has both efficient computing capabilities and low power consumption and strong adaptability. Summary of the Invention
[0005] The present invention aims to solve at least one of the technical problems existing in the related art. To this end, the present invention proposes a processor, a processing method of the processor, and a chip to meet the requirements of efficient computing, low power consumption, and versatility under the condition of limited computing resources of edge devices.
[0006] In a first aspect, the present invention provides a processor, comprising:
[0007] A register file, comprising a plurality of registers, wherein the registers are used to cache data;
[0008] An execution unit, configured to receive instruction information from the kernel and determine a source register, a destination register, and an operator instruction to be executed based on the instruction information; wherein the operator instruction includes at least a data load instruction, a data read instruction, and an artificial intelligence operator instruction;
[0009] The execution unit is further configured to read source data from a source register, execute an instruction operation corresponding to an operator instruction, and store the obtained operation result in a destination register for access by the kernel.
[0010] The processor provided by the present invention supports the efficient execution of artificial intelligence operator instructions by integrating a dedicated execution unit in the processor, reducing the redundancy of general instructions and the overhead of loop control; utilizes the collaborative processing of the register stack and the execution unit to achieve high-speed access and rapid feedback of data, thereby improving the efficiency of the data path; supports multiple artificial intelligence operator operations, and provides a good infrastructure for subsequent more complex instruction scheduling and module expansion; significantly reduces power consumption and latency during algorithm execution, and improves the real-time processing capability of edge AI tasks.
[0011] According to one embodiment of the present invention, the artificial intelligence operator instructions include at least convolution instructions, matrix-vector multiplication instructions, pooling instructions, and activation instructions; wherein:
[0012] The convolution instruction, matrix-vector multiplication instruction, pooling instruction and activation instruction are respectively used to trigger the execution of a convolution operation, a matrix-vector multiplication operation, a pooling operation or an activation function operation on data in at least one source register.
[0013] According to one embodiment of the present invention, each operator instruction has a corresponding instruction format, and the instruction format of the artificial intelligence operator instruction includes a type field for indicating the operation type, wherein:
[0014] When the operator instruction is a convolution instruction, the type field indicates whether to perform regular convolution, dilated convolution, or depthwise separable convolution;
[0015] When the operator instruction is a matrix-vector multiplication instruction, the type field indicates whether inner product calculation or bitwise multiplication is to be performed;
[0016] When the operator instruction is a pooling instruction, the type field indicates whether to perform maximum pooling or average pooling;
[0017] When the operator instruction is an activation function instruction, the type field indicates that nonlinear activation is performed.
[0018] According to one embodiment of the present invention, the execution unit includes:
[0019] Multiple computation modules, including a multiplication array, an addition tree, multiple selectors, and comparators;
[0020] Among them, the computing modules are connected in a preset order to form multiple computing paths; the execution unit is used to select the corresponding computing path according to the operator instruction to implement convolution operation, matrix-vector multiplication operation, pooling operation or activation function operation.
[0021] According to one embodiment of the present invention, when the operator instruction is a convolution instruction, the multiple computing modules include at least:
[0022] A multiplication array, connected to the register file, for performing a multiplication operation on source data read from the register file and outputting a multiplication result;
[0023] a first selector connected to the multiplication array and the register file, respectively, for outputting the multiplication result to the addition tree when the multiplication result is selected;
[0024] an addition tree connected to the first selector, and configured to perform an addition operation on the multiplication results to obtain an addition result;
[0025] The second selector is connected to the multiplication array and the addition tree respectively, and is used to output the depth-separable convolution result when the multiplication result is selected, and to output the normal convolution or the void convolution result when the addition result is selected.
[0026] According to one embodiment of the present invention, when the operator instruction is a matrix-vector multiplication instruction, the multiple computing modules include at least:
[0027] A multiplication array, connected to the register file, for performing a multiplication operation on source data read from the register file and outputting a multiplication result;
[0028] a first selector connected to the multiplication array and the register file, respectively, for outputting the multiplication result to the addition tree when the multiplication result is selected;
[0029] an addition tree connected to the first selector, and configured to perform an addition operation on the multiplication results to obtain an addition result;
[0030] The second selector is connected to the multiplication array and the addition tree respectively, and is used to output the bitwise multiplication result when the multiplication result is selected, and output the inner product calculation result when the multiplication result is selected.
[0031] According to one embodiment of the present invention, when the operator instruction is a pooling instruction, the multiple computing modules include at least:
[0032] a first selector connected to the multiplication array and the register file, respectively, for outputting the source data to the addition tree when the source data read from the register file is selected;
[0033] an addition tree connected to the first selector, for performing an addition operation on the source data to obtain an addition result;
[0034] a second selector connected to the multiplication array and the addition tree, respectively, for outputting the addition result to the third selector when the addition result is selected;
[0035] The third selector is connected to the second selector and the register file respectively, and is used to output the average pooling result when the addition result output by the second selector is selected, and to output the maximum pooling result when the source data read from the register file is selected.
[0036] According to one embodiment of the present invention, when the operator instruction is an activation instruction, the multiple computing modules include at least:
[0037] A comparator is connected to the third selector and is used to perform an activation function operation on the operation result output by the third selector when the convolution instruction or the matrix-vector multiplication instruction performs the activation function operation by default, and output a reference value or the original operation result.
[0038] In a second aspect, the present invention provides a processing method of a processor, the method comprising:
[0039] Receive instruction information, the instruction information including a source register, a destination register, and an operator instruction to be executed; wherein the operator instruction includes at least a data load instruction, a data read instruction, and an artificial intelligence operator instruction, and the artificial intelligence operator instruction is used to trigger an operation of an artificial intelligence operator;
[0040] Read source data from the source register, execute the instruction operation corresponding to the operator instruction, and obtain the operation result;
[0041] The operation result is stored for access by the kernel.
[0042] According to the processing method of the processor provided by the present invention, by receiving instruction information, the instruction information includes a source register, a destination register and an operator instruction to be executed, an instruction-driven automatic operator scheduling mechanism is implemented, the operator instruction is quickly triggered and efficiently executed, and the computing performance of the artificial intelligence algorithm is greatly improved; the source data is read from the source register, the instruction operation corresponding to the operator instruction is executed, the operation result is obtained, and the operation result is stored for kernel access and use. The operator instruction triggers the corresponding operation, which can support hardware accelerated execution of typical artificial intelligence algorithm operations such as convolution, matrix multiplication, and pooling, provides a unified data access and result write-back path, enhances the standardization and modular design of the method, simplifies the kernel's control process of the operation path, improves the overall scheduling efficiency and responsiveness, and supports the parallel utilization and flexible reuse of hardware resources, which is conducive to the deployment of high-efficiency and low-power artificial intelligence computing modules on the edge side.
[0043] In a third aspect, the present invention provides a processing device of a processor, the device comprising:
[0044] A receiving module, configured to receive instruction information, the instruction information including a source register, a destination register, and an operator instruction to be executed; wherein the operator instruction includes at least a data loading instruction, a data reading instruction, and an artificial intelligence operator instruction, and the artificial intelligence operator instruction is used to trigger an operation of an artificial intelligence operator;
[0045] A processing module, configured to read source data from the source register, execute the instruction operation corresponding to the operator instruction, and obtain the operation result;
[0046] The storage module is used to store the operation results for access by the kernel.
[0047] According to the processing device of the processor provided by the present invention, by receiving instruction information, the instruction information includes a source register, a destination register and an operator instruction to be executed, an instruction-driven automatic operator scheduling mechanism is implemented, the operator instruction is quickly triggered and efficiently executed, and the computing performance of the artificial intelligence algorithm is greatly improved; the source data is read from the source register, the instruction operation corresponding to the operator instruction is executed, the operation result is obtained, and the operation result is stored for kernel access and use. The operator instruction triggers the corresponding operation, which can support hardware accelerated execution of typical artificial intelligence algorithm operations such as convolution, matrix multiplication, and pooling, provides a unified data access and result write-back path, enhances the standardization and modular design of the method, simplifies the kernel's control process of the operation path, improves the overall scheduling efficiency and responsiveness, and supports the parallel utilization and flexible reuse of hardware resources, which is conducive to the deployment of high-efficiency and low-power artificial intelligence computing modules on the edge side.
[0048] In a fourth aspect, the present invention provides a chip, comprising a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the processing method of the processor as described in the second aspect above.
[0049] In a fifth aspect, the present invention provides a computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the processor implements the processing method described in the second aspect above.
[0050] In a sixth aspect, the present invention provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the processing method of the processor as described in the second aspect above.
[0051] In a seventh aspect, the present invention provides a computer program product, comprising a computer program, which, when executed by a processor, implements the processing method of the processor as described in the second aspect above.
[0052] Additional aspects and advantages of the present invention will be set forth in part in the description which follows and, in part, will be obvious from the description which follows, or may be learned by practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments with reference to the accompanying drawings, in which:
[0054] Figure 1 is a schematic diagram of the structure of a processor provided in some embodiments of the present invention;
[0055] Figure 2 is a schematic structural diagram of an execution unit provided in some embodiments of the present invention;
[0056] Figure 3 is a schematic diagram of the instruction format of the convolution instruction provided in some embodiments of the present invention;
[0057] Figure 4 1 is a schematic diagram of the instruction format of a matrix-vector multiplication instruction provided in some embodiments of the present invention;
[0058] Figure 5 is a schematic diagram of the instruction format of the pooling instruction provided in some embodiments of the present invention;
[0059] Figure 6 is a schematic diagram of the instruction format of an activation instruction provided in some embodiments of the present invention;
[0060] Figure 7 is a schematic diagram of the instruction format of a data loading instruction provided in some embodiments of the present invention;
[0061] Figure 8 is a schematic diagram of the instruction format of a data storage instruction provided in some embodiments of the present invention;
[0062] Figure 9 is a flowchart of a processing method of a processor provided in some embodiments of the present invention;
[0063] Figure 10 is a schematic structural diagram of a processing device of a processor provided in some embodiments of the present invention;
[0064] Figure 11 It is a schematic diagram of the structure of a computer device provided in some embodiments of the present invention. DETAILED DESCRIPTION
[0065] The following will be combined with the accompanying drawings in the embodiments of the present invention to clearly describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, but not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field are within the scope of protection of the present invention.
[0066] Unless otherwise defined, all technical and scientific terms used in the present invention have the same meanings as commonly understood by those skilled in the art to which the present invention pertains. The terms used in the specification and application of the present invention are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The terms "including" and "having," as well as any variations thereof, in the specification and claims of the present invention and the accompanying drawings are intended to cover non-exclusive inclusions. The terms "first," "second," etc., in the specification and claims of the present invention and the accompanying drawings are used to distinguish between different objects, rather than to describe a specific order or a primary-secondary relationship.
[0067] References to "embodiments" in this disclosure mean that a particular feature, structure, or characteristic described in connection with the embodiment may be included in at least one embodiment of the disclosure. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0068] In the description of the present invention, it should be noted that, unless otherwise expressly specified or limited, the terms "mounted," "connected," "connected," and "attached" should be understood broadly. For example, they may refer to fixed connections, detachable connections, or integral connections; they may refer to direct connections, indirect connections through an intermediate medium, or internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on specific circumstances.
[0069] The term "and / or" in this disclosure simply describes an association between related objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, or B exists alone. Furthermore, the character " / " in this disclosure generally indicates that the related objects are in an "or" relationship.
[0070] The term "multiple" used in the present invention refers to more than two (including two). Similarly, "multiple groups" refers to more than two groups (including two groups), and "multiple sheets" refers to more than two sheets (including two sheets).
[0071] Artificial intelligence algorithms have been widely used in various fields in recent years, such as image recognition, object detection, and language modeling. As AI algorithms become increasingly complex, their demand for computing resources continues to increase. However, in edge computing scenarios, traditional high-computing GPUs or high-performance CPUs are not suitable for directly deploying complex neural network models due to significant limitations on device computing power, power consumption, and size.
[0072] To meet the low power consumption and high efficiency requirements of edge devices for AI inference tasks, the industry has been exploring ways to improve operator computing performance through hardware architecture optimization in recent years. While GPUs possess strong parallel computing capabilities and offer high performance when processing large amounts of data, they are expensive and consume a lot of power. The high cost of deploying algorithms using GPUs across hundreds or thousands of edge nodes makes it difficult to apply GPUs on a large scale to edge devices that are highly sensitive to power consumption and cost. Furthermore, because edge devices typically rely on batteries or low-energy power sources, the high computing power of GPUs leads to high energy consumption, which can lead to practical issues such as insufficient battery life and heat dissipation during long-term operation.
[0073] In comparison, using general-purpose CPUs to deploy artificial intelligence algorithms in software has advantages in terms of versatility and development costs. However, when faced with large-scale parallel computing tasks represented by matrix multiplication and convolution, its instruction execution path is long and relies on a large number of loops and branch judgments, resulting in low overall computing efficiency and large delays.
[0074] Another option is to use ASIC (Application-Specific Integrated Circuit) chips to deploy specific AI algorithms. ASIC chips can provide extreme energy efficiency and execution efficiency for specific tasks, making them suitable for high-performance inference needs. However, ASIC chips require highly customized designs for specific algorithms or tasks. The entire process, from architecture design to chip manufacturing to chip verification, is extremely costly and time-consuming. Furthermore, AI algorithms iterate very quickly, and ASIC chips lack flexibility and scalability, making it difficult to cope with the rapid iteration of AI algorithms and the ever-changing model structures.
[0075] To address these issues, the scalability of the open-source RISC-V instruction set (Reduced Instruction Set Computing-V, the fifth-generation reduced instruction set) provides a new solution for edge AI computing. The RISC-V instruction set architecture allows for flexible expansion and customization of instructions based on application scenarios, making it suitable for introducing AI-related computing modules such as vector computing, matrix computing, and neural network acceleration.
[0076] In view of this, the present invention proposes a processor and a processing method. By designing a new processor based on the RISC-V instruction set, a set of custom instruction sets and their micro-architecture implementations are defined for artificial intelligence operators, supporting calculations such as matrix-vector multiplication, convolution, and pooling. While maintaining the versatility of the chip, it can significantly improve the execution efficiency of operators, reduce power consumption and latency, and thus provide effective support for the deployment of efficient and flexible artificial intelligence algorithms on the edge side.
[0077] In some embodiments, the processor of the present invention can be deployed in edge-side computer equipment. Edge-side computer equipment refers to a computing device that is deployed in a physical location close to the data acquisition source or end user and has localized computing capabilities, data storage capabilities, and network communication capabilities. It is often used to perform low-power, low-latency data processing and artificial intelligence reasoning tasks in a resource-constrained environment. Such devices may include but are not limited to: embedded controllers, edge gateways, single-board computers, smart sensors, AI edge boxes, IoT terminals, and other programmable computing hardware. Compared to cloud servers, these devices emphasize low power consumption, low latency, miniaturization, and high integration.
[0078] Exemplarily, computer devices include but are not limited to embedded system devices, such as industrial controllers, vehicle control units (ECUs), smart cameras, and smart sensors; edge gateway devices, such as industrial edge gateways and smart home gateways; IoT terminal devices, such as smart bracelets, smart speakers, smart home appliances, drones, robots, etc.; AI acceleration devices, such as edge smart boxes equipped with edge AI chips; and mobile terminal devices, such as smartphones, tablets, and wearable devices.
[0079] Faced with the rapid iteration speed of artificial intelligence algorithms, the present invention uses a customized RISC-V instruction set to implement general CPU functions while reusing the CPU's pipeline architecture to accelerate artificial intelligence algorithm calculations, thereby improving flexibility compared to ASIC chips. At the same time, because the algorithm is implemented through fewer customized instructions, the amount of software code is reduced and the execution efficiency is higher.
[0080] The processor and processing method provided by the present invention are described in detail below through specific embodiments and their application scenarios in conjunction with the accompanying drawings.
[0081] The processor is the core component of a computer system, responsible for performing computing tasks and processing data and instructions.
[0082] like Figure 1As shown, the processor provided by the present invention includes at least a processor core and memory. The processor core is the core of the CPU, responsible for executing program instructions, processing data, and controlling computer operations. For example, there may be one or more processor cores. The memory is used to store program code, constant data, or large-scale input / output data.
[0083] In the present invention, the processor core adopts a micro-architecture design based on the RISC-V instruction set to implement various operator operations in artificial intelligence algorithms, and can be regarded as a scheduler of artificial intelligence computing tasks.
[0084] The microarchitecture design includes register files and execution units. In microarchitecture implementation, pipelining is used to improve computing efficiency. For example, a three-stage pipeline can be used, consisting of data loading, operator execution, and data storage.
[0085] The register file is a storage unit used to temporarily store intermediate data or calculation results generated during CPU operation. It consists of multiple readable and writable registers. Registers are the basic storage units in the register file, used to temporarily store data for reading or writing during instruction execution.
[0086] For example, the register file includes N registers, {m0, m1, ..., mN}. The number of registers N can be determined based on the microarchitecture implementation. To implement the calculation of custom operator instructions, the registers can be, for example, long-bitwidth registers that store calculation data. Long-bitwidth registers have a larger bit width than general-purpose registers to meet the data parallel processing requirements of artificial intelligence algorithms.
[0087] The execution unit is responsible for performing arithmetic operations, logical operations, and executing instructions for artificial intelligence operators. Artificial intelligence operators refer to the basic computing units that perform specific mathematical operations in artificial intelligence models, such as convolution, matrix-vector multiplication, pooling, and activation functions.
[0088] In an embodiment of the present invention, the execution unit is used to receive instruction information from the kernel, and based on the instruction information, read the source data from the source register, and trigger the instruction operation corresponding to the operator instruction on the source data to obtain the operation result, and finally store the operation result for access and use by the kernel.
[0089] In some embodiments, the processor core also includes an instruction decoder, a core component of the processor pipeline that converts binary data loaded from memory or cache into control signals recognizable by the execution unit. For example, the instruction information obtained by the execution unit is issued by the core's instruction decoder.
[0090] The source register is the data read register specified in the operator instruction, containing the original data required for calculation. The destination register is the register used to store the calculation result after the operator instruction is executed, for subsequent reading.
[0091] Among them, AI operator instructions are used to trigger the execution of special instructions of a certain AI operator, including but not limited to convolution instructions, matrix-vector multiplication instructions, pooling instructions, and activation instructions. Accordingly, the operator operation of the AI operator includes but is not limited to convolution operations, matrix-vector multiplication operations, pooling operations, or activation function operations.
[0092] Exemplarily, as shown in Table 1 below, the present invention is based on the expansion of the RISC-V instruction set and defines a custom instruction set for a set of artificial intelligence operators, including multiple operator instructions, which can support matrix-vector multiplication, various convolution calculations, pooling calculations, and activation function calculations, and also support batch loading and batch storage instructions to improve data transmission efficiency.
[0093] Table 1 Operator instructions
[0094] Serial number name Assembly code describe 1 LDMM LDMM rd,imm(rs1) Load memory data into the specified register according to the address 2 STMM STMM rs2,imm(rs1) Store the specified register data into the memory according to the address 3 CONV CONV rd,rs1,rs2 Convolution calculation, the result is placed in the destination register 4 MMUL MMUL rd,rs1,rs2 Matrix-vector multiplication, the result is placed in the destination register 5 POOL POOL rd,rs1,rs2 Pooling calculation, the result is placed in the destination register 6 RELU RELU rd,rs1,rs2 Activate the function and put the result into the destination register …… …… …… ……
[0095] Data load instructions are used to read data from a specified destination register and store it in memory for subsequent processing. For example, the data load instruction is LDMM m0,0(s0). This example adds the address stored in general register s0 to the offset 0 as the memory address, and then loads the data at that address into register m0.
[0096] Data transfer instructions are used to transfer data from a destination register to memory. For example, if the data transfer instruction is STMM m1,32(s1), this instruction adds the address of general register s1 plus offset 32 as the memory address, and then transfers the data in register m1 to memory.
[0097] Convolution instructions trigger the execution of convolution operations, including regular convolution, dilated convolution, and depthwise separable convolution. For example, the convolution instruction CONV m2,m0,m1 performs a convolution on the data in registers m0 and m1, and stores the result in register m2. Convolution operations compute a kernel function using a sliding window and the input feature map, and are typically used for feature extraction.
[0098] The matrix-vector multiplication instruction triggers the execution of a matrix-vector multiplication operation and can be used in fully connected layers or attention mechanism calculations. For example, the matrix-vector multiplication instruction is MMUL m3,m0,m1, which performs a matrix-vector multiplication on the data in registers m0 and m1, and stores the result in register m3. Matrix-vector multiplication performs a vector-matrix multiplication operation and is used for transforming, compressing, or aggregating features.
[0099] Pooling instructions trigger the execution of pooling operations, including average pooling and max pooling. For example, the pooling instruction POOL m2,m0,m1 performs a pooling calculation on the data in registers m0 and m1, and then stores the result in register m2. Pooling operations are used to reduce the spatial dimension of feature maps while retaining important features.
[0100] Activation instructions are used to execute activation functions, such as the Rectified Linear Unit (ReLU) and the Sigmoid Function, and are typically executed following a convolution or matrix multiplication operation. For example, an activation instruction such as RELU m3,m0,zero indicates that the activation function is calculated on the data in register m0 and the result is stored in register m3. Activation function operations are nonlinear processing operations used to enhance the expressive power of AI models.
[0101] In the scenario of running artificial intelligence algorithms, the execution efficiency of artificial intelligence algorithms can be improved by directly calling the corresponding operator instructions.
[0102] The processor provided by the present invention supports the efficient execution of artificial intelligence operator instructions by integrating a dedicated execution unit in the processor, reducing the redundancy of general instructions and the overhead of loop control; utilizes the collaborative processing of the register stack and the execution unit to achieve high-speed access and rapid feedback of data, thereby improving the efficiency of the data path; supports multiple artificial intelligence operator operations, and provides a good infrastructure for subsequent more complex instruction scheduling and module expansion; significantly reduces power consumption and latency during algorithm execution, and improves the real-time processing capability of edge AI tasks.
[0103] Among them, the register file supports intermediate data access and fast scheduling during instruction execution, improves operator execution efficiency, and can provide a high-speed, low-latency data access mechanism within the processor, avoiding performance bottlenecks caused by frequent access to main memory.
[0104] Among them, the execution unit receives instruction information from the kernel to determine the source register, destination register and operator instruction to be executed, thereby realizing an instruction-driven automated computing process. The kernel only needs to issue instruction information to trigger the complete data path and operator execution; and supports unified access of multiple types of instructions, which improves the versatility and scalability of operator processing capabilities; by decoupling the specific implementation of the kernel and the hardware execution path, instruction control is made more flexible, which is conducive to the scalability design of the instruction set; in addition, the execution unit reads the source data from the source register, executes the instruction operation corresponding to the operator instruction, and stores the obtained operation result in the destination register for kernel access, supporting automated data flow and operator triggering, avoiding software intervention in multi-step data transfer, reducing control complexity, greatly shortening the data write-back path, improving processor operation efficiency, and supporting rapid feedback of calculation results to the kernel, which helps to realize pipeline processing, result reuse and subsequent scheduling optimization.
[0105] Typically, each operator instruction has a corresponding instruction format. The instruction format of an AI operator instruction includes information such as a type field that indicates the specific operation type. The type field is a coded field in the operator instruction that identifies the operation type corresponding to the current operator instruction. This operation type indicates the specific calculation to be performed by each operator instruction.
[0106] In addition, because RISC-V is a reduced instruction set architecture, each instruction is fixed at 32 bits, so the instruction format of operator instructions is also 32 bits. Each operator instruction has an opcode, which is used to indicate which category the operator instruction belongs to, such as addition, multiplication, or function processing, and determines the instruction format and general functional category.
[0107] In some embodiments, when the operator instruction is a convolution instruction, the type field indicates whether to perform a regular convolution, a dilated convolution, or a depthwise separable convolution. Figure 2 As shown, the instruction format of the convolution instruction CONV is R-type (RegisterType, register type), which is applicable to register-to-register operations. The instruction field funct3 is used to indicate which subcategory under the general category the operator instruction belongs to, such as normal signed multiplication, normal unsigned multiplication, signed and unsigned multiplication, etc. under the multiplication category. The convolution instruction CONV instructs the execution unit to extract the source data data1 from the source register rs1, extract the source data data2 from the source register rs2, and perform conventional convolution, void convolution or depth-separable convolution operations on the source data data1 and the source data data2 according to the operation type sub-type indicated by the type field funct7, obtain the operation result result, and store the operation result result in the destination register rd.
[0108] In some embodiments, when the operator instruction is a matrix-vector multiplication instruction, the type field indicates whether an inner product calculation or a bitwise multiplication is to be performed. Figure 3 As shown, the instruction format of the matrix-vector multiplication MMUL is R-type. The matrix-vector multiplication MMUL instruction instructs the execution unit to extract source data data1 from source register rs1 and source data data2 from source register rs2, perform an inner product or bitwise multiplication operation on source data data1 and data2 according to the operation type sub-type indicated by the type field funct7, and obtain the operation result result, which is then stored in the destination register rd.
[0109] In some embodiments, when the operator instruction is a pooling instruction, the type field indicates whether to perform maximum pooling or average pooling. Figure 4 As shown, the instruction format of the pooling instruction POOL is R-type. The pooling instruction POOL instructs the execution unit to extract source data data1 from source register rs1 and source data data2 from source register rs2, and perform average pooling or max pooling operations on the source data data1 and data2 according to the operation type sub-type indicated by the type field funct7, to obtain the operation result result, and store the operation result result in the destination register rd.
[0110] In some embodiments, when the operator instruction is an activation instruction, the type field indicates that nonlinear activation is to be performed. For example, Figure 5 As shown, the RELU activation instruction has an R-type format. The RELU activation instruction instructs the execution unit to extract source data data1 from source register rs1 and source data data2 from source register rs2, and perform the RELU activation function operation on the source data data1 and source data2 according to the operation type sub-type indicated by the type field funct7, obtaining the operation result result, and storing the operation result result in the destination register rd.
[0111] In some embodiments, when the operator instruction is a data load instruction, since no operation is involved, the instruction format of the data load instruction does not include a type field. Figure 6As shown, the data load instruction LDMM has an I-type (Immediate Type) instruction format. The base address baseaddr typically refers to the value of the address reference register in memory access instructions and is used to calculate the actual address being accessed. Immediate values are typically used in immediate operations, load instructions, or system calls. The immediate field imm[11:0] represents a 12-bit signed immediate value. The data load instruction LDMM instructs the execution unit to load the source data in the base address register rs1 into the destination register rd.
[0112] In some embodiments, when the operator instruction is a data storage instruction, the instruction format of the data storage instruction does not include a type field. Figure 7 As shown in Figure 1, the instruction format for the data storage instruction STMM is S-type (StoreType). In the S-type instruction format, rs1 is used to provide the address reference value baseaddr to calculate the target address, and rs2 is used to provide the data value to be stored. The immediate field is split into two parts, imm[11:5] and imm[4:0], to meet the requirements of the instruction format. After merging, it forms a complete 12-bit immediate value, imm[11:0].
[0113] To implement these defined AI operator instructions in a processor core based on the RISC-V instruction set, a specific microarchitecture is required. This microarchitecture implements three instruction designs: data loading, operator execution, and data storage. The overall process involves the execution unit obtaining instruction-related information, such as the opcode and operands, from the core's instruction decoder. Opcodes include, for example, opcode, funct3, and funct7. Operands include, for example, the numbers of source registers rs1 and rs2, the number of destination registers rd, immediate values, offsets, and base addresses. The opcode determines which operator instruction to execute, and the operands determine which register to access data from. The execution unit loads the corresponding data into registers based on the memory address information, performs calculations based on the corresponding operands, and stores the results in registers. Finally, the results in registers are stored in memory.
[0114] In some embodiments, the execution unit is further configured to send a completion signal back to the core after the execution of an operator instruction is complete, thereby supporting pipeline scheduling and continuous execution of the operator instructions. That is, similar to basic instructions, after each operator instruction is completed, the execution unit generates a complete signal to the core, indicating the completion of the operator instruction.
[0115] Typically, AI algorithms perform convolution, matrix-vector multiplication, or pooling operations repeatedly in a loop. Software code executes in a pipeline, following the order of data loading, operator execution, and data storage. This achieves maximum efficiency, delivering computational results in every cycle. After the computation is complete, the result is written back to the designated destination register for subsequent instruction invocation or kernel readout. This execution is accomplished at the hardware level via a custom computational pipeline, eliminating the need for redundant software calls or loop control, thereby improving data path utilization and execution efficiency.
[0116] To meet the computing requirements of different artificial intelligence operators, the execution unit can flexibly combine multiple modules to form different computing paths, support hardware acceleration of multiple types of artificial intelligence operators, and achieve efficient response to complex instructions.
[0117] To this end, in some embodiments, the execution unit includes: multiple computing modules, including a multiplication array, an addition tree, multiple selectors and comparators; wherein the computing modules are connected in a preset order to form multiple computing paths; the execution unit is used to select the corresponding computing path according to the operator instruction to implement convolution operations, matrix-vector multiplication operations, pooling operations or activation function operations.
[0118] A selector is a module with multiple input channels and one output channel, selecting one input channel for output based on a control signal. Gating refers to a control operation that allows a particular signal to pass through under a specific control signal. A comparator compares the magnitude of two input values and is often used in activation function operations to determine whether an output should be retained. After executing instructions such as convolution or matrix multiplication, the comparator automatically executes the activation function without the need for additional control instructions.
[0119] The computation path is a data processing path composed of computation modules. Different operator types select different paths to complete the corresponding artificial intelligence operator operations.
[0120] Matrix-vector multiplication is divided into inner product calculation and bitwise multiplication. Inner product calculation requires a multiplication array and an adder tree to multiply and accumulate the corresponding data, while bitwise multiplication simply multiplies the two input data sets bit by bit. An activation function is required at the final stage of accumulation in the feedforward layer and the fully connected layer. Therefore, the activation function can be implemented through a comparator or directly bypassed as needed. The specific matrix-vector multiplication operator type is distinguished by the funct7 field.
[0121] Convolution operations are categorized into regular convolution, dilated convolution, and depthwise separable convolution. Similar to matrix-vector multiplication operators, regular and dilated convolutions require a multiplication array and adder tree to multiply and accumulate the corresponding data, while depthwise separable convolution simply multiplies the two input data sets bit by bit. An activation function is required at the final stage of accumulation in the convolutional layer. This activation function can be implemented via a comparator or bypassed as needed. The specific convolution operator type is identified by the funct7 field.
[0122] Pooling operations are categorized into maximum pooling and average pooling. Currently, the algorithm primarily uses maximum pooling, which requires applying a comparator to two input data sets to obtain the maximum value. Average pooling requires applying an adder to two input data sets and right-shifting them by two bits to obtain the average value. Pooling layers do not require an activation function. The specific pooling operator type is identified by the funct7 field.
[0123] The activation function supports the ReLU function. Currently, the vast majority of networks use the ReLU function as the activation function. Furthermore, the ReLU function is relatively simple to implement in circuits. It simply passes the input data through a comparator and outputs the original value or zero based on the positive or negative value of the data.
[0124] In some embodiments, as Figure 8 As shown, the execution unit in the present invention is highly modular and reconfigurable, and its structure includes multiple computing modules: a multiplication array, an addition tree, a first selector, a second selector, a third selector, and a comparator. Among them: the multiplication array is connected to the register stack and is used to perform a multiplication operation on the source data read from the source register to obtain a multiplication result; the first selector is connected to the multiplication array and the register stack respectively and is used to select and output the multiplication result or the source data based on the target operator instruction; the addition tree is connected to the first selector and is used to perform an addition operation on the output of the first selector to obtain an addition result; the second selector is connected to the multiplication array and the addition tree respectively and is used to select and output the multiplication result or the addition result based on the target operator instruction; the third selector is connected to the second selector and the register stack respectively and is used to select and output the output of the second selector or the source data based on the target operator instruction; the comparator is connected to the third selector and is used to perform a comparison operation on the output of the third selector to output the input data or a reference value to obtain the operation result. This structure enables the execution unit to dynamically select different module combinations under instruction drive to form a computing path suitable for convolution operations, matrix-vector multiplication operations, pooling operations and activation function operations.
[0125] The present invention designs a computational path that meets the requirements of multiplication and addition operations and configurable convolution type selection based on the computational characteristics of convolution operations. To this end, in some embodiments, when the operator instruction is a convolution instruction, multiple computational modules include at least: a multiplication array connected to a register file, configured to perform a multiplication operation on source data read from the register file and output a multiplication result; a first selector connected to the multiplication array and the register file, configured to output the multiplication result to an addition tree when the multiplication result is selected; an addition tree connected to the first selector, configured to perform an addition operation on the multiplication result to obtain an addition result; and a second selector connected to the multiplication array and the addition tree, configured to output a depthwise separable convolution result when the multiplication result is selected, and to output a regular convolution or a dilated convolution result when the addition result is selected.
[0126] Combine Figure 8 As shown, the execution flow of the convolution instruction CONV is as follows: the execution unit obtains source data from the register file and sends it to the multiplication array for multiplication of the convolution kernel and the input window data. If the target operator instruction is a convolution instruction and the control signal is signal 1, the first selector MUX (a) selects the multiplication result and outputs the multiplication result to the addition tree. The addition tree receives the multiplication result and performs an addition operation to obtain the addition result, for example, generating a weighted sum of the convolution. The second selector MUX (b) determines whether to output the addition result or directly output the multiplication result based on the convolution type (such as regular convolution, dilated convolution, or depthwise separable convolution). For example, a signal of 1 indicates that the addition result is selected and the regular convolution / dilated convolution result is output; a signal of 0 indicates that the multiplication result is selected and the depthwise separable convolution result is output. The third selector MUX (c) selects the output result of the second selector when the signal is 1. When the signal is 0, the source data directly read from the register file is retained and the selected result is output. If necessary, it is sent to the comparator to perform activation function processing.
[0127] Therefore, fast switching and reuse of different types of convolutions can be achieved through multi-level selectors, reducing duplicate hardware resources and improving operator compatibility.
[0128] Matrix-vector multiplication is a common operation in artificial intelligence algorithms. It needs to support two types of operations: bitwise multiplication and cumulative summation, and requires the hardware path to be reusable. To this end, in some embodiments, when the operator instruction is a matrix-vector multiplication instruction, the multiple computing modules at least include: a multiplication array, connected to the register file, for performing multiplication operations on source data read from the register file and outputting the multiplication result; a first selector, connected to the multiplication array and the register file, respectively, for outputting the multiplication result to the addition tree when the multiplication result is selected; an addition tree, connected to the first selector, for performing addition operations on the multiplication result to obtain an addition result; a second selector, connected to the multiplication array and the addition tree, respectively, for outputting the bitwise multiplication result when the multiplication result is selected, and outputting the inner product calculation result when the multiplication result is selected.
[0129] Combine Figure 8 As shown, the execution flow of the matrix-vector multiplication instruction MMUL is as follows: the processor loads the matrix and vector from the source register into the multiplication array and performs a bitwise multiplication operation; the first selector MUX (a) selects the multiplication result when the signal is 1 and outputs it to the addition tree, and selects the source data when the signal is 0. The first selector MUX (a) outputs the multiplication result to the addition tree to calculate the inner product result. The second selector MUX (b) outputs the multiplication result or the addition result according to the setting, for example, the addition result (inner product) is output when the signal is 1, and the multiplication result (bitwise multiplication result) is output when the signal is 0. The third selector MUX (c) selects the second selector output when the signal is 1, and retains the source data in the register stack when the signal is 0 to output the final result. If necessary, the activation function transformation is performed through the comparator.
[0130] Therefore, execution efficiency is improved through path reuse, while area and power consumption are saved, and integrated support for bitwise multiplication and inner product operations is achieved.
[0131] Pooling operations are often used for feature map downsampling, and their calculation methods include average pooling and maximum pooling, and the hardware path needs to support flexible switching between the two. To this end, in some embodiments, when the operator instruction is a pooling instruction, multiple computing modules include at least: a first selector, respectively connected to the multiplication array and the register stack, for outputting the source data to the addition tree when the source data read from the register stack is selected; an addition tree, connected to the first selector, for performing addition operations on the source data to obtain an addition result; a second selector, respectively connected to the multiplication array and the addition tree, for outputting the addition result to the third selector when the addition result is selected; a third selector, respectively connected to the second selector and the register stack, for outputting the average pooling result when the addition result output by the second selector is selected, and outputting the maximum pooling result when the source data read from the register stack is selected.
[0132] Combine Figure 8As shown, the execution process of the pooling instruction POOL is as follows: the first selector MUX (a) selects the register stack output as the source data when the signal is 0, and performs a summation operation on the selected multiple groups of pooling window data through the addition tree. The second selector MUX (b) selects the output of the addition tree when the signal is 1 for average pooling. The third selector MUX (c) selects the output addition result (for average pooling) when the signal is 1, and selects the output source data (for maximum pooling) when the signal is 0 according to the pooling type to complete the pooling operation. The output data of the third selector MUX (c) does not pass through the comparator and is directly written back or used by the next operator.
[0133] Therefore, by reusing the computing module, two types of pooling operations are completed, which improves path utilization and supports diversified network structure requirements.
[0134] Some operators integrate activation function operations by default, and the activation function calculation depends on the comparator module to perform specific threshold operations. To this end, in some embodiments, when the operator instruction is an activation instruction, the multiple calculation modules include at least: a comparator, connected to the third selector, for performing an activation function operation on the operation result output by the third selector when the convolution instruction or the matrix-vector multiplication instruction performs the activation function operation by default, and outputting a reference value or the original operation result. That is, when the operator instruction is an activation instruction, or the convolution / matrix-vector multiplication instruction is accompanied by an activation function operation by default, the comparator receives the output of the third selector and executes activation function logic such as ReLU, and outputs a reference value or the original value.
[0135] Combine Figure 8 As shown in FIG, the execution process of the activation instruction RELU is as follows: after performing convolution or matrix-vector multiplication, the data is output through the third selector MUX (c) and sent to the comparator; the comparator receives the output, performs conditional judgment on the input value according to the activation function setting (such as ReLU), and outputs the original value or reference value (such as zero); the result is used as the final activation function operation output for subsequent instructions or written back to the destination register.
[0136] Therefore, the hardware fusion calculation of the activation function is realized through the comparator, which reduces the additional instruction overhead and speeds up the overall execution efficiency of the operator.
[0137] In each of the aforementioned execution flows, the execution unit automatically selects the corresponding computation path based on the operator instruction's type field or operation type information, enabling instruction-level operator fusion and hardware acceleration. The entire execution process is driven by operator instructions, with execution paths activated on demand, ensuring computational efficiency while reducing power consumption.
[0138] The execution unit in this invention supports the efficient execution of artificial intelligence operator instructions through the combination and path reconstruction of multiple computing modules. This execution unit dynamically configures its internal computational path based on different types of operator instructions to implement core artificial intelligence operations such as convolution, matrix-vector multiplication, pooling, and activation function operations. This avoids the traditional CPU serial processing model and adopts parallel / pipeline processing to improve the processing speed of artificial intelligence algorithms. This realizes a highly reconfigurable, low-power, and fast-response computing system that is adaptable to a variety of neural network operators and improves the overall computing power of the processor.
[0139] When executing artificial intelligence algorithms, traditional processors typically rely on multiple general instructions for loop control and data processing, resulting in low processing efficiency, high power consumption, and increased complexity in instruction scheduling and register access. The execution unit designed in this invention dynamically connects multiple computing modules through a selector, supporting flexible switching between paths of different operators, improving hardware utilization and making the computing path highly reconfigurable. Furthermore, modules such as the multiplication array and addition tree can be reused in different operators, reducing chip area overhead and achieving high resource reuse. Furthermore, by selecting the computing path on demand, the activation of redundant computing modules can be avoided, reducing dynamic power consumption.
[0140] The execution unit designed in the present invention supports multi-operator fusion computing, such as the default activation function for convolution and optional activation processing for matrix-vector multiplication, to achieve collaborative optimization of software and hardware, reduce instruction interaction and latency, and design basic operator paths for mainstream artificial intelligence models. It can quickly deploy network structures such as CNN and Transformer, and has strong adaptability to cope with the diversity of artificial intelligence operators.
[0141] In order to fully support the execution process of artificial intelligence operators, the processor must have a complete processing path from instruction reading, data reading, operation execution to result writing back. To this end, in some embodiments, such as Figure 9 As shown, the present invention also provides a processing method of a processor, the execution subject is, for example, a processor. The method includes steps 910 to 930:
[0142] Step 910: Receive instruction information, where the instruction information includes a source register, a destination register, and an operator instruction to be executed; wherein the operator instruction includes at least a data load instruction, a data read instruction, and an artificial intelligence operator instruction, and the artificial intelligence operator instruction is used to trigger an operation of an artificial intelligence operator;
[0143] Step 920: Read source data from the source register, execute the instruction operation corresponding to the operator instruction, and obtain the operation result;
[0144] Step 930: Store the operation result for kernel access.
[0145] The processor obtains the instruction information issued by the kernel, determines the source / destination registers and instruction type based on the instruction, executes the corresponding operator (such as convolution, pooling, etc.), and stores the result in the destination register or output buffer for subsequent modules to read. The specific process can be referred to the previous embodiment and will not be repeated here.
[0146] The processing method of the processor provided by the present invention receives instruction information, which includes a source register, a destination register, and an operator instruction to be executed, thereby realizing an instruction-driven automatic operator scheduling mechanism, realizing rapid triggering and efficient execution of operator instructions, and greatly improving the computing performance of the artificial intelligence algorithm; reading source data from the source register, executing the instruction operation corresponding to the operator instruction, obtaining the operation result, and storing the operation result for kernel access and use. The operator instruction triggers the corresponding operation, which can support hardware accelerated execution of typical artificial intelligence algorithm operations such as convolution, matrix multiplication, and pooling, provides a unified data access and result write-back path, enhances the standardization and modular design of the method, simplifies the kernel's control process of the operation path, improves the overall scheduling efficiency and responsiveness, and supports the parallel utilization and flexible reuse of hardware resources, which is conducive to the deployment of high-efficiency and low-power artificial intelligence computing modules on the edge side.
[0147] The present invention is based on the RISC-V instruction set and is deeply integrated with artificial intelligence operators. In view of the large number of calculations in artificial intelligence algorithms, such as matrix multiplication, convolution, pooling and activation functions, special custom instructions are designed to compress the calculations that originally required dozens of general instructions to complete into a few cycles, thereby enabling the execution of artificial intelligence algorithms with low power consumption and low latency. By decomposing the functions of artificial intelligence algorithms into specific operators and defining these operators as RISC-V extended instructions, a tightly coupled architecture of CPU + extended instructions is achieved. This custom instruction set and microarchitecture implementation can be widely used in typical application scenarios of some edge-side smart devices. Such as face detection in smart cameras, heart rate anomaly detection in wearable devices, and fault detection in industrial sensors.
[0148] This CPU design, achieved through instruction extensions, accelerates AI operator computation while maintaining its original versatility. This collaborative approach between hardware and software reduces the amount of code and the conditional branch overhead caused by extensive loop operations, enabling efficient deployment of AI algorithms at the edge. Accelerating AI operators on general-purpose RISC-V processors through custom instruction extensions avoids the cost of ASIC tapeout and reduces development time. It is also compatible with and utilizes existing RISC-V standard development environments (such as GCC compilation and GDB debugging), reducing development costs.
[0149] The processing method of the processor provided in the embodiment of the present invention may be executed by a processing device of the processor. In the embodiment of the present invention, the processing device of the processor provided in the embodiment of the present invention is described by taking the processing method of the processor performed by the processing device of the processor as an example.
[0150] The embodiment of the present invention further provides a processing device of a processor, which is applied to a computer device. Figure 10 As shown, the processing device of the processor includes a receiving module 1001, a processing module 1002 and a storage module 1003.
[0151] A receiving module 1001 is configured to receive instruction information, the instruction information including a source register, a destination register, and an operator instruction to be executed; wherein the operator instruction includes at least a data load instruction, a data read instruction, and an artificial intelligence operator instruction, and the artificial intelligence operator instruction is used to trigger an operation of an artificial intelligence operator;
[0152] Processing module 1002, configured to read source data from a source register, execute an instruction operation corresponding to an operator instruction, and obtain an operation result;
[0153] The storage module 1003 is used to store the operation results for access by the kernel.
[0154] The processing device of the processor provided according to an embodiment of the present invention receives instruction information, which includes a source register, a destination register, and an operator instruction to be executed, thereby implementing an instruction-driven automatic operator scheduling mechanism, realizing rapid triggering and efficient execution of operator instructions, and greatly improving the computing performance of the artificial intelligence algorithm; reading source data from the source register, executing the instruction operation corresponding to the operator instruction, obtaining the operation result, and storing the operation result for kernel access and use. The operator instruction triggers the corresponding operation, which can support hardware accelerated execution of typical artificial intelligence algorithm operations such as convolution, matrix multiplication, and pooling, provides a unified data access and result write-back path, enhances the standardization and modular design of the method, simplifies the kernel's control process of the operation path, improves the overall scheduling efficiency and responsiveness, and supports the parallel utilization and flexible reuse of hardware resources, which is conducive to the deployment of high-efficiency and low-power artificial intelligence computing modules on the edge side.
[0155] The processing device of the processor in the embodiment of the present invention can be a computer device, or a component in the computer device, such as an integrated circuit or a chip. The computer device can be a terminal device or a server. Exemplarily, the computer device can be a mobile phone, a tablet computer, a laptop computer, a PDA, a vehicle-mounted computer device, a mobile Internet device (Mobile Internet Device, MID), an augmented reality (Augmented Reality, AR) / virtual reality (Virtual Reality, VR) device, a robot, a wearable device, an ultra-mobile personal computer (Ultra-mobile Personal Computer, UMPC), a netbook or a personal digital assistant (Personal Digital Assistant, PDA), etc. It can also be a server, a network attached storage (Network Attached Storage, NAS), a personal computer (Personal Computer, PC), a television (Television, TV), an ATM or a self-service machine, etc., and the embodiment of the present invention does not specifically limit it.
[0156] The processing device of the processor in the embodiment of the present invention may be a device having an operating system. The operating system may be a Microsoft (Windows) operating system, an Android operating system, an iOS operating system, or other possible operating systems, which are not specifically limited in the embodiment of the present invention.
[0157] The processing device of the processor provided in the embodiment of the present invention can realize Figure 9 To avoid repetition, the various processes implemented in the method embodiment are not described here.
[0158] In some embodiments, as Figure 11 As shown, an embodiment of the present invention further provides a computer device 1100, including a processor 1101, a memory 1102, and a computer program stored in the memory 1102 and executable on the processor 1101. When the program is executed by the processor 1101, the various processes of the above-mentioned method embodiments are implemented, and the same technical effects can be achieved. To avoid repetition, they will not be described here.
[0159] It should be noted that the computer devices in the embodiments of the present invention include the mobile computer devices and non-mobile computer devices mentioned above.
[0160] An embodiment of the present invention also provides a non-transitory computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the various processes of the processing method embodiment of the above-mentioned processor and can achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0161] The processor is the processor in the computer device described in the above embodiment. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0162] An embodiment of the present invention further provides a computer program product, including a computer program, which implements the processing method of the processor when executed by the processor.
[0163] The processor is the processor in the computer device described in the above embodiment. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0164] An embodiment of the present invention further provides a chip, which includes a processor and a communication interface, the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the various processes of the processing method embodiment of the above-mentioned processor, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0165] It should be understood that the chip mentioned in the embodiment of the present invention can also be called a system-on-chip, a system-on-chip, a chip system, or a system-on-chip chip, etc.
[0166] It should be noted that, in this article, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a..." does not exclude the presence of other identical elements in the process, method, article or device comprising the element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present invention is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the opposite order according to the functions involved. For example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.
[0167] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better embodiment. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the relevant technology can be embodied in the form of a computer software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), including a number of instructions for enabling a terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present invention.
[0168] The embodiments of the present invention are described above in conjunction with the accompanying drawings, but the present invention is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present invention, ordinary technicians in this field can also make many forms without departing from the scope of protection of the present invention and the claims, all of which are protected by the present invention.
[0169] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "illustrative embodiments," "examples," "specific examples," or "some examples" means that a specific feature, structure, material, or characteristic described in conjunction with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, illustrative uses of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.
[0170] Unless otherwise specified, all embodiments and optional embodiments of the present invention can be combined with each other to form new technical solutions.
[0171] Unless otherwise specified, all technical features and optional technical features of the present invention can be combined with each other to form a new technical solution.
[0172] Unless otherwise specified, all steps of the present invention may be performed sequentially or randomly, preferably sequentially. For example, "the method includes steps (a) and (b)" indicates that the method may include steps (a) and (b) performed sequentially, or may include steps (b) and (a) performed sequentially. For example, "the method may further include step (c)" indicates that step (c) may be added to the method in any order, for example, the method may include steps (a), (b), and (c), or may include steps (a), (c), and (b), or may include steps (c), (a), and (b), etc.
[0173] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.
Claims
1. A processor, characterized in that: The processor includes: A register file, comprising a plurality of registers, wherein the registers are used to cache data; An execution unit, configured to receive instruction information from the kernel and determine a source register, a destination register, and an operator instruction to be executed based on the instruction information; wherein the operator instruction includes at least a data load instruction, a data read instruction, and an artificial intelligence operator instruction; The execution unit is further configured to read source data from a source register, execute an instruction operation corresponding to an operator instruction, and store the obtained operation result in a destination register for access by the kernel.
2. The processor according to claim 1, wherein: The artificial intelligence operator instructions include at least convolution instructions, matrix-vector multiplication instructions, pooling instructions, and activation instructions; wherein: The convolution instruction, matrix-vector multiplication instruction, pooling instruction and activation instruction are respectively used to trigger the execution of a convolution operation, a matrix-vector multiplication operation, a pooling operation or an activation function operation on data in at least one source register.
3. The processor according to claim 2, wherein: Each operator instruction has a corresponding instruction format, and the instruction format of the artificial intelligence operator instruction includes a type field for indicating the operation type, wherein: When the operator instruction is a convolution instruction, the type field indicates whether to perform regular convolution, dilated convolution, or depthwise separable convolution; When the operator instruction is a matrix-vector multiplication instruction, the type field indicates whether inner product calculation or bitwise multiplication is to be performed; When the operator instruction is a pooling instruction, the type field indicates whether to perform maximum pooling or average pooling; When the operator instruction is an activation function instruction, the type field indicates that nonlinear activation is performed.
4. The processor according to claim 1, wherein: The execution unit includes: Multiple computation modules, including a multiplication array, an addition tree, multiple selectors, and comparators; Among them, the computing modules are connected in a preset order to form multiple computing paths; the execution unit is used to select the corresponding computing path according to the operator instruction to implement convolution operation, matrix-vector multiplication operation, pooling operation or activation function operation.
5. The processor according to claim 4, wherein: When the operator instruction is a convolution instruction, the multiple computing modules include at least: A multiplication array, connected to the register file, for performing a multiplication operation on source data read from the register file and outputting a multiplication result; a first selector connected to the multiplication array and the register file, respectively, for outputting the multiplication result to the addition tree when the multiplication result is selected; an addition tree connected to the first selector, and configured to perform an addition operation on the multiplication results to obtain an addition result; The second selector is connected to the multiplication array and the addition tree respectively, and is used to output the depth-separable convolution result when the multiplication result is selected, and to output the normal convolution or the void convolution result when the addition result is selected. The processor according to claim 4 , wherein: When the operator instruction is a matrix-vector multiplication instruction, the multiple computing modules include at least: A multiplication array, connected to the register file, for performing a multiplication operation on source data read from the register file and outputting a multiplication result; a first selector connected to the multiplication array and the register file, respectively, for outputting the multiplication result to the addition tree when the multiplication result is selected; an addition tree connected to the first selector, and configured to perform an addition operation on the multiplication results to obtain an addition result; The second selector is connected to the multiplication array and the addition tree respectively, and is used to output the bitwise multiplication result when the multiplication result is selected, and output the inner product calculation result when the multiplication result is selected.
7. The processor according to claim 4, wherein: When the operator instruction is a pooling instruction, the multiple computing modules include at least: a first selector connected to the multiplication array and the register file, respectively, for outputting the source data to the addition tree when the source data read from the register file is selected; an addition tree connected to the first selector, for performing an addition operation on the source data to obtain an addition result; a second selector connected to the multiplication array and the addition tree, respectively, for outputting the addition result to the third selector when the addition result is selected; The third selector is connected to the second selector and the register file respectively, and is used to output the average pooling result when the addition result output by the second selector is selected, and to output the maximum pooling result when the source data read from the register file is selected.
8. The processor according to any one of claims 4 to 6, characterized in that When the operator instruction is an activation instruction, the multiple calculation modules include at least: A comparator is connected to the third selector and is used to perform an activation function operation on the operation result output by the third selector when the convolution instruction or the matrix-vector multiplication instruction performs the activation function operation by default, and output a reference value or the original operation result.
9. A processing method of a processor, characterized in that: The method comprises: Receive instruction information, the instruction information including a source register, a destination register, and an operator instruction to be executed; wherein the operator instruction includes at least a data load instruction, a data read instruction, and an artificial intelligence operator instruction, and the artificial intelligence operator instruction is used to trigger an operation of an artificial intelligence operator; Read source data from the source register, execute the instruction operation corresponding to the operator instruction, and obtain the operation result; The operation result is stored for access by the kernel.
10. A chip, characterized in that: The chip includes a processor and a communication interface, the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the processing method of the processor according to claim 9.