Computing system, method executed by computing system, and storage medium

By combining a general-purpose processor and tensor processor, using a general-purpose processor and tensor processor to process general-purpose computing and in-memory/non-in-memory computing, the problem of mismatch between storage and computing performance in the von Neumann architecture is solved, and efficient neural network computing is achieved.

WO2025124574A1PCT designated stage expired Publication Date: 2025-06-19SUZHOU YIZHU INTELLIGENT TECH CO LTD

Patent Information

Application Number
PCT/CN2024/139346
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-13
Filing Date
2024-12-13
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

In the prior art, when the processors of the von Neumann architecture improve computing speed, the performance mismatch between storage and computing results in low memory access bandwidth, time delay, and high power consumption, making it difficult to deal with complex and changeable neural network computing operators.

Method used

Combining a general-purpose processor and tensor processor, the general-purpose processor is used to process general-purpose computing in neural networks, and tensor processors are used to perform in-memory and non-in-memory computing, reducing data handling and improving computing efficiency.

Benefits of technology

It realizes that on the basis of ensuring universality, improves computing efficiency, supports complex and variable neural network computing operators, and reduces programming complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024139346_19062025_PF_FP_ABST
    Figure CN2024139346_19062025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in the present application is a computing system. The computing system comprises a general processing unit and a tensor processing unit, wherein the general processing unit comprises an instruction scheduling module and at least one computing unit; the instruction scheduling module is used for identifying a computing instruction, and sending, when the computing instruction is identified as a tensor operation instruction, the tensor operation instruction to the tensor processing unit; and the tensor processing unit is used for performing in-memory computing and / or non-in-memory computing on the basis of the tensor operation instruction. Further provided in the present application is a computing method executed by the computing system. The general processing unit and the tensor processing unit are combined together, such that the general processing unit is used to process general computing in a neural network, and the tensor processing unit is used to perform in-memory computing or non-in-memory computing, thereby ensuring generality and also improving computing power, and supporting complex and variable computing operators in the neural network.
Need to check novelty before this filing date? Find Prior Art

Description

Computing system, method executed by the computing system, and storage medium

[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on December 13, 2023, with application number 202311713616.9 and application name “Computing system, method executed by a computing system and storage medium”, the entire contents of which are incorporated into this application by reference. Technical Field

[0002] The present invention relates to the field of computer technology, and more particularly, to a computing system, a method executed by the computing system, and a storage medium. Background Art

[0003] Existing AI chip solutions for detection and computing tasks fall into two main categories: general-purpose processors such as CPUs (Central Processing Units), FPGAs (Field Programmable Gate Arrays), and GPUs (Graphics Processing Units); and accelerators specifically designed for artificial neural networks, such as Google's TPUs (Tensor Processing Units). These chips all utilize the von Neumann architecture, separating computation and storage, which work together to perform data access and computation. However, because processor design primarily focuses on improving computational speed, while storage prioritizes capacity expansion and cost optimization, a performance mismatch between "storage" and "computing" exists, leading to problems such as low memory access bandwidth, extended latency, and high power consumption—commonly known as the "memory wall" and "power wall." The more intensive the memory access, the more severe the "wall" problem, and the more difficult it is to increase computing power.

[0004] Integrated storage and computing, as a new computing architecture, fully integrates storage and computing, effectively overcoming the bottleneck of the von Neumann architecture and significantly improving computing power. However, in-memory computing generally uses an AISC (application-specific integrated circuit) model, which supports a very limited number of operators and has poor programming flexibility and versatility, making it inflexible enough to keep pace with the current development of neural network models. CPUs and GPUs have advantages in data-intensive general-purpose computing, but general-purpose computing struggles with the complex and ever-changing nature of computational operators. Summary of the Invention

[0005] In view of the above problems, the purpose of the present invention is to provide a computing system, a method executed by the computing system and a storage medium, which combine a general-purpose processor and a tensor processor, use the general-purpose processor to handle general-purpose calculations in neural networks, and use the tensor processor to perform in-memory calculations and non-in-memory calculations, thereby reducing data transfer and improving computing efficiency while ensuring versatility.

[0006] According to a first aspect of the present invention, a computing system is provided, comprising a general-purpose processor and a tensor processor; wherein the general-purpose processor comprises an instruction scheduling module and at least one computing unit, the instruction scheduling module being configured to identify computing instructions, and when the computing instructions are identified as tensor operation instructions, sending the tensor operation instructions to the tensor processor; the tensor processor being configured to perform in-memory computing and / or non-memory computing according to the tensor operation instructions.

[0007] Preferably, the instruction scheduling module is further configured to send the general operation instruction to the computing unit when identifying that the computing instruction is a general operation instruction; the computing unit is configured to perform general calculation according to the general operation instruction.

[0008] Preferably, the computing unit includes at least one computing core, and the computing core is used to perform general computing according to general operation instructions.

[0009] Preferably, the computing instruction includes an operation code and an operation field, wherein the operation code includes the instruction type and operation code of the computing instruction, and the operation field includes the source address and target address of the data to be operated.

[0010] Preferably, the instruction types include tensor operations, vector operations, scalar operations and transcendental function operations.

[0011] Preferably, when the instruction type in the computing instruction is a tensor operation, the instruction scheduling module identifies the computing instruction as a tensor operation instruction; when the instruction type in the computing instruction is one of a vector operation, a scalar operation and a transcendental function operation, the instruction scheduling module identifies the computing instruction as a general operation instruction.

[0012] Preferably, the tensor operation instruction includes at least one of a matrix multiplication operation, a matrix multiplication-addition operation, and a convolution operation; the general operation instruction includes at least one of a floating-point multiplication-addition operation, an integer multiplication-addition operation, and a transcendental function operation.

[0013] Preferably, the tensor processor includes an instruction decoding module, a first computing module and a second computing module. The instruction decoding module is used to parse the tensor operation instruction, generate a selection signal according to the source address of the data to be operated and obtain the data to be operated, divide the data to be operated into multiple groups of operation data, and send the operation code and multiple groups of operation data to the first computing module or the second computing module according to the selection signal; the first computing module is used to perform in-memory calculations based on the received operation code, multiple groups of operation data and target address; the second computing module is used to perform non-memory calculations based on the received operation code, multiple groups of operation data and target address.

[0014] Preferably, the instruction decoding module includes: an instruction parsing unit, used to parse the tensor operation instruction to obtain the operation code and the source address and target address of the data to be operated; a data grouping unit, used to obtain the data to be operated according to the source address of the data to be operated, and divide the data to be operated into multiple groups of operation data; a control unit, used to generate a selection signal according to the source address of the data to be operated, and send the operation code, multiple groups of data and target address to the first computing module or the second computing module according to the selection signal.

[0015] Preferably, the first computing module and the second computing module are connected to a general processor using a unified interface.

[0016] Preferably, the unified interface includes one of a PCIE interface and a UCIE interface.

[0017] Preferably, the general purpose processor and the tensor processor are packaged in the same core.

[0018] Preferably, the general-purpose processor and the tensor processor are packaged into different cores and integrated into the same chip or deployed on different chips.

[0019] Preferably, the general-purpose processor is any one of a CPU, a GPU, a DSP, and a GPGPU.

[0020] According to a second aspect of the present invention, there is provided a computing method performed by a computing system, wherein the computing system includes a general-purpose processor and a tensor processor, the general-purpose processor includes an instruction scheduling module and at least one computing unit, and the computing method includes: the instruction scheduling module of the general-purpose processor identifies the computing instruction; when the instruction scheduling module identifies that the computing instruction is a tensor operation instruction, the tensor operation instruction is sent to the tensor processor; the tensor processor performs in-memory computing and / or non-memory computing according to the tensor operation instruction.

[0021] Preferably, the calculation method further includes:

[0022] When the instruction scheduling module identifies that the calculation instruction is a general operation instruction, the general operation instruction is sent to the calculation unit; and the calculation unit performs general calculation according to the general operation instruction.

[0023] Preferably, the computing instruction includes an operation code and an operation field, wherein the operation code includes the instruction type and operation code of the computing instruction, and the operation field includes the source address and target address of the data to be operated.

[0024] Preferably, the instruction types include tensor operations, vector operations and scalar operations.

[0025] Preferably, when the instruction type in the computing instruction is a tensor operation, the instruction scheduling module identifies the computing instruction as a tensor operation instruction; when the instruction type in the computing instruction is a vector operation or a scalar operation, the instruction scheduling module identifies the computing instruction as a general operation instruction.

[0026] Preferably, the tensor processor performs in-memory calculations and / or non-memory calculations according to the tensor operation instructions, including: parsing the tensor operation instructions to obtain the operation code and the source address and target address of the data to be operated; obtaining the data to be operated according to the source address of the data to be operated, and dividing the data to be operated into multiple groups of operation data; generating a selection signal according to the source address of the data to be operated, and performing in-memory calculations and / or non-memory calculations according to the selection signal, the operation code, multiple groups of data and the target address.

[0027] According to a third aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and wherein the computer program implements the above-described method when executed by a processor.

[0028] The computing system, method executed by the computing system, and storage medium provided by the present invention combine a general-purpose processor and a tensor processor, using the general-purpose processor to handle general-purpose calculations in neural networks and using the tensor processor to perform in-memory or non-in-memory calculations, which can not only ensure versatility but also improve computing power and support complex and changeable computing operators in neural networks.

[0029] Furthermore, the computing instructions use a unified instruction format, and one instruction can be used to drive different computing modules in the tensor processor to perform in-memory or non-memory computing, greatly reducing programming complexity. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] The above and other objects, features and advantages of the present invention will become more apparent through the following description of the embodiments of the present invention with reference to the accompanying drawings, in which:

[0031] FIG1 shows a block diagram of a computing system according to an embodiment of the present invention;

[0032] FIG2 is a schematic diagram showing the structure of an instruction decoding module in a tensor processor provided by an embodiment of the present invention;

[0033] FIG3 shows a block diagram of a computing system according to another embodiment of the present invention;

[0034] FIG4 shows a block diagram of a computing system according to another embodiment of the present invention;

[0035] FIG5 is a flowchart of a calculation method provided by an embodiment of the present invention;

[0036] FIG6 shows a flowchart of step S530 provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0037] Various embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. In each of the accompanying drawings, identical elements are represented by identical or similar reference numerals. For the sake of clarity, the various parts in the accompanying drawings are not drawn to scale.

[0038] The specific implementation of the present invention is further described in detail below with reference to the accompanying drawings and examples.

[0039] As mentioned above, general-purpose CPUs and GPUs have advantages in data-intensive general-purpose computing related to AI. However, general-purpose computing struggles to cope with the complex and ever-changing nature of computational operators. In-memory computing, which typically uses an AISC (application-specific integrated circuit) model, supports a very limited range of operators and offers limited programming flexibility and versatility, making it incapable of flexibly matching the current pace of neural network model development.

[0040] In response to the above technical problems, the basic idea of ​​this application is to combine a general-purpose processor and a tensor processor, describe the instruction type in the computing instruction to distinguish between tensor operation instructions and general-purpose operation instructions, send the tensor operation instructions to the tensor processor to perform in-memory calculations or non-memory calculations, and send the general-purpose operation instructions to the general-purpose processor to perform general-purpose calculations. This can not only ensure versatility but also improve computing power, and support complex and changeable computing operators in neural networks.

[0041] FIG1 is a schematic diagram showing the structure of a computing system according to an embodiment of the present application. As shown in FIG1 , the computing system 100 includes a general-purpose processor 110 and a tensor processor 120 .

[0042] The general purpose processor 110 includes an instruction scheduling module 111 and at least one computing unit CU.

[0043] The instruction scheduling module 111 is used to identify a computing instruction, and when the computing instruction is identified as a tensor operation instruction, send the tensor operation instruction to the tensor processor 120.

[0044] In this embodiment, the computing instruction includes an operation code and an operation field, wherein the operation code includes the instruction type and operation code of the computing instruction, and the operation field includes at least the source address and the target address of the data to be operated.

[0045] Specifically, the instruction type is used to describe the type of operation involved in the computing instruction, including tensor operations, vector operations, scalar operations, and transcendental function operations. The opcode is used to describe the operation to be performed by the computing instruction (for example, addition, subtraction, multiplication, division, special functions, etc.), which specifically describes the nature and function of the operation.

[0046] The operation domain includes the source address and target address of the data to be operated on. The source address and target address can be memory addresses or register addresses (i.e., register numbers). The memory or register can be an off-chip memory. Of course, in practical applications, it can also be an on-chip memory, which is used to store data. The data can specifically be a single data value (scalar) or n-dimensional data (vector or tensor), where n is an integer greater than or equal to 1. For example, when n=1, it is a one-dimensional data, i.e., a vector; when n=2, it is a two-dimensional data, i.e., a matrix; when n=3 or more, it is a multi-dimensional tensor.

[0047] In this embodiment, the opcode can be the part of the instruction or field (usually represented by a code) specified in the computer program to perform the operation, which is an instruction sequence number used to inform the device executing the instruction which specific instruction needs to be executed. The operation domain can be the source of all data required to execute the corresponding instruction, such as the corresponding address, etc. All data required to execute the corresponding instruction include the data to be calculated and the corresponding instruction processing method, etc. For a calculation instruction, it must include an opcode and an operation domain, wherein the operation domain includes at least the source address and the target address of the data to be calculated. It should be understood that those skilled in the art can set the instruction format of the calculation instruction and the opcode and operation domain contained therein as needed, and the present disclosure does not limit this.

[0048] In a preferred embodiment, the source address of the data to be operated can be the starting address of the storage space where the data to be operated is located. The general processor 110 or the tensor processor 120 can obtain instructions and data through a data input and output unit, which can be one or more data I / O interfaces or I / O pins. Furthermore, the general processor 110 or the tensor processor 120 can determine the data to be operated based on the source address of the data to be operated and obtain the data to be operated. Of course, in other embodiments, the general processor 110 or the tensor processor 120 can also determine the data required for the operation based on the operation code of the calculation instruction.

[0049] In one possible implementation, the instruction format of the calculation instruction may be as shown in the following table:

[0050] Table 1:

[0051]

[0052] Wherein, type, opcode is the operation code of the computing instruction, dst, scr is the operation domain of the computing instruction, and dst is the target address. src is the source address of the data to be operated on. When there are multiple data to be operated on, src can include multiple data to be operated on src0, src1, ..., srcn, and dst can also include multiple target addresses dst0, dst1, ..., dstn. This disclosure does not limit this.

[0053] In some specific embodiments, the scalar operations and vector operations in the general operation instructions may be shown in the following table:

[0054] Table 2:

[0055] When executing the operation instructions shown in the above table, the general processor 110 reads Num1 floating-point data or integer data from the address specified in the register Reg1, reads Num2 floating-point data or integer data from the address specified in the register Reg2, then performs the multiplication and addition operation, and then stores the calculation result into the address space specified in the register Reg3.

[0056] In some specific embodiments, the transcendental function operation in the general operation instruction may be as shown in the following table:

[0057] Table 3:

[0058] When executing the operation instructions shown in the above table, the general processor 110 reads Num1 data from the address specified in the register Reg1, reads Num2 data from the address specified in the register Reg2, then performs the function operation, and then stores the calculation result in the address space specified in the register Reg3.

[0059] In some specific embodiments, the tensor operation instruction may be as shown in the following table:

[0060] Table 4:

[0061] When executing the operation instructions shown in the above table, the tensor processor 120 reads Num1 tensor data from the address specified in register Reg1, reads Num2 weight data from the address specified in register Reg2, and then performs a convolution operation or a multiplication and addition operation, and then stores the calculation result in the address space specified in register Reg3.

[0062] In a preferred embodiment, if there is no information on the number of data that needs to be read in the operation domain of the operation instruction, it means that the data that needs to be read is a fixed number of data, which can be one data, a row of data, or a column of data, etc., and the embodiment of this application does not make specific restrictions.

[0063] In this embodiment, when the instruction type in the computing instruction is a tensor operation (for example, the instruction type is Tensor), the instruction scheduling module 111 identifies the computing instruction as a tensor operation instruction; when the instruction type in the computing instruction is one of a vector operation, a scalar operation, and a transcendental function operation (for example, the instruction type is Float, INT, SUF), the instruction scheduling module 111 identifies the computing instruction as a general operation instruction. The tensor operation instruction includes at least one of a matrix multiplication operation (MM), a matrix multiply-add operation (MAC), and a convolution operation (Conv); the general operation instruction includes at least one of a floating-point multiplication-add operation, an integer multiplication-add operation, and a transcendental function operation.

[0064] The instruction scheduling module 111 is further configured to send the general operation instruction to the computing unit CU when identifying that the computing instruction is a general operation instruction; the computing unit CU is configured to perform general calculations according to the general operation instruction.

[0065] FIG1 specifically shows two computing units (CUs) as an example, while omitting other possible computing units. Each computing unit (CU) includes an instruction dispatch module, multiple computing cores (Kernels), a register file, a shared L1 cache, and the like. The instruction scheduling module 111 is also used to schedule the execution of computing tasks among multiple computing units (CUs). The general-purpose processor in this embodiment is any one of a CPU, a GPU, a DSP, and a GPGPU.

[0066] The computing system can be used for computing tasks such as matrix calculations, which can be executed in parallel by multiple threads. For example, before execution, these threads are divided into multiple thread blocks in the instruction scheduling module 111, and then these thread blocks are distributed to each computing unit CU (for example, a streaming multiprocessor (SM)). All threads in a thread block are usually assigned to the same computing unit for execution. At the same time, the thread block is split into thread bundles (or simply thread warps), for example, each thread bundle contains a fixed number (or less than this fixed number) of threads, for example, 32 threads. Multiple thread blocks can be executed in the same computing unit, or in different computing units.

[0067] In each computing unit, the instruction dispatch module 112 schedules and allocates thread warps so that the multiple computing cores of the computing unit CU run the corresponding thread warps. Each computing core includes an arithmetic logic unit (ALU), a floating-point computing unit, etc. Depending on the number of computing cores in the computing unit, multiple thread warps in a thread block can be executed simultaneously or in a time-sharing manner. Multiple threads in each thread warp will execute the same instruction, and the results obtained after the instruction execution are updated to the registers corresponding to each thread warp. The instructions and data corresponding to each computing unit CU are sent to the shared cache (e.g., shared L1 cache) in the computing unit or further sent to the unified cache for read and write operations.

[0068] The tensor processor 120 is configured to perform in-memory calculations and / or non-memory calculations according to the tensor operation instructions.

[0069] In this embodiment, the tensor processor 120 includes an instruction decoding module 121 , a first computing module 122 , and a second computing module 123 .

[0070] Among them, the instruction decoding module 121 is used to parse the tensor operation instruction, generate a selection signal according to the source address of the data to be operated and obtain the data to be operated, divide the data to be operated into multiple groups of operation data, and send the operation code and multiple groups of operation data to the first computing module 122 or the second computing module 123 according to the selection signal.

[0071] In this embodiment, the instruction decoding module 121 obtains the data to be operated and generates a selection signal based on the source address of the data to be operated described in the tensor operation instruction. The data to be operated includes activation data and weight data. If the source address of the weight data is a storage-computation integrated unit, the weight data is static data, and the instruction decoding module 121 sends the operation code and multiple sets of operation data to the first computing module 122 based on the selection signal. If the source address of the weight data is a memory, the weight data is dynamic data, and the instruction decoding module 121 sends the operation code and multiple sets of operation data to the second computing module 123 based on the selection signal. In other embodiments, the instruction decoding module 121 can also generate a selection signal based on the operation code.

[0072] Referring to Figure 2, the instruction decoding module 121 includes an instruction parsing unit 1211, a data grouping unit 1212 and a control unit 1213, wherein the instruction parsing unit 1211 is used to parse the tensor operation instruction to obtain the operation code and the source address and target address of the data to be operated; the data grouping unit 1212 is used to obtain the data to be operated according to the source address of the data to be operated, and divide the data to be operated into multiple groups of operation data; the control unit 1213 is used to generate a selection signal according to the source address of the data to be operated, and send the operation code, multiple groups of data and target address to the first computing module 122 or the second computing module 123 according to the selection signal.

[0073] In this embodiment, the operation domain may also include an execution amount. The instruction decoding module 121 is further configured to obtain the execution amount and divide the data to be operated into multiple groups of operation data based on the execution amount. The execution amount is the amount of data that can be processed by the first calculation module 122 or the second calculation module 123 at one time.

[0074] In a preferred embodiment, when the operation domain does not include the execution amount, the data to be operated can be divided into multiple groups of operation data according to a preset default execution amount.

[0075] The first calculation module 122 is used to perform in-memory calculations according to the received operation code, multiple sets of operation data and target addresses.

[0076] In this embodiment, the first computing module 122 can perform computation in memory (CIM) and is composed of an SRAM, ReRAM, or other storage medium integrated computing unit.

[0077] The second calculation module 123 is used to perform non-memory calculations according to the received operation code, multiple sets of operation data and target addresses.

[0078] In this embodiment, the non-in-memory calculation includes near-memory calculation or other calculation, such as general matrix multiplication (GEMM).

[0079] In a preferred embodiment, the first computing module 122 and the second computing module 123 are connected to the general-purpose processor 110 using a unified interface, which can be one of a PCIE interface and a UCIE interface. In this embodiment, the input data and output data of the first computing module 122 and the second computing module 123 have the same data structure. A single instruction can be used to drive the first computing module 122 and the second computing module 123 to perform calculations, and an instruction can also be used to switch between using the first computing module 122 and the second computing module 123 for calculations.

[0080] In this embodiment, the general-purpose processor 110 and the tensor processor 120 can be packaged in the same core. In a preferred embodiment, referring to FIG3 , the general-purpose processor 110 and the tensor processor 120 can also be packaged as different cores and integrated into the same chip or deployed on different chips. The positional relationship between the general-purpose processor 110 and the tensor processor 120 can be set according to actual applications and is not limited thereto.

[0081] The computing system provided by the present invention combines a general-purpose processor and a tensor processor together, using the general-purpose processor to handle general-purpose calculations in neural networks and using the tensor processor to perform in-memory or non-in-memory calculations, which can not only ensure versatility but also improve computing power and support complex and changeable computing operators in neural networks.

[0082] Furthermore, the computing instructions use a unified instruction format, and one instruction can be used to drive different computing modules in the tensor processor to perform in-memory or non-memory computing, greatly reducing programming complexity.

[0083] FIG4 is a schematic diagram showing the structure of a computing system provided by another embodiment of the present invention. As shown in FIG4 , the computing system includes a task distribution module 310 and a plurality of computing devices 320 .

[0084] This embodiment is described by taking two computing devices 320A and 320B as an example, but is not limited thereto.

[0085] In this embodiment, the task distribution module 310 is used to distribute computing instructions to multiple computing devices 320. The computing devices 320 receive the computing instructions and perform corresponding calculations according to the computing instructions. The computing devices 320A and 320B include a general-purpose processor 321 and a tensor processor 322. The general-purpose processor 321 and the tensor processor 322 are the same as those described in the above embodiment and will not be repeated here.

[0086] In other embodiments, the computing devices 320A and 320B only include the general-purpose processor 321 , the tensor processor 322 is located outside the computing devices, and multiple computing devices 320 share the same tensor processor.

[0087] Figure 5 shows a flow chart of a calculation method provided by an embodiment of the present invention. Referring to Figure 5 , the calculation method is executed by the computing system 100 provided by the above embodiment and includes the following steps.

[0088] In step S510 , the instruction dispatch module of the general purpose processor identifies a computing instruction.

[0089] In this embodiment, the computing instruction includes an operation code and an operation field, wherein the operation code includes the instruction type and operation code of the computing instruction, and the operation field includes at least the source address and the target address of the data to be operated.

[0090] Specifically, the instruction type is used to describe the type of operation involved in the computing instruction, including tensor operations, vector operations, scalar operations, and transcendental function operations. The opcode is used to describe the operation to be performed by the computing instruction (for example, addition, subtraction, multiplication, division, special functions, etc.), which specifically describes the nature and function of the operation.

[0091] The operation domain includes the source address and target address of the data to be operated on. The source address and target address can be memory addresses or register addresses (i.e., register numbers). The memory or register can be an off-chip memory. Of course, in practical applications, it can also be an on-chip memory, which is used to store data. The data can specifically be a single data value (scalar) or n-dimensional data (vector or tensor), where n is an integer greater than or equal to 1. For example, when n=1, it is a one-dimensional data, i.e., a vector; when n=2, it is a two-dimensional data, i.e., a matrix; when n=3 or more, it is a multi-dimensional tensor.

[0092] In this embodiment, when the instruction type in the computing instruction is a tensor operation (for example, the instruction type is Tensor), the instruction scheduling module 111 identifies the computing instruction as a tensor operation instruction; when the instruction type in the computing instruction is one of a vector operation, a scalar operation, and a transcendental function operation (for example, the instruction type is Float, INT, SUF), the instruction scheduling module 111 identifies the computing instruction as a general operation instruction. The tensor operation instruction includes at least one of a matrix multiplication operation (MM), a matrix multiply-add operation (MAC), and a convolution operation (Conv); the general operation instruction includes at least one of a floating-point multiplication-add operation, an integer multiplication-add operation, and a transcendental function operation.

[0093] In step S520, when the instruction scheduling module identifies that the computing instruction is a tensor operation instruction, the tensor operation instruction is sent to the tensor processor.

[0094] In step S530, the tensor processor performs in-memory calculations and / or non-in-memory calculations according to the tensor operation instructions.

[0095] In this embodiment, referring to FIG. 6 , step S530 includes steps S531 to S533 .

[0096] In step S531 , the tensor operation instruction is parsed to obtain an operation code and a source address and a target address of data to be operated.

[0097] In step S532, the data to be operated is obtained according to the source address of the data to be operated, and the data to be operated is divided into multiple groups of operation data.

[0098] In this embodiment, the operation domain may also include an execution amount. The instruction decoding module 121 is further configured to obtain the execution amount and divide the data to be operated into multiple groups of operation data based on the execution amount. The execution amount is the amount of data that can be processed by the first calculation module 122 or the second calculation module 123 at one time.

[0099] In a preferred embodiment, when the operation domain does not include the execution amount, the data to be operated can be divided into multiple groups of operation data according to a preset default execution amount.

[0100] In step S533, a selection signal is generated according to the source address of the data to be operated, and in-memory calculation and / or non-in-memory calculation is performed according to the selection signal, the operation code, multiple groups of data and the target address.

[0101] In this embodiment, the data to be calculated includes activation data and weight data; if the source address of the weight data is a storage-computation integrated unit, the weight data is static data, and the instruction decoding module 121 sends the operation code and multiple sets of operation data to the first calculation module 122 according to the selection signal; if the source address of the weight data is a memory, the weight data is dynamic data, and the instruction decoding module 121 sends the operation code and multiple sets of operation data to the second calculation module 123 according to the selection signal. In other embodiments, the instruction decoding module 121 can also generate a selection signal based on the operation code.

[0102] In step S540 , when the instruction scheduling module identifies that the computing instruction is a general operation instruction, the general operation instruction is sent to the computing unit.

[0103] In step S550 , the computing unit performs general calculations according to the general operation instructions.

[0104] The computing method provided by the present invention combines a general-purpose processor and a tensor processor, uses the general-purpose processor to handle general-purpose calculations in neural networks, and uses the tensor processor to perform in-memory or non-in-memory calculations, which can not only ensure versatility but also improve computing power and support complex and changeable computing operators in neural networks.

[0105] Furthermore, the computing instructions use a unified instruction format, and one instruction can be used to drive different computing modules in the tensor processor to perform in-memory or non-memory computing, greatly reducing programming complexity.

[0106] An embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the above-mentioned various method embodiments can be implemented.

[0107] An embodiment of the present application provides a computer program product. When the computer program product is run on an electronic device, the electronic device can implement the steps in the above-mentioned method embodiments when executing the computer program product.

[0108] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, it can implement the steps of each of the above-mentioned method embodiments. The computer program includes computer program code, which can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium can at least include: any entity or device that can carry computer program code to a device / electronic device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electric carrier signal, a telecommunication signal and a software distribution medium. For example, a USB flash drive, a mobile hard drive, a magnetic disk or an optical disk. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electric carrier signals and telecommunication signals.

[0109] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.

[0110] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0111] In the embodiments provided in this application, it should be understood that the disclosed devices / electronic devices and methods can be implemented in other ways. For example, the device / electronic device embodiments described above are merely schematic. For example, the block division of the modules or units is merely a logical function block division. In actual implementation, there may be other block division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0112] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0113] While embodiments of the present invention have been described above, these embodiments do not exhaustively describe all details and do not limit the invention to the specific embodiments described. Obviously, many modifications and variations are possible based on the above description. These embodiments are selected and described in detail in this specification in order to better explain the principles and practical applications of the present invention, thereby enabling those skilled in the art to better utilize the present invention and its modifications. The present invention is limited only by the claims and their full scope and equivalents.

Claims

1. A computing system, characterized in that: Includes general purpose processors and tensor processors; The general purpose processor includes an instruction scheduling module and at least one computing unit. The instruction scheduling module is used to identify the computing instruction, and when the computing instruction is identified as a tensor operation instruction, send the tensor operation instruction to the tensor processor; A tensor processor is used to perform in-memory calculations and / or non-memory calculations according to the tensor operation instructions.

2. The computing system according to claim 1, characterized in that The instruction scheduling module is further configured to send the general operation instruction to the computing unit when identifying that the computing instruction is a general operation instruction; The computing unit is used to perform general computing according to general computing instructions.

3. The computing system according to claim 2, characterized in that The computing unit includes at least one computing core, and the computing core is used to perform general computing according to general operation instructions.

4. The computing system according to claim 1, characterized in that: The computing instruction includes an operation code and an operation field, wherein the operation code includes an instruction type and an operation code of the computing instruction, and the operation field includes a source address and a target address of data to be operated.

5. The computing system according to claim 4, characterized in that: The instruction types include tensor operations, vector operations, scalar operations, and transcendental function operations.

6. The computing system according to claim 5, characterized in that: When the instruction type in the computing instruction is a tensor operation, the instruction scheduling module identifies the computing instruction as a tensor operation instruction; when the instruction type in the computing instruction is one of a vector operation, a scalar operation and a transcendental function operation, the instruction scheduling module identifies the computing instruction as a general operation instruction.

7. The computing system according to claim 1, characterized in that: The tensor operation instructions include at least one of matrix multiplication operation, matrix multiplication and addition operation and convolution operation; the general operation instructions include at least one of floating-point multiplication and addition operation, integer multiplication and addition operation and transcendental function operation.

8. The computing system according to claim 4, characterized in that: The tensor processor includes an instruction decoding module, a first computing module and a second computing module. The instruction decoding module is used to parse the tensor operation instruction, generate a selection signal according to the source address of the data to be operated and obtain the data to be operated, divide the data to be operated into multiple groups of operation data, and send the operation code and the multiple groups of operation data to the first computing module or the second computing module according to the selection signal; The first calculation module is used to perform in-memory calculation according to the received operation code, multiple groups of operation data and target address; The second calculation module is used for performing non-memory calculation according to the received operation code, multiple groups of operation data and target address.

9. The computing system according to claim 8, characterized in that: The instruction decoding module comprises: An instruction parsing unit, used for parsing the tensor operation instruction to obtain an operation code and a source address and a target address of data to be operated; A data grouping unit, configured to obtain the data to be operated according to the source address of the data to be operated, and divide the data to be operated into a plurality of groups of operation data; The control unit is used to generate a selection signal according to the source address of the data to be operated, and send the operation code, multiple groups of data and the target address to the first calculation module or the second calculation module according to the selection signal.

10. The computing system according to claim 8, characterized in that: The first computing module and the second computing module are connected to a general processor using a unified interface.

11. The computing system according to claim 10, characterized in that: The unified interface includes one of a PCIE interface and a UCIE interface.

12. The computing system according to claim 1, characterized in that The general purpose processor and the tensor processor are packaged in the same core.

13. The computing system according to claim 1, characterized in that: The general processor and the tensor processor are packaged into different core particles and integrated into the same chip or deployed on different chips.

14. The computing system according to claim 1, characterized in that: The general-purpose processor is any one of a CPU, a GPU, a DSP, and a GPGPU.

15. A computing method performed by a computing system, characterized in that: The computing system includes a general-purpose processor and a tensor processor, the general-purpose processor includes an instruction scheduling module and at least one computing unit, and the computing method includes: The instruction dispatch module of the general purpose processor identifies the computational instructions; When the instruction scheduling module identifies that the computing instruction is a tensor operation instruction, the tensor operation instruction is sent to the tensor processor; The tensor processor performs in-memory calculations and / or non-in-memory calculations according to the tensor operation instructions.

16. The calculation method according to claim 15, characterized in that: Also includes: When the instruction scheduling module identifies that the computing instruction is a general computing instruction, the general computing instruction is sent to the computing unit; The computing unit performs general calculations according to general operation instructions.

17. The calculation method according to claim 15, characterized in that: The computing instruction includes an operation code and an operation field, wherein the operation code includes an instruction type and an operation code of the computing instruction, and the operation field includes a source address and a target address of data to be operated.

18. The calculation method according to claim 17, characterized in that: The instruction types include tensor operations, vector operations, and scalar operations.

19. The calculation method according to claim 18, characterized in that: When the instruction type in the computing instruction is a tensor operation, the instruction scheduling module identifies the computing instruction as a tensor operation instruction; when the instruction type in the computing instruction is a vector operation or a scalar operation, the instruction scheduling module identifies the computing instruction as a general operation instruction.

20. The calculation method according to claim 17, characterized in that: The tensor processor performs in-memory calculation and / or non-memory calculation according to the tensor operation instruction, including: Parsing the tensor operation instruction to obtain an operation code and a source address and a target address of data to be operated; Acquire the data to be operated according to the source address of the data to be operated, and divide the data to be operated into multiple groups of operation data; A selection signal is generated according to the source address of the data to be operated, and in-memory calculation and / or non-in-memory calculation is performed according to the selection signal, the operation code, multiple groups of data and the target address.

21. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 15 to 20 is implemented.

Citation Information

Patent Citations

  • Instruction processing method and device and related product

    CN111966401A

  • Tensor, vector and scalar calculation acceleration and data scheduling system

    CN115169541A

  • Neural network acceleration device and method, chip, electronic equipment and storage medium

    CN115860079A

  • Computing system, method executed by computing system and storage medium

    CN119088751A

  • Quantization for DNN accelerators

    US20190340499A1

Cited By

  • Instruction scheduling system and instruction scheduling method

    CN121209964A