Computing device, integrated circuit chip, board, electronic device and computing method
By introducing a hardware architecture that supports multi-stage flow computing and efficient processing of tensor data in the computing chip, the problem of instruction set flexibility and tensor computing efficiency in the prior art is solved, and more efficient computing performance and lower power consumption are achieved.
Patent Information
- Application Number
- CN202010619425.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-06-30
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2040-06-30
AI Technical Summary
The instruction sets of existing computing chips are insufficient in terms of flexibility, execution speed, execution efficiency and power consumption, especially when dealing with multidimensional tensor operations.
It provides a hardware architecture, including one or more sets of flow operation circuits, supporting multi-stage flow operation and efficient memory access and processing of tensor data. By analyzing the calculation instructions, descriptors are used to determine the storage address of operands, and efficient execution of multi-stage flow operations and tensor operations are realized.
Improves computing performance, reduces power consumption, improves the execution efficiency of computing operations, and reduces the computational overhead of tensor operations.
Smart Images

Figure CN113867799B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to the field of computing. More specifically, the present disclosure relates to a computing device, an integrated circuit chip, a board, an electronic device, and a computing method. Background Art
[0002] In a computing system, an instruction set is a set of instructions for performing calculations and controlling the computing system, and plays a key role in improving the performance of computing chips (such as processors) in computing systems. Current computing chips (especially chips in the field of artificial intelligence) use associated instruction sets to complete various general or specific control operations and data processing operations. However, the current instruction set still has many defects. For example, the existing instruction set is limited by the hardware architecture and performs poorly in flexibility. Furthermore, many instructions can only complete a single operation, while the execution of multiple operations usually requires multiple instructions, which potentially leads to an increase in the on-chip I / O data throughput. In addition, the current instructions still have room for improvement in terms of execution speed, execution efficiency, and power consumption caused to the chip.
[0003] In addition, the operation instructions of traditional processor CPUs are designed to perform basic single-data scalar operations. Here, single-data operations refer to each operand of the instruction being a scalar data. However, in tasks such as image processing and pattern recognition, the operands are often multi-dimensional vectors (i.e., tensor data) data types, and only using scalar operations cannot enable the hardware to efficiently complete the operation tasks. Therefore, how to efficiently perform multi-dimensional tensor operations is also a problem that needs to be solved in the current computing field. Summary of the invention
[0004] In order to at least solve the problems existing in the above-mentioned prior art, the present disclosure provides a hardware architecture having one or more groups of pipeline operation circuits supporting multi-stage pipeline operations. By utilizing this hardware architecture to execute computing instructions, the scheme disclosed in the present disclosure can obtain technical advantages in multiple aspects including enhancing the processing performance of hardware, reducing power consumption, improving the execution efficiency of computing operations, and avoiding computing overhead. Furthermore, the scheme disclosed in the present disclosure supports efficient memory access and processing of tensor data on the basis of the aforementioned hardware architecture, thereby accelerating tensor operations and reducing the computing overhead caused by tensor operations when multi-dimensional vector operands are included in the computing instructions.
[0005] In a first aspect, the present disclosure provides a computing device, comprising: one or more groups of pipeline operation circuits, which are configured to perform multi-stage pipeline operations according to multiple operation instructions obtained after parsing a computing instruction, wherein each group of the pipeline operation circuits constitutes a multi-stage operation pipeline, and the multi-stage operation pipeline includes multiple operation circuits arranged stage by stage, wherein the operand of the computing instruction includes a descriptor for indicating the shape of a tensor, and the descriptor is used to determine the storage address of the data corresponding to the operand,
[0006] In response to receiving a plurality of operation instructions, at least one stage of operation circuit in the multi-stage operation pipeline is configured to execute a corresponding one of the plurality of operation instructions based on the storage address.
[0007] In a second aspect, the present disclosure provides an integrated circuit chip comprising a computing device as described above and in a number of embodiments below.
[0008] In a third aspect, the present disclosure provides a board comprising an integrated circuit chip as described above and in the following multiple embodiments.
[0009] In a fourth aspect, the present disclosure provides an electronic device comprising an integrated circuit chip as described above and in the following multiple embodiments.
[0010] In the fifth aspect, the present disclosure provides a method for performing calculations using the aforementioned computing device, wherein the computing device includes one or more groups of pipeline operation circuits, and the method includes: configuring each of the one or more groups of pipeline operation circuits to perform multi-stage pipeline operations according to multiple operation instructions obtained after parsing the computing instructions, wherein each group of the pipeline operation circuits constitutes a multi-stage operation pipeline, and the multi-stage operation pipeline includes multiple operation circuits arranged stage by stage, wherein the operand of the computing instruction includes a descriptor for indicating the shape of a tensor, and the descriptor is used to determine the storage address of the data corresponding to the operand; and in response to receiving multiple operation instructions, configuring at least one stage of the operation circuit in the multi-stage operation pipeline to execute a corresponding one of the multiple operation instructions based on the storage address.
[0011] By utilizing the computing device, integrated circuit chip, board, electronic device and method disclosed herein, pipeline operations can be efficiently performed, especially various multi-stage pipeline operations in the field of artificial intelligence. Furthermore, the disclosed solution can achieve efficient computing operations with the help of a unique hardware architecture, thereby improving the overall performance of the hardware and reducing computing overhead. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] By reading the detailed description below with reference to the accompanying drawings, the above and other purposes, features and advantages of the exemplary embodiments of the present disclosure will become readily understood. In the accompanying drawings, several embodiments of the present disclosure are shown in an exemplary and non-limiting manner, and the same or corresponding reference numerals represent the same or corresponding parts, wherein:
[0013] Figure 1a is a block diagram illustrating a computing device according to one embodiment of the present disclosure;
[0014] Figure 1b is a schematic diagram showing a data storage space according to an embodiment of the present disclosure;
[0015] Figure 2 is a block diagram illustrating a computing device according to another embodiment of the present disclosure;
[0016] Figure 3a , 3b and 3c are schematic diagrams showing matrix conversion performed by the data conversion circuit according to an embodiment of the present disclosure;
[0017] Figure 4 is a block diagram illustrating a computing system according to an embodiment of the present disclosure;
[0018] Figure 5 is a simplified flow chart illustrating a method of using a computing device to perform a computing operation according to an embodiment of the present disclosure;
[0019] Figure 6 is a structural diagram showing a combined processing device according to an embodiment of the present disclosure; and
[0020] Figure 7 It is a schematic diagram showing the structure of a board card according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0021] The scheme disclosed herein provides a hardware architecture that supports multi-stage pipeline operations. When the hardware architecture is implemented in a computing device, the computing device includes at least one or more groups of pipeline operation circuits, wherein each group of pipeline operation circuits can constitute a multi-stage operation pipeline disclosed herein. In the multi-stage operation pipeline, multiple operation circuits can be arranged step by step. In one embodiment, when the computing device disclosed herein performs a computing operation involving a tensor, the operand of the computing instruction may include a descriptor for indicating the shape of the tensor, and the descriptor is used to determine the storage direct address of the data corresponding to the operand. Based on this, when multiple operation instructions are received, at least one level of the operation circuit in the aforementioned multi-stage operation pipeline can be configured to execute a corresponding one of the multiple operation instructions according to the aforementioned physical address. With the help of the hardware architecture and operation instructions disclosed herein, parallel pipeline operations can be efficiently performed, the application scenarios of the calculation are expanded, and the calculation overhead is reduced.
[0022] The following will be combined with the drawings in the embodiments of the present disclosure to clearly and completely describe the technical solutions in the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present disclosure.
[0023] Figure 1a is a block diagram showing a computing device 100 according to one embodiment of the present disclosure. Figure 1a As shown in , the computing device 100 may include one or more groups of pipeline operation circuits, such as the first group of pipeline operation circuits 102, the second group of pipeline operation circuits 104 and the third group of pipeline operation circuits 106 shown in the figure, wherein each group of pipeline operation circuits can constitute a multi-stage operation pipeline in the context of the present disclosure. Taking the first group of pipeline operation circuits 102 constituting the first multi-stage operation pipeline as an example, it can perform pipeline operations including 1-1 level pipeline operation, 1-2 level pipeline operation, 1-3 level pipeline operation... 1-N level pipeline operation, a total of N levels of pipeline operation. Similarly, the second group and the third group of pipeline operation circuits also have a structure that supports N levels of pipeline operation. Through such an exemplary architecture, those skilled in the art can understand that the multiple groups of pipeline operation circuits disclosed in the present disclosure can constitute multiple multi-stage operation pipelines, and the multiple multi-stage operation pipelines can execute their respective multiple operation instructions in parallel. In one embodiment, the aforementioned operation instructions can be obtained by parsing the calculation instructions. According to the solution of the present disclosure, the operand of the computing instruction may include a descriptor for indicating the shape of a tensor, and the descriptor is used to determine the storage address of the data corresponding to the operand.
[0024] In order to perform the above-mentioned pipeline operations at each level, an operation circuit including one or more operators can be arranged at each level to execute the corresponding operation instructions to implement the operation operation at this level. When the operation instruction involves the operation of tensors, at least one level of the operation circuit in the multi-stage operation pipeline can be configured to execute a corresponding operation instruction based on the storage address. In one embodiment, in response to receiving multiple operation instructions, one or more groups of pipeline operation circuits disclosed in the present disclosure can be configured to perform multi-data operations, such as executing single instruction multiple data ("SIMD") instructions. In one embodiment, the aforementioned multiple operation instructions can be obtained by parsing the calculation instructions received by the computing device 100, and the operation code of the calculation instruction can represent the multiple operations performed by the multi-stage operation pipeline. In another embodiment, the operation code and the multiple operations represented by it are predetermined according to the functions supported by the multiple operation circuits arranged step by step in the multi-stage operation pipeline.
[0025] In the scheme disclosed herein, each group of pipeline operation circuits can be configured to selectively connect according to multiple operation instructions to complete the corresponding multiple operation instructions in addition to performing the step-by-step operation operation in a multi-stage operation pipeline it constitutes. In an implementation scenario, the multiple multi-stage operation pipelines disclosed herein may include a first multi-stage operation pipeline and a second multi-stage operation pipeline, wherein the output end of the operation circuit of one or more stages of the first multi-stage operation pipeline is configured to be connected to the input end of the operation circuit of one or more stages of the second multi-stage operation pipeline according to the operation instruction. For example, the 1st-2nd stage pipeline operation in the first multi-stage operation pipeline shown in the figure can input its operation result into the 2nd-3rd stage pipeline operation in the second multi-stage operation pipeline according to the operation instruction. Similarly, the 2nd-1st stage pipeline operation in the second multi-stage operation pipeline shown in the figure can input its operation result into the 3rd-3rd stage pipeline operation in the third multi-stage operation pipeline according to the operation instruction. In some scenarios, depending on the operation instructions, two-stage pipeline operations in different pipeline operation lines can realize bidirectional transmission of operation results, such as between the 2-2 stage pipeline operation in the second multi-stage operation pipeline and the 3-2 stage pipeline operation in the third multi-stage operation pipeline shown.
[0026] As can be seen from the above, in order to realize the transfer of data between the same operation pipeline and different operation pipelines, each stage of the operation circuit in the multiple groups of operation pipelines disclosed in the present invention may have an input end and an output end, which is used to receive input data at the operation circuit and output the result of the operation of the operation circuit at this stage. In a multi-stage operation pipeline, the output end of the operation circuit of one or more stages is configured to be connected to the input end of the operation circuit of another stage or multiple stages according to the operation instruction to execute the operation instruction. For example, in the first operation pipeline, the result of the 1-1 stage pipeline operation can be input to the 1-3 stage pipeline operation in the operation pipeline according to the operation instruction.
[0027] In the context of the present disclosure, the aforementioned multiple operation instructions may be microinstructions or control signals running inside a computing device (or processing circuit, processor), which may include (or indicate) one or more operations to be performed by the computing device. According to different operation scenarios, the operation operation may include but is not limited to various operations such as addition operation, multiplication operation, convolution operation, pooling operation, etc. In order to realize multi-stage pipeline operation, each level of operation circuit for performing each level of pipeline operation may include but is not limited to one or more operators or circuits in the following: random number processing circuit, addition and subtraction circuit, subtraction circuit, table lookup circuit, parameter configuration circuit, multiplier, pooler, comparator, absolute value circuit, logic operator, position index circuit or filter. Here, taking the pooler as an example, it can be exemplarily composed of operators such as adders, dividers, comparators, etc., so as to perform pooling operations in neural networks.
[0028] In order to realize multi-stage pipeline operation, the present disclosure can also provide corresponding calculation instructions according to the operations supported by the operation circuit in the multi-stage pipeline operation to realize multi-stage pipeline operation. Depending on the operation scenario, the calculation instruction of the present disclosure may include multiple operation codes, which may represent multiple operations performed by the operation circuit. For example, when N=4 in Figure 1 (i.e., when performing 4-stage pipeline operation), the calculation instruction according to the present disclosure scheme can be expressed as follows (1):
[0029] Result=((((scr0 op0 scr1)op1 src2)op2 src3)op3 src4) (1)
[0030] Among them, scr0~src4 are source operands, and op0~op3 are operation codes. According to different pipeline operation circuit architectures and supported operations, the type, order and number of operation codes of the computing instructions disclosed in the present disclosure may change. In one embodiment, when the computing operation involves the operation of tensors, one of the above-mentioned source operands may include a descriptor for indicating the shape of the tensor, so that the descriptor can be used to determine the storage address of the data corresponding to the operand.
[0031] In some application scenarios, the multi-stage pipeline operation disclosed in the present invention can support unary operations (i.e., a situation where there is only one input data). Taking the operation at the scale layer + relu layer in the neural network as an example, it is assumed that the calculation instruction to be executed is expressed as result = relu (a*ina + b), where ina is the input data (for example, it can be a vector or a matrix), and a and b are both operation constants. For this calculation instruction, a set of three-stage pipeline operation circuits including multipliers, adders, and nonlinear operators disclosed in the present invention can be applied to perform the operation. Specifically, the multiplier of the first-stage pipeline can be used to calculate the product of the input data ina and a to obtain the first-stage pipeline operation result. Then, the adder of the second-stage pipeline can be used to perform an addition operation on the first-stage pipeline operation result (a*ina) and b to obtain the second-stage pipeline operation result. Finally, the relu activation function of the third-stage pipeline can be used to activate the second-stage pipeline operation result (a*ina + b) to obtain the final operation result result.
[0032] In some application scenarios, the multi-stage pipeline operation circuit disclosed in the present invention can support binary operations (for example, convolution calculation instruction result = conv(ina, inb)) or ternary operations (for example, convolution calculation instruction result = conv(ina, inb, bias)), where the input data ina, inb and bias can be either vectors (for example, integer, fixed-point or floating-point data) or matrices. Here, taking the convolution calculation instruction result = conv(ina, inb) as an example, the convolution operation expressed by the calculation instruction can be performed using multiple multipliers, at least one addition tree and at least one nonlinear operator included in the three-stage pipeline operation circuit structure, where the two input data ina and inb can be, for example, neuron data. Specifically, the first-stage pipeline multiplier in the three-stage pipeline operation circuit can be used for calculation first, so that the first-stage pipeline operation result product = ina*inb (regarded as a microinstruction in the operation instruction, which corresponds to the multiplication operation) can be obtained. Then, the addition tree in the second-stage pipeline operation circuit can be used to perform an addition operation on the first-stage pipeline operation result "product" to obtain the second-stage pipeline operation result sum. Finally, the nonlinear operator of the third-stage pipeline operation circuit is used to perform an activation operation on "sum" to obtain the final convolution operation result.
[0033] In some application scenarios, as mentioned above, the disclosed solution can perform bypass operations on one or more stages of pipeline operation circuits that will not be used in the operation, that is, one or more stages of the multi-stage pipeline operation circuit can be selectively used according to the needs of the operation, without requiring the operation to pass through all the multi-stage pipeline operations. Taking the operation of calculating the Euclidean distance as an example, assuming that its calculation instruction is expressed as dis=sum((ina-inb)^2), only a few stages of pipeline operation circuits composed of adders, multipliers, adder trees and accumulators can be used to perform operations to obtain the final operation results, and unused pipeline operation circuits can be bypassed before or during the pipeline operation.
[0034] As mentioned above, the multi-stage pipeline operation disclosed herein also includes using descriptors to obtain information related to the shape of the tensor in order to determine the storage address of the tensor data, thereby obtaining and saving the tensor data through the aforementioned storage address.
[0035] In one possible implementation, a descriptor can be used to indicate the shape of N-dimensional tensor data, where N is a positive integer, such as N = 1, 2 or 3, or zero. Among them, a tensor can contain multiple forms of data composition, and a tensor can be of different dimensions. For example, a scalar can be regarded as a 0-dimensional tensor, a vector can be regarded as a 1-dimensional tensor, and a matrix can be a tensor of 2 dimensions or more. The shape of a tensor includes information such as the dimension of the tensor and the size of each dimension of the tensor. For example, for a tensor:
[0036]
[0037] The shape of the tensor can be described by the descriptor as (2, 4), that is, the two parameters indicate that the tensor is a two-dimensional tensor, and the size of the first dimension (column) of the tensor is 2, and the size of the second dimension (row) is 4. It should be noted that the present application does not limit the way in which the descriptor indicates the shape of the tensor.
[0038] In one possible implementation, the value of N can be determined according to the dimension (order) of the tensor data, or it can be set according to the usage requirements of the tensor data. For example, when the value of N is 3, the tensor data is three-dimensional tensor data, and the descriptor can be used to indicate the shape of the three-dimensional tensor data in three dimensions (e.g., offset, size, etc.). It should be understood that those skilled in the art can set the value of N according to actual needs, and this disclosure does not limit this.
[0039] In a possible implementation, the descriptor may include an identifier of the descriptor and / or content of the descriptor. The identifier of the descriptor is used to distinguish the descriptor, for example, the identifier of the descriptor may be a number for it; the content of the descriptor may include at least one shape parameter representing the shape of the tensor data. For example, the tensor data is 3D data, and among the three dimensions of the tensor data, the shape parameters of two dimensions are fixed, and the content of its descriptor may include the shape parameter representing another dimension of the tensor data.
[0040] In one possible implementation, the identifier and / or content of the descriptor may be stored in a descriptor storage space (internal memory), such as a register, an on-chip SRAM, or other media cache. The tensor data indicated by the descriptor may be stored in a data storage space (internal memory or external memory), such as an on-chip cache or an off-chip memory. The present disclosure does not limit the specific locations of the descriptor storage space and the data storage space.
[0041] In one possible implementation, the identifier, content, and tensor data indicated by the descriptor of the descriptor can be stored in the same area of the internal memory. For example, a continuous area of the on-chip cache can be used to store the relevant content of the descriptor, and its address is ADDR0-ADDR1023. Among them, the address ADDR0-ADDR63 can be used as a descriptor storage space to store the identifier and content of the descriptor, and the address ADDR64-ADDR1023 can be used as a data storage space to store the tensor data indicated by the descriptor. In the descriptor storage space, the identifier of the descriptor can be stored at address ADDR0-ADDR31, and the content of the descriptor can be stored at address ADDR32-ADDR63. It should be understood that the address ADDR is not limited to 1 bit or one byte. It is used here to represent an address and is an address unit. Those skilled in the art can determine the descriptor storage space, data storage space, and their specific addresses according to actual conditions, and this disclosure is not limited to this.
[0042] In one possible implementation, the identifier and content of the descriptor and the tensor data indicated by the descriptor can be stored in different areas of the internal memory. For example, a register can be used as a descriptor storage space to store the identifier and content of the descriptor, and an on-chip cache can be used as a data storage space to store the tensor data indicated by the descriptor.
[0043] In one possible implementation, when a register is used to store the identifier and content of a descriptor, the identifier of the descriptor may be represented by the register number. For example, when the register number is 0, the identifier of the descriptor stored therein is set to 0. When the descriptor in the register is valid, an area may be allocated in the cache space for storing the tensor data according to the size of the tensor data indicated by the descriptor.
[0044] In one possible implementation, the identifier and content of the descriptor may be stored in an internal memory, and the tensor data indicated by the descriptor may be stored in an external memory. For example, the identifier and content of the descriptor may be stored on-chip, and the tensor data indicated by the descriptor may be stored off-chip.
[0045] In one possible implementation, the data address of the data storage space corresponding to each descriptor may be a fixed address. For example, a separate data storage space may be divided for tensor data, and the starting address of each tensor data in the data storage space corresponds to the descriptor one by one. In this case, the control circuit can determine the data address of the data corresponding to the operand in the data storage space according to the descriptor.
[0046] In one possible implementation, when the data address of the data storage space corresponding to the descriptor is a variable address, the descriptor can also be used to indicate the address of N-dimensional tensor data, wherein the content of the descriptor may also include at least one address parameter representing the address of the tensor data. For example, the tensor data is 3D data, and when the descriptor points to the address of the tensor data, the content of the descriptor may include an address parameter representing the address of the tensor data, such as the starting physical address of the tensor data, or may include multiple address parameters of the address of the tensor data, such as the starting address + address offset of the tensor data, or address parameters of the tensor data based on each dimension. Those skilled in the art can set the address parameters according to actual needs, and this disclosure is not limited to this.
[0047] In a possible implementation, the address parameter of the tensor data may include a reference address of the data reference point of the descriptor in the data storage space of the tensor data. The reference address may be different according to the change of the data reference point. The present disclosure does not limit the selection of the data reference point.
[0048] In a possible implementation, the reference address may include the starting address of the data storage space. When the data reference point of the descriptor is the first data block of the data storage space, the reference address of the descriptor is the starting address of the data storage space. When the data reference point of the descriptor is other data other than the first data block in the data storage space, the reference address of the descriptor is the address of the data block in the data storage space.
[0049] In a possible implementation, the shape parameters of the tensor data include at least one of the following: the size of the data storage space in at least one direction of the N dimensional directions, the size of the storage area in at least one direction of the N dimensional directions, the offset of the storage area in at least one direction of the N dimensional directions, the positions of at least two vertices at diagonal positions in the N dimensional directions relative to the data reference point, and the mapping relationship between the data description position of the tensor data indicated by the descriptor and the data address. The data description position is the mapping position of the point or area in the tensor data indicated by the descriptor. For example, when the tensor data is 3D data, the descriptor can use three-dimensional space coordinates (x, y, z) to represent the shape of the tensor data, and the data description position of the tensor data can be the position of the point or area mapped in the three-dimensional space represented by the three-dimensional space coordinates (x, y, z).
[0050] It should be understood that those skilled in the art can select shape parameters representing tensor data according to actual conditions, and this disclosure does not limit this. By using descriptors in the data access process, associations between data can be established, thereby reducing the complexity of data access and improving instruction processing efficiency.
[0051] In one possible implementation, the content of the descriptor of the tensor data can be determined based on the reference address of the data reference point of the descriptor in the data storage space of the tensor data, the size of the data storage space in at least one of the N dimensional directions, the size of the storage area in at least one of the N dimensional directions, and / or the offset of the storage area in at least one of the N dimensional directions.
[0052] Figure 1b FIG. 2 is a schematic diagram showing a data storage space according to an embodiment of the present disclosure. Figure 1b As shown, the data storage space 21 stores two-dimensional data in a row-first manner, which can be represented by (x, y) (where the X axis is horizontal to the right and the Y axis is vertically downward), the size in the X axis direction (the size of each row) is ori_x (not shown in the figure), the size in the Y axis direction (the total number of rows) is ori_y (not shown in the figure), and the starting address PA_start (reference address) of the data storage space 21 is the physical address of the first data block 22. The data block 23 is part of the data in the data storage space 21, and its offset 25 in the X axis direction is represented by offset_x, the offset 24 in the Y axis direction is represented by offset_y, the size in the X axis direction is represented by size_x, and the size in the Y axis direction is represented by size_y.
[0053] In a possible implementation, when a descriptor is used to define the data block 23, the data reference point of the descriptor can use the first data block of the data storage space 21, and the reference address of the descriptor can be agreed to be the starting address PA_start of the data storage space 21. Then, the content of the descriptor of the data block 23 can be determined by combining the size ori_x of the data storage space 21 on the X axis, the size ori_y on the Y axis, the offset offset_y of the data block 23 in the Y axis direction, the offset offset_x in the X axis direction, the size size_x in the X axis direction, and the size size_y in the Y axis direction.
[0054] In a possible implementation, the following formula (2) may be used to represent the content of the descriptor:
[0055]
[0056] It should be understood that although in the above examples, the content of the descriptor represents a two-dimensional space, those skilled in the art may set the specific dimension represented by the content of the descriptor according to actual conditions, and the present disclosure does not limit this.
[0057] In one possible implementation, a reference address of the data reference point of the descriptor in the data storage space can be agreed upon, and based on the reference address, the content of the descriptor of the tensor data is determined according to the positions of at least two vertices at diagonal positions in N dimensional directions relative to the data reference point.
[0058] For example, the reference address PA_base of the data reference point of the descriptor in the data storage space can be agreed upon. For example, a data (e.g., data at position (2, 2)) can be selected in the data storage space 21 as the data reference point, and the physical address of the data in the data storage space is used as the reference address PA_base. The position of the two vertices at the diagonal position relative to the data reference point can be used to determine the reference address PA_base. Figure 1b The content of the descriptor of the data block 23 in the data block 23. First, the positions of at least two vertices at the diagonal positions of the data block 23 relative to the data reference point are determined, for example, the positions of the diagonal vertices from the upper left to the lower right direction relative to the data reference point are used, wherein the relative position of the upper left vertex is (x_min, y_min), and the relative position of the lower right vertex is (x_max, y_max), and then the content of the descriptor of the data block 23 can be determined according to the reference address PA_base, the relative position of the upper left vertex (x_min, y_min), and the relative position of the lower right vertex (x_max, y_max).
[0059] In a possible implementation, the following formula (3) can be used to represent the content of the descriptor (the base address is PA_base):
[0060]
[0061] It should be understood that although the above example uses the vertices at the upper left corner and the lower right corner to determine the content of the descriptor, those skilled in the art can set the specific vertices of at least two vertices at the diagonal positions according to actual needs, and the present disclosure does not limit this.
[0062] In a possible implementation, the content of the descriptor of the tensor data can be determined according to the reference address of the data reference point of the descriptor in the data storage space and the mapping relationship between the data description position and the data address of the tensor data indicated by the descriptor. The mapping relationship between the data description position and the data address can be set according to actual needs. For example, when the tensor data indicated by the descriptor is three-dimensional space data, the function f(x, y, z) can be used to define the mapping relationship between the data description position and the data address.
[0063] In a possible implementation, the following formula (4) may be used to represent the content of the descriptor:
[0064]
[0065] In a possible implementation, the descriptor is also used to indicate the address of the N-dimensional tensor data, wherein the content of the descriptor also includes at least one address parameter representing the address of the tensor data, for example, the content of the descriptor may be:
[0066]
[0067] Wherein PA is an address parameter. The address parameter can be a logical address or a physical address. The descriptor parsing circuit can use PA as any vertex, middle point or preset point of the vector shape, and combine the shape parameters in the X direction and the Y direction to obtain the corresponding data address.
[0068] In a possible implementation, the address parameter of the tensor data includes a reference address of a data reference point of the descriptor in a data storage space of the tensor data, and the reference address includes a starting address of the data storage space.
[0069] In a possible implementation, the descriptor may further include at least one address parameter representing the address of the tensor data. For example, the content of the descriptor may be:
[0070]
[0071] Among them, PA_start is the reference address parameter and will not be repeated here.
[0072] It should be understood that those skilled in the art can set the mapping relationship between the data description location and the data address according to actual conditions, and this disclosure does not limit this.
[0073] In a possible implementation, an agreed reference address can be set in a task, and the descriptors in the instructions under this task all use this reference address, and the descriptor content may include shape parameters based on this reference address. This reference address can be determined by setting the environmental parameters of this task. The relevant description and use of the reference address can be found in the above embodiment. In this implementation, the content of the descriptor can be mapped to a data address more quickly.
[0074] In a possible implementation, the reference address may be included in the content of each descriptor, and the reference address of each descriptor may be different. Compared with the method of setting a common reference address using environmental parameters, each descriptor in this method can describe data more flexibly and use a larger data address space.
[0075] In a possible implementation, the data address of the data corresponding to the operand of the processing instruction in the data storage space can be determined according to the content of the descriptor. The calculation of the data address is automatically completed by hardware, and the calculation method of the data address will be different when the content of the descriptor is expressed in different ways. The present disclosure does not limit the specific calculation method of the data address.
[0076] For example, the content of the descriptor in the operand is expressed using formula (2). The offsets of the tensor data indicated by the descriptor in the data storage space are offset_x and offset_y, respectively, and the size is size_x*size_y. Then, the starting data address PA1 of the tensor data indicated by the descriptor in the data storage space is (x,y) It can be determined using the following formula (5):
[0077] PA1 (x,y) =PA_start+(offset_y-1)*ori_x+offset_x (5)
[0078] The data starting address PA1 is determined according to the above formula (5) (x,y) , combined with the offsets offset_x and offset_y, and the sizes size_x and size_y of the storage area, the storage area of the tensor data indicated by the descriptor in the data storage space can be determined.
[0079] In one possible implementation, when the operand also includes a data description position for the descriptor, the data address of the data corresponding to the operand in the data storage space can be determined according to the content of the descriptor and the data description position. In this way, part of the data (e.g., one or more data) in the tensor data indicated by the descriptor can be processed.
[0080] For example, the content of the descriptor in the operand is expressed using formula (2). The offsets of the tensor data indicated by the descriptor in the data storage space are offset_x and offset_y, respectively, and the size is size_x*size_y. The data description position for the descriptor included in the operand is (x q ,y q ), then the data address PA2 of the tensor data indicated by the descriptor in the data storage space (x,y) It can be determined using the following formula (6):
[0081] PA2 (x,y) =PA_start+(offset_y+y q -1)*ori_x+(offset_x+x q ) (6)
[0082] Combination of the above Figure 1a and Figure 1b The computing device disclosed in the present invention is described. By utilizing one or more groups of pipeline operation circuits in the computing device disclosed in the present invention, computing instructions can be efficiently executed on the computing device to complete multi-stage pipeline operation, thereby improving the efficiency of computing execution and reducing the overhead of computing. In addition, by utilizing computing instructions to perform operations on tensors, the scheme disclosed in the present invention also significantly improves the access and processing efficiency of tensor data and reduces the overhead of tensor operations.
[0083] Figure 2 is a block diagram of a computing device 200 according to another embodiment of the present disclosure. As can be seen from the figure, in addition to having the same two sets of pipeline operation circuits 102 and 104 as the computing device 100, the computing device 200 also additionally includes a control circuit 202 and a data processing circuit 204. In one embodiment, the control circuit 202 can be configured to obtain the computing instructions described above and parse the computing instructions to obtain the multiple operation instructions corresponding to the multiple operations represented by the operation code, such as those represented by formula (1).
[0084] In one embodiment, the data processing unit 204 may include a data conversion circuit 206 and a data splicing circuit 208. When the computing instruction includes a pre-processing operation for a pipeline operation, such as a data conversion operation or a data splicing operation, the data conversion circuit 206 or the data splicing circuit 208 will perform the corresponding conversion operation or splicing operation according to the corresponding computing instruction. The conversion operation and the splicing operation will be described below with examples.
[0085] As far as data conversion operations are concerned, when the data bit width input to the data conversion circuit is relatively high (for example, the data bit width is 1024 bits), the data conversion circuit can convert the input data into data with a lower bit width (for example, the output data bit width is 512 bits) according to the operation requirements. According to different application scenarios, the data conversion circuit can support conversion between multiple data types, for example, FP16 (16-bit floating point number), FP32 (32-bit floating point number), FIX8 (8-bit fixed point number), FIX4 (4-bit fixed point number), FIX16 (16-bit fixed point number) and other data types with different bit widths can be converted. When the data input to the data conversion circuit is a matrix, the data conversion operation can be a transformation of the arrangement position of the matrix elements. The transformation can, for example, include matrix transposition and mirroring (later combined Figure 3a-3c Description), matrix rotation according to a predetermined angle (for example, 90 degrees, 180 degrees or 270 degrees) and conversion of matrix dimensions.
[0086] As for the data splicing operation, the data splicing circuit can perform operations such as parity splicing on the data blocks extracted from the data according to the bit length set in the instruction, for example. For example, when the data bit length is 32 bits, the data splicing circuit can divide the data into 8 data blocks 1 to 8 according to the bit width length of 4 bits, and then splice the data blocks 1, 3, 5 and 7 together, and splice the data blocks 2, 4, 6 and 8 together for operation.
[0087] In other application scenarios, the above data concatenation operation can also be performed on the data M (for example, a vector) obtained after the operation is performed. Assume that the data concatenation circuit can first split the lower 256 bits of the even-numbered rows of data M with an 8-bit bit width as 1 unit data to obtain 32 even-numbered row unit data (represented as M_2i 0 To M_2i 31 ). Similarly, the lower 256 bits of the odd-numbered rows of data M can also be split into 8-bit bit widths as 1 unit data to obtain 32 odd-numbered row unit data (represented as M_(2i+1) 0 to M_(2i+1) 31). Further, the 32 odd-numbered row unit data and the 32 even-numbered row unit data are alternately arranged in the order of data bits from low to high, first even-numbered row and then odd-numbered row. Specifically, the even-numbered row unit data 0 (M_2i 0 ) is arranged in the low position, and then the odd-numbered row unit data 0 (M_(2i+1) is arranged in sequence 0 ). Next, even-numbered row unit data 1 (M_2i 1 )……. And so on, when the odd-numbered row unit data 31 (M_(2i+1) 31 ), 64 unit data are spliced together to form a new data with a bit width of 512 bits.
[0088] According to different application scenarios, the data conversion circuit and the data splicing circuit in the data processing unit can be used in combination to perform pre-processing or post-processing of the data more flexibly. For example, according to the different operations included in the calculation instruction, the data processing unit can only perform data conversion without performing data splicing operations, only perform data splicing operations without performing data conversion, or perform both data conversion and data splicing operations. In some scenarios, when the calculation instruction does not include a pre-processing operation for pipeline operation, the data processing unit can be configured to disable the data conversion circuit and the data splicing circuit. In other scenarios, when the calculation instruction includes a post-processing operation for pipeline operation, the data processing unit can be configured to enable the data conversion circuit and the data splicing circuit to perform post-processing of the intermediate result data, thereby obtaining the final calculation result.
[0089] In order to implement data storage operations, the computing device 200 also includes a storage circuit 210. In an implementation scenario, the storage circuit disclosed herein may include a main storage module and / or a main cache module, wherein the main storage module is configured to store data for performing multi-stage pipeline operations and operation results after the operations are performed, and the main cache module is configured to cache intermediate operation results after the operations are performed in the multi-stage pipeline operations. Furthermore, the storage circuit may also have an interface for data transmission with an off-chip storage medium, so that data transfer between on-chip and off-chip systems can be achieved. When performing tensor operations, the storage circuit 210 can obtain operand corresponding data from the storage address determined by the aforementioned descriptor, and after the tensor operation is completed, the result data is stored in the corresponding storage address using the descriptor.
[0090] Figure 3a , 3b and 3c are schematic diagrams showing matrix conversion performed by the data conversion circuit according to the embodiment of the present disclosure. In order to better understand the conversion operation performed by the data conversion circuit 206, the following will take the transposition operation and horizontal mirror operation of the original matrix as an example for further description.
[0091] like Figure 3a As shown, the original matrix is a matrix of (M+1) rows × (N+1) columns. According to the requirements of the application scenario, the data conversion circuit can Figure 3a The original matrix shown in is transformed by transposing operation to obtain Figure 3b Specifically, the data conversion circuit can exchange the row numbers and column numbers of the elements in the original matrix to form a transposed matrix. Specifically, Figure 3a The coordinates of the original matrix shown are the element "10" at row 1 and column 0. Figure 3b The coordinates in the transposed matrix shown are row 0 and column 1. Similarly, Figure 3a The coordinates of the original matrix shown are the element "M0" at the M+1th row and the 0th column. Figure 3b The coordinates in the transposed matrix shown are then row 0 and column M+1.
[0092] like Figure 3c As shown, the data conversion circuit can Figure 3a The original matrix shown is horizontally mirrored to form a horizontal mirror matrix. Specifically, the data conversion circuit can convert the arrangement order from the first row elements to the last row elements in the original matrix into the arrangement order from the last row elements to the first row elements through the horizontal mirror operation, while the column numbers of the elements in the original matrix remain unchanged. Specifically, Figure 3a The coordinates of the original matrix shown are the element "00" at row 0 and column 0 and the element "10" at row 1 and column 0. Figure 3c The coordinates in the horizontal mirror matrix shown in are the M+1th row, 0th column and the Mth row, 0th column. Figure 3a The coordinates of the original matrix shown are the element "M0" at the M+1th row and the 0th column. Figure 3c The coordinates in the horizontal mirror matrix shown are then row 0, column 0.
[0093] Based on the hardware architecture of FIG. 3 above, the computing device disclosed herein can execute computing instructions including the aforementioned pre-processing and post-processing. Two illustrative examples of computing instructions according to the disclosed solution are given below:
[0094] Example 1: MUAD=(FPMULT)+(FPADD / FPSUB)+(RELU)+(CONVERTFP2FIX) (7)
[0095] A calculation instruction expressed in the above formula (7) is a calculation instruction that inputs a 3-ary operand and outputs a 1-ary operand, and it includes a microinstruction that can be completed by a group of pipeline operation circuits including a three-level pipeline operation (i.e., multiplication + addition / subtraction + activation) according to the present disclosure. Specifically, the ternary operation is A*B+C, wherein the microinstruction of FPMULT completes the floating-point multiplication operation between operands A and B to obtain the product value, i.e., the first-level pipeline operation. Next, the microinstruction of FPADD or FPSUB is executed to complete the floating-point addition or subtraction operation of the aforementioned product value and C to obtain the sum or difference result, i.e., the second-level pipeline operation. Then, the activation operation RELU can be performed on the previous-level result, i.e., the third-level pipeline operation. After the three-level pipeline operation, the microinstruction CONVERTFP2FIX can be finally executed through the type conversion circuit above, so as to convert the type of the result data after the activation operation from a floating-point number to a fixed-point number, so as to be output as the final result or input as an intermediate result to the fixed-point operator for further calculation operations.
[0096] Example 2: SECMUADC=SEARCHC+MULT+ADD (8)
[0097] A calculation instruction expressed in the above formula (8) is a calculation instruction that inputs a 3-ary operand and outputs a 1-ary operand, and it includes a microinstruction that can be completed by a set of pipeline operation circuits including a three-level pipeline operation (i.e., table lookup + multiplication + addition) according to the present disclosure. Specifically, the ternary operation is ST(A)*B+C, wherein the microinstruction of SEARCHC can be completed by the table lookup circuit in the first-level pipeline operation to obtain the table lookup result A. Then, the multiplication operation between operands A and B is completed by the second-level pipeline operation to obtain the product value. Then, the microinstruction of ADD is executed to complete the addition operation of the aforementioned product value and C to obtain the sum result, i.e., the third-level pipeline operation.
[0098] As mentioned above, the computing instructions disclosed in the present invention can be flexibly designed and determined according to the requirements of the calculation, so that the hardware architecture disclosed in the present invention including multiple computing pipelines can be designed and connected based on the computing instructions and the various microinstructions (or microoperations) included therein, so that multiple computing operations can be completed through one computing instruction, thereby improving the execution efficiency of the instructions and reducing computing overhead.
[0099] Figure 4 4 is a block diagram showing a computing system 400 according to an embodiment of the present disclosure. As can be seen from the figure, in addition to the computing device 200, the computing system also includes a plurality of slave processing circuits 402 and an interconnection unit 404 for connecting the computing device 200 and the plurality of slave processing circuits 402.
[0100] In one computing scenario, the slave processing circuit of the present disclosure can operate on the data of the pre-processing operation performed in the computing device according to the computing instruction (implemented as, for example, one or more microinstructions or control signals) to obtain the expected computing result. In another computing scenario, the slave processing circuit can send the intermediate result obtained after its operation (for example, via the interconnection unit) to the data processing unit in the computing device, so that the data conversion circuit in the data processing unit can perform data type conversion on the intermediate result or the data splicing circuit in the data processing unit can perform data splitting and splicing operations on the intermediate result, so as to obtain the final computing result.
[0101] Figure 5 1 is a simplified flow chart showing a method 500 for performing a computing operation using a computing device according to an embodiment of the present disclosure. According to the foregoing description, it can be understood that the computing device here can be a combination of FIG. 1 (including Figure 1a and Figure 1b )- Figure 4 The described computing device has the internal connection relationships shown and supports additional types of operations.
[0102] like Figure 5 As shown, at step 502, the method 500 configures each of the one or more groups of pipeline operation circuits to perform multi-stage pipeline operations according to multiple operation instructions obtained after parsing the calculation instruction, wherein each group of the pipeline operation circuits constitutes a multi-stage operation pipeline, and the multi-stage operation pipeline includes multiple operation circuits arranged stage by stage, wherein the operand of the calculation instruction includes a descriptor for indicating the shape of a tensor, and the descriptor is used to determine the storage address of the data corresponding to the operand. Then, at step 504, in response to receiving multiple operation instructions, the method 500 configures at least one stage of the operation circuit in the multi-stage operation pipeline to execute a corresponding one of the multiple operation instructions based on the storage address.
[0103] For the purpose of simplicity, the above only combines Figure 5 The calculation method of the present disclosure is described. A person skilled in the art can also imagine that the method may include more steps based on the disclosure of the present disclosure, and the execution of these steps can achieve the above-mentioned combination of FIG. Figure 4 The various operations described in this disclosure will not be repeated here.
[0104] Figure 6 is a structural diagram showing a combined processing device 600 according to an embodiment of the present disclosure. Figure 6As shown in FIG. 1-1 , the combined processing device 600 includes a computing processing device 602, an interface device 604, other processing devices 606, and a storage device 608. According to different application scenarios, the computing processing device may include one or more computing devices 610, which may be configured to perform the above-mentioned steps in conjunction with FIG. 1-1 . Figure 5 The operation described.
[0105] In different embodiments, the computing and processing device disclosed herein may be configured to perform user-specified operations. In an exemplary application, the computing and processing device may be implemented as a single-core artificial intelligence processor or a multi-core artificial intelligence processor. Similarly, one or more computing devices included in the computing and processing device may be implemented as an artificial intelligence processor core or a partial hardware structure of an artificial intelligence processor core. When multiple computing devices are implemented as an artificial intelligence processor core or a partial hardware structure of an artificial intelligence processor core, with respect to the computing and processing device disclosed herein, it may be regarded as having a single-core structure or a homogeneous multi-core structure.
[0106] In an exemplary operation, the computing processing device of the present disclosure can interact with other processing devices through an interface device to jointly complete the operation specified by the user. Depending on the implementation, other processing devices of the present disclosure may include one or more types of processors in general and / or special processors such as a central processing unit (CPU), a graphics processing unit (GPU), an artificial intelligence processor, etc. These processors may include but are not limited to a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc., and their number can be determined according to actual needs. As mentioned above, only with respect to the computing processing device of the present disclosure, it can be regarded as having a single-core structure or a homogeneous multi-core structure. However, when the computing processing device and other processing devices are considered together, the two can be regarded as forming a heterogeneous multi-core structure.
[0107] In one or more embodiments, the other processing device can serve as an interface between the computing device disclosed herein (which can be embodied as an artificial intelligence such as a computing device related to neural network computing) and external data and control, and perform basic controls including but not limited to data handling, starting and / or stopping the computing device, etc. In other embodiments, the other processing device can also cooperate with the computing device to jointly complete computing tasks.
[0108] In one or more embodiments, the interface device can be used to transmit data and control instructions between the computing and processing device and other processing devices. For example, the computing and processing device can obtain input data from other processing devices via the interface device and write it into a storage device (or memory) on the computing and processing device chip. Further, the computing and processing device can obtain control instructions from other processing devices via the interface device and write them into a control cache on the computing and processing device chip. Alternatively or optionally, the interface device can also read data from the storage device of the computing and processing device and transmit it to other processing devices.
[0109] Additionally or optionally, the combined processing device of the present disclosure may further include a storage device. As shown in the figure, the storage device is connected to the computing processing device and the other processing device, respectively. In one or more embodiments, the storage device may be used to store data of the computing processing device and / or the other processing device. For example, the data may be data that cannot be fully stored in the internal or on-chip storage device of the computing processing device or other processing device.
[0110] In some embodiments, the present disclosure also discloses a chip (e.g. Figure 7 In one implementation, the chip is a system on chip (SoC) and integrates one or more Figure 6 The chip can be connected to an external interface device (such as Figure 7 The external interface device 706 shown in the figure is connected to other related components. The related components may be, for example, a camera, a display, a mouse, a keyboard, a network card or a wifi interface. In some application scenarios, other processing units (such as a video codec) and / or interface modules (such as a DRAM interface) may be integrated on the chip. In some embodiments, the present disclosure further discloses a chip packaging structure, which includes the above-mentioned chip. In some embodiments, the present disclosure further discloses a board card, which includes the above-mentioned chip packaging structure. The following will be combined with Figure 7 The board is described in detail.
[0111] Figure 7 FIG. 2 is a schematic diagram showing the structure of a board 700 according to an embodiment of the present disclosure. Figure 7 As shown in , the board includes a storage device 704 for storing data, which includes one or more storage units 710. The storage device can be connected and data can be transmitted with the control device 708 and the chip 702 described above by means of, for example, a bus. Further, the board also includes an external interface device 706, which is configured for data relay or switching between the chip (or the chip in the chip packaging structure) and an external device 712 (such as a server or computer, etc.). For example, the data to be processed can be transmitted to the chip by the external device through the external interface device. For another example, the calculation result of the chip can be transmitted back to the external device via the external interface device. According to different application scenarios, the external interface device can have different interface forms, for example, it can adopt a standard PCIE interface, etc.
[0112] In one or more embodiments, the control device in the disclosed board may be configured to regulate the state of the chip. To this end, in an application scenario, the control device may include a microcontroller unit (MCU) to regulate the working state of the chip.
[0113] According to the above combination Figure 6 and Figure 7 Based on the description, those skilled in the art can understand that the present disclosure also discloses an electronic device or apparatus, which may include one or more of the above-mentioned boards, one or more of the above-mentioned chips and / or one or more of the above-mentioned combined processing devices.
[0114] According to different application scenarios, the electronic equipment or device disclosed herein may include servers, cloud servers, server clusters, data processing devices, robots, computers, printers, scanners, tablet computers, smart terminals, PC devices, IoT terminals, mobile terminals, mobile phones, driving recorders, navigators, sensors, cameras, cameras, video cameras, projectors, watches, headphones, mobile storage, wearable devices, visual terminals, automatic driving terminals, transportation, household appliances, and / or medical equipment. The transportation includes airplanes, ships and / or vehicles; the household appliances include televisions, air conditioners, microwave ovens, refrigerators, rice cookers, humidifiers, washing machines, electric lights, gas stoves, and range hoods; the medical equipment includes magnetic resonance imaging, ultrasound machines and / or electrocardiographs. The electronic equipment or device disclosed herein may also be applied to the Internet, IoT, data centers, energy, transportation, public administration, manufacturing, education, power grids, telecommunications, finance, retail, construction sites, and medical fields. Further, the electronic equipment or device disclosed herein may also be used in cloud, edge, and terminal applications related to artificial intelligence, big data, and / or cloud computing. In one or more embodiments, electronic devices or devices with high computing power according to the disclosed solution can be applied to cloud devices (such as cloud servers), while electronic devices or devices with low power consumption can be applied to terminal devices and / or edge devices (such as smart phones or cameras). In one or more embodiments, the hardware information of the cloud device and the hardware information of the terminal device and / or edge device are compatible with each other, so that according to the hardware information of the terminal device and / or edge device, appropriate hardware resources can be matched from the hardware resources of the cloud device to simulate the hardware resources of the terminal device and / or edge device, so as to complete the unified management, scheduling and collaborative work of end-to-end or cloud-edge-to-end.
[0115] It should be noted that, for the purpose of simplicity, the present disclosure describes some methods and embodiments thereof as a series of actions and combinations thereof, but those skilled in the art will appreciate that the scheme of the present disclosure is not limited by the order of the actions described. Therefore, based on the disclosure or teaching of the present disclosure, those skilled in the art will appreciate that some of the steps therein may be performed in other orders or simultaneously. Further, those skilled in the art will appreciate that the embodiments described in the present disclosure may be regarded as optional embodiments, i.e., the actions or modules involved therein may not necessarily be necessary for the implementation of one or some of the schemes of the present disclosure. In addition, depending on the different schemes, the present disclosure also has different focuses on the description of some embodiments. In view of this, those skilled in the art may appreciate the parts that are not described in detail in a certain embodiment of the present disclosure, and may also refer to the relevant descriptions of other embodiments.
[0116] In terms of specific implementation, based on the disclosure and teaching of this disclosure, those skilled in the art can understand that several embodiments disclosed in this disclosure can also be implemented by other methods not disclosed herein. For example, with respect to the various units in the electronic device or device embodiments described above, this article divides them on the basis of considering logical functions, and there may be other division methods in actual implementation. For another example, multiple units or components can be combined or integrated into another system, or some features or functions in a unit or component can be selectively disabled. In terms of the connection relationship between different units or components, the connection discussed in the above text in conjunction with the accompanying drawings can be a direct or indirect coupling between units or components. In some scenarios, the aforementioned direct or indirect coupling involves a communication connection using an interface, wherein the communication interface can support electrical, optical, acoustic, magnetic or other forms of signal transmission.
[0117] In the present disclosure, the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units. The aforementioned components or units may be located in the same location or distributed on multiple network units. In addition, according to actual needs, some or all of the units may be selected to achieve the purpose of the scheme described in the embodiments of the present disclosure. In addition, in some scenarios, multiple units in the embodiments of the present disclosure may be integrated into one unit or each unit may exist physically separately.
[0118] In some implementation scenarios, the above-mentioned integrated unit can be implemented in the form of a software program module. If implemented in the form of a software program module and sold or used as an independent product, the integrated unit can be stored in a computer-readable memory. Based on this, when the scheme of the present disclosure is embodied in the form of a software product (such as a computer-readable storage medium), the software product can be stored in a memory, which may include several instructions to enable a computer device (such as a personal computer, a server or a network device, etc.) to perform some or all of the steps of the method described in the embodiment of the present disclosure. The aforementioned memory may include, but is not limited to, various media that can store program codes, such as a USB flash drive, a flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.
[0119] In some other implementation scenarios, the above-mentioned integrated unit can also be implemented in the form of hardware, that is, a specific hardware circuit, which may include digital circuits and / or analog circuits, etc. The physical implementation of the hardware structure of the circuit may include but is not limited to physical devices, and the physical devices may include but are not limited to devices such as transistors or memristors. In view of this, the various devices described herein (such as computing devices or other processing devices) can be implemented by appropriate hardware processors, such as CPU, GPU, FPGA, DSP and ASIC, etc. Further, the aforementioned storage unit or storage device can be any appropriate storage medium (including magnetic storage medium or magneto-optical storage medium, etc.), which can be, for example, a variable resistive memory (Resistive Random Access Memory, RRAM), a dynamic random access memory (Dynamic Random Access Memory, DRAM), a static random access memory (Static Random Access Memory, SRAM), an enhanced dynamic random access memory (Enhanced Dynamic Random Access Memory, EDRAM), a high bandwidth memory (High Bandwidth Memory, HBM), a hybrid memory cube (Hybrid Memory Cube, HMC), ROM and RAM, etc.
[0120] The foregoing content can be better understood in accordance with the following terms:
[0121] Clause 1. A computing device comprising:
[0122] One or more groups of pipeline operation circuits are configured to perform multi-stage pipeline operations according to multiple operation instructions obtained after parsing the calculation instruction, wherein each group of pipeline operation circuits constitutes a multi-stage operation pipeline, and the multi-stage operation pipeline includes multiple operation circuits arranged stage by stage, wherein the operand of the calculation instruction includes a descriptor for indicating the shape of a tensor, and the descriptor is used to determine the storage address of the data corresponding to the operand,
[0123] In response to receiving the multiple operation instructions, at least one stage of the operation circuit in the multi-stage operation pipeline is configured to execute a corresponding one of the multiple operation instructions based on the storage address.
[0124] Item 2. A computing device according to Item 1, wherein the opcode of the computing instruction represents multiple operations performed by the multi-stage computing pipeline, and the computing device also includes a control circuit configured to obtain the computing instruction and parse it to obtain the multiple computing instructions corresponding to the multiple operations, and in the parsing, the control circuit is also configured to determine the storage address of the data corresponding to the operand according to the descriptor.
[0125] Clause 3. A computing device according to Clause 2, wherein the computing instruction includes an identification of a descriptor and / or the content of the descriptor, and the content of the descriptor includes at least one shape parameter representing the shape of tensor data.
[0126] Clause 4. A computing device according to clause 3, wherein the contents of the descriptor also include at least one address parameter representing an address of tensor data.
[0127] Clause 5. A computing device according to clause 4, wherein the address parameter of the tensor data includes a reference address of a data reference point of the descriptor in a data storage space of the tensor data.
[0128] Clause 6. The computing device of clause 5, wherein the shape parameter of the tensor data comprises at least one of the following:
[0129] The size of the data storage space in at least one direction of the N dimensions, the size of the storage area of the tensor data in at least one direction of the N dimensions, the offset of the storage area in at least one direction of the N dimensions, the positions of at least two vertices at diagonal positions in the N dimensions relative to the data reference point, and the mapping relationship between the data description position of the tensor data indicated by the descriptor and the data address, where N is an integer greater than or equal to zero.
[0130] Clause 7. A computing device according to Clause 2, wherein the operation code and the multiple operations it represents are predetermined according to functions supported by multiple operation circuits arranged stage by stage in a multi-stage operation pipeline.
[0131] Item 8. A computing device according to Item 1, wherein each stage of the operation circuit in the multi-stage operation pipeline is configured to be selectively connected according to the multiple operation instructions so as to execute the multiple operation instructions.
[0132] Item 9. A computing device according to Item 1, wherein the multiple groups of pipeline operation circuits constitute multiple multi-stage operation pipelines, and the multiple multi-stage operation pipelines execute their respective multiple operation instructions in parallel.
[0133] Item 10. A computing device according to Item 1 or 9, wherein each stage of the operation circuit in the multi-stage operation pipeline has an input end and an output end for receiving input data at the operation circuit at that stage and outputting the result of the operation of the operation circuit at that stage.
[0134] Item 11. A computing device according to Item 10, wherein within a multi-stage operation pipeline, the output end of the operation circuit of one or more stages is configured to be connected to the input end of the operation circuit of another stage or multiple stages according to an operation instruction to execute the operation instruction.
[0135] Item 12. A computing device according to Item 10, wherein the plurality of multi-stage operation pipelines include a first multi-stage operation pipeline and a second multi-stage operation pipeline, wherein the output end of the operation circuit of one or more stages of the first multi-stage operation pipeline is configured to be connected to the input end of the operation circuit of one or more stages of the second multi-stage operation pipeline according to the operation instruction.
[0136] Clause 13. The computing device according to clause 1, wherein each stage of the computing circuit comprises one or more of the following operators or circuits:
[0137] Random number processing circuit, addition and subtraction circuit, subtraction circuit, table lookup circuit, parameter configuration circuit, multiplier, pooler, comparator, absolute value circuit, logic operator, position index circuit or filter.
[0138] Item 14. The computing device according to Item 1 further includes a data processing circuit, which includes a type conversion circuit for performing a data type conversion operation and / or a data splicing circuit for performing a data splicing operation.
[0139] Clause 15. A computing device according to Clause 14, wherein the type conversion circuit includes one or more converters for implementing conversion of computing data between multiple different data types.
[0140] Clause 16. A computing device according to clause 14, wherein the data splicing circuit is configured to split the computing data into predetermined bit lengths and to splice the multiple data blocks obtained after the splitting in a predetermined order.
[0141] Clause 17. An integrated circuit chip comprising a computing device according to any one of clauses 1-16.
[0142] Clause 18. A board comprising the integrated circuit chip according to clause 17.
[0143] Clause 19. An electronic device comprising the integrated circuit chip according to clause 17.
[0144] Clause 20. A method of performing a computing operation using a computing device, wherein the computing device includes one or more sets of pipeline operation circuits, the method comprising:
[0145] Each of the one or more groups of pipeline operation circuits is configured to perform multi-stage pipeline operations according to a plurality of operation instructions obtained after parsing the calculation instruction, wherein each group of pipeline operation circuits constitutes a multi-stage operation pipeline, and the multi-stage operation pipeline includes a plurality of operation circuits arranged stage by stage, wherein the operand of the calculation instruction includes a descriptor for indicating the shape of a tensor, and the descriptor is used to determine the storage address of data corresponding to the operand; and
[0146] In response to receiving a plurality of operation instructions, at least one stage of operation circuit in the multi-stage operation pipeline is configured to execute a corresponding one of the plurality of operation instructions based on the storage address.
[0147] Item 21. A method according to Item 20, wherein the opcode of the computing instruction represents multiple operations performed by the multi-stage computing pipeline, the computing device also includes a control circuit, the method includes configuring the control circuit to obtain the computing instruction and parse it to obtain the multiple computing instructions corresponding to the multiple operations, and in the parsing, the method also includes configuring the control circuit to determine the storage address of the data corresponding to the operand according to the descriptor.
[0148] Clause 22. A method according to clause 21, wherein the computation instruction includes an identification of a descriptor and / or content of the descriptor, wherein the content of the descriptor includes at least one shape parameter representing a shape of tensor data.
[0149] Clause 23. A method according to clause 22, wherein the contents of the descriptor also include at least one address parameter representing an address of tensor data.
[0150] Clause 24. A method according to clause 23, wherein the address parameter of the tensor data includes a reference address of a data reference point of the descriptor in a data storage space of the tensor data.
[0151] Clause 25. The method of clause 24, wherein the shape parameter of the tensor data comprises at least one of the following:
[0152] The size of the data storage space in at least one direction of the N dimensions, the size of the storage area of the tensor data in at least one direction of the N dimensions, the offset of the storage area in at least one direction of the N dimensions, the positions of at least two vertices at diagonal positions in the N dimensions relative to the data reference point, and the mapping relationship between the data description position of the tensor data indicated by the descriptor and the data address, where N is an integer greater than or equal to zero.
[0153] Clause 26. A method according to clause 21, wherein the operation code and the multiple operations represented by the operation code are predetermined according to functions supported by multiple operation circuits arranged stage by stage in a multi-stage operation pipeline.
[0154] Clause 27. The method according to clause 20, wherein each stage of the operation circuit in the multi-stage operation pipeline is configured to be selectively connected according to the multiple operation instructions so as to execute the multiple operation instructions.
[0155] Clause 28. The method according to clause 20, wherein the plurality of groups of pipeline operation circuits constitute a plurality of multi-stage operation pipelines, and the plurality of multi-stage operation pipelines execute respective plurality of operation instructions in parallel.
[0156] Item 29. A method according to Item 20 or 28, wherein each stage of the operation circuit in the multi-stage operation pipeline has an input end and an output end for receiving input data at the stage of the operation circuit and outputting the result of the operation of the stage of the operation circuit.
[0157] Item 30. A method according to Item 29, wherein in a multi-stage operation pipeline, the output end of the operation circuit of one or more stages is configured to be connected to the input end of the operation circuit of another stage or multiple stages according to the operation instruction to execute the operation instruction.
[0158] Item 31. A method according to Item 29, wherein the plurality of multi-stage operation pipelines include a first multi-stage operation pipeline and a second multi-stage operation pipeline, wherein the method configures the output end of the operation circuit of one or more stages of the first multi-stage operation pipeline to be connected to the input end of the operation circuit of one or more stages of the second multi-stage operation pipeline according to the operation instruction.
[0159] Clause 32. The method according to clause 20, wherein each stage of the operation circuit comprises one or more of the following operators or circuits:
[0160] Random number processing circuit, addition and subtraction circuit, subtraction circuit, table lookup circuit, parameter configuration circuit, multiplier, pooler, comparator, absolute value circuit, logic operator, position index circuit or filter.
[0161] Clause 33. The method according to Clause 20 further includes a data processing circuit, which includes a type conversion circuit for performing a data type conversion operation and / or a data splicing circuit for performing a data splicing operation.
[0162] Clause 34. A method according to clause 33, wherein the type conversion circuit comprises one or more converters for implementing conversion of computing data between a plurality of different data types.
[0163] Item 35. A method according to Item 33, wherein the data splicing circuit is configured to split the calculation data into predetermined bit lengths and splice the multiple data blocks obtained after the splitting in a predetermined order.
[0164] Although multiple embodiments of the present disclosure have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Those skilled in the art may think of many changes, modifications, and alternatives without departing from the thought and spirit of the present disclosure. It should be understood that in the process of practicing the present disclosure, various alternatives to the embodiments of the present disclosure described herein may be adopted. The attached claims are intended to define the scope of protection of the present disclosure, and therefore cover equivalents or alternatives within the scope of these claims.
Claims
1. A computing device, comprising: One or more groups of pipeline operation circuits are configured to perform multi-stage pipeline operations according to multiple operation instructions obtained after parsing the calculation instruction, wherein each group of pipeline operation circuits constitutes a multi-stage operation pipeline, and the multi-stage operation pipeline includes multiple operation circuits arranged stage by stage, wherein the operand of the calculation instruction includes a descriptor for indicating the shape of a tensor, and the descriptor is used to determine the storage address of the data corresponding to the operand, The computing instruction includes an identifier of a descriptor and / or content of the descriptor, and the content of the descriptor includes shape parameters representing the shape of tensor data, and the shape parameters of the tensor data include: A combination of the size of the data storage space in at least one direction of the N dimensional directions, the size of the storage area of the tensor data in at least one direction of the N dimensional directions, and the offset of the storage area in at least one direction of the N dimensional directions, and / or at least one of the following: The mapping relationship between the positions of at least two vertices at diagonal positions in N dimensional directions relative to the data reference point, the data description position of the tensor data indicated by the descriptor, and the data address, wherein the data description position is the mapping position of a point or region in the tensor data indicated by the descriptor expressed using spatial coordinates, and N is an integer greater than or equal to zero; wherein in response to receiving the plurality of operation instructions, at least one stage of operation circuit in the multi-stage operation pipeline is configured to execute a corresponding one of the plurality of operation instructions based on the storage address; The computing device also includes a data processing circuit including a data splicing circuit for performing a data splicing operation.
2. A computing device according to claim 1, wherein the opcode of the computing instruction represents multiple operations performed by the multi-stage computing pipeline, and the computing device also includes a control circuit, which is configured to obtain the computing instruction and parse it to obtain the multiple computing instructions corresponding to the multiple operations, and in the parsing, the control circuit is also configured to determine the storage address of the data corresponding to the operand according to the descriptor.
3. A computing device according to claim 1, wherein the content of the descriptor also includes at least one address parameter representing an address of tensor data. 4 . The computing device according to claim 3 , wherein the address parameter of the tensor data comprises a reference address of a data reference point of the descriptor in a data storage space of the tensor data.
5. The computing device according to claim 2, wherein the operation code and the multiple operations represented by the operation code are predetermined according to functions supported by multiple operation circuits arranged stage by stage in a multi-stage operation pipeline. 6 . The computing device according to claim 1 , wherein each stage of the operation circuit in the multi-stage operation pipeline is configured to be selectively connected according to the multiple operation instructions so as to execute the multiple operation instructions.
7. The computing device according to claim 1, wherein the plurality of groups of pipeline operation circuits constitute a plurality of multi-stage operation pipelines, and the plurality of multi-stage operation pipelines execute respective plurality of operation instructions in parallel.
8. A computing device according to claim 1 or 7, wherein each stage of the operation circuit in the multi-stage operation pipeline has an input end and an output end, which is used to receive input data at the operation circuit of that stage and output the result of the operation of the operation circuit of that stage.
9. A computing device according to claim 8, wherein in a multi-stage operation pipeline, the output end of the operation circuit of one or more stages is configured to be connected to the input end of the operation circuit of another stage or multiple stages according to the operation instruction to execute the operation instruction.
10. A computing device according to claim 7, wherein the plurality of multi-stage operation pipelines include a first multi-stage operation pipeline and a second multi-stage operation pipeline, wherein the output end of the operation circuit of one or more stages of the first multi-stage operation pipeline is configured to be connected to the input end of the operation circuit of one or more stages of the second multi-stage operation pipeline according to the operation instruction.
11. The computing device according to claim 1, wherein each stage of the computing circuit comprises one or more of the following computing units or circuits: Random number processing circuit, addition and subtraction circuit, subtraction circuit, table lookup circuit, parameter configuration circuit, multiplier, pooler, comparator, absolute value circuit, logic operator, position index circuit or filter.
12. The computing device of claim 1, wherein the data processing circuit further comprises a type conversion circuit for performing a data type conversion operation.
13. The computing device according to claim 12, wherein the type conversion circuit comprises one or more converters for realizing the conversion of computing data between a plurality of different data types.
14. The computing device according to claim 1, wherein the data splicing circuit is configured to split the computing data into predetermined bit lengths and splice the multiple data blocks obtained after the splitting in a predetermined order.
15. An integrated circuit chip comprising a computing device according to any one of claims 1-14.
16. A board comprising the integrated circuit chip according to claim 15.
17. An electronic device comprising the integrated circuit chip according to claim 15.
18. A method for performing a computing operation using a computing device, wherein the computing device comprises one or more sets of pipeline operation circuits and a data processing circuit, the data processing circuit comprises a data splicing circuit for performing a data splicing operation, the method comprising: Each of the one or more groups of pipeline operation circuits is configured to perform multi-stage pipeline operations according to a plurality of operation instructions obtained after parsing the calculation instruction, wherein each group of pipeline operation circuits constitutes a multi-stage operation pipeline, and the multi-stage operation pipeline includes a plurality of operation circuits arranged stage by stage, wherein the operand of the calculation instruction includes a descriptor for indicating the shape of a tensor, and the descriptor is used to determine the storage address of the data corresponding to the operand, The computing instruction includes an identifier of a descriptor and / or content of the descriptor, and the content of the descriptor includes shape parameters representing the shape of tensor data, and the shape parameters of the tensor data include: A combination of the size of the data storage space in at least one direction of the N dimensional directions, the size of the storage area of the tensor data in at least one direction of the N dimensional directions, and the offset of the storage area in at least one direction of the N dimensional directions, and / or at least one of the following: a mapping relationship between the positions of at least two vertices at diagonal positions in N dimensional directions relative to a data reference point, a data description position of the tensor data indicated by the descriptor, and a data address, wherein the data description position is a mapping position of a point or region in the tensor data indicated by the descriptor expressed using spatial coordinates, and N is an integer greater than or equal to zero; and In response to receiving a plurality of operation instructions, at least one stage of operation circuit in the multi-stage operation pipeline is configured to execute a corresponding one of the plurality of operation instructions based on the storage address.
19. A method according to claim 18, wherein the opcode of the computing instruction represents multiple operations performed by the multi-stage computing pipeline, the computing device also includes a control circuit, the method includes configuring the control circuit to obtain the computing instruction and parse it to obtain the multiple computing instructions corresponding to the multiple operations, and in the parsing, the method also includes configuring the control circuit to determine the storage address of the data corresponding to the operand according to the descriptor.
20. The method of claim 18, wherein the content of the descriptor further comprises at least one address parameter representing an address of tensor data. 21 . The method according to claim 20 , wherein the address parameter of the tensor data comprises a reference address of a data reference point of the descriptor in a data storage space of the tensor data.
22. The method according to claim 19, wherein the operation code and the multiple operations represented by the operation code are predetermined according to functions supported by multiple operation circuits arranged stage by stage in a multi-stage operation pipeline.
23. The method according to claim 18, wherein each stage of the operation circuit in the multi-stage operation pipeline is configured to be selectively connected according to the plurality of operation instructions so as to execute the plurality of operation instructions.
24. The method according to claim 18, wherein the plurality of groups of pipeline operation circuits constitute a plurality of multi-stage operation pipelines, and the plurality of multi-stage operation pipelines execute respective plurality of operation instructions in parallel.
25. The method according to claim 18 or 24, wherein each stage of the operation circuit in the multi-stage operation pipeline has an input end and an output end for receiving input data at the operation circuit of that stage and outputting the result of the operation of the operation circuit of that stage.
26. A method according to claim 25, wherein in a multi-stage operation pipeline, the output end of the operation circuit of one or more stages is configured to be connected to the input end of the operation circuit of another stage or multiple stages according to the operation instruction to execute the operation instruction.
27. A method according to claim 24, wherein the plurality of multi-stage operation pipelines include a first multi-stage operation pipeline and a second multi-stage operation pipeline, wherein the method configures the output end of the operation circuit of one or more stages of the first multi-stage operation pipeline to be connected to the input end of the operation circuit of one or more stages of the second multi-stage operation pipeline according to the operation instruction.
28. The method according to claim 18, wherein each stage of the operation circuit comprises one or more of the following operators or circuits: Random number processing circuit, addition and subtraction circuit, subtraction circuit, table lookup circuit, parameter configuration circuit, multiplier, pooler, comparator, absolute value circuit, logic operator, position index circuit or filter.
29. The method of claim 18, wherein the data processing circuit further comprises a type conversion circuit for performing a data type conversion operation.
30. The method according to claim 29, wherein the type conversion circuit comprises one or more converters for implementing conversion of computing data between a plurality of different data types.
31. The method according to claim 18, wherein the data splicing circuit is configured to split the calculation data into predetermined bit lengths and splice the multiple data blocks obtained after the splitting in a predetermined order.
Citation Information
Patent Citations
Device and method for executing forward operation of artificial neural network represented by discrete data
CN107729990A