Runtime Reconfigurable Artificial Intelligence Chip, Processing Unit, Computing Task Execution Method, and Electronic Device

By designing a reconfigurable processing unit at runtime in an artificial intelligence chip, using the controller to analyze computing tasks and configure quantized bit counts, the problems of poor universality of processing units and inflexible computing resource allocation in the prior art are solved, and more efficient computing resource utilization and widespread application of chips are achieved.

CN119576275BActive Publication Date: 2025-07-29BEIJING SMARTCHIP MICROELECTRONICS TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411722029.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-28
Publication Date
2025-07-29
Estimated Expiration
2044-11-28

AI Technical Summary

Technical Problem

Existing artificial intelligence chip processing units are usually designed for fixed quantized bits, resulting in poor chip versatility, increasing manufacturing and usage costs, and being unable to flexibly configure quantized bits based on actual calculation requirements, affecting computing efficiency.

Method used

Design a running-time reconfigurable artificial intelligence chip, which analyzes the calculation task and configures the quantized bit count of the processing unit through the controller, and uses a processing unit composed of intercept registers, multipliers, shifters, adders, splicing registers and multiplexers to realize flexible processing of designated subcomputing tasks.

Benefits of technology

It improves the utilization rate of computing resources and computing efficiency, enhances the versatility of the chip, and is conducive to large-scale promotion and application.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119576275B_ABST
    Figure CN119576275B_ABST
Patent Text Reader

Abstract

The present disclosure relates to the technical field of artificial intelligence chips, and specifically relates to a runtime reconfigurable artificial intelligence chip, a processing unit, a computing task execution method, and an electronic device. The chip includes a controller and multiple processing units. The controller acquires a computing task, parses the computing task, and obtains multiple sub-computing tasks and the quantization bits corresponding to the sub-computing tasks. The processing unit includes an intercept register, a first multiplier, a second multiplier, a shifter, an adder, a splicing register, and a multiplexer, and is used to process a specified sub-computing task. The quantization bits corresponding to the specified sub-computing task are specified quantization bits. The controller configures the quantization bits of the processing unit according to the specified quantization bits, so as to realize the runtime reconfiguration of the artificial intelligence chip.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular, to a runtime reconfigurable artificial intelligence chip, a processing unit, a computing task execution method, and an electronic device. Background Art

[0002] Neural network models often use high-precision parameters (e.g., 32-bit floating-point numbers) to represent weights and activation values. However, this representation may consume a large amount of memory and computing resources, limiting the application of neural network models in resource-constrained devices such as mobile devices and embedded systems. Neural network quantization can be used to reduce the computing and storage resource requirements of neural network models. It converts weights and activation values from high-precision representations (e.g., 32-bit, etc.) to low-precision representations (e.g., 16-bit, 8-bit, 4-bit, 2-bit, 1-bit, etc.), which can significantly reduce the storage resource requirements of the model, reduce the computing cost, and accelerate the model processing speed.

[0003] Quantization includes weight quantization and activation quantization. Weight quantization is to convert the weight parameters in the neural network into low-precision representations, while activation quantization is to convert the activation values (i.e., the outputs of the intermediate layers) in the neural network into low-precision representations. The processing units (PEs) in existing artificial intelligence chips are usually designed for fixed quantization bits and can only calculate weight parameters with fixed quantization bits. For example, they can only calculate 16-bit weight parameters or only calculate 8-bit weight parameters. The chip has poor versatility, increases the manufacturing and use costs of the chip, and is not conducive to the popularization and application of the chip. On the other hand, existing artificial intelligence chips cannot adaptively set the quantization bits according to actual computing requirements, resulting in inflexible computing resource allocation and affecting computing efficiency. Summary of the Invention

[0004] To solve the problems in the related art, embodiments of the present disclosure provide a runtime reconfigurable artificial intelligence chip, a processing unit, a computing task execution method, and an electronic device.

[0005] In a first aspect, embodiments of the present disclosure provide a runtime reconfigurable artificial intelligence chip, including a controller and a plurality of processing units, wherein:

[0006] The controller is configured to obtain a computing task, parse the computing task to obtain a plurality of sub-computing tasks and the quantization bits corresponding to the sub-computing tasks. The sub-computing tasks include calculating the product of a first input data and a second input data, and the quantization bits include the quantization bits of the second input data;

[0007] The processing unit is used to process a specified sub-computation task among the multiple sub-computation tasks, and the quantization bit number corresponding to the specified sub-computation task is a specified quantization bit number. Wherein, when the specified quantization bit number is within the first bit range, the processing unit is used to process one specified sub-computation task; when the specified quantization bit number is within the second bit range, the processing unit is used to process two specified sub-computation tasks. The processing unit includes:

[0008] A truncation register, which is used to split the truncation register input into a high-order input sub-data and a low-order input sub-data. Wherein: when the specified quantization bit number is within the first bit range, the truncation register input includes the second input data of the one specified sub-computation task; when the specified quantization bit number is within the second bit range, the truncation register input includes the concatenation result of the second input data of the two specified sub-computation tasks, and the high-order input sub-data and the low-order input sub-data respectively correspond to the second input data of the two specified sub-computation tasks;

[0009] A first multiplier, which is used to multiply the first input data by the high-order input sub-data to obtain a first multiplier output;

[0010] A second multiplier, which is used to multiply the first input data by the low-order input sub-data to obtain a second multiplier output;

[0011] A shifter, which is used to shift the first multiplier output to obtain a shift result when the specified quantization bit number is within the first bit range;

[0012] An adder, which is used to add the shift result and the second multiplier output to obtain a first calculation result when the specified quantization bit number is within the first bit range;

[0013] A concatenation register, which is used to concatenate the first multiplier output and the second multiplier output to obtain a second calculation result when the specified quantization bit number is within the second bit range;

[0014] A multiplexer, which is used to selectively output the first calculation result and the second calculation result according to the specified quantization bit number.

[0015] According to an embodiment of the present disclosure, the specified sub-computation task includes calculating the product of a first input data and a second input data; when the specified quantization bit number is within the second bit range, the first input data of the two specified sub-computation tasks is the same.

[0016] According to an embodiment of the present disclosure, the minimum value of the first digit range is greater than the maximum value of the second digit range.

[0017] According to an embodiment of the present disclosure, the first digit range is greater than 8 bits and less than or equal to 16 bits, and the second digit range is less than or equal to 8 bits; or, the first digit range is 16 bits, and the second digit range is any one of 8 bits, 4 bits, 2 bits, and 1 bit; the number of bits of the first input data is any one of 32 bits, 16 bits, 8 bits, 4 bits, 2 bits, and 1 bit.

[0018] According to an embodiment of the present disclosure, the truncation register divides the truncation register input into a high-order input sub-data and a low-order input sub-data with equal number of bits.

[0019] According to an embodiment of the present disclosure, the first multiplier output includes a reserved bit set at the highest bit of the first multiplier output; and / or the second multiplier output includes a reserved bit set at the highest bit of the second multiplier output.

[0020] According to an embodiment of the present disclosure, the reserved bit is used to determine whether there is an overflow in the calculation result of the corresponding multiplier.

[0021] According to an embodiment of the present disclosure, the first digit range is greater than 8 bits and less than or equal to 16 bits, the second digit range is less than or equal to 8 bits, and the number of bits of the first input data is 16 bits: the first multiplier is a 16-bit * 8-bit multiplier, and the second multiplier is a 16-bit * 8-bit multiplier; the shifter is used to shift the first multiplier output to the left by 8 bits.

[0022] According to an embodiment of the present disclosure, the first multiplier output is 25 bits, where the highest bit is a reserved bit, and the remaining 24 bits are the product of the first input data and the high-order input sub-data; the second multiplier output is 25 bits, where the highest bit is a reserved bit, and the remaining 24 bits are the product of the first input data and the low-order input sub-data.

[0023] According to an embodiment of the present disclosure, the splicing register is used to splice the first multiplier output and the second multiplier output to obtain a 50-bit second calculation result.

[0024] According to an embodiment of the present disclosure, when the specified quantization number of bits is within the first digit range, the splicing register is disabled; when the specified quantization number of bits is within the second digit range, the shifter and the adder are disabled.

[0025] According to an embodiment of the present disclosure, when the specified quantization bit number is within the first bit number range, the multiplexer outputs the first calculation result; when the specified quantization bit number is within the second bit number range, the multiplexer outputs the second calculation result.

[0026] According to an embodiment of the present disclosure, the processing unit is configured to calculate the product of the input of the convolution kernel and the weight parameter of the convolution kernel, wherein the first input data includes the input of the convolution kernel, and the second input data includes the weight parameter of the convolution kernel; and / or the processing unit is configured to calculate the product of the input of the neuron and the weight parameter of the neuron, wherein the first input data includes the input of the neuron, and the second input data includes the weight parameter of the neuron.

[0027] According to an embodiment of the present disclosure, when the specified quantization bit number is within the first bit number range, the processing unit is configured to calculate the product of the input of one convolution kernel and the weight parameter of the one convolution kernel; when the specified quantization bit number is within the second bit number range, the processing unit is configured to calculate the products of the inputs of two convolution kernels and the weight parameters of the two convolution kernels respectively, and the inputs of the two convolution kernels are equal.

[0028] According to an embodiment of the present disclosure, the controller includes a configuration word pin, and the output level of the configuration word pin is determined according to the specified quantization bit number. When the specified quantization bit number is within the first bit number range, the output level of the configuration word pin is the first level; when the specified quantization bit number is within the second bit number range, the output level of the configuration word pin is the second level, and the first level is different from the second level; the output level of the configuration word pin is used to control the multiplexer to selectively output the first calculation result and the second calculation result, and / or to disable any one or more of the shifter, the adder, and the splicing register.

[0029] A second aspect of the present disclosure provides a processing unit of a runtime reconfigurable artificial intelligence chip, the processing unit is configured to process a specified sub-computation task, the specified sub-computation task includes calculating the product of first input data and second input data, and the quantization bit number corresponding to the specified sub-computation task is a specified quantization bit number. Wherein, when the specified quantization bit number is within the first bit number range, the processing unit is configured to process one specified sub-computation task; when the specified quantization bit number is within the second bit number range, the processing unit is configured to process two specified sub-computation tasks. The processing unit includes:

[0030] A truncation register for splitting the truncation register input into a high-order input sub-data and a low-order input sub-data, where: when the specified quantization bit number is within the first bit range, the truncation register input includes the second input data of the one specified sub-computation task; when the specified quantization bit number is within the second bit range, the truncation register input includes the concatenation result of the second input data of the two specified sub-computation tasks, and the high-order input sub-data and the low-order input sub-data respectively correspond to the second input data of the two specified sub-computation tasks.

[0031] A first multiplier for multiplying the first input data by the high-order input sub-data to obtain a first multiplier output.

[0032] A second multiplier for multiplying the first input data by the low-order input sub-data to obtain a second multiplier output.

[0033] A shifter for shifting the first multiplier output to obtain a shift result when the specified quantization bit number is within the first bit range.

[0034] An adder for adding the shift result and the second multiplier output to obtain a first calculation result when the specified quantization bit number is within the first bit range.

[0035] A concatenation register for concatenating the first multiplier output and the second multiplier output to obtain a second calculation result when the specified quantization bit number is within the second bit range.

[0036] A multiplexer for selectively outputting the first calculation result and the second calculation result according to the specified quantization bit number.

[0037] According to an embodiment of the present disclosure, the specified sub-computation task includes calculating the product of a first input data and a second input data; when the specified quantization bit number is within the second bit range, the first input data of the two specified sub-computation tasks is the same.

[0038] According to an embodiment of the present disclosure, the minimum value of the first bit range is greater than the maximum value of the second bit range.

[0039] According to an embodiment of the present disclosure, the range of the first number of bits is greater than 8 bits and less than or equal to 16 bits, and the range of the second number of bits is less than or equal to 8 bits; or, the range of the first number of bits is 16 bits, and the range of the second number of bits is any one of 8 bits, 4 bits, 2 bits, and 1 bit; the number of bits of the first input data is any one of 32 bits, 16 bits, 8 bits, 4 bits, 2 bits, and 1 bit.

[0040] According to an embodiment of the present disclosure, the truncation register divides the truncation register input into a high-order input sub-data and a low-order input sub-data with equal numbers of bits.

[0041] According to an embodiment of the present disclosure, the output of the first multiplier includes a reserved bit set at the highest bit of the output of the first multiplier; and / or the output of the second multiplier includes a reserved bit set at the highest bit of the output of the second multiplier.

[0042] According to an embodiment of the present disclosure, the reserved bit is used to determine whether there is an overflow in the calculation result of the corresponding multiplier.

[0043] According to an embodiment of the present disclosure, the range of the first number of bits is greater than 8 bits and less than or equal to 16 bits, the range of the second number of bits is less than or equal to 8 bits, and the number of bits of the first input data is 16 bits; the first multiplier is a 16-bit * 8-bit multiplier, and the second multiplier is a 16-bit * 8-bit multiplier; the shifter is used to shift the output of the first multiplier to the left by 8 bits.

[0044] According to an embodiment of the present disclosure, the output of the first multiplier is 25 bits, where the highest bit is a reserved bit, and the remaining 24 bits are the product of the first input data and the high-order input sub-data; the output of the second multiplier is 25 bits, where the highest bit is a reserved bit, and the remaining 24 bits are the product of the first input data and the low-order input sub-data.

[0045] According to an embodiment of the present disclosure, the splicing register is used to splice the output of the first multiplier and the output of the second multiplier to obtain a second calculation result of 50 bits.

[0046] According to an embodiment of the present disclosure, when the specified quantization number of bits is within the range of the first number of bits, the splicing register is disabled; when the specified quantization number of bits is within the range of the second number of bits, the shifter and the adder are disabled.

[0047] When the specified quantization number of bits is within the range of the first number of bits, the multiplexer outputs the first calculation result; when the specified quantization number of bits is within the range of the second number of bits, the multiplexer outputs the second calculation result.

[0048] According to an embodiment of the present disclosure, the processing unit is configured to calculate the product of the input of the convolutional kernel and the weight parameter of the convolutional kernel, wherein the first input data includes the input of the convolutional kernel, and the second input data includes the weight parameter of the convolutional kernel; and / or the processing unit is configured to calculate the product of the input of the neuron and the weight parameter of the neuron, wherein the first input data includes the input of the neuron, and the second input data includes the weight parameter of the neuron.

[0049] According to an embodiment of the present disclosure, when the specified quantization bit number is within the first bit range, the processing unit is configured to calculate the product of the input of one convolutional kernel and the weight parameter of the one convolutional kernel; when the specified quantization bit number is within the second bit range, the processing unit is configured to calculate the products of the inputs of two convolutional kernels and the weight parameters of the two convolutional kernels respectively, and the inputs of the two convolutional kernels are equal.

[0050] According to an embodiment of the present disclosure, the processing unit includes a connection pin, and the connection pin is configured to be connected to the configuration word pin of the controller of the processing unit. The output level of the configuration word pin is determined according to the specified quantization bit number. When the specified quantization bit number is within the first bit range, the output level of the configuration word pin is the first level, and when the specified quantization bit number is within the second bit range, the output level of the configuration word pin is the second level, and the first level is different from the second level; the level of the connection pin is used to control the multiplexer to selectively output the first calculation result and the second calculation result, and / or to disable any one or more of the shifter, the adder, and the splicing register.

[0051] A third aspect of the present disclosure provides a method for executing a computing task of a runtime reconfigurable artificial intelligence chip. The runtime reconfigurable artificial intelligence chip includes a controller and a plurality of processing units. The processing unit includes a truncation register, a first multiplier, a second multiplier, a shifter, an adder, a splicing register, and a multiplexer. The method includes:

[0052] Obtaining a computing task through the controller, parsing the computing task to obtain a plurality of sub-computing tasks and the quantization bit numbers corresponding to the sub-computing tasks, and allocating the sub-computing tasks to corresponding processing units for processing, wherein the sub-computing tasks include calculating the product of first input data and second input data, and the quantization bit numbers include the quantization bit numbers of the second input data;

[0053] Processing a specified sub-computation task among the multiple sub-computation tasks by the processing unit, where the quantization bit number corresponding to the specified sub-computation task is a specified quantization bit number. Among them, when the specified quantization bit number is within the first bit range, the processing unit processes one specified sub-computation task. When the specified quantization bit number is within the second bit range, the processing unit processes two specified sub-computation tasks. Processing the specified sub-computation task among the multiple sub-computation tasks includes:

[0054] Splitting the register input into a high-order input sub-data and a low-order input sub-data by intercepting the register pair, where: when the specified quantization bit number is within the first bit range, the register input includes the second input data of the one specified sub-computation task; when the specified quantization bit number is within the second bit range, the register input includes the concatenation result of the second input data of the two specified sub-computation tasks, and the high-order input sub-data and the low-order input sub-data respectively correspond to the second input data of the two specified sub-computation tasks;

[0055] Multiplying the first input data by the high-order input sub-data by a first multiplier to obtain a first multiplier output;

[0056] Multiplying the first input data by the low-order input sub-data by a second multiplier to obtain a second multiplier output;

[0057] When the specified quantization bit number is within the first bit range, shifting the first multiplier output by a shifter to obtain a shift result, and adding the shift result and the second multiplier output by an adder to obtain a first calculation result;

[0058] When the specified quantization bit number is within the second bit range, concatenating the first multiplier output and the second multiplier output by a concatenation register to obtain a second calculation result;

[0059] Selectively outputting the first calculation result and the second calculation result by a multiplexer according to the specified quantization bit number.

[0060] According to an embodiment of the present disclosure, the controller includes a configuration word pin, and the output level of the configuration word pin is determined according to the specified quantization bit number. When the specified quantization bit number is within the first bit range, the output level of the configuration word pin is a first level. When the specified quantization bit number is within the second bit range, the output level of the configuration word pin is a second level, and the first level is different from the second level. The method further includes:

[0061] The output level of the configuration word pin is used to control the multiplexer to selectively output the first calculation result and the second calculation result, and / or to disable any one or more of the shifter, the adder, and the splicing register.

[0062] According to an embodiment of the present disclosure, the specified sub-computation task includes calculating the product of a first input data and a second input data; when the specified quantization bit number is within the second bit number range, the first input data of the two specified sub-computation tasks is the same.

[0063] According to an embodiment of the present disclosure, the minimum value of the first bit number range is greater than the maximum value of the second bit number range, that is, the first bit number range and the second bit number range have no intersection.

[0064] According to an embodiment of the present disclosure, the first bit number range is greater than 8 bits and less than or equal to 16 bits, and the second bit number range is less than or equal to 8 bits; or, the first bit number range is 16 bits, and the second bit number range is any one of 8 bits, 4 bits, 2 bits, and 1 bit. The bit number of the first input data is any one of 32 bits, 16 bits, 8 bits, 4 bits, 2 bits, and 1 bit.

[0065] According to an embodiment of the present disclosure, the truncation register divides the truncation register input into a high-order input sub-data and a low-order input sub-data with equal bit numbers.

[0066] According to an embodiment of the present disclosure, the first multiplier output includes a reserved bit set at the highest bit of the first multiplier output; and / or the second multiplier output includes a reserved bit set at the highest bit of the second multiplier output, and the reserved bit is used to determine whether there is an overflow in the calculation result of the corresponding multiplier.

[0067] According to an embodiment of the present disclosure, the first bit number range is greater than 8 bits and less than or equal to 16 bits, the second bit number range is less than or equal to 8 bits, and the bit number of the first input data is 16 bits: the first multiplier is a 16-bit * 8-bit multiplier, and the second multiplier is a 16-bit * 8-bit multiplier; the shifter is used to shift the first multiplier output left by 8 bits.

[0068] According to an embodiment of the present disclosure, the first multiplier output is 25 bits, where the highest bit is a reserved bit, and the remaining 24 bits are the product of the first input data and the high-order input sub-data; the second multiplier output is 25 bits, where the highest bit is a reserved bit, and the remaining 24 bits are the product of the first input data and the low-order input sub-data. The splicing register is used to splice the first multiplier output and the second multiplier output to obtain a 50-bit second calculation result.

[0069] According to an embodiment of the present disclosure, the method further includes: when the specified quantization bit number is within the first bit range, disabling the stitching register; when the specified quantization bit number is within the second bit range, disabling the shifter and the adder.

[0070] According to an embodiment of the present disclosure, when the specified quantization bit number is within the first bit range, the multiplexer outputs the first calculation result; when the specified quantization bit number is within the second bit range, the multiplexer outputs the second calculation result.

[0071] According to an embodiment of the present disclosure, the processing unit is configured to calculate the product of the input of the convolution kernel and the weight parameter of the convolution kernel, wherein the first input data includes the input of the convolution kernel, and the second input data includes the weight parameter of the convolution kernel; and / or the processing unit is configured to calculate the product of the input of the neuron and the weight parameter of the neuron, wherein the first input data includes the input of the neuron, and the second input data includes the weight parameter of the neuron.

[0072] According to an embodiment of the present disclosure, when the specified quantization bit number is within the first bit range, the processing unit is configured to calculate the product of the input of one convolution kernel and the weight parameter of the one convolution kernel; when the specified quantization bit number is within the second bit range, the processing unit is configured to calculate the product of the input of each of the two convolution kernels and the weight parameter of each of the two convolution kernels, and the inputs of the two convolution kernels are equal.

[0073] A fourth aspect of the present disclosure provides an electronic device, including the runtime reconfigurable artificial intelligence chip according to any one of the above, or including the processing unit according to any one of the above.

[0074] According to an embodiment of the present disclosure, during the execution of an artificial intelligence computing task, the quantization bit numbers of the respective processing units in the artificial intelligence chip can be flexibly configured according to the needs of the computing task, solving the technical problem in the prior art that the quantization bit numbers of the processing units cannot be adaptively set according to the actual computing requirements, improving the computing resource utilization rate and computing efficiency, enhancing the chip versatility, and being conducive to the wide promotion and application of the chip.

[0075] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0076] In combination with the drawings, through the following detailed description of non-limiting embodiments, other features, objects, and advantages of the present disclosure will become more apparent. In the drawings:

[0077] Figure 1Shows a structural block diagram of a runtime reconfigurable artificial intelligence chip according to an embodiment of the present disclosure.

[0078] Figure 2 Shows a structural block diagram of one processing unit in a runtime reconfigurable artificial intelligence chip according to an embodiment of the present disclosure.

[0079] Figure 3 Shows a flowchart of a method for executing a computing task for a runtime reconfigurable artificial intelligence chip according to an embodiment of the present disclosure.

[0080] Figure 4 Shows a flowchart of processing a specified sub-computing task among the multiple sub-computing tasks according to an embodiment of the present disclosure.

[0081] Figure 5 Shows a specific example of a method for executing a computing task for a runtime reconfigurable artificial intelligence chip according to an embodiment of the present disclosure.

[0082] Figure 6 Shows another specific example of a method for executing a computing task for a runtime reconfigurable artificial intelligence chip according to an embodiment of the present disclosure. Detailed implementation manners

[0083] In the following, exemplary embodiments of the present disclosure will be described in detail with reference to the accompanying drawings so that those skilled in the art can easily implement them. In addition, for clarity, parts irrelevant to the description of the exemplary embodiments are omitted in the drawings.

[0084] In the present disclosure, it should be understood that terms such as "including" or "having" are intended to indicate the existence of features, numbers, steps, actions, components, parts, or combinations thereof disclosed in this specification, and do not intend to exclude the possibility of the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.

[0085] In addition, it should be noted that, without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other. The present disclosure will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0086] As mentioned above, the processing units in existing artificial intelligence chips are usually designed for fixed quantization bits and can only perform calculations on weight parameters with fixed quantization bits. For example, they can only calculate 16-bit weight parameters or only calculate 8-bit weight parameters. The universality of the chips is poor, which increases the manufacturing and use costs of the chips and is not conducive to the popularization and application of the chips. On the other hand, existing artificial intelligence chips cannot adaptively set the quantization bits according to actual computing requirements, resulting in inflexible computing resource allocation and affecting computing efficiency.

[0087] The present disclosure provides a runtime reconfigurable artificial intelligence chip, which is characterized in that it includes a controller and a plurality of processing units, wherein: the controller is configured to obtain a computing task, parse the computing task to obtain a plurality of sub-computing tasks and the quantization bits corresponding to the sub-computing tasks, the sub-computing tasks include calculating the product of a first input data and a second input data, and the quantization bits include the quantization bits of the second input data; the processing unit is configured to process a specified sub-computing task among the plurality of sub-computing tasks, and the quantization bits corresponding to the specified sub-computing task are specified quantization bits, wherein when the specified quantization bits are within a first bit range, the processing unit is configured to process one specified sub-computing task, and when the specified quantization bits are within a second bit range, the processing unit is configured to process two specified sub-computing tasks, and the processing unit includes: an intercept register, configured to split the intercept register input into a high-order input sub-data and a low-order input sub-data, wherein: when the specified quantization bits are within the first bit range, the intercept register input includes the second input data of the one specified sub-computing task; when the specified quantization bits are within the second bit range, the intercept register input includes the concatenation result of the second input data of the two specified sub-computing tasks, and the high-order input sub-data and the low-order input sub-data respectively correspond to the second input data of the two specified sub-computing tasks; a first multiplier, the first multiplier is configured to multiply the first input data by the high-order input sub-data to obtain a first multiplier output; a second multiplier, the second multiplier is configured to multiply the first input data by the low-order input sub-data to obtain a second multiplier output; a shifter, the shifter is configured to, when the specified quantization bits are within the first bit range, shift the first multiplier output to obtain a shift result; an adder, the adder is configured to, when the specified quantization bits are within the first bit range, add the shift result and the second multiplier output to obtain a first calculation result; a concatenation register, the concatenation register is configured to, when the specified quantization bits are within the second bit range, concatenate the first multiplier output and the second multiplier output to obtain a second calculation result; a multiplexer, the multiplexer is configured to selectively output the first calculation result and the second calculation result according to the specified quantization bits.

[0088] According to the embodiments of the present disclosure, when performing artificial intelligence computing tasks, the quantization bits of each processing unit in the artificial intelligence chip can be flexibly configured according to the needs of the tasks, solving the technical problem in the prior art that the quantization bits of the processing unit cannot be adaptively set according to the actual computing requirements, improving the computing resource utilization rate and computing efficiency, enhancing the chip versatility, and being conducive to the wide-range popularization and application of the chip.

[0089] Figure 1 A structural block diagram of a runtime reconfigurable artificial intelligence chip according to an embodiment of the present disclosure is shown.

[0090] As Figure 1 shown, the runtime reconfigurable artificial intelligence chip according to an embodiment of the present disclosure includes a controller and a plurality of processing units (PEs). The controller is configured to obtain a computing task, parse the computing task to obtain a plurality of sub-computing tasks and the quantization bits corresponding to the sub-computing tasks. The sub-computing tasks include calculating the product of a first input data and a second input data, and the quantization bits include the quantization bits of the second input data.

[0091] According to an embodiment of the present disclosure, the computing task may be an artificial intelligence computing task, such as a computing task involved in the training or inference of a neural network, such as a convolution operation, a neuron calculation, etc., but not limited thereto. The controller parses the computing task to obtain a plurality of sub-computing tasks and the quantization bits corresponding to the sub-computing tasks.

[0092] For example, when the computing task is a convolution operation, the sub-computing tasks obtained by parsing the computing task may include multiplying the input of the convolution kernel by the weight parameters of the convolution kernel. The first input data of the sub-computing task includes the input of the convolution kernel, the second input data includes the weight parameters of the convolution kernel, and the quantization bits corresponding to the sub-computing task are the quantization bits of the weight parameters of the convolution kernel. When the computing task is a neuron calculation, the sub-computing tasks obtained by parsing the computing task may include multiplying the input of the neuron by the weight parameters of the neuron. The first input data of the sub-computing task includes the input of the neuron, the second input data includes the weight parameters of the neuron, and the quantization bits corresponding to the sub-computing task are the quantization bits of the weight parameters of the neuron. It can be understood that the processing units in the artificial intelligence chip according to an embodiment of the present disclosure can be used to perform various multiplication operations, not limited to the multiplication operations in the above examples.

[0093] According to an embodiment of the present disclosure, the processing unit is configured to process a specified sub-computing task among the plurality of sub-computing tasks. The quantization bits corresponding to the specified sub-computing task are specified quantization bits. When the specified quantization bits are within a first bit range, the processing unit is configured to process one specified sub-computing task. When the specified quantization bits are within a second bit range, the processing unit is configured to process two specified sub-computing tasks with the same first input data. The minimum value of the first bit range is greater than the maximum value of the second bit range.

[0094] According to an embodiment of the present disclosure, the range of the first number of bits can be greater than 8 bits and less than or equal to 16 bits, and the range of the second number of bits can be less than or equal to 8 bits; alternatively, the range of the first number of bits can be 16 bits, and the range of the second number of bits can be any one of 8 bits, 4 bits, 2 bits, and 1 bit. According to an embodiment of the present disclosure, the number of bits of the first input data is any one of 32 bits, 16 bits, 8 bits, 4 bits, 2 bits, and 1 bit.

[0095] In a specific example, the number of bits of the first input data is 16 bits, the range of the first number of bits is greater than 8 bits and less than or equal to 16 bits, and the range of the second number of bits is less than or equal to 8 bits. In this case, when the specified quantization number of bits is 16 bits, the corresponding processing unit processes one specified sub-computation task, and when the specified quantization number of bits is any one of 8 bits, 4 bits, 2 bits, and 1 bit, the corresponding processing unit processes two specified sub-computation tasks, and the first input data of the two specified sub-computation tasks is the same.

[0096] In Figure 1 it, the controller includes a plurality of configuration word pins respectively connected to each processing unit. After the controller parses to obtain a plurality of sub-computation tasks, it distributes the plurality of sub-computation tasks to the processing unit for execution, and configures the quantization number of bits of the corresponding processing unit through the configuration word pins according to the quantization number of bits corresponding to each sub-computation task, so as to achieve adaptive configuration of the quantization number of bits of the processing unit according to the needs of the computation task during runtime. According to an embodiment of the present disclosure, the output level of the configuration word pin is determined according to the quantization number of bits of the corresponding processing unit. When the quantization number of bits is within the range of the first number of bits, the output level of the configuration word pin is the first level, and when the quantization number of bits is within the range of the second number of bits, the output level of the configuration word pin is the second level, and the first level is different from the second level.

[0097] Figure 2 The structural block diagram of a processing unit in a runtime reconfigurable artificial intelligence chip according to an embodiment of the present disclosure is shown.

[0098] As Figure 2 shown, the processing unit according to an embodiment of the present disclosure includes: a truncation register, a first multiplier, a second multiplier, a shifter, an adder, a stitching register, and a multiplexer.

[0099] The truncation register divides the truncation register input to obtain a high-order input sub-data and a low-order input sub-data. Among them, when the specified quantization bit number is within the first bit range, the truncation register input includes the second input data of the one specified sub-computation task; when the specified quantization bit number is within the second bit range, the truncation register input includes the concatenation result of the second input data of the two specified sub-computation tasks, and the high-order input sub-data and the low-order input sub-data respectively correspond to the second input data of the two specified sub-computation tasks.

[0100] The first multiplier is used to multiply the first input data by the high-order input sub-data to obtain a first multiplier output, and the second multiplier is used to multiply the first input data by the low-order input sub-data to obtain a second multiplier output. The number of bits of the first input data and the high-order input sub-data matches the input bit number of the first multiplier, and the number of bits of the first input data and the low-order input sub-data matches the input bit number of the second multiplier.

[0101] Taking the first input data as 16 bits, the first bit range as greater than 8 bits and less than or equal to 16 bits, the second bit range as less than or equal to 8 bits, the high-order input sub-data and the low-order input sub-data both as 8 bits, and the first multiplier and the second multiplier both as 16-bit * 8-bit multipliers as an example, the operations of the truncation register and the multiplier are described below.

[0102] When the specified quantization bit number is 16 bits, the specified quantization bit number is within the first bit range, the truncation register input is the 16-bit second input data, and the truncation register divides the second input data into an 8-bit high-order input sub-data and an 8-bit low-order input sub-data.

[0103] When the specified quantization bit number is greater than 8 bits and less than 16 bits, the specified quantization bit number is within the first bit range, the truncation register input is the 16-bit data obtained by padding "0" in front of the second input data, and the truncation register divides the 16-bit data into an 8-bit high-order input sub-data and an 8-bit low-order input sub-data.

[0104] When the specified quantization bit number is 8 bits, the specified quantization bit number is within the second bit range, the truncation register input is the concatenation result of two second input data, totaling 16 bits, and the truncation register divides the 16-bit concatenation result into an 8-bit high-order input sub-data and an 8-bit low-order input sub-data.

[0105] When the specified quantization bit number is less than 8 bits, the specified quantization bit number is within the second bit range, the truncation register input is the 16-bit concatenation result obtained by padding "0" to 8 bits in front of the two second input data respectively and then concatenating them, and the truncation register divides the 16-bit concatenation result into an 8-bit high-order input sub-data and an 8-bit low-order input sub-data.

[0106] The first multiplier multiplies 16-bit first input data and 8-bit high-order input sub-data to obtain a first multiplier output, and the second multiplier multiplies 16-bit first input data and 8-bit low-order input sub-data to obtain a second multiplier output.

[0107] When the specified quantization bit number is within the first bit range, the shifter shifts the first multiplier output to obtain a shift result, and the adder adds the shift result and the second multiplier output to obtain a first calculation result. The shift number of the shifter is set such that the first calculation result output by the adder is equal to the product of the first input data and the second input data.

[0108] When the first input data is 16 bits, the first bit range is greater than 8 bits and less than or equal to 16 bits, the second bit range is less than or equal to 8 bits, the high-order input sub-data and the low-order input sub-data are both 8 bits, and the first multiplier and the second multiplier are both 16-bit * 8-bit multipliers, the first multiplier output is 24 bits, the second multiplier output is 24 bits, and the shifter shifts the first multiplier output left by 8 bits and then adds it to the second multiplier output to obtain a 32-bit first calculation result.

[0109] According to an embodiment of the present disclosure, a reserved bit can be set at the highest bit of each of the first multiplier output and the second multiplier output to determine whether there is an overflow in the calculation result of the corresponding multiplier. For example, when there is an overflow in the calculation result of the multiplier, the reserved bit is 1, indicating a calculation error, otherwise the reserved bit is 0. In this case, each of the first multiplier output and the second multiplier output is 25 bits, and the first calculation result is 33 bits.

[0110] When the specified quantization bit number is within the second bit range, the stitching register stitches the first multiplier output and the second multiplier output to obtain a second calculation result. The second calculation result contains the products of the first input data and the second input data of the respective two specified calculation tasks.

[0111] When the first input data is 16 bits, the first bit range is greater than 8 bits and less than or equal to 16 bits, the second bit range is less than or equal to 8 bits, the high-order input sub-data and the low-order input sub-data are both 8 bits, and the first multiplier and the second multiplier are both 16-bit * 8-bit multipliers, the first multiplier output is 24 bits, the second multiplier output is 24 bits, and the stitching result is 48 bits.

[0112] When a reserved bit is set at the highest bit of each of the first multiplier output and the second multiplier output, each of the first multiplier output and the second multiplier output is 25 bits, and the second calculation result is 50 bits.

[0113] The multiplexer selectively outputs the first calculation result and the second calculation result according to the specified quantization bits. When the specified quantization bits are within the first bit range, the multiplexer outputs the first calculation result. When the specified quantization bits are within the second bit range, the multiplexer outputs the second calculation result. The output level of the configuration word pin of the controller can be used to control the multiplexer to selectively output the first calculation result and the second calculation result.

[0114] According to the embodiments of the present disclosure, not only can the processing unit be flexibly configured at runtime for the quantization bits of the task, but also when the quantization bits are within the second bit range, two multiplication operations with the same first input data can be executed simultaneously, improving the versatility and computing efficiency of the chip.

[0115] According to the embodiments of the present disclosure, to reduce the chip power consumption, when the specified quantization bits are within the first bit range, the stitching register is disabled; when the specified quantization bits are within the second bit range, the shifter and the adder are disabled. The output level of the configuration word pin of the controller can be used to disable any one or more of the shifter, the adder, and the stitching register.

[0116] Figure 3 The flowchart of the calculation task execution method for a runtime reconfigurable artificial intelligence chip according to an embodiment of the present disclosure is shown. The runtime reconfigurable artificial intelligence chip includes a controller and multiple processing units. The processing unit includes an intercept register, a first multiplier, a second multiplier, a shifter, an adder, a stitching register, and a multiplexer. According to the embodiments of the present disclosure, the runtime reconfigurable artificial intelligence chip can be, for example, the runtime reconfigurable artificial intelligence chip referred to above in Figure 1 and Figure 2 described.

[0117] As Figure 3 shown, the calculation task execution method includes steps S1 - S2.

[0118] In step S1, the controller obtains a calculation task, parses the calculation task to obtain multiple sub - calculation tasks and the quantization bits corresponding to the sub - calculation tasks, and assigns the sub - calculation tasks to the corresponding processing units for processing. Among them, the sub - calculation task includes calculating the product of the first input data and the second input data, and the quantization bits include the quantization bits of the second input data.

[0119] In step S2, the processing unit processes a specified sub-computation task among the multiple sub-computation tasks, and the quantization bit number corresponding to the specified sub-computation task is a specified quantization bit number. Among them, when the specified quantization bit number is within the first bit range, the processing unit processes one specified sub-computation task; when the specified quantization bit number is within the second bit range, the processing unit processes two specified sub-computation tasks.

[0120] Figure 4 The flowchart shows the processing of a specified sub-computation task among the multiple sub-computation tasks according to an embodiment of the present disclosure. As Figure 4 shown, the processing of the specified sub-computation task among the multiple sub-computation tasks includes steps S21 - S24.

[0121] In step S21, the register input is segmented by intercepting the register pair to obtain a high-order input sub-data and a low-order input sub-data. Among them: when the specified quantization bit number is within the first bit range, the register input includes the second input data of the one specified sub-computation task; when the specified quantization bit number is within the second bit range, the register input includes the concatenation result of the second input data of the two specified sub-computation tasks, and the high-order input sub-data and the low-order input sub-data respectively correspond to the second input data of the two specified sub-computation tasks.

[0122] In step S22, the first multiplier multiplies the first input data by the high-order input sub-data to obtain a first multiplier output; the second multiplier multiplies the first input data by the low-order input sub-data to obtain a second multiplier output.

[0123] In step S23, when the specified quantization bit number is within the first bit range, the shifter shifts the first multiplier output to obtain a shift result, and the adder adds the shift result to the second multiplier output to obtain a first calculation result; when the specified quantization bit number is within the second bit range, the concatenation register concatenates the first multiplier output and the second multiplier output to obtain a second calculation result.

[0124] In step S24, the multiplexer selectively outputs the first calculation result and the second calculation result according to the specified quantization bit number.

[0125] According to an embodiment of the present disclosure, the controller includes a configuration word pin, and an output level of the configuration word pin is determined according to a specified quantization bit number. When the specified quantization bit number is within the first bit number range, the output level of the configuration word pin is a first level. When the specified quantization bit number is within the second bit number range, the output level of the configuration word pin is a second level. The first level is different from the second level. The method further includes: controlling, by the output level of the configuration word pin, the multiplexer to selectively output the first calculation result and the second calculation result, and / or disabling any one or more of the shifter, the adder, and the stitching register, so as to reduce the chip operation power consumption.

[0126] According to an embodiment of the present disclosure, the specified sub-computation task includes calculating a product of a first input data and a second input data; when the specified quantization bit number is within the second bit number range, the first input data of the two specified sub-computation tasks is the same.

[0127] According to an embodiment of the present disclosure, a minimum value of the first bit number range is greater than a maximum value of the second bit number range, that is, the first bit number range and the second bit number range have no intersection.

[0128] According to an embodiment of the present disclosure, the first bit number range is greater than 8 bits and less than or equal to 16 bits, and the second bit number range is less than or equal to 8 bits; or, the first bit number range is 16 bits, and the second bit number range is any one of 8 bits, 4 bits, 2 bits, and 1 bit. The number of bits of the first input data is any one of 32 bits, 16 bits, 8 bits, 4 bits, 2 bits, and 1 bit.

[0129] According to an embodiment of the present disclosure, the truncation register divides the truncation register input into a high-order input sub-data and a low-order input sub-data with equal number of bits.

[0130] According to an embodiment of the present disclosure, the first multiplier output includes a reserved bit set at the highest bit of the first multiplier output; and / or the second multiplier output includes a reserved bit set at the highest bit of the second multiplier output. The reserved bit is used to determine whether there is an overflow in the calculation result of the corresponding multiplier. For example, when there is an overflow in the calculation result of the multiplier, the reserved bit is 1, indicating a calculation error, otherwise the reserved bit is 0.

[0131] According to an embodiment of the present disclosure, the first bit number range is greater than 8 bits and less than or equal to 16 bits, the second bit number range is less than or equal to 8 bits, and the number of bits of the first input data is 16 bits: the first multiplier is a 16-bit * 8-bit multiplier, and the second multiplier is a 16-bit * 8-bit multiplier; the shifter is used to shift the first multiplier output left by 8 bits.

[0132] According to an embodiment of the present disclosure, the output of the first multiplier is 25 bits, where the highest bit is a reserved bit, and the remaining 24 bits are the product of the first input data and the high-order input sub-data; the output of the second multiplier is 25 bits, where the highest bit is a reserved bit, and the remaining 24 bits are the product of the first input data and the low-order input sub-data. The splicing register is used to splice the output of the first multiplier and the output of the second multiplier to obtain a 50-bit second calculation result.

[0133] According to an embodiment of the present disclosure, the method further includes: when the specified quantization bit number is within the first bit range, disabling the splicing register; when the specified quantization bit number is within the second bit range, disabling the shifter and the adder.

[0134] According to an embodiment of the present disclosure, when the specified quantization bit number is within the first bit range, the multiplexer outputs the first calculation result; when the specified quantization bit number is within the second bit range, the multiplexer outputs the second calculation result.

[0135] According to an embodiment of the present disclosure, the processing unit is used to calculate the product of the input of the convolution kernel and the weight parameter of the convolution kernel, where the first input data includes the input of the convolution kernel, and the second input data includes the weight parameter of the convolution kernel; and / or the processing unit is used to calculate the product of the input of the neuron and the weight parameter of the neuron, where the first input data includes the input of the neuron, and the second input data includes the weight parameter of the neuron.

[0136] According to an embodiment of the present disclosure, when the specified quantization bit number is within the first bit range, the processing unit is used to calculate the product of the input of one convolution kernel and the weight parameter of the one convolution kernel; when the specified quantization bit number is within the second bit range, the processing unit is used to calculate the product of the inputs of the two convolution kernels and the weight parameters of the two convolution kernels respectively, and the inputs of the two convolution kernels are equal.

[0137] Next, refer to Figure 5 and Figure 6 to describe a specific example of a method for executing a computing task of a runtime reconfigurable artificial intelligence chip according to an embodiment of the present disclosure.

[0138] In this specific example, the first input data is 16 bits, the first bit range is greater than 8 bits and less than or equal to 16 bits, the second bit range is less than or equal to 8 bits, the high-order input sub-data and the low-order input sub-data are both 8 bits, and the first multiplier and the second multiplier are both 16-bit * 8-bit multipliers.

[0139] In Figure 5In it, assume that the first input data assigned to the processing unit for processing the specified sub-computation task is 16 bits, the second input data is 16 bits, and the specified quantization bit number is 16 bits. Since the specified quantization bit number is within the first bit range, the processing unit processes one specified sub-computation task and calculates the product of the first input data and the second input data of the specified computation task.

[0140] As Figure 5 shown, the 16-bit second input data is split into an 8-bit high-order input sub-data and an 8-bit low-order input sub-data, and are respectively input into the first multiplier and the second multiplier.

[0141] The first multiplier calculates the product of the 16-bit first input data and the 8-bit high-order input sub-data, obtaining a 25-bit first multiplier output, where the highest bit is a reserved bit for detecting overflow, and the remaining 24 bits are the product result. When an overflow occurs in the multiplication calculation, the reserved bit is 1, otherwise the reserved bit is 0.

[0142] The second multiplier calculates the product of the 16-bit first input data and the 8-bit low-order input sub-data, obtaining a 25-bit second multiplier output, where the highest bit is a reserved bit for detecting overflow, and the remaining 24 bits are the product result. When an overflow occurs in the multiplication calculation, the reserved bit is 1, otherwise the reserved bit is 0.

[0143] The 25-bit first multiplier output is shifted left by 8 bits by a shifter, obtaining a 33-bit shift result, which is added to the 25-bit second multiplier output at an adder to obtain a 33-bit first calculation result.

[0144] The selector selects and outputs according to the value of the configuration word. When the value of the configuration word is 1, the first calculation result is output as the output result of the specified sub-computation task.

[0145] In Figure 6 it, assume that the first input data assigned to the processing unit for processing the specified sub-computation task is 16 bits, the second input data is 8 bits, and the specified quantization bit number is 8 bits. Since the specified quantization bit number is within the second bit range, the processing unit processes two specified sub-computation tasks and calculates the product of the first input data and the second input data of the two specified computation tasks.

[0146] As Figure 6 shown, after the two 8-bit second input data are concatenated, they are split by an intercept register into an 8-bit high-order input sub-data and an 8-bit low-order input sub-data, corresponding to the two 8-bit second input data respectively. The 8-bit high-order input sub-data and the 8-bit low-order input sub-data are respectively input into the first multiplier and the second multiplier.

[0147] The first multiplier calculates the product of 16-bit first input data and 8-bit high-order input sub-data, obtaining a 25-bit first multiplier output, where the highest bit is a reserved bit for detecting overflow, and the remaining 24 bits are the product result. When an overflow occurs in the multiplication calculation, the reserved bit is 1; otherwise, the reserved bit is 0.

[0148] The second multiplier calculates the product of 16-bit first input data and 8-bit low-order input sub-data, obtaining a 25-bit second multiplier output, where the highest bit is a reserved bit for detecting overflow, and the remaining 24 bits are the product result. When an overflow occurs in the multiplication calculation, the reserved bit is 1; otherwise, the reserved bit is 0.

[0149] The splicer splices the 25-bit first multiplier output and the 25-bit second multiplier output to obtain a 50-bit second calculation result.

[0150] The selector makes a selection and outputs according to the value of the configuration word. When the value of the configuration word is 0, it outputs the second calculation result as the output result of the specified sub-computation task.

[0151] The following refers to Figure 5 to illustrate the case where the specified quantization bit number is within the first bit range but less than 16 bits.

[0152] In Figure 5 it is assumed that the first input data of the specified sub-computation task assigned to the processing unit for processing is 16 bits, the second input data is 10 bits, and the specified quantization bit number is 10 bits. Since the specified quantization bit number is within the first bit range, the processing unit processes a specified sub-computation task to calculate the product of the first input data and the second input data of the specified computation task.

[0153] As Figure 5 shown, first fill 0s at the high-order part of the 10-bit second input data to obtain a 16-bit filled second input data, and then divide the 16-bit filled second input data into an 8-bit high-order input sub-data and an 8-bit low-order input sub-data, which are respectively input into the first multiplier and the second multiplier.

[0154] The first multiplier calculates the product of 16-bit first input data and 8-bit high-order input sub-data, obtaining a 25-bit first multiplier output, where the highest bit is a reserved bit for detecting overflow, and the remaining 24 bits are the product result. When an overflow occurs in the multiplication calculation, the reserved bit is 1; otherwise, the reserved bit is 0.

[0155] The second multiplier calculates the product of 16-bit first input data and 8-bit low-order input sub-data, obtaining a 25-bit second multiplier output, where the highest bit is a reserved bit for detecting overflow, and the remaining 24 bits are the product result. When an overflow occurs in the multiplication calculation, the reserved bit is 1; otherwise, the reserved bit is 0.

[0156] The 25-bit output of the first multiplier is shifted left by 8 bits by the shifter to obtain a 33-bit shifted result, which is added to the 25-bit output of the second multiplier at the adder to obtain a 33-bit first calculation result.

[0157] The selector makes a selection and outputs according to the value of the configuration word. When the value of the configuration word is 1, the first calculation result is output as the output result of the specified sub-computation task.

[0158] The following refers to Figure 6 to illustrate the case where the specified quantization bit number is within the second bit range but less than 8 bits.

[0159] In Figure 6 it is assumed that the first input data of the specified sub-computation task assigned to the processing unit for processing is 16 bits, the second input data is 4 bits, and the specified quantization bit number is 4 bits. Since the specified quantization bit number is within the second bit range, the processing unit processes two specified sub-computation tasks and calculates the product of the first input data and the second input data of the two specified computation tasks.

[0160] As Figure 6 shown, the two 4-bit second input data are first filled with 0s at the high bits respectively to obtain two 8-bit filled second input data. After the two 8-bit filled second input data are concatenated, they are divided into an 8-bit high-bit input sub-data and an 8-bit low-bit input sub-data by the truncation register, corresponding to the two 8-bit filled second input data respectively. The 8-bit high-bit input sub-data and the 8-bit low-bit input sub-data are respectively input into the first multiplier and the second multiplier.

[0161] The first multiplier calculates the product of the 16-bit first input data and the 8-bit high-bit input sub-data to obtain a 25-bit output of the first multiplier, where the highest bit is a reserved bit for detecting overflow, and the remaining 24 bits are the product result. When an overflow occurs in the multiplication calculation, the reserved bit is 1, otherwise the reserved bit is 0.

[0162] The second multiplier calculates the product of the 16-bit first input data and the 8-bit low-bit input sub-data to obtain a 25-bit output of the second multiplier, where the highest bit is a reserved bit for detecting overflow, and the remaining 24 bits are the product result. When an overflow occurs in the multiplication calculation, the reserved bit is 1, otherwise the reserved bit is 0.

[0163] The splicer concatenates the 25-bit output of the first multiplier and the 25-bit output of the second multiplier to obtain a 50-bit second calculation result.

[0164] The selector makes a selection and outputs according to the value of the configuration word. When the value of the configuration word is 0, the second calculation result is output as the output result of the specified sub-computation task.

[0165] According to an embodiment of the present disclosure, the maximum value W1 of the first digit range and the maximum value W2 of the second digit range satisfy the following relationship: W1 ≤ 2W2. The number of digits of the high-order input sub-data and the number of digits of the low-order input sub-data are both equal to the maximum value W2 of the second digit range. The number of digits of the first input of the first multiplier and the second multiplier are both equal to the maximum value of the digit range of the first input data, and the number of digits of the second input of the first multiplier and the second multiplier are both equal to W2.

[0166] When the number of digits of the second input data is within the first digit range but less than 2W2, first fill 0 at the high position to make the number of digits of the second input data up to 2W2, and then input it into the truncation register for subsequent calculation.

[0167] When the number of digits of the second input data is within the second digit range but less than W2, first fill 0 at the high position to make the number of digits of the second input data up to W2 respectively, and then splice the two second input data and input them into the truncation register for subsequent calculation.

[0168] According to an embodiment of the present disclosure, the processing units in the chip can be flexibly configured into multiple different quantization bits during runtime. For example, they can be configured to 16 bits, 8 bits, 4 bits, 2 bits, 1 bit. According to different computing task requirements, the processing capabilities of each processing unit can be fully utilized, avoiding the situation in the prior art where some processing units are task congested while some processing units are idle due to the fixed quantization bits of the processing units, and significantly improving the computing efficiency.

[0169] The embodiment of the present disclosure also provides an electronic device, which includes the above-mentioned runtime reconfigurable artificial intelligence chip, or includes the above-mentioned processing unit.

[0170] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the present disclosure is not limited to the technical solution formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the inventive concept. For example, the technical solutions formed by mutually replacing the above features with the (but not limited to) technical features with similar functions disclosed in the present disclosure.

Claims

1. A runtime reconfigurable artificial intelligence chip, characterized in that, It includes a controller and multiple processing units, where: The controller is used to obtain a computing task, parse the computing task to obtain multiple sub-computing tasks and the quantization bits corresponding to the sub-computing tasks. The sub-computing tasks include calculating the product of a first input data and a second input data, and the quantization bits include the quantization bits of the second input data; The processing unit is used to process a specified sub-computing task among the multiple sub-computing tasks. The quantization bits corresponding to the specified sub-computing task are specified quantization bits. Wherein, when the specified quantization bits are within a first bit range, the processing unit is used to process one specified sub-computing task. When the specified quantization bits are within a second bit range, the processing unit is used to process two specified sub-computing tasks. The processing unit includes: A truncation register, which is used to split the truncation register input into a high-order input sub-data and a low-order input sub-data. Wherein: when the specified quantization bits are within the first bit range, the truncation register input includes the second input data of the one specified sub-computing task; when the specified quantization bits are within the second bit range, the truncation register input includes the concatenation result of the second input data of the two specified sub-computing tasks, and the high-order input sub-data and the low-order input sub-data respectively correspond to the second input data of the two specified sub-computing tasks; A first multiplier, which is used to multiply the first input data by the high-order input sub-data to obtain a first multiplier output; A second multiplier, which is used to multiply the first input data by the low-order input sub-data to obtain a second multiplier output; A shifter, which is used to shift the first multiplier output to obtain a shift result when the specified quantization bits are within the first bit range; An adder, which is used to add the shift result and the second multiplier output to obtain a first calculation result when the specified quantization bits are within the first bit range; A concatenation register, which is used to concatenate the first multiplier output and the second multiplier output to obtain a second calculation result when the specified quantization bits are within the second bit range; A multiplexer, which is used to selectively output the first calculation result and the second calculation result according to the specified quantization bits.

2. The runtime reconfigurable artificial intelligence chip according to claim 1, wherein: The specified sub-computing task includes calculating the product of a first input data and a second input data; When the specified quantization bits are within the second bit range, the first input data of the two specified sub-computing tasks is the same.

3. The runtime reconfigurable artificial intelligence chip according to claim 1, wherein: The minimum value of the first bit range is greater than the maximum value of the second bit range.

4. The runtime reconfigurable artificial intelligence chip according to claim 1, wherein: The range of the first number is greater than 8 bits and less than or equal to 16 bits, and the range of the second number is less than or equal to 8 bits; or, the range of the first number is 16 bits, and the range of the second number is any one of 8 bits, 4 bits, 2 bits, and 1 bit; The number of bits of the first input data is any one of 32 bits, 16 bits, 8 bits, 4 bits, 2 bits, and 1 bit.

5. The runtime reconfigurable artificial intelligence chip according to claim 1, wherein: The truncation register divides the truncation register input into a high-order input sub-data and a low-order input sub-data with equal number of bits.

6. The runtime reconfigurable artificial intelligence chip according to claim 1, wherein: The output of the first multiplier includes a reserved bit set at the highest bit of the output of the first multiplier; and / or The output of the second multiplier includes a reserved bit set at the highest bit of the output of the second multiplier.

7. The runtime reconfigurable artificial intelligence chip according to claim 6, characterized in that: The reserved bit is used to determine whether there is an overflow in the calculation result of the corresponding multiplier.

8. The runtime reconfigurable artificial intelligence chip according to claim 1, wherein: The range of the first number is greater than 8 bits and less than or equal to 16 bits, the range of the second number is less than or equal to 8 bits, and the number of bits of the first input data is 16 bits: The first multiplier is a 16-bit * 8-bit multiplier, and the second multiplier is a 16-bit * 8-bit multiplier; The shifter is used to shift the output of the first multiplier left by 8 bits.

9. The runtime reconfigurable artificial intelligence chip according to claim 8, wherein: The output of the first multiplier is 25 bits, where the highest bit is a reserved bit, and the remaining 24 bits are the product of the first input data and the high-order input sub-data; The output of the second multiplier is 25 bits, where the highest bit is a reserved bit, and the remaining 24 bits are the product of the first input data and the low-order input sub-data.

10. The runtime reconfigurable artificial intelligence chip according to claim 9, wherein: The stitching register is used to stitch the output of the first multiplier and the output of the second multiplier to obtain a 50-bit second calculation result.

11. The runtime reconfigurable artificial intelligence chip according to claim 1, wherein: When the specified quantization number of bits is within the range of the first number, the stitching register is disabled; When the specified quantization number of bits is within the range of the second number, the shifter and the adder are disabled.

12. The runtime reconfigurable artificial intelligence chip according to claim 1, wherein: When the specified quantization number of bits is within the range of the first number, the multiplexer outputs the first calculation result; When the specified quantization number of bits is within the range of the second number, the multiplexer outputs the second calculation result.

13. The runtime reconfigurable artificial intelligence chip according to claim 1, wherein: The processing unit is used to calculate the product of the input of the convolution kernel and the weight parameter of the convolution kernel, wherein the first input data includes the input of the convolution kernel, and the second input data includes the weight parameter of the convolution kernel; and / or The processing unit is used to calculate the product of the input of the neuron and the weight parameter of the neuron, wherein the first input data includes the input of the neuron, and the second input data includes the weight parameter of the neuron.

14. The runtime reconfigurable artificial intelligence chip according to claim 13, wherein: When the specified quantization bit number is within the first bit number range, the processing unit is used to calculate the product of the input of one convolution kernel and the weight parameter of the one convolution kernel; When the specified quantization bit number is within the second bit number range, the processing unit is used to calculate the products of the inputs of two convolution kernels and the weight parameters of the two convolution kernels respectively, and the inputs of the two convolution kernels are equal.

15. The runtime reconfigurable artificial intelligence chip according to claim 1, wherein: The controller includes a configuration word pin, and the output level of the configuration word pin is determined according to the specified quantization bit number. When the specified quantization bit number is within the first bit number range, the output level of the configuration word pin is the first level, and when the specified quantization bit number is within the second bit number range, the output level of the configuration word pin is the second level, and the first level is different from the second level; The output level of the configuration word pin is used to control the multiplexer to selectively output the first calculation result and the second calculation result, and / or to disable any one or more of the shifter, the adder, and the splicing register.

16. A processing unit of a runtime reconfigurable artificial intelligence chip, characterized in that, The processing unit is used to process a specified sub-computation task, and the specified sub-computation task includes calculating the product of the first input data and the second input data. The quantization bit number corresponding to the specified sub-computation task is the specified quantization bit number. Among them, when the specified quantization bit number is within the first bit number range, the processing unit is used to process one specified sub-computation task, and when the specified quantization bit number is within the second bit number range, the processing unit is used to process two specified sub-computation tasks. The processing unit includes: An intercept register, which is used to split the intercept register input to obtain a high-order input sub-data and a low-order input sub-data, wherein: when the specified quantization bit number is within the first bit number range, the intercept register input includes the second input data of the one specified sub-computation task; when the specified quantization bit number is within the second bit number range, the intercept register input includes the splicing result of the second input data of the two specified sub-computation tasks, and the high-order input sub-data and the low-order input sub-data respectively correspond to the second input data of the two specified sub-computation tasks; A first multiplier, which is used to multiply the first input data by the high-order input sub-data to obtain a first multiplier output; A second multiplier, which is used to multiply the first input data by the low-order input sub-data to obtain a second multiplier output; A shifter, which is used to shift the first multiplier output to obtain a shift result when the specified quantization bit number is within the first bit number range; An adder, which is used to add the shifted result and the output of the second multiplier to obtain a first calculation result when the specified quantization bit number is within the first bit range; A splicing register, which is used to splice the output of the first multiplier and the output of the second multiplier to obtain a second calculation result when the specified quantization bit number is within the second bit range; A multiplexer, which is used to selectively output the first calculation result and the second calculation result according to the specified quantization bit number.

17. The processing unit according to claim 16, wherein: The specified sub-computation task includes calculating the product of a first input data and a second input data; When the specified quantization bit number is within the second bit range, the first input data of the two specified sub-computation tasks is the same.

18. The processing unit according to claim 16, wherein: The minimum value of the first bit range is greater than the maximum value of the second bit range.

19. The processing unit according to claim 16, wherein: The first bit range is greater than 8 bits and less than or equal to 16 bits, and the second bit range is less than or equal to 8 bits; or, the first bit range is 16 bits, and the second bit range is any one of 8 bits, 4 bits, 2 bits, and 1 bit; The number of bits of the first input data is any one of 32 bits, 16 bits, 8 bits, 4 bits, 2 bits, and 1 bit.

20. The processing unit according to claim 16, wherein: The truncation register divides the truncation register input into a high-order input sub-data and a low-order input sub-data with equal number of bits.

21. The processing unit according to claim 16, wherein: The output of the first multiplier includes a reserved bit set at the highest bit of the output of the first multiplier; and / or The output of the second multiplier includes a reserved bit set at the highest bit of the output of the second multiplier.

22. The processing unit according to claim 21, wherein: The reserved bit is used to determine whether there is an overflow in the calculation result of the corresponding multiplier.

23. The processing unit according to claim 16, wherein: The first bit range is greater than 8 bits and less than or equal to 16 bits, the second bit range is less than or equal to 8 bits, and the number of bits of the first input data is 16 bits: The first multiplier is a 16-bit * 8-bit multiplier, and the second multiplier is a 16-bit * 8-bit multiplier; The shifter is used to shift the output of the first multiplier left by 8 bits.

24. The processing unit according to claim 23, wherein: The output of the first multiplier is 25 bits, where the highest bit is a reserved bit, and the remaining 24 bits are the product of the first input data and the high-order input sub-data; The output of the second multiplier is 25 bits, where the highest bit is a reserved bit, and the remaining 24 bits are the product of the first input data and the low-order input sub-data.

25. The processing unit according to claim 24, wherein: The splicing register is used to splice the output of the first multiplier and the output of the second multiplier to obtain a 50-bit second calculation result.

26. The processing unit according to claim 16, wherein: when the specified quantization bit number is within the first bit range, the stitching register is disabled; when the specified quantization bit number is within the second bit range, the shifter and the adder are disabled.

27. The processing unit according to claim 16, wherein: when the specified quantization bit number is within the first bit range, the multiplexer outputs the first calculation result; when the specified quantization bit number is within the second bit range, the multiplexer outputs the second calculation result.

28. The processing unit according to claim 16, wherein: the processing unit is used to calculate the product of the input of the convolution kernel and the weight parameter of the convolution kernel, wherein the first input data includes the input of the convolution kernel, and the second input data includes the weight parameter of the convolution kernel; and / or the processing unit is used to calculate the product of the input of the neuron and the weight parameter of the neuron, wherein the first input data includes the input of the neuron, and the second input data includes the weight parameter of the neuron.

29. The processing unit according to claim 28, wherein: when the specified quantization bit number is within the first bit range, the processing unit is used to calculate the product of the input of one convolution kernel and the weight parameter of the one convolution kernel; when the specified quantization bit number is within the second bit range, the processing unit is used to calculate the products of the inputs of two convolution kernels and the weight parameters of the two convolution kernels respectively, and the inputs of the two convolution kernels are equal.

30. The processing unit according to claim 16, wherein: the processing unit includes a connection pin, and the connection pin is used to connect to the configuration word pin of the controller of the processing unit. The output level of the configuration word pin is determined according to the specified quantization bit number. When the specified quantization bit number is within the first bit range, the output level of the configuration word pin is the first level. When the specified quantization bit number is within the second bit range, the output level of the configuration word pin is the second level, and the first level is different from the second level; the level of the connection pin is used to control the multiplexer to selectively output the first calculation result and the second calculation result, and / or to disable any one or more of the shifter, the adder, and the stitching register.

31. A method for executing a computing task of a runtime reconfigurable artificial intelligence chip, characterized in that, The runtime reconfigurable artificial intelligence chip includes a controller and a plurality of processing units. The processing unit includes a truncation register, a first multiplier, a second multiplier, a shifter, an adder, a stitching register, and a multiplexer. The method includes: obtaining a calculation task through the controller, parsing the calculation task to obtain a plurality of sub-calculation tasks and the quantization bit numbers corresponding to the sub-calculation tasks, and allocating the sub-calculation tasks to the corresponding processing units for processing, wherein the sub-calculation tasks include calculating the product of the first input data and the second input data, and the quantization bit numbers include the quantization bit numbers of the second input data; Processing a specified sub-computation task among the multiple sub-computation tasks by the processing unit, where the quantization bit number corresponding to the specified sub-computation task is a specified quantization bit number. Among them, when the specified quantization bit number is within the first bit range, the processing unit processes one specified sub-computation task. When the specified quantization bit number is within the second bit range, the processing unit processes two specified sub-computation tasks. Processing the specified sub-computation task among the multiple sub-computation tasks includes: Splitting the register input into a high-order input sub-data and a low-order input sub-data by intercepting the register pair. Among them: when the specified quantization bit number is within the first bit range, the register input includes the second input data of the one specified sub-computation task. When the specified quantization bit number is within the second bit range, the register input includes the concatenation result of the second input data of the two specified sub-computation tasks. The high-order input sub-data and the low-order input sub-data respectively correspond to the second input data of the two specified sub-computation tasks; Multiplying the first input data by the high-order input sub-data through a first multiplier to obtain a first multiplier output; Multiplying the first input data by the low-order input sub-data through a second multiplier to obtain a second multiplier output; When the specified quantization bit number is within the first bit range, shifting the first multiplier output through a shifter to obtain a shift result, and adding the shift result to the second multiplier output through an adder to obtain a first calculation result; When the specified quantization bit number is within the second bit range, concatenating the first multiplier output and the second multiplier output through a concatenation register to obtain a second calculation result; Selectively outputting the first calculation result and the second calculation result through a multiplexer according to the specified quantization bit number.

32. The method according to claim 31, wherein The controller includes a configuration word pin, and the output level of the configuration word pin is determined according to the specified quantization bit number. When the specified quantization bit number is within the first bit range, the output level of the configuration word pin is a first level. When the specified quantization bit number is within the second bit range, the output level of the configuration word pin is a second level. The first level is different from the second level. The method further includes: Controlling the multiplexer to selectively output the first calculation result and the second calculation result through the output level of the configuration word pin, and / or disabling any one or more of the shifter, the adder, and the concatenation register.

33. An electronic device, characterized in that, Including the runtime reconfigurable artificial intelligence chip according to any one of claims 1-15, or including the processing unit according to any one of claims 16-30.

Citation Information

Patent Citations

  • Reconfigurable processing unit for deep learning

    CN114780481A

  • Neural network calculation-oriented multi-bit-width reconstruction approximate tensor multiplication and addition method and system

    CN117170623A