A computing device, integrated circuit chip, board, device, and computing method
By decomposing the large bit width data into small bit width components and using these components for calculation, the problem of low computing efficiency caused by the limitation of the processor bit width is solved, and the effect of efficient processing of large bit width data is achieved.
Patent Information
- Application Number
- CN202010610807.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-06-29
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2040-09-07
AI Technical Summary
Due to the limited processing bit width, existing processors are difficult to efficiently process data with large bit widths, resulting in low computing efficiency.
By decomposing the large bit width data into multiple small bit width components, these components are used to perform calculations instead of the large bit width data, thereby realizing the calculation of the large bit width data in scenarios where the processor bit width is limited.
Under the condition that the processor bit width is limited, it can efficiently process large bit width data, simplify neural network computing, improve computing efficiency, and is suitable for a variety of application scenarios, including neural network computing.
Smart Images

Figure CN113934678B_ABST
Abstract
Description
Technical Field
[0001] This disclosure generally relates to the field of data processing. More specifically, this disclosure relates to a computing device, an integrated circuit chip, a board, a device, and a computing method. Background Art
[0002] Currently, the data bit widths processed by different types of processors may vary. For a processor that performs operations on a specific data type, the data bit width it processes is often limited. For example, for a fixed-point arithmetic unit, the data bit width it can usually process does not exceed 16 bits, such as 16-bit integer data. However, in order to save computing costs and overheads and improve computing efficiency, how to enable a processor with a limited bit width to process more bit-width data has become a technical problem to be solved. Summary of the Invention
[0003] To at least solve the above-mentioned technical problems, this disclosure proposes a solution in multiple aspects that uses small-bit-width components (i.e., data with fewer bits) of large-bit-width data (i.e., data with more bits) to participate in calculations instead of the large-bit-width data. Through the calculation solution of this disclosure, at least two small-bit-width data can be used to represent the large-bit-width data and perform arithmetic processing instead of the large-bit-width data, so that in a scenario where the processing bit width of the processor is limited, the processor can still be used to complete the calculation for the large-bit-width data.
[0004] In a first aspect, this disclosure provides a computing device, including: an arithmetic circuit configured to: receive a plurality of data to be operated on associated with an operation instruction, where at least one data to be operated on is represented by two or more components, the at least one data to be operated on has a source data bit width, each component has its own target data bit width, and the target data bit width is less than the source data bit width; and use the two or more components to perform the operation specified by the operation instruction instead of the represented data to be operated on to obtain two or more intermediate results. The computing device further includes: a combination circuit configured to: combine the above intermediate results to obtain a final result; and a storage circuit configured to store the above intermediate results and / or the final result.
[0005] In a second aspect, this disclosure provides an integrated circuit chip including the computing device of the first aspect.
[0006] In a third aspect, this disclosure provides an integrated circuit board including the integrated circuit chip of the second aspect.
[0007] In a fourth aspect, this disclosure provides a computing device including the board of the third aspect.
[0008] In a fifth aspect, the present disclosure provides a method executed by a computing device. The method includes: receiving a plurality of data to be operated on associated with an operation instruction, where at least one of the data to be operated on is characterized by two or more components, the at least one data to be operated on has a source data bit width, each component has its own target data bit width, and the target data bit width is less than the source data bit width; performing the operation specified by the operation instruction using the two or more components instead of the characterized data to be operated on to obtain two or more intermediate results; and combining the intermediate results to obtain a final result.
[0009] Through the computing device, integrated circuit chip, board, computing device, and method provided as above, the solution of the present disclosure utilizes the small-bit-width components of large-bit-width data to participate in calculations instead of the large-bit-width data, so as to be unrestricted by the processing bit width of the processor and give full play to the computing power of the processor in artificial intelligence application scenarios such as neural network operations or other general scenarios. Further, in scenarios such as neural network operations, the solution of the present disclosure can also simplify the calculations of the neural network and improve the calculation efficiency by using at least two small-bit-width components to perform calculations instead of large-bit-width data. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] By reading the following detailed description with reference to the accompanying drawings, the above and other objects, features, and advantages of the exemplary embodiments of the present disclosure will become readily understood. In the drawings, several embodiments of the present disclosure are shown in an exemplary rather than restrictive manner, and the same or corresponding reference numerals represent the same or corresponding parts, where:
[0011] Figure 1 is a simplified block diagram showing a computing device according to an embodiment of the present disclosure;
[0012] Figure 2 is a detailed block diagram showing a computing device according to an embodiment of the present disclosure;
[0013] Figure 3 is a detailed block diagram showing a computing device according to an embodiment of the present disclosure;
[0014] Figure 4 is a detailed block diagram showing a computing device according to an embodiment of the present disclosure;
[0015] Figure 5 is a detailed block diagram showing a computing device according to an embodiment of the present disclosure;
[0016] Figure 6 is a flowchart showing the computing method of a computing device according to an embodiment of the present disclosure;
[0017] Figure 7 is a structural diagram showing a combination processing device according to an embodiment of the present disclosure; and
[0018] Figure 8 It is a schematic structural diagram of a board card according to an embodiment of the present disclosure. Detailed implementation manners
[0019] Next, the technical solutions in the embodiments of the present disclosure will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, rather than all of the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those skilled in the art without creative efforts shall fall within the protection scope of the present disclosure.
[0020] It should be understood that the terms "first" and "second" in the claims, the description and the drawings of the present disclosure are used to distinguish different objects, rather than to describe a specific order. The terms "comprising" and "including" used in the description and claims of the present disclosure indicate the presence of the described features, wholes, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or their combinations.
[0021] It should also be understood that the terms used in the description of the present disclosure herein are only for the purpose of describing specific embodiments, and are not intended to limit the present disclosure. As used in the description and claims of the present disclosure, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include the plural forms. It should be further understood that the term "and / or" used in the description and claims of the present disclosure refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0022] As used in this specification and the claims, the term "if" can be interpreted as "when" or "once" or "in response to determining" or "in response to detecting" according to the context. Similarly, the phrase "if determined" or "if detecting [the described condition or event]" can be interpreted as meaning "once determined" or "in response to determining" or "once detecting [the described condition or event]" or "in response to detecting [the described condition or event]" according to the context.
[0023] As mentioned above, in view of the problem of limited processing bit width of the processor, the present disclosure proposes a solution in multiple aspects to use small-bit-width components of large-bit-width data to participate in calculations instead of the large-bit-width data. Since at least two small-bit-width data are used to represent the large-bit-width data and perform arithmetic processing instead of the large-bit-width data, it is necessary to perform combined processing on the arithmetic results obtained using the small-bit-width data to obtain the final result. The solution of the present disclosure overcomes the obstacle of limited processor bit width by enabling a large-bit-width (e.g., 24-bit) data to be represented by at least two small-bit-width (e.g., 16-bit and 8-bit) data (or components). Further, by using small-bit-width data / components to replace the operations, the computational complexity is simplified, thereby improving the computational efficiency of, for example, neural network calculations. Even further, since the large-bit-width data operations are decomposed into multiple small-bit-width data operations, the processing circuits can perform corresponding processing in parallel, further improving the computational efficiency. The solution of the present disclosure is particularly suitable for arithmetic processing involving multiplication operations, such as multiplication or multiply-accumulate operations, and the multiply-accumulate operations may include, for example, convolution operations. Therefore, the solution of the present disclosure can be used to perform neural network operations, particularly to process weight data and neuron data to obtain the desired arithmetic results. For example, when the neural network is a convolutional neural network for images, the weight data may be convolutional kernel data, and the neuron data may be, for example, pixel data of the image or output data after the previous layer of arithmetic operations.
[0024] The specific embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.
[0025] Figure 1 FIG. 7 is a simplified block diagram showing a computing device 100 according to an embodiment of the present disclosure. In one or more embodiments, the computing device 100 can be used for arithmetic processing of large-bit-width data for various application scenarios, such as artificial intelligence applications including neural network operations or general scenarios where large-bit-width data needs to be decomposed into small-bit-width data for calculations.
[0026] As Figure 1 shown, the computing device 100 includes an arithmetic circuit 110, a combination circuit 120, and a storage circuit 130.
[0027] In some embodiments, the arithmetic circuit 110 can be configured to: receive a plurality of data to be operated on associated with an operation instruction, where at least one data to be operated on is represented by two or more components. The at least one data to be operated on has a source data bit width, and each component has its own target data bit width, where the target data bit width is less than the source data bit width.
[0028] As described above, the data bit width of the data to be operated on may exceed the processing bit width of the operation circuit. Based on this, the data to be operated on with a large bit width (source data bit width) can be decomposed into two or more components with a small bit width (target data bit width) for representation.
[0029] The decomposition of the data to be operated on can be achieved based on various existing and / or future-developed data decomposition techniques to be decomposed into two or more components.
[0030] In some embodiments, the number of components used to represent the data to be operated on can be determined at least in part based on the source data bit width of the data to be operated on and the data bit width supported by the operation circuit. In still other embodiments, the target data bit width can be determined at least in part based on the data bit width supported by the operation circuit. For example, when the data to be operated on has a 24-bit data bit width and the operation circuit supports a maximum of 16-bit data bit width, in one example, the data to be operated on can be decomposed into 2 components with unequal target data bit widths, namely: an 8-bit high-order digit component and a 16-bit low-order digit component, or a 16-bit high-order digit component and an 8-bit low-order digit component. In another example, the data to be operated on can be decomposed into 3 components with equal target data bit widths, namely: an 8-bit high-order digit component, an 8-bit middle-order digit component, and an 8-bit low-order digit component. The present disclosure has no limitation in this regard, as long as the target data bit widths of the decomposed components satisfy the processing bit width limit of the operation circuit.
[0031] The data to be operated on can be decomposed into multiple components according to the required number of components and the target data bit width of each component, where each component has a corresponding component value and a component scaling factor. Taking the decomposition of a large-bit-width data into two small-bit-width components as an example below, a possible data decomposition method will be briefly described, but those skilled in the art can understand that the present disclosure has no limitation in this regard.
[0032] In one example, the large-bit-width data is decomposed into two components: a first component and a second component. The first component can be a high-order digit component or a low-order digit component; correspondingly, the second component can be a low-order digit component or a high-order digit component.
[0033] First, based on the target data bit width of each component and / or the digit position of each component in the data before decomposition (large-bit-width data), the component scaling factor of each component can be determined. For example, when the target data bit width of the first component (the high-order digit component in this example) is n1 and the target data bit width of the second component (the low-order digit component in this example) is n2, when not considering that n2 includes a sign bit, the component scaling factor of the first component can be 2 n2-1 . In comparison, when considering that n2 includes a sign bit, the component scaling factor of the first component can be 2 n2 . Generally, the component scaling factor of the low-order digit component is defaulted to 1.
[0034] Next, using the component scaling factor of the first component, the large-bitwidth data to be decomposed can be calculated to obtain the component value of the first component. The first component is characterized by the component value and the corresponding component scaling factor.
[0035] Then, based on the large-bitwidth data to be decomposed and the component value of the first component obtained previously, the component value of the second component can be calculated. In one example, when the data bitwidth of the second component does not include a sign bit, for example, the highest bit of the data bitwidth of the second component is not a sign bit, the value of the second component can be obtained by subtracting the first component from the large-bitwidth data to be decomposed. Here, the first component is the product value of the component value of the first component and the corresponding component scaling factor.
[0036] In the above manner, the large-bitwidth data can be decomposed into two components, and each component has a corresponding component value and component scaling factor. When it is necessary to decompose the data into more than two components, the above method can be iteratively executed until the required number of components is obtained. For example, for data with a 24-bit data bitwidth, when it is determined that it needs to be decomposed into 3 components and all three components have an 8-bit data bitwidth, the data with a 24-bit data bitwidth can be first decomposed into a first component with an 8-bit data bitwidth and an intermediate second component with a 16-bit data bitwidth through the above steps. Then, the above steps are repeatedly executed for the intermediate second component with a 16-bit data bitwidth to further decompose it into a second component with an 8-bit data bitwidth and a third component with an 8-bit data bitwidth.
[0037] Those skilled in the art can understand that various processes can also be adopted to optimize the data decomposition method, and this disclosure has no limitation in this regard, as long as the decomposed components are received for the specified operation.
[0038] Continue Figure 1 , in some embodiments, the arithmetic circuit 110 can be further configured to perform the operation specified by the operation instruction using the two or more received components instead of the to-be-operated data represented, to obtain two or more intermediate results.
[0039] Specifically, the arithmetic circuit 110 can be configured to perform the specified operation on two or more components of one to-be-operated data respectively with the corresponding data of other to-be-operated data, and output the corresponding operation results to the combination circuit 120.
[0040] Depending on the specific arithmetic instruction, the other data to be operated on may include one or more data to be operated on. Each data to be operated on may have a different data bit width. When the data bit width of the data to be operated on meets the processing bit width limit of the arithmetic circuit, it may not be necessary to decompose it, but rather use the original data for the operation. On the other hand, although some data to be operated on are decomposed into multiple components, in some cases, only a certain or some of the components may be required for the operation. Therefore, in this case, the corresponding data of these other data to be operated on may include any of the following: the original data of the data to be operated on, or at least one component representing the data to be operated on.
[0041] The arithmetic circuit 110 can use the received data to perform the operation specified by the arithmetic instruction, thereby obtaining two or more intermediate results and outputting them to the combinational circuit 120. Those skilled in the art can understand that the arithmetic circuit 110 can perform the specified operation in the order of receiving these two or more components, thereby obtaining each intermediate result in sequence and outputting it to the combinational circuit 120. The order of these components can include, for example: from the high digit to the low digit, or from the low digit to the high digit.
[0042] In some embodiments, the combinational circuit 120 can be configured to: combine the intermediate results input from the arithmetic circuit 110 to obtain the final result. As described above, since at least one data to be operated on uses two or more of its components to replace the operation, therefore, the intermediate results are obtained by performing the operation using each component, and they need to be combined to obtain the final result.
[0043] In some embodiments, the combinational circuit can be further configured to perform a weighted combination of these operation results as intermediate results to obtain the final result. Since the components that participate in the operation instead of the original data to be operated on have corresponding component values and component scaling factors, and the arithmetic circuit 110 can perform the operation only using the component values to obtain the intermediate results, therefore, in the combinational circuit 120, the component scaling factors of the components participating in the operation can be considered to perform a weighted combination of the intermediate results. Subsequently, various implementations of the combinational circuit will be described in detail based on several embodiments.
[0044] The computing device 100 may further include a storage circuit 130 configured to store the above intermediate results and / or final results. As described above, since the results obtained by the arithmetic circuit 110 performing operations using the components are intermediate results, these intermediate results need to be combined. During the combination, cyclic combination can be performed according to the generation of the intermediate results, such as weighted accumulation. Therefore, the storage circuit can be used to store these intermediate results temporarily or permanently. Preferably, in some embodiments, the intermediate results and the final results can share the storage space in the storage circuit, thereby saving storage space. Those skilled in the art can understand that the storage circuit 130 can also be used to store other data and information. For example, the intermediate data that needs to be stored generated during the operation of the arithmetic circuit 110 is not limited in this disclosure.
[0045] Figure 2 FIG. 4 is a detailed block diagram showing a computing device 200 according to an embodiment of the present disclosure. As described above, the solution of the present disclosure is particularly suitable for arithmetic processing involving multiplication operations. Therefore, in this embodiment, the arithmetic circuit 210 of the computing device 200 can be specifically implemented as a multiplication circuit 211 or a multiply-accumulate circuit 212. The multiply-accumulate circuit 212 can be used, for example, to implement a convolution operation.
[0046] Since the components participating in the operation instead of the original data to be operated have corresponding component values and component scaling factors, and the component scaling factors are associated with the digit positions of the components in the data to be operated they represent, when it comes to operations such as multiplication, for example, multiplication operations or multiply-accumulate operations, the multiplication circuit 211 or the multiply-accumulate circuit 212 can perform the operation only using the component values to obtain an operation result as an intermediate result. The influence of the component scaling factors can then be processed by the combination circuit 220.
[0047] As Figure 2 shown, the combination circuit 220 may include a weighting circuit 221 and an addition circuit 222. The weighting circuit 221 can be configured to use a weighting factor to perform weighting processing on the current operation result of the arithmetic circuit 210, such as the product result of the multiplication circuit 211 or the multiply-accumulate result of the multiply-accumulate circuit 212, or the previous combination result of the combination circuit 220. Depending on the different weighting objects, the weighting factors can also be different. In some embodiments, the weighting factor is at least partially determined based on the component scaling factors of the components generating the corresponding operation results. The addition circuit 222 can be configured to accumulate the weighted result with other intermediate results to obtain a final result.
[0048] The following describes the possible implementation manners of the weighting circuit 221 in the combination circuit 220 Figure 2 for different weighting object cases respectively.
[0049] Figure 3is a detailed block diagram showing a computing device 300 according to an embodiment of the present disclosure. In this embodiment, an implementation of a weighted circuit 221 is further shown. Figure 2 In this implementation, the object of weighting is the current operation result of the operation circuit 210.
[0050] As Figure 3 shown, the weighted circuit 321 can be configured to multiply the operation result of the operation circuit 310 by a first weighting factor to obtain a weighted result. When the operation of the operation circuit 310 is a multiplication operation or a multiply-accumulate operation, the first weighting factor can be the product of the component scaling factors corresponding to the components of the operation result. Those skilled in the art can understand that for different operation results, the first weighting factor may also be different. At this time, the addition circuit 322 can be configured to accumulate the obtained weighted result with the previous addition result of the addition circuit 322.
[0051] The following takes the operation of two data as an example to further describe Figure 3 the specific implementation of the shown embodiment.
[0052] In one example, assume that the operation instruction specifies to perform a multiplication operation on data A and data B with a large bit width. Each of data A and data B has been previously decomposed into two components. For example, data A and data B can be respectively expressed as:
[0053] A = a1 * scaleA1 + a0 * scaleA0
[0054] B = b1 * scaleB1 + b0 * scaleB0
[0055] where a1 and a0 are the component values of the high-order digit component and the low-order digit component of data A respectively; scaleA1 and scaleA0 are the corresponding component scaling factors. Similarly, b1 and b0 are the component values of the high-order digit component and the low-order digit component of data B respectively; scaleB1 and scaleB0 are the corresponding component scaling factors. In this example, when using components to perform a multiplication operation instead of data A and B, 4 multiplication operations need to be performed. No matter in what order these 4 multiplication operations are performed, only by adjusting the weighting factor accordingly can the final operation result be obtained.
[0056] For example, the above multiplication operation can be expressed as:
[0057] A * B = (a1 * scaleA1 + a0 * scaleA0) * (b1 * scaleB1 + b0 * scaleB0)
[0058] =(a1 * b1) * (scaleA1 * scaleB1)+(a1 * b0) * (scaleA1 * scaleB0)+
[0059] (a0 * b1) * (scaleA0 * scaleB1)+(a0 * b0) * (scaleA0 * scaleB0)
[0060] As can be seen from the above description, the arithmetic circuit 310 can perform multiplication operations between the numerical values of each component. In this example, they are a1 * b1, a1 * b0, a0 * b1, and a0 * b0 respectively. Additionally, in this example, the corresponding first weighting factors for the above four intermediate results are: scaleA1 * scaleB1, scaleA1 * scaleB0, scaleA0 * scaleB1, and scaleA0 * scaleB0. The weighting circuit 321 weights the above intermediate results respectively using the corresponding first weighting factors. The addition circuit 322 can add the weighted intermediate results to obtain the final result.
[0061] In some embodiments, the values of some component scaling factors are 1. For example, scaleA0 or scaleB0 may be 1. In this case, when calculating the first weighting factor, the corresponding multiplication can be omitted. For example, the calculations of scaleA1 * scaleB0, scaleA0 * scaleB1, and scaleA0 * scaleB0 can be omitted.
[0062] In another example, assume that the operation instruction specifies to perform a convolution operation on data A and data B with a large bit width, where data A can be, for example, a neuron in a neural network operation, and data B can be a weight value in a neural network operation. Each of data A and data B has been pre-decomposed into two components. For example, data A and data B can be respectively expressed as:
[0063] A = a1 * scaleA1 + a0 * scaleA0
[0064] B = b1 * scaleB1 + b0 * scaleB0
[0065] Among them, a1 and a0 are respectively the component numerical values of the high-order digit component and the low-order digit component of data A; scaleA1 and scaleA0 are respectively the corresponding component scaling factors. Similarly, b1 and b0 are respectively the component numerical values of the high-order digit component and the low-order digit component of data B; scaleB1 and scaleB0 are respectively the corresponding component scaling factors. In this example, when using components to perform the convolution operation instead of data A and B, 4 convolution operations need to be performed. No matter in what order these 4 convolution operations are performed, only the weighting factors need to be adjusted accordingly to obtain the final operation result.
[0066] In one example, the operation process is shown with the operation order from the lower digits to the higher digits:
[0067] a0 (conv) b0 -> tmp0, tmp0 * W00 -> p0
[0068] a1 (conv) b0 -> tmp1, tmp1 * W10 + p0 –> p1
[0069] a0 (conv) b1 -> tmp2, tmp2 * W01 + p1 –> p2
[0070] a1 (conv) b1 -> tmp3, tmp3 * W11 + p2 –> p3
[0071] Where conv represents the convolution operation, tmp0, tmp1, tmp2, and tmp3 are the convolution results of the four convolution operations respectively, W00, W10, W01, and W11 are the corresponding weighting factors respectively, and p0, p1, p2, and p3 are the combined results after weighted combination. It can be understood that p0 is the first combined result because there is no previous combined data, so p0 directly corresponds to the weighted result.
[0072] In another example, the operation process is shown with the operation order from the higher digits to the lower digits:
[0073] a1 (conv) b1 -> tmp3, tmp3 * W11 -> p3
[0074] a0 (conv) b1 -> tmp2, tmp2 * W01 + p3 –> p2
[0075] a1 (conv) b0 -> tmp1, tmp1 * W10 + p2 –> p1
[0076] a0 (conv) b0 -> tmp0, tmp0 * W00 + p1 –> p0
[0077] Where the meanings of the symbols are the same as those above. It can be understood that p3 is the first combined result because there is no previous combined data, so p3 directly corresponds to the weighted result.
[0078] In the above two examples, the weighting factor can be the product of the component scaling factors of the corresponding convolution result components. For example,
[0079] W00 = scaleA0 * scaleB0;
[0080] W10 = scaleA1 * scaleB0;
[0081] W01 = scaleA0 * scaleB1;
[0082] W11 = scaleA1 * scaleB1.
[0083] Similarly to the above, in some embodiments, the values of some component scaling factors are 1. For example, scaleA0 or scaleB0 may be 1. In this case, when calculating the first weighting factor, the corresponding multiplication can be omitted. For example, the calculations of W00, W10, and W01 can be omitted, thereby improving the calculation efficiency.
[0084] Any one or both of data A and data B in the above two examples can be a scalar or a vector. When the data is a vector, each element in the vector is decomposed into two or more components and participates in the operation instead of these elements. Since the elements of the vector do not affect each other, the operations involving each element can be processed in parallel, thereby improving the operation efficiency.
[0085] In addition, as can be seen from the above operation process, regardless of the order of operations between the components, since the first weighting factor directly corresponds to the product of the component scaling factors of the components of each intermediate result / operation result, the final result can be obtained by directly accumulating the weighted results. Figure 3 The implementation manner is not limited by the operation order of the operation circuit 310 and / or the output order of the intermediate results.
[0086] Figure 4 shows Figure 2 Another implementation of the weighting circuit 221 is shown. In this implementation, the operation order from high digits to low digits is optimized. In this case, the weighting object of the weighting circuit is the previous operation result of the combinational circuit.
[0087] As Figure 4 shown, the weighting circuit 421 can be configured to multiply the previous addition result of the addition circuit 422 by a second weighting factor to obtain a weighted result. At this time, the second weighting factor is the ratio of the scaling factor of the previous operation result of the operation circuit 410 to the scaling factor of the current operation result, where the scaling factor of the operation result is determined by the component scaling factors corresponding to the components of the operation result. Those skilled in the art can understand that the second weighting factor may be different for each combination. At this time, the addition circuit 422 can be configured to accumulate the weighted result of the weighting circuit 421 and the current operation result of the operation circuit 410.
[0088] Taking the convolution operation of the above data A and data B as an example, the specific implementation of the Figure 4 shown embodiment is further described.
[0089] According to the operation order from high digits to low digits, the operation process is as follows:
[0090] a1(conv)b1 -> tmp3, tmp3 -> p3
[0091] a0(conv)b1 -> tmp2, tmp2 + p3 * H33 –> p2
[0092] a1(conv)b0 -> tmp1, tmp1 + p2 * H22 –> p1
[0093] a0(conv)b0 -> tmp0, tmp0 + p1 * H11 –> p0
[0094] p0 = p0 * H00
[0095] Wherein, the meanings of the symbols are the same as before, and H00, H11, H22, and H33 are the corresponding weighting factors respectively. In this example, the weighting factors can be determined as follows:
[0096] H33 = (scaleA1 * scaleB1) / (scaleA0 * scaleB1);
[0097] H22 = (scaleA0 * scaleB1) / (scaleA1 * scaleB0);
[0098] H11 = (scaleA1 * scaleB0) / (scaleA0 * scaleB0);
[0099] H00 = scaleA0 * scaleB0.
[0100] It can be seen from the above operation process that the combined result needs to be weighted again finally, and the weighting factor H00 corresponds to the scaling factor of the last operation result tmp0. At this time, the operation of the operation circuit 410 for this operation instruction has ended, and there is no current operation result. In order to unify the calculation of the weighting factor, the scaling factor of the current operation result can be set to 1, so that the weighting factor of the last weighting still corresponds to the ratio of the scaling factor of the previous operation result to the scaling factor of the current operation result.
[0101] Similarly, in some embodiments, the values of some component scaling factors are 1. For example, scaleA0 or scaleB0 may be 1. At this time, when calculating the second weighting factor, the corresponding multiplication can be omitted. For example, the calculations of scaleA1 * scaleB0, scaleA0 * scaleB1, and scaleA0 * scaleB0 can be omitted, and at the same time, the weighting of the last combined result can also be omitted, that is, p0 = p0 * H00, thereby improving the calculation efficiency.
[0102] As can be seen from the above operation process, the operation results of the component values of the high-order components gradually increase through multiple weightings. Therefore, it is possible to avoid the loss of precision that may occur during the alignment step when adding two numbers with a large difference, such as adding a very large number and a very small number.
[0103] In some embodiments, when the operation instruction is a multiplication operation or a multiply-accumulate operation instruction, if any of the data participating in the operation is zero, the result must be zero. At this time, this zero data can be omitted from participating in the calculation, and correspondingly, the current operation circuit can be turned off without performing the operation and directly output the result, thereby saving operation power consumption and also saving computing and / or storage resources.
[0104] Figure 5 A detailed block diagram of the computing device 500 according to an embodiment of the present disclosure is shown. In this embodiment, a first comparison circuit 513 is added to the operation circuit 510, and the comparison circuit 513 can be configured to determine whether any of the data to which a specified operation is to be performed is zero. It can be understood that this data can include any of the following: the original data of the data to be operated, or the component representing the data to be operated. If the data is zero, the specified operation on this data can be omitted and the operation can be directly skipped to the next data. Otherwise, continue to perform the specified operation on this data as described above.
[0105] Alternatively or additionally, a second comparison circuit 523 can be provided in the combinational circuit 520. The second comparison circuit 523 can be configured to: determine whether the received intermediate result is zero; and if the intermediate result is zero, omit the combinational processing on the intermediate result; otherwise, continue to perform the combinational processing on the intermediate result as described above. Similarly to the above, this processing method can save operation power consumption and also save computing and / or storage resources.
[0106] Figure 6 A flowchart of a computing method 600 executed by a computing device according to an embodiment of the present disclosure is shown. As described above, in one or more embodiments, the computing method 600 can be used for the operation processing of large-bitwidth data for various application scenarios, such as artificial intelligence applications including neural network operations or general scenarios that require decomposing large-bitwidth data into small-bitwidth data for calculation.
[0107] As Figure 6 shown, in step S610, a plurality of data to be operated associated with the operation instruction are received, where at least one data to be operated is represented by two or more components. The at least one data to be operated has a source data bitwidth, each component has its own target data bitwidth, and the target data bitwidth is less than the source data bitwidth.
[0108] Optionally, in some embodiments, when the arithmetic instruction involves a multiplication operation or a multiply-add operation (e.g., a convolution operation), method 600 may further include step S615. In step S615, for example, by Figure 5 the first comparison circuit 513 in to determine whether any of the data to which the operation is about to be performed is zero. The data may include any of the following: the original data of the data to be operated on, or the components representing the data to be operated on.
[0109] If none of the data is zero, method 600 proceeds to step S620, where the operation specified by the arithmetic instruction is performed using the two or more received components instead of the data to be operated on represented, to obtain two or more intermediate results.
[0110] If any of the data is zero, method 600 may skip step S620, that is, the operation specified by the arithmetic instruction is not performed using this data, and directly continue with the next operation. Because when the operation is a multiplication operation or a multiply-add operation, if any one of the operands is zero, the result will be zero. Therefore, the operation specified for the zero data can be omitted, thus saving computing resources and reducing power consumption.
[0111] Continuing with step S620, performing the specified operation may include: respectively performing the specified operation on two or more components of one data to be operated on and the corresponding data of other data to be operated on to obtain corresponding operation results. As mentioned earlier, the other data to be operated on may include one or more data to be operated on. And the corresponding data of these other data to be operated on may include any of the following: the original data of the data to be operated on, or at least one component representing the data to be operated on.
[0112] Optionally, in some embodiments, method 600 may further include step S625. In step S625, for example, by Figure 5 the second comparison circuit 523 in to determine whether the intermediate result to which the combined processing is about to be performed is zero. If the intermediate result is zero, method 600 may skip step S630, that is, the combined processing is not performed using this intermediate result, and directly continue with the combination of the next intermediate result, thereby saving computing resources and reducing power consumption.
[0113] Finally, in step S630, the intermediate results obtained in step S620 may be combined to obtain a final result. In some embodiments, combining the intermediate results may include: performing a weighted combination on the operation results output in step S620 to obtain a final result.
[0114] The calculation method 600 of the embodiments of the present disclosure is particularly applicable to operation processing involving multiplication operations, such as multiplication or multiply-accumulate operations. The multiply-accumulate operation may include, for example, a convolution operation. Since the components participating in the operation instead of the original data to be operated have corresponding component values and component scaling factors, and the component scaling factors are associated with the digit positions of the components in the data to be operated, when performing operations such as multiplication, for example, multiplication operations or multiply-accumulate operations, only the component values can be used for the operation to obtain the operation result as an intermediate result. The influence of the component scaling factors can be processed in the subsequent result combination.
[0115] For example, in some embodiments, in step S620, performing the specified operation may include: using the component values to perform the operation to obtain the operation result. Further, in step S630, performing the weighted combination may include: using a weighting factor to perform a weighted combination of the current operation result and the previous combined result, where the weighting factor is determined at least in part based on the component scaling factor of the component corresponding to the current operation result.
[0116] As mentioned above, different weighted combination methods can be adopted based on the operation order of the components, such as from low digits to high digits, or from high digits to low digits.
[0117] In some embodiments, performing the weighted combination in step S630 may include: multiplying the operation result of step S620 by a first weighting factor to obtain a weighted result, where the first weighting factor is the product of the component scaling factors of the components corresponding to the current operation result; and accumulating the weighted result with the previous combined result.
[0118] In other embodiments, performing the weighted combination in step S630 may include: multiplying the previous combined result by a second weighting factor to obtain a weighted result, where the second weighting factor is the ratio of the scaling factor of the previous operation result to the scaling factor of the current operation result, and the scaling factor of the operation result is determined by the component scaling factors of the components corresponding to the operation result; and accumulating the weighted result with the current operation result of step S620.
[0119] The calculation method performed by the calculation device of the embodiments of the present disclosure has been described above with reference to the flowchart. Those skilled in the art can understand that since the operation of large-bitwidth data is decomposed into multiple operations of small-bitwidth data, corresponding processing can be performed in parallel between the steps of the above method, further improving the calculation efficiency.
[0120] Figure 7 is a structural diagram showing a combination processing device 700 according to an embodiment of the present disclosure. As Figure 7As shown in [description not provided], the combined processing device 700 includes a computing processing device 702, an interface device 704, other processing devices 706, and a storage device 708. According to different application scenarios, the computing processing device may include one or more computing devices 710, which may be configured to perform the operations described herein in connection with the attached Figures 1 - 6 operations described.
[0121] In different embodiments, the computing processing device of the present disclosure may be configured to perform user-specified operations. In an exemplary application, the computing processing device may be implemented as a single-core artificial intelligence processor or a multi-core artificial intelligence processor. Similarly, one or more computing devices included in the computing processing device may be implemented as artificial intelligence processor cores or partial hardware structures of artificial intelligence processor cores. When multiple computing devices are implemented as artificial intelligence processor cores or partial hardware structures of artificial intelligence processor cores, for the computing processing device of the present disclosure, it may be regarded as having a single-core structure or a homogeneous multi-core structure.
[0122] In an exemplary operation, the computing processing device of the present disclosure may interact with other processing devices through the interface device to jointly complete user-specified operations. Depending on the implementation, the other processing devices of the present disclosure may include one or more types of general-purpose and / or special-purpose processors such as a central processing unit (CPU), a graphics processing unit (GPU), and an artificial intelligence processor. These processors may include, but are not limited to, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., and the number thereof may be determined according to actual needs. As described above, only for the computing processing device of the present disclosure, it may be regarded as having a single-core structure or a homogeneous multi-core structure. However, when the computing processing device and other processing devices are considered together, the two may be regarded as forming a heterogeneous multi-core structure.
[0123] In one or more embodiments, the other processing device may serve as an interface between the computing processing device of the present disclosure (which may be specifically implemented as an operation device related to artificial intelligence such as neural network operations) and external data and control, and perform basic controls including but not limited to data transfer, start-up and / or stop of the computing device. In other embodiments, the other processing device may also cooperate with the computing processing device to jointly complete computing tasks.
[0124] In one or more embodiments, the interface device can be used to transfer data and control instructions between a computing processing device and other processing devices. For example, the computing processing device can obtain input data from other processing devices via the interface device and write it into the storage device (or memory) on the chip of the computing processing device. Further, the computing processing device can obtain control instructions from other processing devices via the interface device and write them into the control cache on the chip of the computing processing device. Alternatively or optionally, the interface device can also read the data in the storage device of the computing processing device and transfer it to other processing devices.
[0125] Additionally or optionally, the combined processing device of the present disclosure may further include a storage device. As shown in the figure, the storage device is respectively connected to the computing processing device and the other processing devices. In one or more embodiments, the storage device can be used to store the data of the computing processing device and / or the other processing devices. For example, the data may be data that cannot be fully stored in the internal or on-chip storage device of the computing processing device or other processing devices.
[0126] In some embodiments, the present disclosure also discloses a chip (such as Figure 8 the chip 802 shown in Figure 7 ). In one implementation, the chip is a System on Chip (SoC) and integrates one or more combined processing devices as shown in Figure 8 . The chip can be connected to other related components through an external interface device (such as the external interface device 806 shown in Figure 8 ). The related components can be, for example, a camera, a display, a mouse, a keyboard, a network card, or a wifi interface. In some application scenarios, other processing units (such as a video codec) and / or interface modules (such as a DRAM interface) etc. can be integrated on the chip. In some embodiments, the present disclosure also discloses a chip package structure that includes the above chip. In some embodiments, the present disclosure also discloses a board that includes the above chip package structure. The following will be described in detail with reference to Figure 8 the board.
[0127] Figure 8 is a schematic structural diagram of a board 800 according to an embodiment of the present disclosure. As shown in Figure 8As shown in [the figure], the board card includes a storage device 804 for storing data, which includes one or more storage units 810. The storage device can be connected to the control device 808 and the chip 802 described above and perform data transmission through means such as a bus. Further, the board card also includes an external interface device 806, which is configured to perform data relay or transfer functions between the chip (or the chip in the chip package structure) and an external device 812 (such as a server or a computer, etc.). For example, the data to be processed can be transmitted from the external device to the chip through the external interface device. Another example is that the calculation result of the chip can be transmitted back to the external device via the external interface device. According to different application scenarios, the external interface device can have different interface forms. For example, it can adopt a standard PCIE interface, etc.
[0128] In one or more embodiments, the control device in the board card disclosed in the present disclosure can be configured to regulate the state of the chip. For this purpose, in one application scenario, the control device can include a microcontroller (Micro Controller Unit, MCU) for regulating the working state of the chip.
[0129] According to the above combination Figure 7 and Figure 8 the description of, those skilled in the art can understand that the present disclosure also discloses an electronic device or apparatus, which can include one or more of the above board cards, one or more of the above chips, and / or one or more of the above combined processing devices.
[0130] According to different application scenarios, the electronic devices or apparatuses of the present disclosure may include servers, cloud servers, server clusters, data processing devices, robots, computers, printers, scanners, tablet computers, smart terminals, PC devices, Internet of Things terminals, mobile terminals, mobile phones, dash cams, navigators, sensors, cameras, cameras, video cameras, projectors, watches, headphones, mobile storage, wearable devices, vision terminals, autonomous driving terminals, transportation means, household appliances, and / or medical devices. The transportation means include airplanes, ships, and / or vehicles; the household appliances include televisions, air conditioners, microwave ovens, refrigerators, rice cookers, humidifiers, washing machines, electric lights, gas stoves, range hoods; the medical devices include nuclear magnetic resonance instruments, B-ultrasound instruments, and / or electrocardiographs. The electronic devices or apparatuses of the present disclosure can also be applied to fields such as the Internet, Internet of Things, data centers, energy, transportation, public management, manufacturing, education, power grids, telecommunications, finance, retail, construction sites, and medical care. Further, the electronic devices or apparatuses of the present disclosure can also be used in application scenarios related to artificial intelligence, big data, and / or cloud computing, such as the cloud, edge, and terminal. In one or more embodiments, the electronic devices or apparatuses with high computing power according to the solution of the present disclosure can be applied to cloud devices (such as cloud servers), while the electronic devices or apparatuses with low power consumption can be applied to terminal devices and / or edge devices (such as smartphones or cameras). In one or more embodiments, the hardware information of cloud devices is compatible with the hardware information of terminal devices and / or edge devices, so that appropriate hardware resources can be matched from the hardware resources of cloud devices according to the hardware information of terminal devices and / or edge devices to simulate the hardware resources of terminal devices and / or edge devices, so as to complete the unified management, scheduling, and collaborative work of end-cloud integration or cloud-edge-end integration.
[0131] It should be noted that for the purpose of simplicity, the present disclosure describes some methods and their embodiments as a series of actions and combinations thereof. However, those skilled in the art can understand that the solution of the present disclosure is not limited by the order of the described actions. Therefore, based on the disclosure or teachings of the present disclosure, those skilled in the art can understand that some of the steps can be executed in other orders or simultaneously. Further, those skilled in the art can understand that the embodiments described in the present disclosure can be regarded as optional embodiments, that is, the actions or modules involved are not necessarily required for the implementation of a certain or certain solutions of the present disclosure. In addition, according to different solutions, the present disclosure also focuses on the descriptions of some embodiments. In view of this, those skilled in the art can understand the parts not detailed in a certain embodiment of the present disclosure by referring to the relevant descriptions of other embodiments.
[0132] In terms of specific implementation, based on the disclosure and teachings of the present disclosure, those skilled in the art can understand that several embodiments disclosed in the present disclosure can also be implemented in other ways not disclosed herein. For example, for each unit in the foregoing embodiments of the electronic device or apparatus, it is divided herein based on consideration of logical functions, and there may be other division methods in actual implementation. For another example, multiple units or components can be combined or integrated into another system, or some features or functions of the units or components can be selectively disabled. Regarding the connection relationship between different units or components, the connections discussed in conjunction with the accompanying drawings herein can be direct or indirect couplings between the units or components. In some scenarios, the foregoing direct or indirect couplings involve communication connections using interfaces, where the communication interfaces can support signal transmission in electrical, optical, acoustic, magnetic, or other forms.
[0133] In the present disclosure, the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units. The foregoing components or units may be located at the same location or distributed to multiple network units. Additionally, according to actual needs, some or all of the units can be selected to achieve the objectives of the solutions described in the embodiments of the present disclosure. Further, in some scenarios, multiple units in the embodiments of the present disclosure can be integrated into one unit or each unit exists physically separately.
[0134] In some implementation scenarios, the above-mentioned integrated units can be implemented in the form of software program modules. If implemented in the form of software program modules and sold or used as an independent product, the integrated units can be stored in a computer-readable memory. Based on this, when the solution of the present disclosure is embodied in the form of a software product (such as a computer-readable storage medium), the software product can be stored in the memory, which may include several instructions for causing a computer device (such as a personal computer, a server, or a network device, etc.) to execute some or all of the steps of the method described in the embodiments of the present disclosure. The foregoing memory may include, but is not limited to, various media that can store program codes, such as USB flash drives, flash memory drives, read-only memory (ROM), random access memory (RAM), mobile hard disks, magnetic disks, or optical discs.
[0135] In some other implementation scenarios, the above-integrated units can also be implemented in the form of hardware, i.e., a specific hardware circuit, which can include digital circuits and / or analog circuits, etc. The physical implementation of the hardware structure of the circuit can include, but is not limited to, physical devices, and the physical devices can include, but are not limited to, devices such as transistors or memristors. In view of this, various devices described herein (such as computing devices or other processing devices) can be implemented by appropriate hardware processors, such as CPUs, GPUs, FPGAs, DSPs, and ASICs, etc. Further, the aforementioned storage unit or storage device can be any appropriate storage medium (including magnetic storage media or magneto-optical storage media, etc.), which can be, for example, a resistive random access memory (RRAM), a dynamic random access memory (DRAM), a static random access memory (SRAM), an enhanced dynamic random access memory (EDRAM), a high bandwidth memory (HBM), a hybrid memory cube (HMC), a ROM, and a RAM, etc.
[0136] Although multiple embodiments of the present disclosure have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Many changes, variations, and alternative forms may be contemplated by those skilled in the art without departing from the spirit and scope of the present disclosure. It should be understood that various alternatives to the embodiments of the present disclosure described herein may be employed in practicing the present disclosure. The appended claims are intended to define the scope of the present disclosure and thus cover equivalents or alternatives within the scope of these claims.
[0137] The foregoing can be better understood in accordance with the following clauses:
[0138] Clause 1. A computing device, comprising:
[0139] An arithmetic circuit configured to:
[0140] Receive a plurality of data to be operated on associated with an operation instruction, wherein at least one of the data to be operated on is characterized by two or more components, the at least one data to be operated on has a source data bit width, each of the components has its own target data bit width, and the target data bit width is less than the source data bit width; and
[0141] Performing the operation specified by the operation instruction using the two or more components instead of the represented data to be operated on to obtain two or more intermediate results;
[0142] A combinational circuit configured to:
[0143] Combine the intermediate results to obtain a final result; and
[0144] A storage circuit configured to store the intermediate result and / or the final result.
[0145] Clause 2. The computing device according to Clause 1, wherein
[0146] The operation circuit is configured to perform the operation on the two or more components of one data to be operated on respectively with the corresponding data of other data to be operated on, and output the corresponding operation result to the combinational circuit; and
[0147] The combinational circuit is configured to perform a weighted combination of the operation results to obtain a final result.
[0148] Clause 3. The computing device according to Clause 2, wherein the other data to be operated on includes one or more data to be operated on, and its corresponding data includes any one of the following: the original data of the data to be operated on, or at least one component representing the data to be operated on.
[0149] Clause 4. The computing device according to any one of Clauses 2-3, wherein the operation instruction includes an instruction involving a multiplication operation or a multiply-accumulate operation, and the operation circuit includes a multiplication operation circuit or a multiply-accumulate operation circuit.
[0150] Clause 5. The computing device according to Clause 4, wherein each component has a component value and a component scaling factor, and the component scaling factor is associated with the digit position of the corresponding component in the represented data to be operated on;
[0151] wherein the operation circuit is configured to perform the operation using the component value to obtain an operation result; and
[0152] The combinational circuit is configured to perform a weighted combination of the current operation result of the operation circuit and the previous combined result of the combinational circuit using a weighting factor, wherein the weighting factor is determined at least in part based on the component scaling factor of the component corresponding to the operation result.
[0153] Clause 6. The computing device according to Clause 5, wherein the combinational circuit includes a weighting circuit and an addition circuit,
[0154] The weighted circuit configuration is used to multiply the operation result of the operation circuit by a first weighting factor to obtain a weighted result, where the first weighting factor is the product of the component scaling factors corresponding to the components of the operation result; and
[0155] The addition circuit configuration is used to accumulate the weighted result with the previous addition result of the addition circuit.
[0156] Clause 7. The computing device according to clause 5, wherein the combination circuit includes a weighted circuit and an addition circuit,
[0157] The weighted circuit configuration is used to multiply the previous addition result of the addition circuit by a second weighting factor to obtain a weighted result, where the second weighting factor is the ratio of the scaling factor of the previous operation result of the operation circuit to the scaling factor of the current operation result, and the scaling factor of the operation result is determined by the component scaling factors corresponding to the components of the operation result; and
[0158] The addition circuit configuration is used to accumulate the weighted result with the current operation result of the operation circuit.
[0159] Clause 8. The computing device according to any one of clauses 4-6, wherein the operation circuit further includes a first comparison circuit, and the first comparison circuit is configured to:
[0160] Determine whether any of the data to which the operation is about to be performed is zero, where the data includes any of the following: the original data of the data to be operated, or the components representing the data to be operated; and
[0161] If the data is zero, omit performing the operation specified by the operation instruction for the data;
[0162] Otherwise, use the data to perform the operation specified by the operation instruction.
[0163] Clause 9. The computing device according to any one of clauses 1-8, wherein the combination circuit further includes a second comparison circuit, and the second comparison circuit is configured to:
[0164] Determine whether the received intermediate result is zero; and
[0165] If the intermediate result is zero, omit performing the combination for the intermediate result;
[0166] Otherwise, use the intermediate result for the combination.
[0167] Clause 10. The computing device according to any one of clauses 1-9, wherein,
[0168] The number of components for characterizing the at least one data to be operated on is determined at least in part based on the source data bit width and the data bit width supported by the arithmetic circuit; and / or
[0169] The target data bit width is determined at least in part based on the data bit width supported by the arithmetic circuit.
[0170] Clause 11. The computing device according to any one of Clauses 1-10, wherein
[0171] The arithmetic circuit is further configured to perform the operation specified by the operation instruction in the order of receiving the two or more components, where the order includes: from high significant bits to low significant bits, or from low significant bits to high significant bits.
[0172] Clause 12. The computing device according to any one of Clauses 1-11, wherein the data to be operated on is a vector, and performing the operation specified by the operation instruction includes
[0173] Performing the operation in parallel among the elements in the vector.
[0174] Clause 13. An integrated circuit chip, comprising the computing device according to any one of Clauses 1-12.
[0175] Clause 14. An integrated circuit board, comprising the integrated circuit chip according to Clause 13.
[0176] Clause 15. A computing device, comprising the board according to Clause 14.
[0177] Clause 16. A method performed by a computing device, the method comprising
[0178] Receiving a plurality of data to be operated on associated with an operation instruction, wherein at least one data to be operated on is characterized by two or more components, the at least one data to be operated on has a source data bit width, each of the components has its own target data bit width, and the target data bit width is less than the source data bit width;
[0179] Using the two or more components to replace the characterized data to be operated on to perform the operation specified by the operation instruction to obtain two or more intermediate results; and
[0180] Combining the intermediate results to obtain a final result.
[0181] Clause 17. The method according to Clause 16, wherein
[0182] Performing the operation specified by the operation instruction includes
[0183] Performing the operation on each of the two or more components of a data to be operated on with corresponding data of other data to be operated on to obtain corresponding operation results; and
[0184] Combining the intermediate results includes:
[0185] Performing a weighted combination on the operation results to obtain a final result.
[0186] Clause 18. The computing device according to Clause 17, wherein the operation instruction includes an instruction involving a multiplication operation or a multiply-accumulate operation.
[0187] Clause 19. The method according to Clause 18, wherein each of the components has a component value and a component scaling factor, and the component scaling factor is associated with the digit position of the corresponding component in the data to be operated on represented;
[0188] Performing the operation specified by the operation instruction includes:
[0189] Performing the operation using the component value to obtain an operation result; and
[0190] Performing the weighted combination includes:
[0191] Using a weighting factor to perform a weighted combination of the current operation result and the previous combined result, wherein the weighting factor is determined at least in part based on the component scaling factor of the component corresponding to the operation result.
[0192] Clause 20. The method according to Clause 19, wherein performing the weighted combination includes:
[0193] Multiplying the operation result by a first weighting factor to obtain a weighted result, wherein the first weighting factor is the product of the component scaling factors of the components corresponding to the operation result; and
[0194] Accumulating the weighted result and the previous combined result.
[0195] Clause 21. The method according to Clause 19, wherein performing the weighted combination includes:
[0196] Multiplying the previous combined result by a second weighting factor to obtain a weighted result, wherein the second weighting factor is the ratio of the scaling factor of the previous operation result to the scaling factor of the current operation result, and the scaling factor of the operation result is determined by the component scaling factors of the components corresponding to the operation result; and
[0197] Accumulating the weighted result and the current operation result.
[0198] Clause 22. The method according to any one of Clauses 16 - 21 further includes:
[0199] judging whether any of the data to which the operation is to be performed is zero, where the data includes any of the following: the original data of the data to be operated, or the component representing the data to be operated; and
[0200] if the component is zero, then do not perform the operation specified by the operation instruction using the component.
[0201] If the data is zero, then omit performing the operation specified by the operation instruction for the data;
[0202] otherwise, perform the operation specified by the operation instruction using the data.
[0203] The embodiments of the present disclosure have been introduced in detail above. Specific examples are used herein to elaborate on the principle and implementation manner of the present disclosure. The description of the above embodiments is only used to help understand the method and its core idea of the present disclosure. At the same time, changes or deformations made by those skilled in the art based on the idea of the present disclosure, within the specific implementation manner and application scope of the present disclosure, all belong to the scope protected by the present disclosure. In summary, the content of this specification should not be construed as a limitation to the present disclosure.
Claims
1. A computing device, comprising: An arithmetic circuit configured to: Receive a plurality of data to be operated on associated with an operation instruction, wherein at least one data to be operated on is characterized by two or more components, the at least one data to be operated on has a source data bit width, each of the components has a respective target data bit width, and the target data bit width is less than the source data bit width, wherein the number of components for characterizing the at least one data to be operated on is determined at least in part based on the source data bit width and the data bit widths supported by the arithmetic circuit; and / or the target data bit width is determined at least in part based on the data bit widths supported by the arithmetic circuit; and Use the two or more components to perform the operation specified by the operation instruction instead of the characterized data to be operated on to obtain two or more intermediate results; A combination circuit configured to: Combine the intermediate results to obtain a final result; and A storage circuit configured to store the intermediate results and / or the final result, wherein each of the components has a component value and a component scaling factor, and the component scaling factor is associated with the digit position of the corresponding component in the characterized data to be operated on; wherein the arithmetic circuit is configured to perform the operation using the component values to obtain an operation result; and The combination circuit is configured to perform weighted combination of the current operation result of the arithmetic circuit and the previous combination result of the combination circuit using a weighting factor, wherein the weighting factor is determined at least in part based on the component scaling factor of the component corresponding to the operation result.
2. The computing device according to claim 1, wherein The arithmetic circuit is configured to separately perform the operation on the two or more components of one data to be operated on with the corresponding data of other data to be operated on, and output the corresponding operation results to the combination circuit; and The combination circuit is configured to perform weighted combination of the operation results to obtain a final result.
3. The computing device according to claim 2, wherein the other data to be operated on includes one or more data to be operated on, and the corresponding data thereof includes any one of the following: the original data of the data to be operated on, or at least one component characterizing the data to be operated on.
4. The computing device according to claim 3, wherein, The operation instruction includes an instruction involving a multiplication operation or a multiply-accumulate operation, and the arithmetic circuit includes a multiplication arithmetic circuit or a multiply-accumulate arithmetic circuit.
5. The computing device according to claim 1, wherein the combination circuit includes a weighting circuit and an addition circuit, The weighting circuit is configured to multiply the operation result of the arithmetic circuit by a first weighting factor to obtain a weighted result, wherein the first weighting factor is the product of the component scaling factors of the components corresponding to the operation result; and The addition circuit is configured to accumulate the weighted result with the previous addition result of the addition circuit.
6. The computing device according to claim 1, wherein the combination circuit includes a weighting circuit and an addition circuit, The weighted circuit configuration is used to multiply the previous addition result of the addition circuit by a second weighting factor to obtain a weighted result, where the second weighting factor is the ratio of the scaling factor of the previous operation result of the operation circuit to the scaling factor of the current operation result, and the scaling factor of the operation result is determined by the component scaling factor corresponding to the component of the operation result; and The addition circuit configuration is used to accumulate the weighted result and the current operation result of the operation circuit.
7. The computing device according to any one of claims 4-5, wherein, The operation circuit further includes a first comparison circuit, and the first comparison circuit is configured to: Determine whether any of the data to which the operation is about to be performed is zero, where the data includes any of the following: the original data of the data to be operated, or the component representing the data to be operated; And If the data is zero, omit performing the operation specified by the operation instruction for the data; Otherwise, use the data to perform the operation specified by the operation instruction.
8. The computing device according to any one of claims 1-6, wherein, The combination circuit further includes a second comparison circuit, and the second comparison circuit is configured to: Determine whether the received intermediate result is zero; and If the intermediate result is zero, omit performing the combination for the intermediate result; Otherwise, use the intermediate result for the combination.
9. The computing device according to any one of claims 1-6, wherein The operation circuit is further configured to perform the operation specified by the operation instruction in the order of receiving two or more components, where the order includes: from high significant bits to low significant bits, or from low significant bits to high significant bits.
10. The computing device according to claim 7, wherein, The data to be operated is a vector, and performing the operation specified by the operation instruction includes: Performing the operation in parallel among the elements in the vector.
11. An integrated circuit chip, comprising the computing device according to any one of claims 1-10.
12. An integrated circuit board, comprising the integrated circuit chip according to claim 11.
13. A computing device, comprising the board according to claim 12.
14. A method performed by a computing device, the method comprising: Receiving a plurality of data to be operated associated with an operation instruction, where at least one data to be operated is represented by two or more components, the at least one data to be operated has a source data width, each of the components has its own target data width, and the target data width is less than the source data width, and the number of components used to represent the at least one data to be operated is at least partially determined based on the source data width and the data width supported by the operation circuit; And / or the target data width is at least partially determined based on the data width supported by the operation circuit; Using the two or more components to replace the represented data to be operated to perform the operation specified by the operation instruction to obtain two or more intermediate results; And Combining the intermediate results to obtain a final result, Wherein, each of the components has a component value and a component scaling factor, and the component scaling factor is associated with the digit position of the corresponding component in the data to be operated represented. The operations specified by the operation instruction include: Performing the operation using the component value to obtain an operation result; and Using a weighting factor to perform weighted combination of the current operation result of the operation circuit and the previous combined result of the combination circuit, wherein the weighting factor is determined at least in part based on the component scaling factor of the component corresponding to the operation result.
15. The method according to claim 14, wherein Performing the operations specified by the operation instruction includes: Performing the operation on the two or more components of a data to be operated respectively with the corresponding data of other data to be operated to obtain corresponding operation results; and Combining the intermediate results includes: Performing weighted combination on the operation results to obtain a final result.
16. The method according to claim 15, wherein, The operation instruction includes an instruction involving a multiplication operation or a multiply-accumulate operation.
17. The method according to claim 14, wherein performing the weighted combination includes: Multiplying the operation result by a first weighting factor to obtain a weighted result, wherein the first weighting factor is the product of the component scaling factors of the components corresponding to the operation result; And Accumulating the weighted result and the previous combined result.
18. The method according to claim 14, wherein performing the weighted combination includes: Multiplying the previous combined result by a second weighting factor to obtain a weighted result, wherein the second weighting factor is the ratio of the scaling factor of the previous operation result to the scaling factor of the current operation result, and the scaling factor of the operation result is determined by the component scaling factors of the components corresponding to the operation result; and Accumulating the weighted result and the current operation result.
19. The method according to any one of claims 14-18, further comprising: Judging whether any of the data to which the operation is about to be performed is zero, wherein the data includes any of the following: the original data of the data to be operated, or the component representing the data to be operated; And If the component is zero, not performing the operations specified by the operation instruction using the component; If the data is zero, omitting to perform the operations specified by the operation instruction for the data; Otherwise, performing the operations specified by the operation instruction using the data.
Citation Information
Patent Citations
Computer data processing method and device
CN110262773A