Transform network-oriented mixed precision multiplier

By recoding and parallel computing of mixed-precision multipliers, the resource waste problem of low-bit-width multiplication operations of Transformer networks on FPGAs is solved, improving the efficiency and resource utilization of high-resolution image recognition and classification tasks.

CN120803399APending Publication Date: 2025-10-17XIDIAN UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510893659.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

When deploying Transformer networks on FPGAs, multiplication operations on low-bit-width and extremely low-bit-width data lead to wasted DSP resources and slow computation speed, making it difficult to meet the efficiency requirements of high-resolution image recognition and classification tasks.

Method used

A mixed-precision multiplier for Transformer networks is provided. The precision of weights and feature maps is obtained through a control module, the data is re-encoded using a re-encoding module, multiplication is performed using a multiplication module, and feature map data is recovered through a decoding module. This enables parallel fusion computation of multiple data sets, improving resource utilization.

Benefits of technology

It improves the efficiency of high-resolution image recognition and classification tasks, reduces hardware overhead, adapts to various precision combinations, and enhances the computational density and resource utilization of edge execution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120803399A_ABST
    Figure CN120803399A_ABST
Patent Text Reader

Abstract

The invention discloses a Tranformer network-oriented hybrid precision multiplier, which is applied to an FPGA (Field Programmable Gate Array), and is used for allocating a target high-resolution image according to the weight precision of the weight data of the current layer of the target Tranformer network, the feature map precision of the feature map data of the target high-resolution image and a preset data allocation rule. Determining a first quantity of weight data and a second quantity of feature map data of a single calculation cycle; according to the first quantity of weight data and the second quantity of feature map data, recoding the weight data and the feature map data after determining a target recoding algorithm to obtain weight coded data and feature map coded data; performing multiplication calculation on the weight coding data and the feature map coding data in a single calculation period to obtain an initial product result of the single calculation period; and decoding the initial product result, and merging the multiple sub-feature map data of all the calculation periods corresponding to the current layer of the target Transform network to obtain target feature map data. According to the method, the resource utilization rate is improved on the basis of ensuring the model precision.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of image processing, and particularly relates to a mixed-precision multiplier for a Transformer network. BACKGROUND

[0002] At present, the demand for high resolution and high performance of mobile image recognition and classification tasks is increasing, and such tasks often require fast processing of high-definition images on low-power terminal devices, especially in applications such as smart phones, drones, and security cameras. In these application scenarios, the Transformer network has stronger global modeling capability and has gradually replaced the traditional CNN network to become the mainstream network architecture.

[0003] However, the Transformer network model has large parameter size and high computational complexity, especially in its structure, there are a large number of matrix multiplication operations, and the computational complexity is much higher than that of the traditional convolutional neural network. When implemented in hardware, there is a power and resource bottleneck, which poses a serious challenge to the deployment on end-side devices. Therefore, in order to realize the efficient deployment of the Transformer network in the high-resolution image recognition and classification scene, the industry generally adopts quantization compression means, especially low-bit or extremely low-bit quantization methods such as W8A8 (weight quantization is INT8, feature map quantization is INT8), W4A8, W8A4, W4A4, W2A4, W2A2, etc. to reduce the storage and operation requirements of the network model. When using extremely low bit width such as INT4, INT2 (INT represents data as an integer variable, and the number represents the bit width, such as INT4 represents data as a 4-bit wide integer variable) for model quantization to reduce storage and computing overhead, although the resources and power consumption are significantly reduced, the network model accuracy is severely lost, and it is difficult to meet the actual application requirements, therefore, the mixed-precision quantization method has gradually become a mainstream trend, that is, different quantization precisions are used in different layers or different channels of the network model, and by using high-precision quantization for sensitive layers and low-precision quantization for relatively insensitive layers, the purpose of improving the hardware execution efficiency while ensuring the accuracy of the network model is achieved.

[0004] Based on this, in the process of deploying the Transformer network on the FPGA, the DSP48E2 unit is the most common DSP unit on the Xilinx FPGA chip, which has excellent timing optimization and can quickly respond to the huge advantage of the chip internal circuit structure, so using the DSP48E2 unit to complete the multiplication operation in the FPGA development can maximize the use of the resources of the chip itself and speed up the speed of neural network inference. However, for standardized design, the input bit width of a single DSP48E2 unit is limited by the chip structure, which is 27bitx18bit, and a DSP48E2 multiplication unit is needed to complete a single multiplication operation, but when used for multiplication of low bit width and very low bit width data, such as input weight data and feature map data bit width of 8bitx4bit, 4bitx4bit, 2bitx4bit, etc., a DSP48E2 unit is still needed to perform 27bitx18bit multiplication operation, which will cause waste of DSP resources on the FPGA and slow down the matrix multiplication operation speed. Most of the operations in the neural network inference of high-resolution image recognition and classification tasks are multiplication calculations, so the efficiency of the high-resolution image recognition and classification task will be low.

[0005] In summary, for the mixed precision quantization of the Transformer network for high-resolution image recognition and classification tasks, how to perform data scheduling when multiplying different precision data on the FPGA, while also maximizing the use of DSP48E2 to perform multiple multiplication operations simultaneously, to improve the efficiency of high-resolution image recognition and classification tasks, has become a problem to be solved. SUMMARY

[0006] In order to solve the above problems existing in the prior art, the application provides a mixed precision multiplier for a Transformer network.

[0007] The technical problem to be solved by the application is solved by the following technical scheme:

[0008] The application provides a mixed precision multiplier for a Transformer network, which is applied to a field programmable gate array (FPGA), and the mixed precision multiplier comprises:

[0009] A control module is configured to obtain a weight precision of weight data of a current layer of a target Transformer network and a feature map precision of feature map data of a target high-resolution image, and determine a first number of weight data and a second number of feature map data of a single calculation period according to the weight precision, the feature map precision and a preset data allocation rule.

[0010] a recoding module, configured to determine a target recoding algorithm from preset recoding algorithms according to the first quantity of weight data and the second quantity of feature map data, and recode the first quantity of weight data and the second quantity of feature map data by using the target recoding algorithm to obtain corresponding weight coded data and feature map coded data;

[0011] a multiplication calculation module, configured to perform multiplication calculation on the weight coded data and the feature map coded data in a single calculation period to obtain an initial product result of the single calculation period;

[0012] a decoding module, configured to decode the initial product result to obtain a plurality of sub-feature map data calculated in the single calculation period;

[0013] The control module is further configured to merge the plurality of sub-feature map data of all calculation periods corresponding to the current layer of the target Transformer network to obtain target feature map data.

[0014] The application provides a mixed-precision multiplier for a Transformer network, which is applied to a field programmable gate array (FPGA). A control module is configured to obtain weight precision of weight data of a current layer of a target Transformer network and feature map precision of feature map data of a target high-resolution image, and determine a first quantity of weight data and a second quantity of feature map data of a single calculation period according to the weight precision, the feature map precision and a preset data distribution rule. A recoding module is configured to determine a target recoding algorithm from preset recoding algorithms according to the first quantity of weight data and the second quantity of feature map data, and recode the first quantity of weight data and the second quantity of feature map data by using the target recoding algorithm to obtain corresponding weight coded data and feature map coded data. A multiplication calculation module is configured to perform multiplication calculation on the weight coded data and the feature map coded data in a single calculation period to obtain an initial product result of the single calculation period. A decoding module is configured to decode the initial product result to obtain a plurality of sub-feature map data calculated in the single calculation period. The control module is further configured to merge the plurality of sub-feature map data of all calculation periods corresponding to the current layer of the target Transformer network to obtain target feature map data.

[0015] The application can flexibly adapt to various precision data combinations by using a dynamic scheduling and unified control mechanism of the control module, and ensures the model precision. The recoding module and the multiplication calculation module are used to realize multiple data parallel fusion calculation, which greatly improves the resource utilization, thereby improving the recognition and classification task efficiency of the high-resolution image and reducing the hardware overhead.

[0016] The application will be further described in detail below with reference to the accompanying drawings and the application. BRIEF DESCRIPTION OF DRAWINGS

[0017] Figure 1 is a structural schematic diagram of a mixed precision multiplier for a Transformer network provided by an embodiment of the present application.

[0018] Figure 2 is a schematic diagram of a preset data allocation rule in a mixed precision multiplier for a Transformer network provided by an embodiment of the present application.

[0019] Figures 3A-3B is a schematic diagram of a preset recoding algorithm corresponding to a weight precision of INT2 in a mixed precision multiplier for a Transformer network provided by an embodiment of the present application.

[0020] Figures 4A-4B is a schematic diagram of a preset recoding algorithm corresponding to a weight precision of INT4 in a mixed precision multiplier for a Transformer network provided by an embodiment of the present application.

[0021] Figures 5A-5B is a schematic diagram of a preset recoding algorithm corresponding to a weight precision of INT8 in a mixed precision multiplier for a Transformer network provided by an embodiment of the present application.

[0022] Figure 6 is an exemplary data flow scheduling schematic diagram of a mixed precision multiplier for a Transformer network provided by an embodiment of the present application. DETAILED DESCRIPTION

[0023] The present application will be further described in detail below in conjunction with specific embodiments, but the embodiments of the present application are not limited thereto.

[0024] An embodiment of the present application provides a mixed precision multiplier for a Transformer network, applied to a field programmable gate array (FPGA), referring to Figure 1 , comprising the following modules:

[0025] The control module 10 is configured to obtain a weight precision of weight data of a current layer of a target Transformer network and a feature map precision of feature map data of a target high-resolution image, and determine a first quantity of weight data and a second quantity of feature map data of a single calculation period according to the weight precision, the feature map precision and a preset data allocation rule.

[0026] Exemplarily, the weight precision of the weight data corresponding to each layer in the target Transformer network and the output feature map precision of the output feature map can be preset in advance, and the feature map precision of the feature map data of the target high-resolution image corresponding to the current layer of the target Transformer network is the output feature map precision of the last layer. A single calculation period refers to a period of completing one calculation by sequentially calling the recoding module, the multiplication calculation module, the decoding module and the optional re-quantization module for the first quantity of weight data and the second quantity of feature map data. The multiplication calculation module of the embodiment is implemented by using a DSP (Digital Signal Processing) unit, which can be a DSP48E2 unit, and the bit width of the input data is limited to 27bitx18bit, and the input bit width is configured in the range of 2-27bitx2-18bit when used.

[0027] Optionally, the preset data allocation rule includes the number of weight data and the number of feature map data that can be calculated in a single calculation period corresponding to different weight feature map precision combinations; the weight feature map precision combination is a combination formed by an optional weight precision and an optional feature map precision; the optional weight precision is one of INT8, INT4 and INT2, and the optional feature map precision is one of INT8 and INT4.

[0028] The preset data allocation rule is determined by the following method:

[0029] Based on the data calculation bit width of the DSP unit in the FPGA, the optional weight precision and the optional feature map precision in each weight feature map precision combination, the number of weight data and the number of feature map data that can be calculated in a single calculation period corresponding to each weight feature map precision combination are determined.

[0030] Exemplarily, the optional weight precision can be one of INT8, INT4 and INT2, and the weight precision variable corresponding to the optional weight precision can be set to 00 (representing INT8), 01 (representing INT4) or 11 (representing INT2). The optional feature map precision is one of INT8 and INT4, and the feature map precision variable corresponding to the optional feature map precision can be set to 0 (representing INT8) or 1 (representing INT4).

[0031] The data calculation bit width of the DSP unit in the FPGA can be 27bitx18bit of the data calculation bit width of the DSP48E2 unit, and the preset data allocation rule is determined by the following manner: in the recoding process, according to the selectable weight precision and the selectable feature map precision in each weight feature map precision combination, the arrangement combination of the weight data and the feature map data is carried out on the basis of the data calculation bit width of the DSP unit in the FPGA, when the product of the number of weight data and the number of feature map data that can be calculated in a single calculation period is maximum, and the positions of the different weight data and feature map data product results are not repeated when the product results are decoded, the number of weight data and the number of feature map data that can be calculated in a single calculation period corresponding to the weight feature map precision combination can be determined.

[0032] Referring to Figure 2 The preset data allocation rule can include five cases: when the selectable weight precision is INT8, whether the selectable feature map precision is INT8 or INT4, two weight data and one feature map data can participate in calculation in a single calculation period; when the selectable weight precision is INT4 and the selectable feature map data is INT8, three weight data and one feature map data can participate in calculation in a single calculation period; when the selectable weight precision is INT4 and the selectable feature map data is INT4, two weight data and three feature map data can participate in calculation in a single calculation period; when the selectable weight precision is INT2 and the selectable feature map data is INT8, two weight data and two feature map data can participate in calculation in a single calculation period; and when the selectable weight precision is INT2 and the selectable feature map data is INT4, four weight data and two feature map data can participate in calculation in a single calculation period.

[0033] In the embodiment, after the first number of weight data and the second number of feature map data in a single calculation period are determined according to the weight precision, the feature map precision and the preset data allocation rule, the first number of weight data and the second number of feature map data are packaged and delivered to the recoding module 20, and the weight precision variable corresponding to the weight data and the feature map precision variable corresponding to the feature map data are also delivered.

[0034] The embodiment dynamically controls the data input packaging number and the path according to the input weight precision and the feature map precision, flexibly schedules the calculation process, improves the universality and the reuse rate of the module, and is deployed on an image terminal device supporting multiple models and cross tasks.

[0035] The recoding module 20 is configured to determine a target recoding algorithm from preset recoding algorithms according to the first number of weight data and the second number of feature map data, and recode the first number of weight data and the second number of feature map data by using the target recoding algorithm to obtain corresponding weight coding data and feature map coding data.

[0036] Exemplarily, the first quantity of weight data and the second quantity of feature map data are re-encoded to form multiplication input data suitable for a 27bitx18bit DSP48E2 unit. Since the weight data is often positive and negative, and the feature map data representing image pixels is positive, the weight data is usually symmetrically quantized, the INT8 weight data range is (-128, 127), the INT4 weight data range is (-8, 7), and the INT2 weight data range is (-2, 1), and the feature map data is asymmetrically quantized, the INT8 precision feature map data range is (0, 255), and the INT4 precision weight data range is (0, 15). In order to fully utilize the bit width resources of the DSP48E2 unit and also simplify the processing of the subsequent multiplication results, the absolute values of the input weight data of different precisions are taken and then re-encoded, and the sign bit variables of the weight data are retained. Thus, the sign bit variables of the weight data can be used to process the subsequent multiplication results to obtain the product results of each weight data and feature map data.

[0037] Optionally, the re-encoding module 20 is further configured to:

[0038] take absolute values of the first quantity of weight data and extract sign bits to obtain a first quantity of weight absolute value data and corresponding sign bit variables; determine a target re-encoding algorithm from preset re-encoding algorithms according to the weight absolute value data and the feature map data; and re-encode the first quantity of weight absolute value data and the second quantity of feature map data by using the target re-encoding algorithm to obtain weight encoding data and feature map encoding data.

[0039] Optionally, the re-encoding module 20 is further configured to:

[0040] determine, for each weight absolute value data, whether the weight precision of the weight absolute value data is INT8, INT4 or INT2;

[0041] determine a target re-encoding algorithm from preset re-encoding algorithms according to the determination result; and re-encode the first quantity of weight absolute value data and the second quantity of feature map data by using the target re-encoding algorithm to obtain weight encoding data and feature map encoding data.

[0042] Exemplarily, the preset re-encoding algorithms include re-encoding algorithms corresponding to different weight precisions.

[0043] Optionally, the re-encoding module 20 is further configured to:

[0044] If the weight precision corresponding to the weight absolute value data is INT2, the first quantity of weight absolute value data and the second quantity of feature map data are re-encoded to obtain weight encoding data and feature map encoding data.

[0045] Exemplarily, when the weight precision corresponding to the weight absolute value data is INT2, the first quantity of weight absolute value data and the second quantity of feature map data are directly re-encoded to obtain weight encoding data and feature map encoding data. Here, the bit width of each weight absolute value data is 2 bits.

[0046] The specific re-encoding process is as follows:

[0047] Referring to Figure 3A , the process of re-encoding for the weight feature map precision combination W2 (weight precision INT2) A8 (feature map precision INT8) is as follows: the weight absolute value data participating in re-encoding is 2 (the first quantity), and the feature map data is 2 (the second quantity). The 2 2-bit weight absolute value data A and B are re-encoded with the 2 8-bit feature map data a and b as W2A8. The weight absolute value data A and B are re-encoded as 11-bit weight encoding data after an interval of 7-bit width (here, the interval of 7 bits is because the weight absolute value data is 2 bits, but the maximum value is only 2, and after multiplication with the 8-bit feature map data, the maximum expansion of the product result is 1 bit width, so the product result of A and B can be distinguished by ensuring that the high-bit weight absolute value data B is in the 10th bit position), and then transmitted to the multiplication calculation module. The feature map data a and b are re-encoded as 26-bit feature map encoding data after an interval of 10-bit width (here, the interval of 10 bits is also because the re-encoded weight absolute value data is 11 bits, but limited by the maximum value, so that the product result of the 8-bit feature map data is expanded by 10 bits at most, so the interval bit number is 10 bits, which can distinguish the product result of a and b), and then transmitted to the multiplication calculation module.

[0048] Referring to Figure 3B, the re-encoding process for the weight feature map precision combination W2 (weight precision INT2) A4 (feature map precision INT4) is as follows: 4 weight absolute value data (first quantity) and 2 feature map data (second quantity) participate in re-encoding. The 4 2-bit weight absolute value data A, B, C, and D are re-encoded with the 2 4-bit feature map data a and b to obtain W2A4 re-encoded data, and then the re-encoded data is transmitted to the multiplication calculation module. The 2 4-bit feature map data a and b are re-encoded with a 16-bit space in the middle to obtain 24-bit feature map encoding data, and then the feature map encoding data is transmitted to the multiplication calculation module.

[0049] Optionally, the re-encoding module 20 is further configured to:

[0050] If the weight precision corresponding to the weight absolute value data is INT4, it is determined whether the first quantity of weight absolute value data is all not equal to the first target value. If yes, the first quantity of weight absolute value data and the second quantity of feature map data are re-encoded to obtain weight encoding data and feature map encoding data. If there is a first target weight absolute value data equal to the first target value, each bit of the corresponding first bit width is filled with a preset value when re-encoding the first quantity of weight absolute value data, the first bit width is the bit width corresponding to the first target weight absolute value data minus one, and the second quantity of feature map data is re-encoded to obtain weight encoding data and feature map encoding data. The first target value is determined according to the weight precision INT4.

[0051] Exemplarily, the first target value is 8, which is the maximum value of the weight absolute value data corresponding to INT4. If the first quantity of weight absolute value data is all not equal to 8, the first quantity of weight absolute value data and the second quantity of feature map data are directly re-encoded. If there is a first target weight absolute value data equal to 8 in the first quantity of weight absolute value data, each bit of the corresponding first bit width is filled with a preset value 0 when re-encoding the first quantity of weight absolute value data, the first bit width is the bit width corresponding to the first target weight absolute value data minus one, that is, 3 bits, to obtain the weight absolute value data to be encoded, and the weight absolute value data to be encoded and the second quantity of feature map data are re-encoded to obtain weight encoding data and feature map encoding data.

[0052] It should be noted that the first bit width is set to the bit width corresponding to the first target weight absolute value data minus one because the re-encoding design bit width length for the weight precision INT4 is the bit width length corresponding to INT4 minus one, that is, 3 bits.

[0053] Here, by filling each bit corresponding to a bit width of 3 bits of the first target weight absolute value data with a preset value 0, the subsequent 3-bit first target weight absolute value data can be prevented from participating in the operation of the multiplication calculation module, and the product result corresponding to the first target weight absolute value data is directly left shifted by 3 bits to obtain the feature map data that should be multiplied.

[0054] The specific re-encoding process is as follows:

[0055] Referring to Figure 4A , the process of re-encoding for the weight feature map precision combination W4 (weight precision INT4) A8 (feature map precision INT8) is as follows: the weight absolute value data participating in re-encoding is 3 (the first quantity), and the feature map data is 1 (the second quantity). The 3-bit weight absolute value data A, B, and C are re-encoded as 25-bit weight encoding data after each interval of 8-bit empty positions in the middle of the weight absolute value data A, B, and C, and then transmitted to the multiplication calculation module. The feature map data a is 8-bit and does not need to be re-encoded, and is directly transmitted to the multiplication calculation module as feature map encoding data.

[0056] Referring to Figure 4B , the process of re-encoding for the weight feature map precision combination W4 (weight precision INT4) A4 (feature map precision INT4) is as follows: the weight absolute value data participating in re-encoding is 2 (the first quantity), and the feature map data is 3 (the second quantity). The 3-bit weight absolute value data A and B are re-encoded as 24-bit weight encoding data after each interval of 18-bit empty positions in the middle of the weight absolute value data A and B, and then transmitted to the multiplication calculation module. The feature map data a, b, and c are re-encoded as 18-bit feature map encoding data after each interval of 3-bit empty positions in the middle of the feature map data a, b, and c, and then transmitted to the multiplication calculation module.

[0057] Optionally, the re-encoding module 20 is further configured to:

[0058] If the weight precision corresponding to the weight absolute value data is INT8, it is determined whether the absolute value of the weight data is not equal to the second target value. If yes, the weight absolute value data and the feature map data are re-encoded to obtain the first quantity of weight encoded data and the second quantity of feature map encoded data. If there is second target weight absolute value data equal to the second target value, each bit corresponding to the second bit width is padded with a preset value when re-encoding the first quantity of weight absolute value data, the second bit width is the bit width corresponding to the second target weight absolute value data minus one, and the second quantity of feature map data is re-encoded to obtain the weight encoded data and the feature map encoded data. The second target value is determined according to the weight precision INT8.

[0059] Exemplarily, the second target value is 128, which is the maximum value of the weight absolute value data corresponding to INT8. If none of the first quantity of weight absolute value data is equal to 128, the first quantity of weight absolute value data and the second quantity of feature map data are directly re-encoded. If there is second target weight absolute value data equal to 128 in the first quantity of weight absolute value data, each bit corresponding to the second bit width is padded with a preset value 0 when re-encoding the first quantity of weight absolute value data, the second bit width is the bit width corresponding to the second target weight absolute value data minus one, that is, 7 bits, to obtain the weight absolute value data to be encoded, and the weight absolute value data to be encoded and the second quantity of feature map data are re-encoded to obtain the weight encoded data and the feature map encoded data.

[0060] It should be noted that the second bit width is set to the bit width corresponding to the second target weight absolute value data minus one because the re-encoding design for the weight precision INT8 has a bit width length of INT8 corresponding bit width length minus one, that is, 7 bits.

[0061] Here, by padding each bit corresponding to the bit width 7 bits of the second target weight absolute value data with the preset value 0, the subsequent 7-bit second target weight absolute value data can be prevented from participating in the operation of the multiplication calculation module, and the product result corresponding to the second target weight absolute value data is directly left shifted by 7 bits to obtain the feature map data that should be multiplied with it after a period of time.

[0062] The specific re-encoding process is as follows:

[0063] Referring to Figure 5AThe re-encoding process for the weight feature map precision combination W8 (weight precision INT8) A8 (feature map precision INT8) is as follows: the weight absolute value data involved in the re-encoding is 2 (first quantity), and the feature map data is 1 (second quantity). The two 7-bit weight absolute value data A and B are re-encoded with the 8-bit feature map data a using W8A8. The weight absolute value data A and B are separated by an 8-bit width space and re-encoded into 22-bit weighted encoded data, which is then passed to the multiplication calculation module. The feature map data a is 8 bits and does not need to be re-encoded. It is directly passed to the multiplication calculation module as feature map encoded data.

[0064] Reference Figure 5B , the re-encoding process for the weight feature map precision combination W8 (weight precision INT8) A4 (feature map precision INT4) is as follows: the weight absolute value data involved in the re-encoding is 2 (first quantity), and the feature map data is 1 (second quantity). The two 7-bit weight absolute value data A and B are re-encoded with the 4-bit feature map data a by W8A4. The weight absolute value data A and B are separated by a 4-bit width space and re-encoded into 18-bit weighted encoded data, and then passed to the multiplication calculation module. The feature map data a is 4 bits and does not need to be re-encoded. It is directly passed to the multiplication calculation module as feature map encoded data.

[0065] This embodiment uses different recoding rules for the six weight feature map precision combinations, adapting multiple small-bit-width data to the input bit width of the DSP48E2 unit and enabling parallel merging of multiplication operations. This allows a single multiplier to support parallel computation of multiple precision combinations, enabling hardware sharing and significantly improving computational density and resource utilization when executing high-resolution image recognition tasks on the client side.

[0066] The multiplication calculation module 30 is used to perform multiplication calculation on the weighted encoded data and the feature map encoded data in a single calculation cycle to obtain an initial product result of the single calculation cycle.

[0067] Exemplarily, the multiplication calculation module 30 can implement the multiplication operation by calling a DSP Macro. The DSP Macro provides an easy-to-use interface for implementing the multiplication operation by abstracting the configuration of the DSP48E2 unit and implementing the operation by a user-defined arithmetic expression, thereby simplifying the dynamic operation of the DSP48E2 unit. The multiplication operation of the re-encoded weight encoding data and the feature map encoding data after the re-encoding is completed is implemented by using the DSP48E2 unit by configuring the multiplication expression on a single port of the generated DSP Macro IP. In the multiplication calculation module 30, the two multipliers of the input port are the weight encoding data and the feature map encoding data, respectively, and the multiplication calculation result, i.e., the initial product result, of the output port is directly passed to the decoding module.

[0068] The decoding module 40 is configured to decode the initial product result to obtain a plurality of sub-feature map data calculated in a single calculation period.

[0069] Exemplarily, the decoding module 40 is configured to process the initial product result of a single calculation period to obtain the product result of each weight data and each feature map data, i.e., a plurality of sub-feature map data of a single calculation period. The number of sub-feature map data in a single calculation period is determined by the product of the first number of weight data and the second number of feature map data.

[0070] Corresponding to the re-encoding module 20, the decoding module 40 is also configured to decode the initial product result based on the sign bit variable corresponding to each weight absolute value data to obtain a plurality of sub-feature map data calculated in a single calculation period.

[0071] Optionally, the sign bit variable is used to indicate whether the weight data corresponding to the weight absolute value data is positive or negative, and the decoding module 40 is also configured to:

[0072] determine the bit width information and the position information of a plurality of product results contained in the initial product result in a single calculation period according to the first number, the second number, the weight encoding data, and the feature map encoding data;

[0073] if the sign bit variable indicates that the weight data is positive, decode the product result corresponding to the weight data based on the bit width information and the position information of the product result corresponding to the weight data by using a first decoding algorithm to obtain the sub-feature map data corresponding to the product result;

[0074] If the sign bit variable indicates that the weight data is negative, the second decoding algorithm is used to decode the product result corresponding to the weight data based on the bit width information and the position information of the product result corresponding to the weight data, to obtain the sub-feature map data corresponding to the product result.

[0075] The first decoding algorithm and the second decoding algorithm are different.

[0076] Exemplarily, according to the first quantity of weight data, the second quantity of feature map data, the weight encoding data and the feature map encoding data, when the weight precision is INT8 and the feature map precision is INT8 or INT4 precision, two product results are contained in the initial product result of a single calculation period, from low bit to high bit, A*a and B*a respectively, and the corresponding bit width information and position information can be obtained, as shown in Figure 5A and Figure 5B When the weight precision is INT4 and the feature map precision is INT8, three product results are contained in the initial product result of a single calculation period, from low bit to high bit, A*a, B*a and C*a respectively, and the corresponding bit width information and position information can be obtained, as shown in Figure 4A When the weight precision is INT4 and the feature map precision is INT4, six product results are contained in the initial product result of a single calculation period, from low bit to high bit, A*a, A*b, A*c, B*a, B*b and B*c respectively, and the corresponding bit width information and position information can be obtained, as shown in Figure 4B When the weight precision is INT2 and the feature map precision is INT8, four product results are contained in the initial product result of a single calculation period, from low bit to high bit, A*a, B*a, A*b and B*b respectively, and the corresponding bit width information and position information can be obtained, as shown in Figure 3A When the weight precision is INT2 and the feature map precision is INT4, eight product results are contained in the initial product result of a single calculation period, from low bit to high bit, A*a, B*a, C*a, D*a, A*b, B*b, C*b and D*b respectively, and the corresponding bit width information and position information can be obtained, as shown in Figure 3B .

[0077] After obtaining the bit width information and position information of multiple product results in a single calculation period, the product results need to be processed for sign recovery. A sign bit variable obtained in advance is used as an enable signal. When the sign bit variable is 0, it indicates that the weight data corresponding to the weight absolute value data is positive. Then, the first decoding algorithm is as follows: 1 bit of 0 is added in front of the highest bit of the corresponding product result to obtain the sub-feature map data corresponding to the product result. When the weight sign bit variable is 1, it indicates that the weight data corresponding to the weight absolute value data is negative. Then, the second decoding algorithm is as follows: 1 is subtracted from the corresponding product result, and then the result is inverted bit by bit. Then, 1 bit of 1 is added in front of the highest bit of the processing result to obtain the sub-feature map data corresponding to the product result. The sub-feature map data corresponding to all product results in a single period are combined to obtain the sub-feature map data calculated in a single calculation period. Finally, the sub-feature map data of a single calculation period is transmitted to the re-quantization module.

[0078] Here, for the case that the first target weight absolute value data corresponding to the bit width of 3 bits is filled with a preset value 0 in the re-encoding module 20, the 3-bit first target weight absolute value data does not participate in the operation of the multiplication calculation module. The product result corresponding to the first target weight absolute value data is directly processed by left shifting the feature map data that should be multiplied by 3 bits after a number of periods.

[0079] Similarly, for the case that the second target weight absolute value data corresponding to the bit width of 7 bits is filled with a preset value 0 in the re-encoding module 20, the 7-bit second target weight absolute value data does not participate in the operation of the multiplication calculation module. The product result corresponding to the second target weight absolute value data is directly processed by left shifting the feature map data that should be multiplied by 7 bits after a number of periods.

[0080] The embodiment uses the sign bit variable of the weight data to realize bit width restoration and sign recovery of multiple product results, ensures calculation correctness, and saves the consumption of multiplication resources occupied by the sign bit. It is especially suitable for high-resolution image analysis tasks with high feature fidelity requirements in models such as Transformer.

[0081] The control module 10 is also used to combine the multiple sub-feature map data of all calculation periods corresponding to the current layer of the target Transformer network to obtain target feature map data.

[0082] For example, for the feature map data input into the current layer of the target Transformer network, multiple calculation periods are needed for calculation. After obtaining the multiple sub-feature map data of each calculation period, the multiple sub-feature map data of all calculation periods corresponding to the current layer of the target Transformer network are combined to obtain the target feature map data output by the current layer.

[0083] Optionally, the mixed-precision multiplier of the embodiment can further include:

[0084] The re-quantization module 50 is configured to determine a target quantization algorithm from preset quantization algorithms according to an output feature map precision corresponding to a current layer of the target Transformer network, perform linear quantization on the target feature map data by using the target quantization algorithm, and obtain quantized target feature map data and a corresponding target feature map precision; the preset quantization algorithms include quantization algorithms corresponding to different output feature map precisions, and the target feature map precision is INT8 or INT4.

[0085] For example, the re-quantization module 50 is configured to linearly quantize the sub-feature map data (with an INT32 feature map precision) obtained by the decoding module 40 into feature map data with an INT8 or INT4 precision, so as to directly perform calculation when the next calculation is needed. The output feature map precision corresponding to the current layer of the target Transformer network can be INT8 or INT4. When the output feature map precision is INT8, the sub-feature map data is right shifted by 24 bits to obtain quantized sub-feature map data and a corresponding sub-feature map precision INT8; when the output feature map precision is INT4, the sub-feature map data is right shifted by 28 bits to obtain quantized sub-feature map data and a corresponding sub-feature map precision INT4.

[0086] After the linear quantization operation is completed, the quantized sub-feature map data and a corresponding sub-feature map precision variable can be output to the control module 10, and the sub-feature map precision variable includes 0 (indicating INT8) and 1 (indicating INT4).

[0087] Here, reference is made to Figure 1 Before the re-quantization module 50 is processed, it can be determined whether the preset output feature map precision is obtained. When the output feature map precision is not obtained, the re-quantization module 50 does not need to be processed, and the decoding module 40 directly outputs the sub-feature map data to the control module 10.

[0088] The re-quantization module of the embodiment integrates the re-quantization logic supporting INT8 and INT4, so that the output of the re-quantization module directly meets the requirements of the next layer quantization network, simplifies the overall system design, and improves the integration and energy efficiency ratio of the overall system in the image recognition pipeline.

[0089] The hybrid precision multiplier for the Transformer network provided by the application realizes that a plurality of low-bit multiplications share a group of multiplier logic by adopting a plurality of data parallel fusion calculation modes after re-encoding the weight data and the feature map data, greatly improves the resource utilization rate, and reduces the hardware overhead. Moreover, the control module adopts a dynamic scheduling and unified control mechanism, can flexibly adapt to a plurality of precision combinations, has stronger adaptability and expandability, and meets the dynamic adjustment demand of the feature processing precision of the high-resolution image in different task stages. In addition, the embodiment integrates product decoding, symbol recovery and re-quantization processing, reduces the data transfer and logic complexity, improves the overall calculation efficiency, and significantly improves the practicability and deployability of the high-resolution Transformer image classification model deployed at the end side.

[0090] The processing process of the hybrid precision multiplier for the Transformer network provided by the application is further described below through an example.

[0091] As Figure 6As shown, the data flow scheduling diagram of the multiplier designed by the application is shown, taking the weight precision of the input weight data as INT8, the feature map precision of the input feature map data as INT4, and the output feature map precision as INT8 as an example, to complete the process of reading, recoding, multiplication calculation, decoding processing, and re-quantization of the weight data of the target Transformer network and the feature map data of the target high-resolution image. The control module receives the weight precision variable of the input weight data of the current layer of the target Transformer network from the previous stage as 00, indicating that the weight precision of the input weight data is INT8, the feature map precision variable of the input feature map data of the current layer is 1, indicating that the feature map precision of the input feature map data is INT4, and the output feature map precision variable of the current layer is 0, indicating that the output feature map precision of the current layer is INT8. Then read 2 weight data and 1 feature map data. Then pass the input precision variable and data to the recoding module and the decoding module, and pass the output feature map precision variable to the re-quantization module. In the recoding module, first extract the sign bit variable of the weight data, then recode the weight data and the feature map data to obtain 18-bit weight coding data and 4-bit feature map coding data, and then pass the recoded weight coding data and feature map coding data to the multiplication calculation module, and pass the sign bit variable to the decoding module. The multiplication calculation module completes the calculation of 2 weight data x 1 feature map data once by calling DSP Macro, and then passes the 22-bit initial product result to the decoding module. The decoding module processes the initial product result according to the sign bit variable to obtain two 12-bit sub-feature map data, and then passes the sub-feature map data to the re-quantization module, and the output feature map precision variable of the current layer is 0, indicating that re-quantization processing is needed, the type of re-quantization processing is INT8, and the re-quantized sub-feature map data and the corresponding sub-feature map precision variable are output.

[0092] It should be noted that the terms "first", "second" and the like are used to distinguish similar objects, not necessarily describing a particular order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the application described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the application. Rather, they are merely examples of devices and methods consistent with some aspects of the application.

[0093] In the description of the specification, the description of the terms "one embodiment", "some embodiments", "an example", "a specific example", or "some examples" etc. means that the specific features or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In the description of the specification, the illustrative description of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in the specification.

[0094] Although the present application is described herein in conjunction with various embodiments, those skilled in the art, with the benefit of the description and drawings presented herein, can understand and appreciate other variations and modifications in the disclosed embodiments. In the description of the present application, the word "comprising" does not exclude other components or steps, "a" or "one" does not exclude a plurality, and "plurality" means two or more, unless otherwise expressly specified. In addition, some measures are described in different embodiments, but this does not mean that these measures cannot be combined to produce good results.

[0095] The above is a further detailed description of the present application in conjunction with specific preferred embodiments, and cannot be considered as limiting the specific implementation of the present application to these descriptions. For those skilled in the art, without departing from the concept of the present application, a number of simple deductions or substitutions can be made, which should be considered as falling within the scope of protection of the present application.

Claims

1. A mixed-precision multiplier for Transformer networks, characterized in that: Applied to a field programmable gate array (FPGA), the mixed-precision multiplier includes: a control module, configured to obtain weight accuracy of weight data of a current layer of a target Transformer network and feature map accuracy of feature map data of a target high-resolution image, and determine a first quantity of weight data and a second quantity of feature map data for a single computation cycle based on the weight accuracy, the feature map accuracy, and a preset data allocation rule; a recoding module, configured to determine a target recoding algorithm from preset recoding algorithms based on the first amount of weight data and the second amount of feature map data; and recode the first amount of weight data and the second amount of feature map data using the target recoding algorithm to obtain corresponding weight-coded data and feature map-coded data; a multiplication calculation module, configured to perform a multiplication calculation on the weighted encoded data and the feature map encoded data within a single calculation cycle to obtain an initial product result of the single calculation cycle; A decoding module, configured to decode the initial product result to obtain a plurality of sub-feature map data calculated in a single calculation cycle; The control module is further used to merge multiple sub-feature map data of all calculation cycles corresponding to the current layer of the target Transformer network to obtain target feature map data.

2. The mixed-precision multiplier for Transformer networks according to claim 1, wherein: The mixed precision multiplier further includes: A requantization module is used to determine a target quantization algorithm from a preset quantization algorithm based on the preset output feature map accuracy corresponding to the current layer of the target Transformer network, and use the target quantization algorithm to perform a linear quantization operation on the sub-feature map data to obtain the quantized sub-feature map data and the corresponding sub-feature map accuracy; the preset quantization algorithm includes quantization algorithms corresponding to different output feature map accuracy, and the sub-feature map accuracy is INT8 or INT4.

3. The mixed-precision multiplier for Transformer networks according to claim 2, wherein: The preset data allocation rule includes the number of weight data and the number of feature map data that can be calculated in a single calculation cycle corresponding to different weight feature map precision combinations; the weight feature map precision combination is a combination of an optional weight precision and an optional feature map precision; the optional weight precision is one of INT8, INT4 and INT2, and the optional feature map precision is one of INT8 and INT4; The preset data allocation rule is determined in the following manner: Based on the data calculation bit width of the DSP unit in the FPGA, the optional weight precision and optional feature map precision in each weight feature map precision combination, the number of weight data and the number of feature map data that can be calculated in a single calculation cycle corresponding to each weight feature map precision combination are determined.

4. The mixed-precision multiplier for Transformer networks according to claim 3, wherein: The recoding module is further configured to take the absolute value of the first amount of weight data and extract the sign bit to obtain the first amount of weight absolute value data and the corresponding sign bit variable; determine a target recoding algorithm from the preset recoding algorithms based on the weight absolute value data and the feature map data; and recode the first amount of weight absolute value data and the second amount of feature map data using the target recoding algorithm to obtain the weight coded data and the feature map coded data; The decoding module is further used to decode the initial product result based on the sign bit variable corresponding to each weight absolute value data to obtain multiple sub-feature map data of a single calculation cycle.

5. The mixed-precision multiplier for Transformer networks according to claim 4, wherein: The re-encoding module is further used to: For each of the weight absolute value data, determine whether the weight precision of the weight absolute value data is INT8, INT4 or INT2; According to the judgment result, a target recoding algorithm is determined from the preset recoding algorithms; the target recoding algorithm is used to recode the first quantity of the weight absolute value data and the second quantity of the feature map data to obtain the weight coding data and the feature map coding data.

6. The mixed-precision multiplier for Transformer networks according to claim 5, wherein: The re-encoding module is further used to: If the weight accuracy corresponding to the weight absolute value data is INT2, the first number of the weight absolute value data and the second number of the feature map data are re-encoded to obtain the weight encoded data and the feature map encoded data.

7. The mixed-precision multiplier for Transformer networks according to claim 5, wherein: The re-encoding module is further used to: If the weight precision corresponding to the weight absolute value data is INT4, determine whether the first number of the weight absolute value data are not equal to the first target value. If so, re-encode the first number of the weight absolute value data and the second number of the feature map data to obtain the weight encoded data and the feature map encoded data; if there is a first target weight absolute value data equal to the first target value, then when re-encoding the first number of the weight absolute value data, use a preset value to fill each bit corresponding to the first bit width, and the first bit width is the bit width corresponding to the first target weight absolute value data minus one, and re-encode the second number of the feature map data to obtain the weight encoded data and the feature map encoded data; the first target value is determined according to the weight precision INT4.

8. The mixed-precision multiplier for Transformer networks according to claim 5, wherein: The re-encoding module is further used to: If the weight precision corresponding to the weight absolute value data is INT8, determine whether the absolute value of the weight data is not equal to the second target value. If so, re-encode the weight absolute value data and the feature map data to obtain the corresponding first quantity of weight encoded data and second quantity of feature map encoded data; if there is a second target weight absolute value data equal to the second target value, use a preset value to fill each bit corresponding to the second bit width when re-encoding the first quantity of the weight absolute value data, and the second bit width is the bit width corresponding to the second target weight absolute value data minus one, and re-encode the second quantity of the feature map data to obtain the weight encoded data and the feature map encoded data; the second target value is determined according to the weight precision INT8.

9. The mixed-precision multiplier for a Transformer network according to any one of claims 6 to 8, characterized in that: The sign bit variable is used to indicate whether the weight data corresponding to the weight absolute value data is a positive number or a negative number. The decoding module is further used to: Determining bit width information and position information of a plurality of product results included in the initial product result in a single calculation cycle according to the first number, the second number, the weight encoded data, and the feature map encoded data; If the sign bit variable indicates that the weight data is a positive number, decoding the product result corresponding to the weight data using a first decoding algorithm based on the bit width information and position information of the product result corresponding to the weight data to obtain sub-feature map data corresponding to the product result; If the sign bit variable indicates that the weight data is a negative number, decoding the product result corresponding to the weight data using a second decoding algorithm based on the bit width information and position information of the product result corresponding to the weight data to obtain sub-feature map data corresponding to the product result; The first decoding algorithm and the second decoding algorithm are different.

Citation Information

Cited By

  • Multi-input-output Transform model chip architecture and calculation method

    CN121072630A