Transformer Quantization Acceleration Method in the Text Domain Based on Dictionary Index Fixed-Point Arithmetic
The benchmark dictionary is generated through aggregation hierarchical clustering, and linear transformation and exponential function fitting are combined with the statistical features of the transformer tensor, which solves the problem of efficient operation and activation value quantification of the Transformer model on resource-constrained devices, achieving efficient and real-time quantitative acceleration effect.
Patent Information
- Application Number
- CN202510322794.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-03-19
AI Technical Summary
The Transformer model is difficult to operate efficiently on resource-constrained devices in the prior art, and the quantification of activation values cannot meet the real-time requirements due to the difficulty in processing dynamic range.
By performing aggregation hierarchical clustering of the numerical values of a random Gaussian distribution, a reference dictionary is generated, and a reference dictionary is linearly transformed according to the statistical characteristics of each tensor in the transformer, a localized dictionary is generated that is suitable for the tensor. Finally, the dictionary value of the localized dictionary is fitted exponentially and the dictionary index is generated to achieve efficient fixed-point operation.
It realizes efficient quantization acceleration of the Transformer model, reduces the requirements of computing resources and storage space, reduces training costs and deployment complexity, and quantifies weights and activation values at the same time, meeting real-time requirements.
Smart Images

Figure CN119848171B_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of quantization acceleration, and specifically relates to a method, device, storage medium, equipment and computer program product for quantization acceleration of a transformer in the text field based on dictionary index fixed-point operations. Background Art
[0002] With the rapid development of artificial intelligence technology, Transformer models have been widely used in the field of natural language processing. They can efficiently process large-scale text data and are applied to many fields such as machine translation, text generation, and semantic understanding. However, Transformer models are usually large in scale and have high computational complexity, making it difficult to run efficiently on resource-constrained devices. Therefore, quantizing and accelerating Transformer models has become an important research direction.
[0003] At present, the quantization acceleration technologies for Transformer models mainly include integer or fixed-point value-based quantization methods and dictionary-based quantization methods. Integer or fixed-point value-based quantization methods improve computational efficiency by quantizing floating-point values to lower-precision integers or fixed-point values. Dictionary-based quantization methods map weights or activation values to a predefined dictionary whose dictionary items are fixed-precision values.
[0004] However, methods that rely on fine-tuning require additional computing resources and storage space, increasing training costs and deployment complexity, while the quantization of activation values cannot meet real-time requirements due to the difficulty in handling the dynamic range. Summary of the invention
[0005] The present application aims to provide a method, apparatus, storage medium, device and computer program product for accelerating transformer quantization in the text field based on dictionary index fixed-point operations, which at least solves the problem that the scope of application and real-time performance of text transformer quantization compression cannot be achieved at the same time.
[0006] In a first aspect, the embodiment of the present application discloses a transformer quantization acceleration method in the text field based on dictionary index fixed-point operation, including:
[0007] Determine a reference dictionary for characterizing clustering relationships of a plurality of values of random Gaussian distribution according to aggregate hierarchical clustering;
[0008] Determine a first localized dictionary of the first tensor according to the first dictionary value determined by linear transformation of each first tensor in the transformer to the reference dictionary; the first tensor is a tensor in the transformer that conforms to a standard Gaussian distribution; the first dictionary value is floating point data;
[0009] Determine a first dictionary index for the first localized dictionary by performing an exponential function fitting on the first dictionary value; the first index value in the first dictionary index is integer data; the first index value is used to enable the transducer to obtain an execution result of the first target operation command according to a call of a target first tensor corresponding to a target first index value in the first index value when the first target operation command is obtained; the first target operation command is an operation command in which the target first tensor participates in the operation.
[0010] In a second aspect, an embodiment of the present application further discloses a transducer quantization acceleration device in the text field based on dictionary index fixed-point arithmetic, including:
[0011] A reference module, configured to determine a reference dictionary for characterizing a clustering relationship of the values according to an agglomerative hierarchical clustering of a plurality of values of a random Gaussian distribution;
[0012] A localization module, configured to determine a first localized dictionary of each first tensor according to a first dictionary value of each first tensor determined by a linear transformation of the reference dictionary by each first tensor in the transducer; the first tensor is a tensor in the transducer that conforms to a standard Gaussian distribution; the first dictionary value is floating-point data;
[0013] An index module, configured to determine a first dictionary index for the first localized dictionary by performing an exponential function fitting on the first dictionary value; the first index value in the first dictionary index is integer data; the first index value is used to enable the transducer to obtain an execution result of the first target operation command according to a call of a target first tensor corresponding to a target first index value in the first index value when the first target operation command is obtained; the first target operation command is an operation command in which the target first tensor participates in the operation.
[0014] In a third aspect, an embodiment of the present application further discloses a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps described in the first aspect are implemented.
[0015] In a fourth aspect, an embodiment of the present application further discloses an electronic device, including a processor, a memory, and a computer program stored on the memory and executable on the processor, and when the computer program is executed by the processor, the steps described in the first aspect are implemented.
[0016] In a fifth aspect, an embodiment of the present application further discloses a computer program product, on which a computer program is stored, and when the computer program is executed by a processor, the steps described in the first aspect are implemented.
[0017] In summary, in the embodiment of the present application, by aggregating and hierarchically clustering multiple values of a random Gaussian distribution to generate a reference dictionary, the numerical characteristics of the Gaussian distribution can be effectively characterized, so that subsequent quantization operations have higher accuracy, reducing the need to regenerate the dictionary each time, and improving computational efficiency; then, according to the statistical characteristics of each first tensor in the converter, the reference dictionary is linearly transformed to generate a first localized dictionary adapted to the tensor, and the characteristics of each tensor are combined with the reference dictionary, making the quantization process more flexible and efficient. Through linear transformation, it is ensured that the generated localized dictionary can accurately represent the data characteristics of each tensor, reducing quantization errors; finally, an exponential function is fitted to the dictionary value of the first localized dictionary to generate a first dictionary index. Using the exponential function fitting, complex floating-point operations are simplified to integer index operations, so that when the converter obtains the first target operation command, the corresponding target tensor can be quickly called according to the index value, reducing computational complexity and significantly improving computational efficiency. Therefore, based on the method of the embodiment of the present application, through the generation of a benchmark dictionary, linear transformation of tensors, exponential function fitting and generation of dictionary indexes, efficient quantization acceleration of the transformer model is achieved without fine-tuning, reducing the need for additional computing resources and storage space, reducing training costs and deployment complexity, and quantizing weights and activation values at the same time, effectively solving the problem of the difficult-to-handle dynamic range of activation values in the prior art, meeting real-time requirements, and solving the problem that the scope of application and real-time performance of text transformer quantization compression cannot be achieved at the same time. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Various other advantages and benefits will become apparent to those of ordinary skill in the art by reading the detailed description of the preferred embodiments below. The accompanying drawings are only for the purpose of illustrating the preferred embodiments and are not to be considered as limiting the present application. Also, the same reference symbols are used throughout the accompanying drawings to represent the same components. In the accompanying drawings:
[0019] Figure 1 It is a flowchart of the steps of a transformer quantization acceleration method in the text field based on dictionary index fixed-point operation provided by an embodiment of the present application;
[0020] Figure 2 It is a flowchart of another method for accelerating transformer quantization in the text field based on dictionary index fixed-point operation provided by an embodiment of the present application;
[0021] Figure 3 It is a data transfer process in the embodiment of the present application;
[0022] Figure 4 It is a structural schematic diagram of a transformer quantization acceleration device in the text field based on dictionary index fixed-point operation provided by an embodiment of the present application;
[0023] Figure 5 It is a block diagram of an electronic device provided by an embodiment of the present application;
[0024] Figure 6 It is a block diagram of another electronic device provided by an embodiment of the present application. Detailed implementation manners
[0025] The exemplary embodiments of the present application will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present application can be more thoroughly understood and the scope of the present application can be fully communicated to those skilled in the art.
[0026] Transformer models in the field of text, such as the Bidirectional Encoder Representations from Transformers (BERT) model, mainly rely on their internal multi-layer attention mechanisms during their inference operations. Each layer of the attention mechanism performs weighted calculations on the input data to capture long-range dependencies, thereby generating richer context representations. These attention mechanisms generate attention weights by calculating the similarity between query, key, and value matrices, and perform weighted summation on the input data based on these weights. During this process, the outputs and intermediate activation values of each layer of the model are continuously updated to reflect the complex relationships in the input data. Therefore, the Transformer model can effectively process and understand the complex semantic structures in natural language, and achieve efficient inference for tasks such as text classification, question answering systems, and machine translation.
[0027] The present application utilizes a rule discovered in a large number of experiments: the values in the Transformer model driven by the attention mechanism generally exhibit a Gaussian distribution, and the distribution of each layer is the same. Based on this, a benchmark dictionary can be generated by performing hierarchical clustering on these Gaussian-distributed values, and the statistical features of the first tensor in each Transformer can be linearly transformed to generate a localized dictionary to ensure that the dictionary can accurately represent the data characteristics of each tensor. Then, by fitting an exponential function to the localized dictionary to generate a dictionary index that is convenient for calculation, efficient fixed-point operations can be achieved. Specifically, through exponential function fitting, complex floating-point operations are simplified to integer index operations, which can greatly reduce the resources and time required for calculation, enabling the Transformer to quickly call the corresponding target tensor according to the integer index value during inference, significantly reducing the computational complexity, and improving the computational efficiency and real-time performance of the model. Finally, using these dictionary indexes, when an operation command is obtained, the corresponding target tensor can be quickly called to replace the complex attention operation mechanism in the original Transformer model.
[0028] Based on this, Figure 1 A text domain transducer quantization acceleration method based on dictionary index fixed-point operation provided by this embodiment specifically includes the following steps:
[0029] Step 101, determine a reference dictionary for characterizing the clustering relationship of numerical values according to the agglomerative hierarchical clustering of multiple numerical values of a random Gaussian distribution.
[0030] In some embodiments of the present application, a reference dictionary for characterizing the clustering relationship of numerical values is determined according to the agglomerative hierarchical clustering of multiple numerical values of a random Gaussian distribution. The reason for performing this step is to generate a reference dictionary that can effectively characterize the numerical characteristics of the Gaussian distribution. The process of performing this step includes generating random Gaussian distribution samples and applying the agglomerative hierarchical clustering (AC) method to cluster these samples into several clusters, and the center point value of each cluster is used as a dictionary entry. Agglomerative hierarchical clustering is a bottom-up clustering method that first treats each data point as an independent cluster and then gradually merges the clusters with the smallest distance until a predetermined goal is reached. After performing this step, the generated reference dictionary can be reused in subsequent quantization operations, reducing the need to regenerate the dictionary each time and improving the calculation efficiency.
[0031] In a specific example, first generate random Gaussian distribution data containing 50,000 samples, and these samples satisfy the conditions of a mean of 0 and a standard deviation of 1. Then, apply the agglomerative hierarchical clustering method to cluster the generated Gaussian distribution samples, and finally obtain 16 clusters. The center point value of each cluster is used as a fixed-point value (dictionary entry) to form a reference dictionary. According to the above execution process, the generated reference dictionary can effectively characterize the numerical characteristics of the Gaussian distribution and provide a reliable basis for subsequent quantization operations.
[0032] Step 102, determine the first local dictionary of the first tensor according to the first dictionary value of each first tensor determined by the linear transformation of each first tensor in the transducer with respect to the reference dictionary.
[0033] Among them, the first tensor is a tensor in the transducer that conforms to the standard Gaussian distribution; the first dictionary value is floating-point data.
[0034] In some embodiments of the present application, the first localized dictionary of the first tensors is determined based on the linear transformation of each first tensor in the transducer on the reference dictionary, and the first dictionary value of each first tensor is determined. The reason for performing this step is to generate a localized dictionary that adapts to the characteristics of each tensor, making the quantization process more flexible and efficient. The process of performing this step includes calculating the mean and standard deviation of each first tensor, and performing a linear transformation on the fixed-point values of the reference dictionary to generate a localized dictionary that adapts to the tensor. The fixed-point value refers to the floating-point data in the dictionary, and through linear transformation, it is ensured that the generated localized dictionary can accurately represent the data characteristics of each tensor. After performing this step, the generated localized dictionary can reduce the quantization error and improve the quantization accuracy.
[0035] In a specific example, the mean m and standard deviation s of a certain weight tensor in the transducer can be calculated first. Then, the fixed-point values in the reference dictionary are subjected to a linear transformation operation, that is, using the formula Dict' = s * Dict + m, where Dict' represents the first localized dictionary and Dict represents the reference dictionary. The reference dictionary is scaled and offset to generate a localized dictionary that adapts to the weight tensor. The localized dictionary generated according to the above execution process can accurately represent the data characteristics of the weight tensor, reduce the quantization error, and improve the quantization accuracy.
[0036] Step 103, determine the first dictionary index of the first localized dictionary by fitting the exponential function of the first dictionary value.
[0037] Among them, the first index value in the first dictionary index is integer data; the first index value is used to enable the transducer to obtain the execution result of the first target operation command according to the call of the target first tensor corresponding to the target first index value in the first index value when the first target operation command is obtained; the first target operation command is an operation command in which the target first tensor participates in the operation.
[0038] In some embodiments of the present application, the first dictionary index of the first localized dictionary is determined by fitting the exponential function of the first dictionary value. The reason for performing this step is to simplify complex floating-point operations into integer index operations, thereby improving the calculation efficiency. The process of performing this step includes fitting the exponential function to the first dictionary value to generate an integer index value. The exponential function fitting is a mathematical method that maps floating-point fixed-point values to integer index values. After performing this step, the generated first dictionary index can quickly call the corresponding target tensor according to the index value when the transducer obtains the first target operation command, reducing the calculation complexity.
[0039] In a specific example, an exponential function fitting method is applied to the floating-point fixed-point values in the localized dictionary of a weight tensor to map them to integer index values. These integer index values include a sign bit and an integer index bit, the sign bit indicates the positive or negative value, and the integer index bit indicates the quantization level. For example, 1011 represents the third quantization level of a negative number. According to the above execution process, the generated first dictionary index can quickly call the corresponding target tensor according to the index value when the converter obtains the first target operation command, thereby reducing the computational complexity and improving the computational efficiency.
[0040] In summary, in the embodiment of the present application, by aggregating and hierarchically clustering multiple values of a random Gaussian distribution to generate a reference dictionary, the numerical characteristics of the Gaussian distribution can be effectively characterized, so that subsequent quantization operations have higher accuracy, reducing the need to regenerate the dictionary each time, and improving computational efficiency; then, according to the statistical characteristics of each first tensor in the converter, the reference dictionary is linearly transformed to generate a first localized dictionary adapted to the tensor, and the characteristics of each tensor are combined with the reference dictionary, making the quantization process more flexible and efficient. Through linear transformation, it is ensured that the generated localized dictionary can accurately represent the data characteristics of each tensor, reducing quantization errors; finally, an exponential function is fitted to the dictionary value of the first localized dictionary to generate a first dictionary index. Using the exponential function fitting, complex floating-point operations are simplified to integer index operations, so that when the converter obtains the first target operation command, the corresponding target tensor can be quickly called according to the index value, reducing computational complexity and significantly improving computational efficiency. Therefore, based on the method of the embodiment of the present application, through the generation of a benchmark dictionary, linear transformation of tensors, exponential function fitting and generation of dictionary indexes, efficient quantization acceleration of the transformer model is achieved without fine-tuning, reducing the need for additional computing resources and storage space, reducing training costs and deployment complexity, and quantizing weights and activation values at the same time, effectively solving the problem of the difficult-to-handle dynamic range of activation values in the prior art, meeting real-time requirements, and solving the problem that the scope of application and real-time performance of text transformer quantization compression cannot be achieved at the same time.
[0041] Figure 2 This is another transformer quantization acceleration method in the text field based on dictionary index fixed-point operation provided by this embodiment, which specifically includes the following steps:
[0042] Step 201 : Determine a reference dictionary for characterizing the clustering relationship of the values based on the aggregated hierarchical clustering of multiple values of random Gaussian distribution.
[0043] The method shown in this step has been described in step 101 and will not be repeated here.
[0044] Optionally, step 201 includes the following sub-steps:
[0045] Sub-step 2011: Generate multiple groups of values according to a preset sample size so that each group of values follows a random Gaussian distribution respectively.
[0046] In some embodiments of the present application, multiple groups of values are generated according to a preset sample size so that each group of values follows a random Gaussian distribution respectively. The reason for performing this step is to ensure that the generated samples can accurately reflect the characteristics of the Gaussian distribution and provide a basis for subsequent clustering operations. The process of performing this step includes generating multiple groups of values according to the preset sample size, and each group of values is distributed according to the standard Gaussian distribution. The sample size refers to the number of generated values, ensuring that each group of values is representative. After performing this step, the multiple groups of generated values can effectively represent the Gaussian distribution and provide a reliable data basis for subsequent quantization processing and dictionary generation.
[0047] In a specific example, first, according to the preset sample size, five groups of values are generated, with each group containing 50,000 data points. Each group of values is generated according to the standard Gaussian distribution with a mean of 0 and a standard deviation of 1. These generated value sets can accurately reflect the characteristics of the Gaussian distribution and ensure representativeness in subsequent agglomerative hierarchical clustering operations. Multiple groups of values that conform to the random Gaussian distribution are generated according to the above execution process, laying a foundation for subsequent dictionary generation.
[0048] Sub-step 2012: Generate multiple reference dictionaries corresponding to each group of values according to the agglomerative hierarchical clustering operation on the multiple groups of values, so that each reference dictionary contains multiple value clusters respectively.
[0049] Among them, the number of value clusters contained in different reference dictionaries is the same; each value cluster has a corresponding central value respectively.
[0050] In some embodiments of the present application, multiple reference dictionaries corresponding to each group of values are generated according to the agglomerative hierarchical clustering operation on the multiple groups of values, so that each reference dictionary contains multiple value clusters respectively. The reason for performing this step is to utilize the numerical characteristics of the Gaussian distribution and generate multiple reference dictionaries through clustering operations. These dictionaries can effectively represent the characteristics of the value clusters. The process of performing this step includes performing an agglomerative hierarchical clustering operation on the multiple groups of values, dividing each group of values into several value clusters, and each value cluster has a corresponding central value. The agglomerative hierarchical clustering method is a bottom-up clustering method that generates a hierarchical clustering result by gradually merging the clusters with the smallest distance. After performing this step, the multiple generated reference dictionaries can accurately reflect the clustering relationship of each group of values and provide a basis for subsequent benchmark dictionary generation.
[0051] In a specific example, first, the agglomerative hierarchical clustering method is applied to the five generated sets of Gaussian distribution values. Each set of values is gradually merged with the cluster having the smallest distance according to the agglomerative hierarchical clustering method, and finally 16 numerical clusters are generated for each set. Each numerical cluster has a corresponding central value, and these central values are used as dictionary entries in the reference dictionary. Five reference dictionaries are generated according to the above execution process, and each reference dictionary contains 16 numerical clusters, laying a foundation for the generation of the benchmark dictionary.
[0052] Sub-step 2013: Obtain the benchmark dictionary according to the averaging operation on the central values of multiple reference dictionaries.
[0053] In some embodiments of the present application, the benchmark dictionary is obtained according to the averaging operation on the central values of multiple reference dictionaries. The reason for performing this step is to generate a benchmark dictionary that can represent the characteristics of the Gaussian distribution by averaging multiple reference dictionaries. The process of performing this step includes averaging the central values of multiple previously generated reference dictionaries, and using the average value of these central values as the dictionary entry of the benchmark dictionary. The central value refers to the center point of each numerical cluster in the reference dictionary, representing the statistical characteristics of the values within the cluster. After performing this step, the generated benchmark dictionary can accurately reflect the characteristics of the Gaussian distribution, providing a reliable data basis for subsequent quantization operations.
[0054] In a specific example, first, the central values in the five previously generated reference dictionaries are extracted. Each of these reference dictionaries contains 16 central values. Then, an averaging operation is performed on the central values at each position, that is, the average value of the central values at the corresponding position in the five reference dictionaries is calculated, and finally a benchmark dictionary containing 16 central values is generated. According to the above execution process, by averaging the central values of multiple reference dictionaries, a benchmark dictionary that can accurately characterize the characteristics of the Gaussian distribution is successfully generated, providing reliable data support for subsequent quantization processing.
[0055] Step 202: Determine the dictionary value of each tensor according to the linear transformation of the benchmark dictionary by each tensor in the transformer.
[0056] In some embodiments of the present application, the dictionary value of each tensor is determined according to the linear transformation of the benchmark dictionary by each tensor in the transformer. The reason for performing this step is to generate dictionary values adapted to the characteristics of each tensor, making the quantization process more accurate and efficient. The process of performing this step includes calculating the statistical characteristics of each tensor, such as the mean and standard deviation, and performing a linear transformation on the fixed-point values of the benchmark dictionary to generate dictionary values adapted to the tensor. The fixed-point value refers to the floating-point data in the dictionary, and through the linear transformation, it is ensured that the generated dictionary values can accurately represent the data characteristics of each tensor. After performing this step, the generated dictionary values can reduce the quantization error and improve the quantization accuracy.
[0057] In a specific example, first calculate the mean m and standard deviation s of a certain weight tensor in the transformer. Then, through a linear transformation operation on the fixed-point values in the reference dictionary, that is, using the formula Dict' = s * Dict + m, scale and offset the reference dictionary to generate dictionary values adapted to this weight tensor. The dictionary values generated according to the above execution process can accurately represent the data characteristics of this weight tensor, reduce quantization errors, and improve quantization accuracy.
[0058] Step 203, compare the statistical characteristics of each tensor with the dictionary values according to the preset statistical matching conditions to determine the first tensor and the second tensor in the tensor.
[0059] Among them, the second tensor is a tensor other than the tensor that conforms to the standard Gaussian distribution in the transformer.
[0060] In some embodiments of the present application, compare the statistical characteristics of each tensor with the dictionary values according to the preset statistical matching conditions to determine the first tensor and the second tensor in the tensor. The reason for performing this step is to distinguish tensors that conform to the standard Gaussian distribution from tensors that do not conform to the standard Gaussian distribution, so as to perform appropriate processing on different types of tensors. The process of performing this step includes calculating the statistical characteristics (such as mean and standard deviation) of each tensor and comparing them with the dictionary values to determine the type of the tensor. Statistical characteristics refer to the data characteristics of the tensor, including parameters such as mean and standard deviation. After performing this step, the first tensor and the second tensor can be accurately distinguished, providing a basis for quantization and processing in subsequent steps.
[0061] In a specific example, the mean or standard deviation of each tensor in the transformer can be calculated first. Then, compare the statistical characteristics of each tensor with the dictionary values in the reference dictionary. According to the preset matching conditions, the tensor that conforms to the standard Gaussian distribution is determined as the first tensor, and the tensor that does not conform to the standard Gaussian distribution is determined as the second tensor. According to the above execution process, the first tensor and the second tensor can be accurately distinguished, providing a basis for subsequent quantization and processing.
[0062] Optionally, step 203 includes the following sub-steps:
[0063] Sub-step 2031, calculate the mean of each tensor.
[0064] In some embodiments of the present application, the mean of each tensor is calculated. The reason for performing this step is to obtain the statistical characteristics of each tensor, which serves as the basis for distinguishing the first tensor and the second tensor in subsequent steps. The process of performing this step includes performing statistical calculations on all the values in each tensor to obtain its mean. The mean refers to the arithmetic average of all the element values in the tensor and represents the central tendency of the tensor. After performing this step, the mean of each tensor can be accurately obtained, providing data support for subsequent quantization and classification operations.
[0065] In a specific example, it is necessary to calculate the mean of a certain weight tensor. First, perform a statistical sum of all the elements in the tensor, and then divide the sum result by the total number of elements to obtain the mean of the tensor. According to the above execution process, the mean of the weight tensor can be accurately calculated, providing reliable statistical data for determining the tensor type in subsequent steps.
[0066] Sub-step 2032: When the ratio between the mean of the tensor and the dictionary value of the tensor is less than a preset scale threshold, determine the tensor as the first tensor; when the ratio between the mean of the tensor and the dictionary value of the tensor is greater than or equal to the preset scale threshold, determine the tensor as the second tensor.
[0067] In some embodiments of the present application, when the ratio between the mean of the tensor and the dictionary value of the tensor is less than a preset scale threshold, determine the tensor as the first tensor; when the ratio between the mean of the tensor and the dictionary value of the tensor is greater than or equal to the preset scale threshold, determine the tensor as the second tensor. The reason for performing this step is to distinguish the first tensor and the second tensor according to the ratio of the statistical characteristics of the tensor to the dictionary value, so as to appropriately process different types of tensors. The process of performing this step includes calculating the mean of each tensor, then calculating the ratio of the mean to the dictionary value of the tensor, and comparing it with the preset scale threshold. If the ratio is less than the threshold, determine the tensor as the first tensor; if the ratio is greater than or equal to the threshold, determine the tensor as the second tensor. The ratio refers to the proportional relationship between the mean of the tensor and the dictionary value and is used to judge the type of the tensor. After performing this step, the tensor type can be accurately distinguished, providing a basis for subsequent quantization and processing.
[0068] In a specific example, first calculate the mean of a certain tensor as m and extract the dictionary value of the tensor as d. Next, calculate the ratio of m to d and compare it with the preset scale threshold. For example, the preset scale threshold is 1.5. If the ratio m / d is less than 1.5, determine the tensor as the first tensor; if the ratio m / d is greater than or equal to 1.5, determine the tensor as the second tensor. According to the above execution process, the first tensor and the second tensor can be accurately distinguished, providing a reliable basis for quantization and processing in subsequent steps.
[0069] Step 204: Generate a localized dictionary for the first tensor according to the first dictionary values corresponding to the first tensor.
[0070] In some embodiments of the present application, a localized dictionary for the first tensor is generated according to the first dictionary values corresponding to the first tensor. The reason for performing this step is to utilize the specific dictionary values of each first tensor to generate a localized dictionary that matches its data characteristics, thereby further improving the quantization accuracy. The process of performing this step includes applying the first dictionary values of each first tensor to the generated reference dictionary, and through a linear transformation step, generating a localized dictionary adapted to the tensor. Dictionary values refer to the floating-point data in the reference dictionary, and after these data are linearly transformed, they can better represent the corresponding first tensor. After performing this step, the generated localized dictionary can accurately describe the characteristics of each first tensor, reduce quantization errors, and improve the model performance.
[0071] In a specific example, first, according to the statistical characteristics of a certain weight tensor, the first dictionary values are extracted from the reference dictionary. Then, linear transformation and exponential function fitting operations are performed on these dictionary values to convert them into data characteristics adapted to the weight tensor. Finally, a localized dictionary for the weight tensor is generated to ensure that in subsequent quantization operations, the dictionary can accurately represent the numerical characteristics of the tensor. The localized dictionary generated according to the above execution process can accurately describe the corresponding first tensor, reduce quantization errors, and improve the accuracy and efficiency of model quantization.
[0072] Step 205: Determine the first dictionary index for the first localized dictionary by performing an exponential function fitting on the first dictionary values.
[0073] Among them, the first index value in the first dictionary index is integer data; the first index value is used to enable the transducer to obtain the execution result of the first target operation command according to the call of the target first tensor corresponding to the target first index value in the first index value; the first target operation command is an operation command in which the target first tensor participates in the operation.
[0074] The method shown in this step has been described in step 103 and will not be elaborated here.
[0075] Optionally, step 205 includes the following sub-steps:
[0076] Sub-step 2051: Perform an exponential function fitting on the first dictionary values in the form of to determine the binary array and binary parameters corresponding to each first dictionary value.
[0077] Among them, the first variable of the binary array is used to represent the magnitude relationship between each first dictionary value and the average value of the first dictionary values by a one-bit integer data; the second variable of the binary array is used to represent the exponential value of the first parameter a in the binary parameter after each first dictionary value is decomposed according to the exponential function; b is the second parameter in the binary parameter.
[0078] In some embodiments of the present application, the exponential function fitting of the first dictionary values is performed in the form of to determine the binary array and the binary parameter corresponding to each first dictionary value. The reason for performing this step is to map each first dictionary value into a binary array and a binary parameter that are convenient for calculation through exponential function fitting. The process of performing this step includes performing exponential function fitting on each first dictionary value to determine the fitting result, which includes a binary array and a binary parameter. The first variable of the binary array is used to represent the magnitude relationship between each first dictionary value and the average value of the first dictionary values by a one-bit integer data, and the second variable is used to represent the exponential value of the first parameter a in the binary parameter after each first dictionary value is decomposed according to the exponential function. b is the second parameter in the binary parameter. After performing this step, the generated binary array and binary parameter can be used for subsequent quantization and calculation, improving the calculation efficiency of the model.
[0079] In a specific example, the exponential function fitting method is applied to the first dictionary values of a certain tensor, and the fitting is performed in the form of After fitting, the binary array and the binary parameter corresponding to each first dictionary value are determined. The first variable of the binary array represents the magnitude relationship between each first dictionary value and its average value, and the second variable represents the exponential value of the exponential function. a and b in the binary parameter are the parameters in the fitting respectively. According to the above execution process, the binary array and the binary parameter for subsequent quantization and calculation are generated, improving the calculation efficiency of the model.
[0080] Sub-step 2052: Combine and record the binary parameters corresponding to each first dictionary value respectively to obtain the first index value corresponding to each first dictionary value respectively.
[0081] In some embodiments of the present application, the binary parameters corresponding to each first dictionary value are merged and recorded respectively to obtain the first index value corresponding to each first dictionary value. The reason for performing this step is to integrate the binary parameters obtained by exponential function fitting and generate the first index value for quick lookup and calculation. The process of performing this step includes merging the binary parameters in the binary array corresponding to each first dictionary value and recording the merged result as the first index value. The binary parameters refer to two parameters obtained by exponential function fitting, which are used to represent the numerical relationship and the exponential value respectively. After performing this step, the generated first index value can be used for the quick lookup and calculation of the transducer, improving the calculation efficiency of the model.
[0082] In a specific example, the first dictionary value of a certain tensor is integrated with its corresponding binary parameters. The binary parameters include the first variable representing the numerical relationship and the second variable representing the exponential value. These binary parameters are merged into a complete record and used as the first index value of this first dictionary value. According to the above execution process, the complete first index values are generated, and these index values can be used for the quick lookup and calculation of the transducer, improving the calculation efficiency of the model.
[0083] The following will explain the operation simplification principle brought by exponential function fitting:
[0084] For a dataset conforming to the Gaussian distribution, assume that 4-bit quantization indices are obtained and utilized in model inference to simplify the multiplication and addition operations and achieve model quantization acceleration (at this time, the calculation of the outlier subset part is not considered for the time being). Assume that there is a weight tensor, which can be expressed as , and similarly, the activation value tensor is expressed as . Among them, n is the dimension of A and W, A i , and W i are the i-th components of A and W respectively, a is the first parameter used for fitting, represents the sign bit, represents its corresponding standard deviation, is the bias value related to the mean, standard deviation, and the second parameter b used for fitting. In the inference process of the model, each output activation is the sum of the product of the previous activation value and the weight, which is also the step with the highest computational complexity in the Transformer. Given the form of the input value, the output activation is expressed as:
[0085]
[0086] It can be seen that the output activation consists of the sum of four terms. These four terms are analyzed separately:
[0087] The focus of the first term is the sum of integers in the exponent. When using the obtained 4-bit quantization values, the range of this sum only contains 15 unique values. Therefore, for the addition operation of this term, we can directly obtain the main value of this part by counting the number of occurrences of the sum of quantization values on each exponent, and then multiplying and accumulating the number of occurrences and the corresponding values. And the multiplication and addition operation here is the addition of integer multiples of a series of known numbers, and we can obtain the values of these numbers before model inference without additional operations;
[0088] The second term can be decomposed into a constant and a part similar to the steps of the first term. The only difference is that here, instead of the sum of quantization values, individual quantization values need to be counted, which contains 8 unique values. Similar to part b, by multiplying and accumulating the number of occurrences and the corresponding values, and then adding the constant bias, the main value of this part can be obtained;
[0089] The third term is the processing of the activation value vector, which is completely similar to the processing steps of the second term and will not be elaborated here;
[0090] The fourth term can be disassembled into some fixed constants regarding the fitting parameters a and b, and the sum of the multiplication results of symbolic values. The former of this part is some constants that can be calculated in advance, and the latter is completely an operation between integers, so its computational complexity is very low.
[0091] Thus, it can be shown that by adding up each term through the above steps, the quantization processing of the output activation value is completed. By successfully using the counting of index values, a large number of floating-point multiplication and addition operations are significantly reduced, thereby achieving the quantization acceleration of the overall model.
[0092] Step 206: According to the hierarchical clustering of the second tensor, determine a second localization dictionary for characterizing the clustering relationship of the second tensor, and generate a second dictionary index for the second localization dictionary.
[0093] Among them, the second index value in the second dictionary index is used to enable the transducer to obtain the execution result of the second target operation command according to the call of the target second tensor corresponding to the target second index value in the second index value when the second target operation command is obtained; the second target operation command is the operation command in which the target second tensor participates in the operation.
[0094] In some embodiments of the present application, according to the agglomerative hierarchical clustering of the second tensor, a second localization dictionary for characterizing the clustering relationship of the second tensor is determined, and a second dictionary index for the second localization dictionary is generated. The reason for performing this step is to cluster the second tensor that does not conform to the standard Gaussian distribution, so as to generate a localization dictionary and a dictionary index that can effectively characterize the characteristics of these tensors. The process of performing this step includes performing agglomerative hierarchical clustering on the second tensor, generating a second localization dictionary representing its clustering relationship, and generating a second dictionary index by exponential function fitting or other methods. The second dictionary index is used to quickly call the corresponding target second tensor according to the target second index value when the transducer obtains the second target operation command, and execute the operation command. After performing this step, the generated second localization dictionary and second dictionary index can effectively characterize the second tensor, improving the computational efficiency and performance of the model.
[0095] In a specific example, first, agglomerative hierarchical clustering is performed on some tensors (second tensors) in the transducer that do not conform to the standard Gaussian distribution to generate a second localization dictionary representing their clustering relationship. Then, a second dictionary index is generated for the second localization dictionary to characterize the target second tensor. According to the above execution process, when the transducer obtains the second target operation command, it can quickly call the corresponding target second tensor according to the target second index value and execute the operation command, further improving the computational efficiency and performance.
[0096] Figure 3 This is the data flow process in the next embodiment of the present application:
[0097] S1. Agglomerative hierarchical classification: Perform agglomerative hierarchical classification on the basis of the Gaussian distribution to generate a reference dictionary. In this step, through the clustering operation, the Gaussian distribution samples are divided into multiple clusters, and the central point values of each cluster form the dictionary entries of the reference dictionary, laying a foundation for subsequent quantization operations;
[0098] S2. Classify tensors by calculating variance and mean: Classify tensors according to variance and mean to generate a Gaussian value subset and an outlier subset. In this step, by calculating the variance and mean of the tensors, it is judged whether they conform to the standard Gaussian distribution, and the tensors that conform are used as the Gaussian value subset, and the rest are used as the outlier subset;
[0099] S3.1. Linear transformation: Perform a linear transformation on the Gaussian value subset to generate a quantization dictionary. In this step, according to the characteristics of the Gaussian value subset, the reference dictionary is linearly transformed to adapt to the data characteristics of the Gaussian value subset, generating a quantization dictionary to improve the quantization accuracy;
[0100] S3.2, separate clustering: cluster the outlier subset separately and generate an outlier dictionary. This step clusters the outlier subset and generates a dictionary that adapts to the data characteristics of these outliers to ensure the accuracy of the quantization process.
[0101] Finally, in the actual model calculation process (S0), the weight / embedding tensor and activation value tensor of the Transformer model are processed in the quantization process to generate a 4-bit quantization index for index value calculation, which converts the operation form during model inference into a form suitable for quantization, thereby improving the computational efficiency and performance of the model.
[0102] In summary, in the embodiment of the present application, by aggregating and hierarchically clustering multiple values of a random Gaussian distribution to generate a reference dictionary, the numerical characteristics of the Gaussian distribution can be effectively characterized, so that subsequent quantization operations have higher accuracy, reducing the need to regenerate the dictionary each time, and improving computational efficiency; then, according to the statistical characteristics of each first tensor in the converter, the reference dictionary is linearly transformed to generate a first localized dictionary adapted to the tensor, and the characteristics of each tensor are combined with the reference dictionary, making the quantization process more flexible and efficient. Through linear transformation, it is ensured that the generated localized dictionary can accurately represent the data characteristics of each tensor, reducing quantization errors; finally, an exponential function is fitted to the dictionary value of the first localized dictionary to generate a first dictionary index. Using the exponential function fitting, complex floating-point operations are simplified to integer index operations, so that when the converter obtains the first target operation command, the corresponding target tensor can be quickly called according to the index value, reducing computational complexity and significantly improving computational efficiency. Therefore, based on the method of the embodiment of the present application, through the generation of a benchmark dictionary, linear transformation of tensors, exponential function fitting and generation of dictionary indexes, efficient quantization acceleration of the transformer model is achieved without fine-tuning, reducing the need for additional computing resources and storage space, reducing training costs and deployment complexity, and quantizing weights and activation values at the same time, effectively solving the problem of the difficult-to-handle dynamic range of activation values in the prior art, meeting real-time requirements, and solving the problem that the scope of application and real-time performance of text transformer quantization compression cannot be achieved at the same time.
[0103] like Figure 4 As shown, the embodiment of the present application also discloses a text field transformer quantization acceleration device 30 based on dictionary index fixed-point operation, including:
[0104] A reference module 301 is used to determine a reference dictionary for characterizing the clustering relationship of the values according to the aggregated hierarchical clustering of multiple values of random Gaussian distribution;
[0105] A localization module 302, configured to determine a first localization dictionary of a first tensor according to a first dictionary value of each first tensor determined by a linear transformation of a reference dictionary by each first tensor in a transducer; the first tensor is a tensor in the transducer that conforms to a standard Gaussian distribution; the first dictionary value is floating-point data;
[0106] An indexing module 303, configured to determine a first dictionary index of the first localization dictionary by fitting an exponential function to the first dictionary value; a first index value in the first dictionary index is integer data; the first index value is used to enable the transducer to obtain an execution result of a first target operation command according to a call of a target first tensor corresponding to a target first index value in the first index value when the first target operation command is obtained; the first target operation command is an operation command in which the target first tensor participates in an operation.
[0107] Optionally, the reference module 301 includes:
[0108] A numerical generation sub-module, configured to generate multiple groups of numerical values according to a preset sample size, so that each group of numerical values respectively follows a random Gaussian distribution;
[0109] A clustering sub-module, configured to generate multiple reference dictionaries respectively corresponding to each group of numerical values according to an agglomerative hierarchical clustering operation on the multiple groups of numerical values, so that each reference dictionary respectively includes multiple numerical clusters; the number of numerical clusters included in different reference dictionaries is the same; each numerical cluster respectively has a corresponding central numerical value;
[0110] An averaging sub-module, configured to obtain a reference dictionary according to an averaging operation on the central numerical values of the multiple reference dictionaries.
[0111] Optionally, the localization module 302 includes:
[0112] A linear transformation sub-module, configured to determine a dictionary value of each tensor according to a linear transformation of the reference dictionary by each tensor in the transducer;
[0113] A first screening sub-module, configured to compare a statistical feature of each tensor with the dictionary value according to a preset statistical matching condition to determine a first tensor in the tensors;
[0114] A first dictionary sub-module, configured to generate a localization dictionary of the first tensor according to the first dictionary value corresponding to the first tensor.
[0115] Optionally, the first screening sub-module includes:
[0116] A mean unit, configured to calculate a mean of each tensor;
[0117] A screening unit, configured to determine a tensor as the first tensor when a ratio between the mean of the tensor and the dictionary value of the tensor is less than a preset scale threshold.
[0118] Optionally, the transducer quantization acceleration device 30 in the text field based on dictionary index fixed-point operation further includes:
[0119] A second screening module, configured to compare the statistical characteristics of each tensor with dictionary values according to a preset statistical matching condition to determine a second tensor in the tensor; the second tensor is a tensor other than the tensor conforming to the standard Gaussian distribution in the transducer;
[0120] A second dictionary module, configured to determine a second localization dictionary for characterizing the clustering relationship of the second tensor according to the hierarchical clustering of the second tensor, and generate a second dictionary index for the second localization dictionary; the second index value in the second dictionary index is used to enable the transducer to obtain the execution result of the second target operation command according to the call of the target second tensor corresponding to the target second index value in the second index value when obtaining the second target operation command; the second target operation command is an operation command in which the target second tensor participates in the operation.
[0121] Optionally, the index module 303 includes:
[0122] A fitting sub-module, configured to fit the exponential function of the first dictionary value in the form of to determine a binary array and binary parameters corresponding to each first dictionary value; the first variable of the binary array is used to represent the magnitude relationship between each first dictionary value and the average value of the first dictionary values through a one-bit integer data; the second variable of the binary array is used to represent the exponential value of the first parameter a in the binary parameters after each first dictionary value is decomposed according to the exponential function; b is the second parameter in the binary parameters.
[0123] An index sub-module, configured to merge and record the binary parameters corresponding to each first dictionary value respectively to obtain a first index value corresponding to each first dictionary value.
[0124] In summary, in the embodiment of the present application, by aggregating and hierarchically clustering multiple values of a random Gaussian distribution to generate a reference dictionary, the numerical characteristics of the Gaussian distribution can be effectively characterized, so that subsequent quantization operations have higher accuracy, reducing the need to regenerate the dictionary each time, and improving computational efficiency; then, according to the statistical characteristics of each first tensor in the converter, the reference dictionary is linearly transformed to generate a first localized dictionary adapted to the tensor, and the characteristics of each tensor are combined with the reference dictionary, making the quantization process more flexible and efficient. Through linear transformation, it is ensured that the generated localized dictionary can accurately represent the data characteristics of each tensor, reducing quantization errors; finally, an exponential function is fitted to the dictionary value of the first localized dictionary to generate a first dictionary index. Using the exponential function fitting, complex floating-point operations are simplified to integer index operations, so that when the converter obtains the first target operation command, the corresponding target tensor can be quickly called according to the index value, reducing computational complexity and significantly improving computational efficiency. Therefore, based on the method of the embodiment of the present application, through the generation of a benchmark dictionary, linear transformation of tensors, exponential function fitting and generation of dictionary indexes, efficient quantization acceleration of the transformer model is achieved without fine-tuning, reducing the need for additional computing resources and storage space, reducing training costs and deployment complexity, and quantizing weights and activation values at the same time, effectively solving the problem of the difficult-to-handle dynamic range of activation values in the prior art, meeting real-time requirements, and solving the problem that the scope of application and real-time performance of text transformer quantization compression cannot be achieved at the same time.
[0125] The embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, each process of the above-mentioned transformer quantization acceleration method embodiment in the text field based on dictionary index fixed-point operation is implemented, and the same technical effect can be achieved. To avoid repetition, it is not repeated here. The computer-readable storage medium is, for example, a read-only memory (ROM), a random access memory (RAM), a disk or an optical disk, etc.
[0126] Figure 5 700 is a block diagram of an electronic device 700 provided in an embodiment of the present application. For example, the electronic device 700 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0127] Reference Figure 5, the electronic device 700 may include one or more of the following components: a processing component 702, a memory 704, a power component 706, a multimedia component 708, an audio component 710, an input / output (I / O) interface 712, a sensor component 714, and a communication component 716.
[0128] The processing component 702 generally controls the overall operation of the electronic device 700, such as operations associated with display, telephone calls, data communications, camera operations, and recording operations. The processing component 702 may include one or more processors 720 to execute instructions to complete all or part of the steps of the above-mentioned text domain transformer quantization acceleration method based on dictionary index fixed-point arithmetic. In addition, the processing component 702 may include one or more modules to facilitate the interaction between the processing component 702 and other components. For example, the processing component 702 may include a multimedia module to facilitate the interaction between the multimedia component 708 and the processing component 702.
[0129] The memory 704 is used to store various types of data to support the operation of the electronic device 700. Examples of such data include instructions for any application or method operating on the electronic device 700, contact data, phone book data, messages, pictures, multimedia, etc. The memory 704 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.
[0130] The power component 706 provides power to various components of the electronic device 700. The power component 706 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the electronic device 700.
[0131] The multimedia component 708 includes a screen that provides an output interface between the electronic device 700 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operations. In some embodiments, the multimedia component 708 includes a front camera and / or a rear camera. When the electronic device 700 is in an operating mode, such as a shooting mode or a multimedia mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capabilities.
[0132] The audio component 710 is used to output and / or input audio signals. For example, the audio component 710 includes a microphone (MIC) that is used to receive external audio signals when the electronic device 700 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 704 or transmitted via the communication component 716. In some embodiments, the audio component 710 further includes a speaker for outputting audio signals.
[0133] The I / O interface 712 provides an interface between the processing component 702 and a peripheral interface module, and the peripheral interface module can be a keyboard, a click wheel, buttons, etc. These buttons can include but are not limited to: a home button, a volume button, a power button, and a lock button.
[0134] The sensor component 714 includes one or more sensors for providing status assessments of various aspects of the electronic device 700. For example, the sensor component 714 can detect the on / off state of the electronic device 700, the relative positioning of components, such as the display and the keypad of the electronic device 700. The sensor component 714 can also detect a change in the position of the electronic device 700 or a component of the electronic device 700, the presence or absence of user contact with the electronic device 700, the orientation or acceleration / deceleration of the electronic device 700, and the temperature change of the electronic device 700. The sensor component 714 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor component 714 can also include a light sensor, such as a CMOS or a CCD image sensor, for use in imaging applications. In some embodiments, the sensor component 714 can further include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0135] The communication component 716 is used to facilitate communication between the electronic device 700 and other devices in a wired or wireless manner. The electronic device 700 can access a communication standard-based wireless network, such as WiFi, a carrier network (such as 2G, 3G, 4G, or 7G), or a combination thereof. In an exemplary embodiment, the communication component 716 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 716 further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra Wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0136] In an exemplary embodiment, the electronic device 700 can be implemented by one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components, for implementing the text domain transformer quantization acceleration method based on dictionary index fixed-point arithmetic provided in the embodiments of the present application.
[0137] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 704 including instructions, and the above instructions can be executed by a processor 720 of the electronic device 700 to complete the above text domain transformer quantization acceleration method based on dictionary index fixed-point arithmetic. For example, the non-transitory storage medium can be a ROM, Random Access Memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.
[0138] Figure 6 is a block diagram of an electronic device 800 shown according to an exemplary embodiment. For example, the electronic device 800 can be provided as a server. Referring to Figure 6 , the electronic device 800 includes a processing component 822, which further includes one or more processors, and memory resources represented by a memory 832 for storing instructions executable by the processing component 822, such as application programs. The application programs stored in the memory 832 can include one or more modules each corresponding to a set of instructions. In addition, the processing component 822 is configured to execute instructions to perform the text domain transformer quantization acceleration method based on dictionary index fixed-point arithmetic provided in the embodiments of the present application.
[0139] The electronic device 800 may further include a power supply component 826 configured to perform power management of the electronic device 800, a wired or wireless network interface 850 configured to connect the electronic device 800 to a network, and an input / output (I / O) interface 858. The electronic device 800 may operate based on an operating system stored in the memory 832, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSD TM or the like.
[0140] The embodiments of the present application also provide a computer program product, including a computer program, which when executed by a processor implements a method for accelerating the quantization of a transducer in the text field based on dictionary-indexed fixed-point arithmetic.
[0141] Those skilled in the art will readily conceive of other embodiments of the present application after considering the specification and practicing the application disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include known common knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only to be considered as exemplary, and the true scope and spirit of the present application are pointed out by the following claims.
[0142] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.
[0143] Each embodiment in this specification is described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other.
[0144] It is easy for those skilled in the art to think that any combination application of the above various embodiments is feasible. Therefore, any combination of the above various embodiments is an implementation scheme of the present application. However, due to space limitations, this specification will not elaborate on each of them here.
[0145] The method for accelerating the quantization of a transducer in the text field based on dictionary-indexed fixed-point arithmetic provided herein is not inherently related to any specific computer, virtual system, or other device. Various general-purpose systems can also be used in conjunction with the teachings herein. It is obvious from the above description how to construct the required structure of a system with the solution of the present application. In addition, the present application is not directed to any specific programming language. It should be understood that the content of the present application described herein can be implemented using various programming languages, and the description of the specific language above is for disclosing the best implementation mode of the present application.
[0146] In the specification provided herein, a large number of specific details are set forth. However, it will be understood that embodiments of the present application may be practiced without these specific details. In some instances, well-known methods, structures and techniques have not been shown in detail so as not to obscure the understanding of this specification.
[0147] Similarly, it should be understood that, in order to streamline the present application and assist in understanding one or more of the various aspects of the application, in the above description of the exemplary embodiments of the present application, the various features of the present application are sometimes grouped together into a single embodiment, figure, or description thereof. However, the disclosed method should not be construed as reflecting an intention that the claimed present application requires more features than are expressly recited in each claim. Rather, as reflected by the claims, the aspects of the application lie in less than all of the features of the single foregoing disclosed embodiment. Thus, the claims following the detailed description are hereby expressly incorporated into the detailed description, with each claim standing on its own as a separate embodiment of the present application.
[0148] Those skilled in the art will appreciate that the modules in the devices in the embodiments can be adaptively changed and disposed in one or more devices different from the embodiments. The modules or units or components in the embodiments can be combined into one module or unit or component, and in addition, they can be divided into multiple sub-modules or sub-units or sub-components. Except that at least some of such features and / or processes or units are mutually exclusive, any combination can be used to combine all the features disclosed in this specification (including the accompanying claims, abstract and drawings) and all the processes or units of any method or device so disclosed. Unless otherwise expressly stated, each feature disclosed in this specification (including the accompanying claims, abstract and drawings) can be replaced by an alternative feature that provides the same, equivalent or similar purpose.
[0149] In addition, those skilled in the art will be able to understand that although some of the embodiments described herein include certain features included in other embodiments rather than other features, the combination of the features of different embodiments means that it is within the scope of the present application and forms different embodiments. For example, in the claims, any one of the claimed embodiments can be used in any combination.
[0150] Each component embodiment of the present application can be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. Those skilled in the art should understand that a microprocessor or a digital signal processor (DSP) can be used in practice to implement some or all of the functions of some or all of the components in the text domain transducer quantization acceleration method based on dictionary index fixed-point arithmetic according to the embodiments of the present application. The present application can also be implemented as a device or device program (such as a computer program and a computer program product) for executing part or all of the methods described herein. Such a program for implementing the present application can be stored on a computer-readable medium, or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, or provided on a carrier signal, or provided in any other form.
[0151] In another embodiment provided by the present invention, there is also provided a computer program product containing instructions, which when run on a computer, causes the computer to execute the transducer quantization acceleration method in the text domain based on dictionary index fixed-point arithmetic according to the embodiments of the present application.
[0152] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server, data center, etc. that integrates one or more available media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).
[0153] It should be noted that the above embodiments are illustrative of the present application rather than restrictive of the present application, and those skilled in the art can design alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. The word "comprising" does not exclude the presence of elements or steps not listed in the claim. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present application can be implemented by means of hardware including several different elements and by means of a suitably programmed computer. In a unit claim listing several devices, several of these devices may be embodied by the same item of hardware. The use of the words first, second, and third, etc. does not denote any order. These words can be interpreted as names.
[0154] It should be noted that, for the method embodiments of the present application, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should understand that the embodiments of the present application are not limited by the described order of actions, because according to the embodiments of the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also understand that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily essential to the embodiments of the present application.
[0155] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the embodiments of the system or device, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiments.
[0156] The above are only the preferred embodiments of the present invention and are not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention are included in the protection scope of the present invention.
Claims
1. A transformer quantization acceleration method in the text field based on dictionary index fixed-point operation, characterized in that: include: Determine a reference dictionary for characterizing clustering relationships of a plurality of values of random Gaussian distribution according to aggregate hierarchical clustering; Determine a first localized dictionary of each first tensor according to the first dictionary value determined by performing a linear transformation of each first tensor in the transformer on the reference dictionary; The first tensor is a tensor in the converter that conforms to a standard Gaussian distribution; the first dictionary value is floating point data; Determine a first dictionary index for the first localized dictionary by fitting an exponential function to the first dictionary value; a first index value in the first dictionary index is an integer data; The first index value is used to enable the converter to obtain the execution result of the first target operation command by calling the target first tensor corresponding to the target first index value in the first index value when the first target operation command is obtained; the first target operation command is an operation command in which the target first tensor participates in the operation; The determining the first dictionary index for the first localized dictionary by fitting an exponential function to the first dictionary value includes: according to The exponential function of the first dictionary value is fitted in the form of to determine a binary array corresponding to each of the first dictionary values and a binary parameter; the first variable of the binary array is used to represent the magnitude relationship between each of the first dictionary values and the average value of the first dictionary values through one bit of integer data; the second variable of the binary array is used to represent the exponential value of each of the first dictionary values after decomposition according to the exponential function to the first parameter a in the binary parameter through multi-bit integer data; b is the second parameter in the binary parameter; The binary parameters corresponding to each of the first dictionary values are merged and recorded to obtain first index values corresponding to each of the first dictionary values.
2. The transformer quantization acceleration method in the text field based on dictionary index fixed-point operation as claimed in claim 1, characterized in that: The step of determining a reference dictionary for characterizing the clustering relationship of the values according to the aggregated hierarchical clustering of the multiple values of the random Gaussian distribution includes: Generate multiple groups of values according to a preset sample capacity, so that each group of values obeys a random Gaussian distribution; According to the aggregated hierarchical clustering operation on the multiple groups of the values, a plurality of reference dictionaries corresponding to each group of the values are generated, so that each of the reference dictionaries contains a plurality of value clusters; different reference dictionaries contain the same number of value clusters; each of the value clusters has a corresponding central value; The reference dictionary is obtained by averaging central values of a plurality of the reference dictionaries.
3. The transformer quantization acceleration method in the text field based on dictionary index fixed-point operation as claimed in claim 1, characterized in that: The method of determining a first localized dictionary of each first tensor according to the linear transformation of the reference dictionary by each first tensor in the transformer and the first dictionary value of each first tensor determined comprises: Determining a dictionary value of each tensor according to a linear transformation of the reference dictionary by each tensor in the transformer; According to a preset statistical matching condition, compare the statistical features of each of the tensors with the dictionary value to determine a first tensor in the tensors; Generate a localized dictionary for the first tensor according to a first dictionary value corresponding to the first tensor.
4. The transformer quantization acceleration method in the text field based on dictionary index fixed-point operation as claimed in claim 3, characterized in that: The step of comparing the statistical features of each of the tensors with the dictionary value according to a preset statistical matching condition to determine the first tensor in the tensors includes: computing the mean of each of said tensors; When a ratio between the mean value of the tensor and the dictionary value of the tensor is less than a preset scale threshold, the tensor is determined as a first tensor.
5. The transformer quantization acceleration method in the text field based on dictionary index fixed-point operation as claimed in claim 3, characterized in that: The method further comprises: According to a preset statistical matching condition, the statistical feature of each tensor is compared with the dictionary value to determine a second tensor in the tensor; the second tensor is a tensor in the converter other than the tensor that conforms to the standard Gaussian distribution; According to the aggregated hierarchical clustering of the second tensor, a second localized dictionary is determined to characterize the clustering relationship of the second tensor, and a second dictionary index for the second localized dictionary is generated; the second index value in the second dictionary index is used to enable the converter to obtain the execution result of the second target operation command by calling the target second tensor corresponding to the target second index value in the second index value when obtaining the second target operation command; the second target operation command is an operation command in which the target second tensor participates in the operation.
6. A transformer quantization acceleration device in the text field based on dictionary index fixed-point operation, characterized in that: include: A benchmark module, for determining a benchmark dictionary for characterizing the clustering relationship of the values according to the aggregated hierarchical clustering of a plurality of values of random Gaussian distribution; A localization module, configured to determine a first localized dictionary of each first tensor according to a first dictionary value determined by performing a linear transformation of each first tensor in a transformer on the reference dictionary; The first tensor is a tensor in the converter that conforms to a standard Gaussian distribution; the first dictionary value is floating point data; An indexing module, configured to determine a first dictionary index for the first localized dictionary by fitting an exponential function to the first dictionary value; a first index value in the first dictionary index is an integer data; The first index value is used to enable the converter to obtain the execution result of the first target operation command by calling the target first tensor corresponding to the target first index value in the first index value when the first target operation command is obtained; the first target operation command is an operation command in which the target first tensor participates in the operation; The index module comprises: The fitting submodule is used to The exponential function of the first dictionary value is fitted in the form of to determine a binary array corresponding to each of the first dictionary values and a binary parameter; the first variable of the binary array is used to represent the magnitude relationship between each of the first dictionary values and the average value of the first dictionary values through one bit of integer data; the second variable of the binary array is used to represent the exponential value of each of the first dictionary values after decomposition according to the exponential function to the first parameter a in the binary parameter through multi-bit integer data; b is the second parameter in the binary parameter; The index submodule is used to merge and record the binary parameters corresponding to each of the first dictionary values to obtain the first index values corresponding to each of the first dictionary values.
7. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the transformer quantization acceleration method in the text field based on dictionary index fixed-point operation as described in any one of claims 1 to 5 are implemented.
8. An electronic device, characterized in that: The invention comprises a processor, a memory and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of a transformer quantization acceleration method in a text field based on dictionary index fixed-point operations as described in any one of claims 1 to 5.
9. A computer program product, characterized in that The computer program product stores a computer program, and when the computer program is executed by a processor, the steps of the transformer quantization acceleration method in the text field based on dictionary index fixed-point operation as described in any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Super-division model processing method and device, computer equipment and storage medium
CN116012228A
Model compression method, terminal deployment method, system and electronic equipment
CN118502709A