Processing method and electronic equipment
By separating and processing feature channels with outliers in the machine learning model, and using hardware acceleration units with different precision processing, the problem of high computational volume in the hybrid precision quantization method is solved, and the model inference rate is improved and the accuracy is guaranteed.
Patent Information
- Application Number
- CN202510541299.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-07-22
AI Technical Summary
In the field of machine learning, there is a problem of high computational cost in the hybrid precision quantization method, which affects the model inference rate.
By obtaining the current input data, based on the set of outliers feature channel indexes in the current layer of the preset layer of the target model, the first target column and the second target column with outliers are separated from the input activation value, and the hardware acceleration unit with different precision is processed until the inference result of the target model is obtained.
It improves the efficiency and accuracy of the model inference process, reduces unnecessary calculations, and reduces computing resource consumption.
Smart Images

Figure CN120354948A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to a processing method and an electronic device. Background Art
[0002] In the field of machine learning, the model inference process can be accelerated by means of mixed-precision quantization (i.e., different quantization precisions are assigned to different layers or different data blocks in a machine learning model). However, there is still a problem of large computational volume during the quantization process, which affects the rate of model inference. Summary of the Invention
[0003] The technical solutions provided by this application are as follows:
[0004] In a first aspect of this application, a processing method is provided, including:
[0005] Obtain current input data;
[0006] Process the current input data based on a target model to obtain the current input activation value of a preset layer of the target model;
[0007] Based on the set of in-layer outlier feature channel indexes of the preset layer of the target model, separate a first target column with outliers and a second target column other than the first target column from the current input activation value of the preset layer; the set of in-layer outlier feature channel indexes is used to identify at least one outlier feature channel of an input source layer that provides input to the preset layer; the set of in-layer outlier feature channel indexes of the preset layer is obtained by updating a pre-constructed set of in-layer outlier feature channel indexes of the preset layer based on a set scheduling strategy before obtaining the current input data;
[0008] Process the first target column and the second target column until the inference result of the target model is obtained.
[0009] The pre-constructed set of in-layer outlier feature channel indexes of the preset layer is obtained through the following method:
[0010] During the loading process of the target model, generate a global set of outlier feature channel indexes corresponding to the target model based on multiple verification samples; the global set of outlier feature channel indexes is used to identify the outlier feature channels of each layer;
[0011] Before performing inference based on the target model, construct a set of in-layer outlier feature channel indexes of a preset layer of the target model based on the global set of outlier feature channel indexes.
[0012] The process of updating the outlier feature channel index set within the current layer through the set scheduling strategy includes:
[0013] During the inference process immediately before obtaining the current input data, detect the columns with outliers in the input activation values of the preset layer during the previous inference process;
[0014] By comparing the columns with outliers and the pre-constructed in-layer outlier feature channel index set of the preset layer, obtain new index entries;
[0015] Based on the set scheduling strategy and the new index entries, perform scheduling management on the pre-constructed in-layer outlier feature channel index set to obtain the outlier feature channel index set within the current layer.
[0016] The pre-constructed in-layer outlier feature channel index set includes: at least one index entry; the index entry includes: at least one of the contribution value and priority of each outlier feature channel in the input source layer providing input to the preset layer and the identifier of the outlier feature channel;
[0017] The performing scheduling management on the pre-constructed in-layer outlier feature channel index set based on the set scheduling strategy and the new index entries includes:
[0018] If the number of index entries in the pre-constructed in-layer outlier feature channel index set does not meet the set quantity threshold, add the new index entries to the pre-constructed in-layer outlier feature channel index set;
[0019] If the number of index entries in the pre-constructed in-layer outlier feature channel index set meets the set quantity threshold, based on at least one of the contribution values and priorities in each index entry, eliminate the corresponding index entries from the pre-constructed in-layer outlier feature channel index set, and add the new index entries to the in-layer outlier feature channel index set after the eliminated index entries.
[0020] The process of processing the first target column and the second target column until obtaining the inference result of the target model includes:
[0021] Allocate different storage spaces for the first target column and the second target column;
[0022] By processing the first target column and the second target column in the different storage spaces until obtaining the inference result of the target model.
[0023] The allocating different storage spaces for the first target column and the second target column includes:
[0024] Copy the first target column from the storage space of the current input activation values of the preset layer to a new storage space, and set to zero the position in the storage space corresponding to the first target column.
[0025] The processing of the first target column and the second target column in different storage spaces includes:
[0026] Process the first target column based on a first hardware acceleration unit to obtain a first processing result; the first hardware acceleration unit supports a first data precision;
[0027] Quantize and process the second target column based on a second hardware acceleration unit to obtain a second processing result; the second hardware acceleration unit supports a second data precision; the second data precision is lower than the first data precision;
[0028] Combine the first processing result and the second processing result to obtain a third processing result.
[0029] The generating of the global outlier feature channel index set corresponding to the target model based on multiple verification samples includes:
[0030] Process each verification sample in the multiple verification samples based on the target model to obtain the output activation values of each layer of the target model corresponding to each verification sample;
[0031] By detecting the columns with outliers in the output activation values of each layer of the target model corresponding to each verification sample, determine the index items corresponding to the outlier feature channels corresponding to the columns with outliers in each layer, to obtain the global outlier feature channel index set corresponding to the target model.
[0032] The determining of the index items corresponding to the outlier feature channels corresponding to the columns with outliers in each layer by detecting the columns with outliers in the output activation values of each layer of the target model corresponding to each verification sample includes:
[0033] By detecting the columns with outliers in the output activation values of each layer of the target model corresponding to each verification sample, accumulate the hit times of the outlier feature channels corresponding to the columns with outliers in each layer to obtain the contribution value of the outlier feature channel;
[0034] Assign priorities to each of the outlier feature channels;
[0035] Form index items with at least one of the contribution value and the priority of the outlier feature channel and the identifier of the outlier feature channel.
[0036] On the other hand, the present application provides an electronic device, including:
[0037] A memory for storing a computer program;
[0038] A processor for executing the computer program so that the electronic device can implement the processing method described in any one of the above. Description of the Drawings
[0039] Combined with the drawings and referring to the following specific embodiments, the above and other features, advantages and aspects of the embodiments of the present disclosure will become more obvious. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and the original elements and elements are not necessarily drawn to scale.
[0040] Figure 1 A schematic flowchart of a processing method provided in Embodiment 1 of the present application;
[0041] Figure 2 A schematic flowchart of a processing method provided in Embodiment 5 of the present application;
[0042] Figure 3 A schematic flowchart of a processing method provided in Embodiment 6 of the present application;
[0043] Figure 4 A schematic flowchart of a processing method provided in Embodiment 7 of the present application;
[0044] Figure 5 A schematic structural diagram of a processing device provided by the present application. Detailed Description of the Embodiments
[0045] The following describes the embodiments of the present application with reference to the drawings in the embodiments of the present application. The terms used in the embodiments of the present application are only used to explain the specific embodiments of the present application, and are not intended to limit the present application.
[0046] The following describes the embodiments of the present application with reference to the drawings. Those skilled in the art know that with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of the present application are equally applicable to similar technical problems.
[0047] In the description, claims and the above drawings of this application, terms such as "first" and "second" are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances, which is only a way of distinguishing objects with the same attributes when describing embodiments of this application. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, so that a process, method, system, product or device including a series of units does not have to be limited to those units, but may include other units not clearly listed or inherent to these processes, methods, products or devices.
[0048] To make the above objects, features and advantages of this application more obvious and understandable, the following further describes this application in detail with reference to the drawings and specific embodiments.
[0049] Refer to Figure 1 , which is a schematic flowchart of a processing method provided in Embodiment 1 of this application. As Figure 1 shown, the method may include but is not limited to the following steps:
[0050] Step S101, obtain the current input data.
[0051] In this embodiment, the current input data may be data such as images, text, or audio.
[0052] Step S102, process the current input data based on the target model to obtain the current input activation value of a preset layer of the target model.
[0053] In this embodiment, the preset layer of the target model may include layers in the target model that contain matrix multiplication, such as the Q / K / V projection layer, MLP layer, etc. in the Transformer model; or, the convolutional layer in the convolutional neural network, etc.
[0054] The preset layer of the target model may also include: layers specified by the user that are different from the layers containing matrix multiplication according to the characteristics or deployment requirements of the target model.
[0055] In this embodiment, the current input data is processed layer by layer starting from the input layer of the target model. Each layer performs inference based on its input and outputs an activation value. The output activation value can be used as the input for the subsequent layer. When the calculation reaches the input source layer of the preset layer of the target model, the activation value output by its input source layer can be used as the current input activation value of the preset layer.
[0056] Step S103: Based on the outlier feature channel index set of the current layer in the preset layer of the target model, separate the first target column with outliers and the second target column other than the first target column from the current input activation values of the preset layer; the outlier feature channel index set in the current layer is used to identify at least one outlier feature channel of the input source layer that provides input to the preset layer.
[0057] The outlier feature channel index set in the current layer is obtained by updating the pre-constructed in-layer outlier feature channel index set of the preset layer based on a set scheduling policy before obtaining the current input data.
[0058] The generation of outliers may be directly related to the model parameters (e.g., weight matrix) of the input source layer, resulting in outliers in the columns of the output activation values. Therefore, once the in-layer outlier feature channel index set is pre-constructed (i.e., at least one outlier feature channel of the input source layer that provides input to the preset layer is pre-detected), as long as the model parameters of the input source layer of the target model are not updated, the outlier feature channels can be considered to stably generate outliers. Therefore, in the target model inference stage, there is no need to re-detect outliers for the current input data, and the in-layer outlier feature channel index set of the current layer can be directly reused to separate the first target column with outliers and the second target column other than the first target column from the current input activation values of the preset layer.
[0059] Step S104: Process the first target column and the second target column until the inference result of the target model is obtained.
[0060] In this embodiment, there are outliers in the first target column. If the first target column is quantized from high precision to low precision, the error of the outliers may be amplified. Therefore, the first target column can be processed while retaining high precision to ensure the accuracy of the processing.
[0061] There are no outliers in the second target column, and the quantization error caused by quantizing from high precision to low precision has little impact on the second target column. Therefore, the second target column can be processed by quantizing it from high precision to low precision to improve the processing efficiency.
[0062] In this embodiment, the first target column and the second target column are processed with different precisions to obtain two independent intermediate results. Subsequently, the two intermediate results can be merged to obtain the output result of the preset layer. This output result can be used as the input of the next layer and iterated in turn until the final inference of the target model is completed.
[0063] In this embodiment, by obtaining the current input data, processing the current input data based on the target model to obtain the current input activation value of the preset layer of the target model, and separating the first target column with outliers and the second target column other than the first target column from the current input activation value of the preset layer based on the set of inlier feature channel indices within the current layer of the target model, it is possible to avoid performing real-time outlier detection on the current input activation value during the inference process, thereby improving the efficiency of separating the first target column and the second target column. On this basis, by processing the first target column and the second target column until the inference result of the target model is obtained, this process is not only accelerated by sampling and processing the first target column and the second target column with different precisions, but also unnecessary calculations are reduced due to the pre-determined set of inlier feature channel indices within the current layer, further accelerating the inference process while ensuring the accuracy of the inference.
[0064] As another optional embodiment of the present application, a processing method provided for Embodiment 2 of the present application. This embodiment is mainly an implementation manner of the method for obtaining the pre-constructed set of inlier feature channel indices within the preset layer, and may specifically include but is not limited to the following steps:
[0065] Step S11: During the loading process of the target model, generate a global set of inlier feature channel indices corresponding to the target model based on multiple validation samples; the global set of inlier feature channel indices is used to identify the inlier feature channels of each layer.
[0066] In this embodiment, a publicly available dataset can be directly used as multiple validation samples. Publicly available datasets usually have a wide range of application scenarios and representativeness, and can provide relatively comprehensive feature distribution information for the target model.
[0067] Of course, multiple samples can also be randomly collected as validation samples according to specific business scenarios. This method can ensure that the feature distribution of the validation samples is more consistent with the actual business data, so that the generated global set of inlier feature channel indices is more in line with the business requirements.
[0068] During the loading process of the target model, the parameter can be set to determine whether to enable the pre-computation function. The setting method of the parameter is flexible. It can be predefined in the configuration file in advance so that it can be automatically read and applied when the target model is loaded; it can also be input in real time through an external command during the loading process of the target model to meet the dynamic requirements in different scenarios.
[0069] If the pre-computation function is selected to be turned off, then during the subsequent inference initialization, the global outlier feature channel index set will be initialized to be empty. This means that during the subsequent inference process, the target model will detect columns with outliers in real time in the conventional manner and perform corresponding processing.
[0070] If the pre-computation function is selected to be turned on, then based on multiple validation samples, the global outlier feature channel index set corresponding to the target model can be generated.
[0071] Step S12, before performing inference based on the target model, based on the global outlier feature channel index set, construct the in-layer outlier feature channel index set of the preset layer of the target model.
[0072] Before performing inference based on the target model, when constructing the in-layer outlier feature channel index set of the preset layer, index items directly related to the input source layer of the preset layer can be screened out from the global index set. For example, determine whether there are index items corresponding to the feature channels of the input source layer of the preset layer in the global outlier feature channel index set. The screened relevant index items can be copied (or initialized in other ways) to the in-layer outlier feature channel index set of the preset layer to ensure that the in-layer outlier feature channel index set only contains index items that can identify the outlier feature channels of the input source layer of the preset layer.
[0073] In this embodiment, the loading process of the target model and the inference process of the target model can be independent of each other. Therefore, generating the global outlier feature channel index set during the loading process will not interfere with the inference process of the target model, ensuring the normal progress of the inference process.
[0074] Moreover, generating the global outlier feature channel index set during the target model loading process can associate the global outlier feature channel index set with the target model structure information. When the target model structure is adjusted subsequently, such as adding new layers, modifying layer parameters, etc., the global outlier feature channel index set can be updated conveniently.
[0075] Also, based on the global outlier feature channel index set, constructing the in-layer outlier feature channel index set of the preset layer of the target model can ensure that during the inference stage, only the outlier feature channels of the preset layer need to be concerned, avoiding unnecessary searches and judgments in the global scope, and further improving the inference efficiency.
[0076] As another optional embodiment of the present application, a processing method provided in Embodiment 3 of the present application, this embodiment is mainly an implementation manner of updating the current in-layer outlier feature channel index set through the set scheduling strategy, and specifically may include but is not limited to the following steps:
[0077] Step S21: During the inference process immediately before obtaining the current input data, detect the columns with outliers in the input activation values of the preset layer during the previous inference process.
[0078] Detecting the columns with outliers in the input activation values of the preset layer during the previous inference process may include, but is not limited to: checking each column of the input activation values of the preset layer during the previous inference process one by one. If the absolute value of at least one element in a certain column exceeds a preset threshold, then determine that this column is a column with outliers.
[0079] Step S22: By comparing the column with outliers and the pre-constructed in-layer outlier feature channel index set of the preset layer, obtain a new index item.
[0080] In this embodiment, the feature channels of the preset layer and the identifiers of the columns in the input activation values are in one-to-one correspondence in the target model. Therefore, it is possible to check whether there is an identifier corresponding to the column with outliers in the pre-constructed in-layer outlier feature channel index set. If not, based on this column with outliers, determine a new index item. For example, the identifier of this column with outliers can be used as one of the items in the new index item.
[0081] The new index item can be used to identify the outlier feature channels of the input source layer that provides input to the preset layer newly detected under the current input data.
[0082] Step S23: Based on the set scheduling policy and the new index item, perform scheduling management on the pre-constructed in-layer outlier feature channel index set to obtain the in-layer outlier feature channel index set of the current layer.
[0083] In this embodiment, based on the set scheduling policy, index items can be selected from the pre-constructed in-layer outlier feature channel index set for elimination, and the new index item is added to the in-layer outlier feature channel index set after elimination to obtain the in-layer outlier feature channel index set of the current layer.
[0084] During the previous inference process, steps S21 - S23 can be executed in parallel with the process of separating the columns with outliers from the input activation values of the preset layer during the previous inference process based on the pre-constructed in-layer outlier feature channel index set of the preset layer.
[0085] Similarly, in the current inference phase, there is a process that can be executed in parallel with the above step S103. This process may include: detecting columns with outliers in the current input activation values; obtaining new index entries by comparing the columns with outliers in the current input activation values and the set of in-layer outlier feature channel indices of the preset layer; based on the set scheduling policy and the new index entries, performing scheduling management on the set of in-layer outlier feature channel indices of the current layer to obtain a new set of in-layer outlier feature channel indices, and the new set of in-layer outlier feature channel indices can be used as the set of in-layer outlier feature channel indices of the preset layer in the next inference process.
[0086] In this embodiment, during the inference phase, when there are differences in the feature patterns between the input data and the verification samples, one or more new outlier feature channels may be detected. The new outlier feature channels are different from the outlier feature channels identified by the pre-constructed set of in-layer outlier feature channel indices. Based on the set scheduling policy and the new index entries, performing scheduling management on the pre-constructed set of in-layer outlier feature channel indices can ensure that the set of in-layer feature channel indices of the current layer focuses on the feature channels that actually generate outliers under the input data in the inference phase, continuously and effectively separate the columns with outliers, and improve the inference speed.
[0087] Moreover, during the inference process, the update of the set of in-layer outlier feature channel indices and the separation of the columns with outliers according to the set of in-layer outlier feature channel indices can be performed simultaneously, ensuring that the latest set of in-layer outlier feature channel indices can be immediately used in the next inference, avoiding the missed detection of the columns with outliers caused by the lag in the update of the set of in-layer outlier feature channel indices, thereby ensuring the accuracy of the separated columns with outliers and improving the inference speed.
[0088] As another optional embodiment of the present application, a processing method provided for Embodiment 4 of the present application. This embodiment is mainly an implementation manner of the above step S23. In this embodiment, the pre-constructed set of in-layer outlier feature channel indices may include: at least one index entry; the index entry may include: at least one of the contribution value and the priority of each outlier feature channel in the input source layer that provides input to the preset layer and the identifier of the outlier feature channel; the contribution value represents the risk degree of the outlier feature channel generating outliers; the higher the priority of the outlier feature channel, the more important it is to the performance of the target model. Step S23 may include but is not limited to the following steps:
[0089] Step S231: If the number of index items in the pre-constructed intra-layer outlier feature channel index set does not meet the set number threshold, add the new index item to the pre-constructed intra-layer outlier feature channel index set to obtain the current intra-layer outlier feature channel index set.
[0090] The specific value of the set quantity threshold can be flexibly adjusted according to actual needs, and this application does not impose constraints on its specific value range or setting method. However, the following conditions must be met: the upper limit of the set quantity threshold shall not exceed the total number of feature channels of the input source layer of the preset layer.
[0091] Step S232: If the number of index items in the pre-constructed intra-layer outlier feature channel index set meets the set number threshold, based on at least one of the contribution value and priority in each of the index items, the corresponding index item is eliminated from the pre-constructed intra-layer outlier feature channel index set, and the new index item is added to the intra-layer outlier feature channel index set after the index item is eliminated, so as to obtain the current intra-layer outlier feature channel index set.
[0092] In this embodiment, the corresponding index items can be eliminated from the pre-constructed intra-layer outlier feature channel index set based only on the contribution value in each of the index items. Specifically, the method may include: sorting the index items in the pre-constructed intra-layer outlier feature channel index set in order of contribution value from small to large, and selecting at least one index item with a top ranking from the sorting result for elimination based on the number of new index items. The smaller the contribution value, the lower the risk of the corresponding outlier feature channel generating outliers. By eliminating index items with smaller contribution values, the interference of the outlier feature channels with lower risk on the separation of columns with outliers can be reduced, thereby ensuring the accuracy of separating columns with outliers.
[0093] In this embodiment, the corresponding index items can also be eliminated from the pre-constructed intra-layer outlier feature channel index set based only on the priority of each index item. Specifically, it can include: sorting the index items in the pre-constructed intra-layer outlier feature channel index set in order of priority from low to high, and selecting at least one index item with a higher ranking from the sorting results for elimination based on the number of new index items. By eliminating index items with lower priorities, it means that outlier feature channels with higher importance to the performance of the target model are retained preferentially, which can reduce the risk of mis-separation of high-precision calculation requirement columns (i.e., avoid misclassifying these columns as low-precision calculation columns), and help maintain the calculation accuracy and performance stability of the target model when separating columns with outliers from columns without outliers.
[0094] In this embodiment, it is also possible to eliminate the corresponding index items from the pre-constructed in-layer outlier feature channel index set based on the contribution values and priorities in each of the index items. Specifically, it may include:
[0095] Step S31: Determine the weight factor of the index item according to the priority in the index item. The higher the priority, the greater the corresponding weight factor.
[0096] Step S32: Perform a weighted calculation on the weight factor of the index item and the contribution value in the index item to obtain the weighted result of the index item.
[0097] Step S33: Sort each index item in the pre-constructed in-layer outlier feature channel index set in ascending order according to the weighted result, and select at least one index item with a forward arrangement order from the sorting result for elimination according to the number of new index items.
[0098] In this embodiment, through the joint constraint of the contribution value and the priority, it is possible to avoid misjudgments of eliminating high-priority index items only by the contribution value (for example, the contribution value in a certain index item is extremely low, but the priority is extremely high. If eliminated only based on the contribution value, this index item may be miseliminated), and it is also possible to avoid misjudgments of eliminating high-contribution index items only by the priority (for example, if the priority in a certain index item is extremely low, but the contribution value is extremely high. If eliminated only based on the priority, this index item may be miseliminated), ensuring the accuracy of the columns with outliers during separation.
[0099] As another alternative embodiment of the present application, refer to Figure 2 , which is a schematic flowchart of a processing method provided in Embodiment 5 of the present application. This embodiment is mainly an implementation manner of the above step S104. As Figure 2 shown, step S104 may include but is not limited to the following steps:
[0100] Step S1041: Allocate different storage spaces for the first target column and the second target column.
[0101] By allocating different storage spaces for the first target column and the second target column, the first target column and the second target column can be respectively stored in two independent storage spaces, which can ensure that there is no storage overlap between the first target column and the second target column.
[0102] During allocation, the storage space of the first target column or the second target column can be continuous or discontinuous.
[0103] Step S1042: Process the first target column and the second target column in the different storage spaces until the inference result of the target model is obtained.
[0104] By allocating different storage spaces for the first target column and the second target column, when processing the first target column and the second target column, different precisions can be used to process the first target column and the second target column in parallel, and two intermediate results can be obtained synchronously.
[0105] Subsequently, the two intermediate results can be merged to obtain the output result of the preset layer. This output result can be used as the input of the next layer, and the iteration is performed sequentially until the final inference of the target model is completed.
[0106] In this embodiment, by allocating different storage spaces for the first target column and the second target column, when processing the first target column and the second target column, it is not necessary to move or copy the first target column and the second target column, thereby avoiding the performance consumption caused by memory transfer.
[0107] Moreover, the independent storage space allows parallel processing of the first target column and the second target column, improving the inference efficiency.
[0108] As another optional embodiment of the present application, referring to Figure 3 , it is a schematic flowchart of a processing method provided in Embodiment 6 of the present application. This embodiment is mainly an implementation manner of the above step S1041. As Figure 3 shown, step S1041 may include but is not limited to the following steps:
[0109] Step S10411: Copy the first target column from the storage space of the current input activation value of the preset layer to a new storage space, and set the position corresponding to the first target column in the storage space to zero.
[0110] In this embodiment, a new storage space can be allocated for the first target column according to the data volume of the first target column, and the size of the new storage space can be not less than the data volume of the first target column.
[0111] By setting the position corresponding to the first target column in the storage space to zero, it can be ensured that the valid data in the storage space is only the second target column. When processing the second target column in the storage space, the continuity of memory access can be ensured, and the calculation efficiency can be improved.
[0112] In this embodiment, by copying the first target column from the storage space of the current input activation value of the preset layer to a new storage space and setting the position corresponding to the first target column in the storage space to zero, only one additional memory allocation operation can be introduced. Compared with other methods that may involve multiple memory allocations, the memory management overhead can be reduced. In addition, since direct data transfer within the original storage space or reallocation of multiple storage spaces, which may cause frequent memory access, is avoided, the performance consumption caused by memory transfer is effectively reduced, and the overall processing efficiency is improved.
[0113] As another optional embodiment of the present application, a processing method provided in Embodiment 7 of the present application. This embodiment is mainly an implementation manner for processing the first target column and the second target column in different storage spaces in Embodiment 5, and specifically may include but is not limited to the following steps:
[0114] Step S41: Process the first target column based on the first hardware acceleration unit to obtain a first processing result; the first hardware acceleration unit supports the first data precision.
[0115] In this embodiment, when the target model is loaded, the first hardware acceleration unit can load the weight matrix of the target model. The weight matrix of the target model can at least include the weight matrix of the preset layer.
[0116] In the inference stage, based on the first hardware acceleration unit, operations (such as matrix multiplication or convolution operation, etc.) can be performed on the first target column and the weight matrix of the preset layer to obtain a first processing result.
[0117] During the operation process, the first hardware acceleration unit can utilize its high-precision computing ability to ensure the accuracy of the first processing result.
[0118] Step S42: Quantize and process the second target column based on the second hardware acceleration unit to obtain a second processing result; the second hardware acceleration unit supports the second data precision; the second data precision is lower than the first data precision.
[0119] In this embodiment, when the target model is loaded, the second hardware acceleration unit can load the weight matrix of the target model. The weight matrix of the target model can at least include: the weight matrix of the preset layer.
[0120] In the inference stage, first, the second target column can be quantized based on the second hardware acceleration unit. The quantization process can include: converting the second target column from high precision (such as FP16) to low precision (such as INT8 or INT4, etc.).
[0121] After quantization is completed, the second hardware acceleration unit can perform operations (such as matrix multiplication or convolution operations, etc.) on the quantized second target column and the weight matrix of the preset layer to obtain a second processing result.
[0122] Step S43: Combine the first processing result and the second processing result to obtain a third processing result.
[0123] In this embodiment, the first hardware acceleration unit supports the first data precision and can ensure accurate processing of the first target column containing outliers, avoiding error magnification caused by insufficient precision. The second hardware acceleration unit supports the second data precision, which can improve the calculation efficiency and reduce the consumption of computing resources on the premise of ensuring that the processing result of the second target column meets certain precision requirements.
[0124] Moreover, by using hardware acceleration units that support different data precisions to process the first target column and the second target column, the advantages of hardware resources can be fully utilized, avoiding waste of hardware resources and improving the overall inference performance.
[0125] In this embodiment, for the specific implementation process, the entire processing flow is described in detail in combination with different hardware acceleration units and preset layer 1 and preset layer 2. For example, as Figure 4 shown, during the model loading process, if the pre-computation function is enabled by setting parameters, the judgment module 1 can generate a global outlier feature channel index set corresponding to the target model based on multiple verification samples.
[0126] Before reasoning based on the target model, a within-layer outlier feature channel index set of preset layer 1 can be constructed based on the global outlier feature channel index set to obtain a pre-constructed within-layer outlier feature channel index set 1.
[0127] During the current reasoning process, the processing module can separate the first target column with outliers and the second target column other than the first target column from the input activation value A with a precision of FP16 based on the pre-constructed within-layer outlier feature channel index set 1. The processing module can pass the first target column to the first hardware acceleration unit and the second target column to the second hardware acceleration unit. At the same time, the judgment module 2 can compare the columns with outliers in the input activation value A with the pre-constructed within-layer outlier feature channel index set 1 to obtain new index items, and pass the new index items to the management module. The management module can perform scheduling management on the pre-constructed within-layer outlier feature channel index set 1 based on the set scheduling policy and the new index items to obtain an updated within-layer outlier feature channel index set 1, and the updated within-layer outlier feature channel index set 1 can be used for the next reasoning process.
[0128] The first hardware acceleration unit supports the data precision FP16 and directly operates on the first target column and the weight matrix of the preset layer 1 to obtain the first processing result.
[0129] The second hardware acceleration unit can first quantize the second target column, and the quantization process may include: converting the second target column from high precision (e.g., FP16) to low precision (e.g., INT8).
[0130] After completing the quantization, the second hardware acceleration unit can operate on the quantized second target column and the weight matrix of the preset layer 1 to obtain the second processing result.
[0131] Then, the first processing result and the second processing result can be merged to obtain the third processing result.
[0132] In the subsequent inference process, the judgment module 2 can use the set of in-layer outlier feature channel indexes updated by the management module in the previous inference process, that is, for the new inference process, this updated set of in-layer outlier feature channel indexes can be directly called as the current set of in-layer outlier feature channel indexes without reloading the global set of outlier feature channel indexes, because before performing inference based on the target model, the initialization operation of constructing the set of in-layer outlier feature channel indexes of the preset layer 1 based on the global set of outlier feature channel indexes can be executed only once.
[0133] Regarding the related processing process of the preset layer 2, reference can be made to the processing method of the preset layer 1, which will not be elaborated here.
[0134] As another optional embodiment of the present application, a processing method provided for Embodiment 8 of the present application, this embodiment is mainly an implementation manner of generating the global set of outlier feature channel indexes corresponding to the target model based on multiple verification samples in Embodiment 2, and may specifically include but not be limited to the following steps:
[0135] Step S51: Process each verification sample in the multiple verification samples based on the target model to obtain the output activation values of each layer of the target model corresponding to each verification sample.
[0136] In this embodiment, each verification sample can be input into the target model in sequence, and for each verification sample, the output activation value is calculated layer by layer.
[0137] Step S52: By detecting the columns with outliers in the output activation values of each layer of the target model corresponding to each verification sample, determine the index items corresponding to the outlier feature channels corresponding to the columns with outliers in each layer, and obtain the global set of outlier feature channel indexes corresponding to the target model.
[0138] The detailed process of detecting the columns with outliers in the output activation values of each layer of the target model corresponding to each of the verification samples can refer to the detection method in step S21 of Embodiment 3, which will not be elaborated here.
[0139] In this embodiment, the feature channels of the layer and the columns in the output activation values are in one-to-one correspondence. If a column with an outlier is determined, the feature channel corresponding to the column with the outlier can be regarded as an outlier feature channel, and at least the identifier of the column in the input activation value can be used as the index item corresponding to the outlier feature channel.
[0140] In this embodiment, by processing each verification sample in multiple verification samples based on the target model respectively, the output activation values of each layer of the target model corresponding to each verification sample are obtained, which can reduce the risk of misjudging outliers due to data characteristics based on a single verification sample, ensure the accuracy of the detected columns with outliers, and further ensure the accuracy of the global outlier feature channel index set corresponding to the target model.
[0141] As another optional embodiment of the present application, a processing method provided in Embodiment 9 of the present application. This embodiment is mainly an implementation manner of determining the index items corresponding to the outlier feature channels corresponding to the columns with outliers in the output activation values of each layer of the target model corresponding to each verification sample in Embodiment 8, and may specifically include but are not limited to the following steps:
[0142] Step S521: By detecting the columns with outliers in the output activation values of each layer of the target model corresponding to each verification sample, accumulate the hit times of the outlier feature channels corresponding to the columns with outliers in each layer to obtain the contribution value of the outlier feature channel.
[0143] In this embodiment, for each layer of the target model, as long as an outlier is detected in the output activation value of a certain feature channel of this layer in any verification sample, this feature channel can be regarded as an outlier feature channel. At the same time, the hit times of this outlier feature channel can be incremented by 1. That is, as multiple verification samples are detected, the hit times of this outlier feature channel can be continuously accumulated based on the actual detection results, and finally the total hit times (i.e., the contribution value) of the outlier feature channel are obtained.
[0144] Step S522: Assign priorities to each of the outlier feature channels.
[0145] In this embodiment, the priorities can be assigned, but are not limited to, according to data precision and inference precision. For example, if an outlier feature channel can reduce data precision and the loss of inference precision is smaller, the assigned priority is higher.
[0146] Step S523: Compose an index entry from at least one of the contribution value and the priority of the outlier feature channel and the identifier of the outlier feature channel.
[0147] In this embodiment, the contribution value can reflect the degree of frequent occurrence of outliers in multiple verification samples in the outlier feature channel. Including the contribution value in the index entry can serve as a reference for scheduling and managing the pre-constructed in-layer outlier feature channel index set. During scheduling and management, according to the contribution value, the high and low differences in the degree of outlier occurrence in the outlier feature channel can be identified, and the index entry corresponding to the outlier feature channel with a low degree of outlier occurrence is preferentially selected for elimination, improving the accuracy of scheduling and management.
[0148] The priority can represent the importance of different feature channels to the performance of the target model. Including the priority in the index entry can serve as a reference for scheduling and managing the pre-constructed in-layer outlier feature channel index set. During scheduling and management, the elimination order of the index entries can be determined according to the priority, and the index entries with low priority are preferentially selected for elimination, improving the accuracy of scheduling and management.
[0149] Of course, the contribution value and the priority can also be used simultaneously as a reference for scheduling and managing the pre-constructed in-layer outlier feature channel index set, which can avoid the mis-elimination of index entries caused by only using the contribution value or the priority, and further improve the accuracy of scheduling and management.
[0150] Next, the processing device provided by the present application will be introduced. The processing device described below can be correspondingly referred to the processing method described above.
[0151] Refer to Figure 5 , the processing device includes: a processing module 100, a first judgment module 200, a construction module 300, a second judgment module 400, and a management module 500.
[0152] The processing module 100 is used for:
[0153] Obtain the current input data;
[0154] Process the current input data based on the target model to obtain the current input activation value of the preset layer of the target model;
[0155] Based on the set of outlier feature channel indices in the current layer of the preset layer of the target model, separate from the current input activation values of the preset layer a first target column with outliers and a second target column other than the first target column; the set of outlier feature channel indices in the current layer is used to identify at least one outlier feature channel of the input source layer that provides input to the preset layer; the set of outlier feature channel indices in the current layer is obtained by updating, based on a set scheduling policy before obtaining the current input data, the pre-constructed set of in-layer outlier feature channel indices of the preset layer;
[0156] By processing the first target column and the second target column until the inference result of the target model is obtained.
[0157] The first judgment module 200 is used to generate, during the loading process of the target model, a global set of outlier feature channel indices corresponding to the target model based on multiple verification samples; the global set of outlier feature channel indices is used to identify the outlier feature channels of each layer.
[0158] The construction module 300 is used to construct, before inferring based on the target model, a set of in-layer outlier feature channel indices of the preset layer of the target model based on the global set of outlier feature channel indices.
[0159] In this embodiment, the second judgment module 400 is used for:
[0160] During the inference process immediately before obtaining the current input data, detect the column with outliers in the input activation values of the preset layer during the previous inference process;
[0161] By comparing the column with outliers and the pre-constructed set of in-layer outlier feature channel indices of the preset layer, obtain new index entries.
[0162] The management module 500 is used to perform scheduling management on the pre-constructed set of in-layer outlier feature channel indices based on the set scheduling policy and the new index entries to obtain the set of outlier feature channel indices in the current layer.
[0163] In this embodiment, the pre-constructed set of in-layer outlier feature channel indices may include: at least one index entry; the index entry includes: at least one of the contribution value and priority of each outlier feature channel in the input source layer that provides input to the preset layer and the identifier of the outlier feature channel.
[0164] The management module 500 may specifically be used for:
[0165] If the number of index entries in the pre-constructed in-layer outlier feature channel index set does not meet the set quantity threshold, add the new index entry to the pre-constructed in-layer outlier feature channel index set;
[0166] If the number of index entries in the pre-constructed in-layer outlier feature channel index set meets the set quantity threshold, based on at least one of the contribution values and priorities in each index entry, eliminate the corresponding index entry from the pre-constructed in-layer outlier feature channel index set, and add the new index entry to the in-layer outlier feature channel index set after the eliminated index entry.
[0167] The process by which the processing module 100 processes the first target column and the second target column until the inference result of the target model is obtained may specifically include:
[0168] Allocate different storage spaces for the first target column and the second target column;
[0169] Process the first target column and the second target column in the different storage spaces until the inference result of the target model is obtained.
[0170] The process by which the processing module 100 allocates different storage spaces for the first target column and the second target column may specifically include:
[0171] Copy the first target column from the storage space of the current input activation value of the preset layer to a new storage space, and set the position corresponding to the first target column in the storage space to zero.
[0172] The process by which the processing module 100 processes the first target column and the second target column in the different storage spaces may specifically include:
[0173] Process the first target column based on the first hardware acceleration unit to obtain a first processing result; the first hardware acceleration unit supports the first data precision;
[0174] Quantize and process the second target column based on the second hardware acceleration unit to obtain a second processing result; the second hardware acceleration unit supports the second data precision; the second data precision is lower than the first data precision;
[0175] Merge the first processing result and the second processing result to obtain a third processing result.
[0176] The process by which the first processing module 200 generates the global outlier feature channel index set corresponding to the target model based on multiple verification samples may specifically include:
[0177] Process each verification sample among a plurality of verification samples based on the target model to obtain the output activation values of each layer of the target model corresponding to each verification sample;
[0178] By detecting the columns with outliers in the output activation values of each layer of the target model corresponding to each verification sample, determine the index items corresponding to the outlier feature channels corresponding to the columns with outliers in each layer, and obtain the global outlier feature channel index set corresponding to the target model.
[0179] The process by which the first processing module 200 determines the index items corresponding to the outlier feature channels corresponding to the columns with outliers in each layer by detecting the columns with outliers in the output activation values of each layer of the target model corresponding to each verification sample may specifically include:
[0180] By detecting the columns with outliers in the output activation values of each layer of the target model corresponding to each verification sample, accumulate the hit times of the outlier feature channels corresponding to the columns with outliers in each layer to obtain the contribution value of the outlier feature channel;
[0181] Assign priorities to each outlier feature channel;
[0182] Form index items by using at least one of the contribution value and priority of the outlier feature channel and the identifier of the outlier feature channel.
[0183] In another embodiment of the present application, an electronic device is provided. The electronic device may include:
[0184] A memory for storing a computer program.
[0185] A processor for executing the computer program so that the electronic device can implement the processing method introduced in any one of Embodiments 1-9.
[0186] In another embodiment of the present application, another electronic device is provided. The electronic device may include:
[0187] A first processor for obtaining current input data and processing the current input data based on a target model to obtain the current input activation value of a preset layer of the target model.
[0188] The first processor may include but is not limited to: GPU (Graphics Processing Unit) or NPU (Neural Network Processor), etc.
[0189] A second processor, configured to separate, from the current input activation values of the current layer of a preset layer of the target model, a first target column having outliers and a second target column other than the first target column based on a set of outlier feature channel indices within the current layer of the preset layer; the set of outlier feature channel indices within the current layer is used to identify at least one outlier feature channel of an input source layer that provides input to the preset layer; the set of outlier feature channel indices within the current layer is obtained by updating a pre-constructed set of outlier feature channel indices within the preset layer based on a set scheduling policy before the current input data is acquired.
[0190] The second processor may include, but is not limited to, a CPU (Central Processing Unit).
[0191] The first processor is further configured to process the first target column and the second target column until an inference result of the target model is obtained.
[0192] The first processor may include:
[0193] A first hardware acceleration unit and a second hardware acceleration unit;
[0194] The first hardware acceleration unit is configured to quantize and process the first target column to obtain a first processing result; the first hardware acceleration unit supports a first data precision.
[0195] The second hardware acceleration unit is configured to process the second target column to obtain a second processing result; the second hardware acceleration unit supports a second data precision; the second data precision is higher than the first data precision;
[0196] The second processor may also be configured to merge the first processing result and the second processing result to obtain a third processing result.
[0197] It should be further noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the attached drawings of the device embodiments provided in this application, the connection relationships between the modules indicate that they have communication connections, which can be specifically implemented as one or more communication buses or signal lines.
[0198] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus necessary general hardware. Of course, it can also be implemented by dedicated hardware including application-specific integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions accomplished by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures for implementing the same function can also be diverse, such as analog circuits, digital circuits, or dedicated circuits. However, for the present application, in more cases, software program implementation is a better embodiment. Based on such understanding, the technical solution of the present application, in essence, or the part that makes a contribution to the prior art, can be embodied in the form of a software product. The computer software product is stored in a readable storage medium, such as a floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disc of a computer, and includes several instructions for causing a computer device (which can be a personal computer, training device, or network device, etc.) to execute the methods described in various embodiments of the present application.
[0199] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.
[0200] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a dedicated computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, training device, or data center to another website, computer, training device, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can store, or a data storage device such as a training device or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).
Claims
1. A processing method, comprising: Obtaining current input data; Processing the current input data based on a target model to obtain current input activation values of a preset layer of the target model; Based on a set of in-layer outlier feature channel indices of the preset layer of the target model, separating a first target column with outliers and a second target column other than the first target column from the current input activation values of the preset layer; the set of in-layer outlier feature channel indices is used to identify at least one outlier feature channel of an input source layer that provides input to the preset layer; the set of in-layer outlier feature channel indices of the current layer is obtained by updating a pre-constructed set of in-layer outlier feature channel indices of the preset layer based on a set scheduling strategy before obtaining the current input data; Processing the first target column and the second target column until an inference result of the target model is obtained.
2. The processing method according to claim 1, wherein the pre-constructed set of in-layer outlier feature channel indices of the preset layer is obtained by the following method: During the loading process of the target model, a global outlier feature channel index set corresponding to the target model is generated based on multiple validation samples; The global set of outlier feature channel indices is used to identify outlier feature channels of each layer; Before performing inference based on the target model, based on the global set of outlier feature channel indices, constructing a set of in-layer outlier feature channel indices of a preset layer of the target model.
3. The processing method according to claim 1, wherein the process of updating the set of in-layer outlier feature channel indices of the current layer by the set scheduling strategy includes: During an inference process immediately before obtaining the current input data, detecting columns with outliers in the input activation values of the preset layer during the previous inference process; Obtaining new index entries by comparing the columns with outliers and the pre-constructed set of in-layer outlier feature channel indices of the preset layer; Based on the set scheduling strategy and the new index entries, performing scheduling management on the pre-constructed set of in-layer outlier feature channel indices to obtain a set of in-layer outlier feature channel indices of the current layer.
4. The processing method according to claim 3, wherein the pre-constructed in-layer outlier feature channel index set includes: At least one index entry; the index entry includes at least one of a contribution value and a priority of each outlier feature channel in an input source layer that provides input to the preset layer and an identifier of the outlier feature channel; The performing scheduling management on the pre-constructed set of in-layer outlier feature channel indices based on the set scheduling strategy and the new index entries includes: If the number of index entries in the pre-constructed set of in-layer outlier feature channel indices does not meet a set quantity threshold, adding the new index entries to the pre-constructed set of in-layer outlier feature channel indices; If the number of index entries in the pre-constructed set of in-layer outlier feature channel indices meets the set quantity threshold, based on at least one of the contribution values and priorities in each index entry, eliminating corresponding index entries from the pre-constructed set of in-layer outlier feature channel indices, and adding the new index entries to the set of in-layer outlier feature channel indices after the eliminated index entries.
5. The processing method according to claim 1, wherein the processing of the first target column and the second target column until the inference result of the target model is obtained includes: Allocating different storage spaces for the first target column and the second target column; Processing the first target column and the second target column in the different storage spaces until the inference result of the target model is obtained.
6. The processing method according to claim 5, wherein the allocating different storage spaces for the first target column and the second target column includes: Copying the first target column from the storage space of the current input activation value of the preset layer to a new storage space, and setting the position corresponding to the first target column in the storage space to zero.
7. The processing method according to claim 5, wherein the processing of the first target column and the second target column in the different storage spaces includes: Processing the first target column based on a first hardware acceleration unit to obtain a first processing result; The first hardware acceleration unit supports a first data precision; Quantifying and processing the second target column based on a second hardware acceleration unit to obtain a second processing result; The second hardware acceleration unit supports a second data precision; The second data precision is lower than the first data precision; Combining the first processing result and the second processing result to obtain a third processing result.
8. The processing method according to claim 2, wherein the generating a global outlier feature channel index set corresponding to the target model based on a plurality of verification samples includes: Processing each verification sample in the plurality of verification samples based on the target model respectively to obtain the output activation values of each layer of the target model corresponding to each verification sample; Determining the index items corresponding to the outlier feature channels corresponding to the columns with outliers in the output activation values of each layer of the target model corresponding to each verification sample by detecting the columns with outliers in the output activation values of each layer of the target model corresponding to each verification sample, so as to obtain the global outlier feature channel index set corresponding to the target model.
9. The processing method according to claim 8, wherein the determining the index items corresponding to the outlier feature channels corresponding to the columns with outliers in the output activation values of each layer of the target model corresponding to each verification sample by detecting the columns with outliers in the output activation values of each layer of the target model corresponding to each verification sample includes: By detecting the columns with outliers in the output activation values of each layer of the target model corresponding to each verification sample, accumulating the hit times of the outlier feature channels corresponding to the columns with outliers in each layer to obtain the contribution value of the outlier feature channel; Assigning priorities to each outlier feature channel; Forming index items by at least one of the contribution value and the priority of the outlier feature channel and the identifier of the outlier feature channel.
10. An electronic device, comprising: A memory for storing a computer program; A processor for executing the computer program so that the electronic device can implement the processing method according to any one of claims 1-9.