Business processing method and device, equipment, medium and program product
By quantizing and dequantizing the weights of large models, the problem of high computing resource consumption of large models is solved, storage resources and bandwidth consumption are reduced, and processing speed is improved.
Patent Information
- Application Number
- CN202510900208.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-09-26
AI Technical Summary
Large models have a huge number of weights, which leads to high consumption of computing resources and affects computing efficiency and processing speed.
The model weights are quantized into a quantization parameter group and loaded into the second storage unit. The calculation unit extracts multiple quantization parameters for inverse quantization to form inverse quantization weights to process business data, reducing storage and bandwidth resource consumption.
Through quantization and dequantization processing, storage resource occupancy and bandwidth consumption are reduced, and business processing speed is improved.
Smart Images

Figure CN120706513A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, in particular to technical fields such as large models and data processing, and in particular to business processing methods, devices, electronic devices, storage media, and program products. Background Art
[0002] With the rapid development of computer technology, models, especially large models such as large language models, large visual models, and large multimodal models, have attracted increasing attention and research due to their outstanding data processing capabilities. However, due to the large number of weights in the models, accessing the weights requires significant resources. Therefore, reducing the resource consumption of model operations has become a major concern. Summary of the Invention
[0003] The present disclosure provides a service processing method, device, electronic device, storage medium, and program product.
[0004] According to one aspect of the present disclosure, a business processing method is provided, comprising: in response to receiving a business processing request for a target business, loading the business data of the target business and a quantization parameter group of a model for processing the target business from a first storage unit to a second storage unit, the quantization parameter group being obtained by quantizing the weights of the model; using a computing unit to perform the following operations: extracting multiple quantization parameters from the quantization parameter group acquired from the second storage unit to obtain multiple quantization weights, the multiple quantization parameters having different numbers of bits and the multiple quantization weights having the same number of bits; respectively dequantizing the multiple quantization weights to obtain multiple dequantization weights; and processing the business data using a target model formed based on the multiple dequantization weights to obtain a business processing result for the target business.
[0005] According to another aspect of the present disclosure, a business processing device is provided, including: a loading module, for loading the business data of the above-mentioned target business and a quantization parameter group of a model for processing the above-mentioned target business from a first storage unit to a second storage unit in response to receiving a business processing request for a target business, the above-mentioned quantization parameter group being obtained by quantizing the weights of the above-mentioned model; a calculation module, for using the calculation unit to perform the following operations: an extraction submodule, for extracting multiple quantization parameters from the above-mentioned quantization parameter group obtained from the above-mentioned second storage unit, to obtain multiple quantization weights, the multiple quantization parameters have different numbers of bits, and the multiple quantization weights have the same number of bits; an inverse quantization submodule, for respectively inverse quantizing the multiple quantization weights to obtain multiple inverse quantization weights; and a calculation submodule, for processing the above-mentioned business data using a target model formed based on the multiple inverse quantization weights, to obtain a business processing result for the above-mentioned target business.
[0006] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method described above.
[0007] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the method described above.
[0008] According to another aspect of the present disclosure, a computer program product is provided, comprising a computer program, which implements the method described above when executed by a processor.
[0009] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.
[0011] Figure 1A Schematically illustrates an exemplary system architecture to which the service processing method and apparatus according to an embodiment of the present disclosure can be applied;
[0012] Figure 1B Schematically shows a structural block diagram of an electronic device suitable for implementing a service processing method according to an embodiment of the present disclosure;
[0013] Figure 2 The flowchart of the service processing method according to the embodiment of the present disclosure is schematically shown;
[0014] Figure 3 The following schematically illustrates a flow chart of obtaining multiple quantization weights according to an embodiment of the present disclosure;
[0015] Figure 4 The following schematically illustrates a flow chart of determining multiple quantization weights according to another embodiment of the present disclosure;
[0016] Figure 5 A schematic diagram schematically illustrates redundant bits of a quantization parameter group according to an embodiment of the present disclosure;
[0017] Figure 6A A schematic diagram schematically illustrates a method of performing operations on service data and multiple inverse quantization weights according to an embodiment of the present disclosure;
[0018] Figure 6B Schematically illustrates a schematic diagram of operating service data and multiple inverse quantization weights according to another embodiment of the present disclosure;
[0019] Figure 7 A block diagram schematically shows a service processing device according to an embodiment of the present disclosure; and
[0020] Figure 8 A block diagram schematically shows an electronic device suitable for implementing a service processing method according to another embodiment of the present disclosure. DETAILED DESCRIPTION
[0021] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0022] With the rapid development of computer technology, models, especially large models such as large language models, large visual models, and large multimodal models, have attracted increasing attention and in-depth research due to their outstanding data processing capabilities. However, due to the large number of weights in large models, the amount of computation required is very large, and thus requires relatively large computing resources. Therefore, how to reduce the operating costs of large models has become a major concern.
[0023] Specifically, large models involve many network layers, and correspondingly, a large number of weights in the weight matrix. When performing inference operations based on large models, multiple memory access operations for the weight matrix are often required, significantly reducing computational efficiency and slowing down the processing of large models.
[0024] In order to solve this technical problem, an embodiment of the present disclosure provides a business processing method, including: in response to receiving a business processing request for a target business, loading the business data of the target business and a quantization parameter group of a model for processing the target business from a first storage unit to a second storage unit, the quantization parameter group being obtained by quantizing the weights of the model; using a computing unit to perform the following operations: parsing multiple quantization parameters from the quantization parameter group obtained from the second storage unit to obtain multiple quantization weights, the multiple quantization parameters having different numbers of bits and the multiple quantization weights having the same number of bits; respectively dequantizing the multiple quantization weights to obtain multiple dequantized weights; and using a target model formed based on the multiple dequantized weights to process the business data to obtain a business processing result for the target business.
[0025] By utilizing the service processing method provided by the embodiment of the present disclosure, multiple quantization parameters can be stored in the form of unit bytes. On the basis of quantization, the storage amount of data can be further reduced, thereby reducing the storage resource occupancy of the first storage unit and the second storage unit. In addition, the bandwidth resource consumption of loading from the first storage unit to the second storage unit can also be reduced, thereby improving the processing speed of the target service.
[0026] Figure 1A An exemplary system architecture to which the service processing method and apparatus according to an embodiment of the present disclosure can be applied is schematically shown.
[0027] It should be noted that Figure 1A The examples shown are merely examples of system architectures to which the embodiments of the present disclosure may be applied, to help those skilled in the art understand the technical content of the present disclosure. This does not mean that the embodiments of the present disclosure cannot be applied to other devices, systems, environments, or scenarios. For example, in another embodiment, the exemplary system architecture to which the service processing method and apparatus may be applied may include a terminal device, but the terminal device may implement the service processing method and apparatus provided by the embodiments of the present disclosure without interacting with a server.
[0028] like Figure 1A As shown, the system architecture 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used as a medium for providing communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.
[0029] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, and 103, such as knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software (for example only).
[0030] The terminal devices 101 , 102 , and 103 may be various electronic devices having a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, and desktop computers.
[0031] Server 105 may be a server that provides various services, such as a background management server (for example only) that supports content browsed by users using terminal devices 101, 102, and 103. The background management server may analyze and process received data such as user requests, and feed back processing results (e.g., web pages, information, or data obtained or generated based on user requests) to the terminal device.
[0032] For example, the server can be a cloud server, also known as a cloud computing server or cloud host. This is a host product within the cloud computing service system that addresses the management difficulties and poor scalability of traditional physical hosts and VPS services ("Virtual Private Servers," or simply "VPS"). The server can also be a server in a distributed system or a server integrated with blockchain.
[0033] It should be noted that the service processing method provided in the embodiment of the present disclosure can generally be executed by the terminal device 101, 102, or 103. Accordingly, the service processing apparatus provided in the embodiment of the present disclosure can also be provided in the terminal device 101, 102, or 103.
[0034] Alternatively, the service processing method provided by the embodiment of the present disclosure may also be generally executed by the server 105. Accordingly, the service processing apparatus provided by the embodiment of the present disclosure may generally be provided in the server 105. The service processing method provided by the embodiment of the present disclosure may also be performed by a server or server cluster that is different from the server 105 and that is capable of communicating with the terminal devices 101, 102, 103 and / or the server 105. Accordingly, the service processing apparatus provided by the embodiment of the present disclosure may also be provided in a server or server cluster that is different from the server 105 and that is capable of communicating with the terminal devices 101, 102, 103 and / or the server 105.
[0035] It should be understood that Figure 1A The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.
[0036] The following will be Figure 1B right Figure 1A The internal structure of the terminal device or server is described in detail to assist in explaining that the main purpose of the business processing method provided in the embodiment of the present disclosure is to improve the internal performance of the electronic device.
[0037] Figure 1B The structural block diagram of an electronic device suitable for implementing the service processing method according to an embodiment of the present disclosure is schematically shown.
[0038] Electronic equipment is intended to refer to various forms of digital computers, such as Figure 1A The terminal device or server in the
[0039] like Figure 1BAs shown, the electronic device includes a main processor 1601, a first storage unit 1602, a second storage unit 1603, a computing unit 1604, and a bus 1605. The main processor 1601, the first storage unit 1602, the second storage unit 1603, and the computing unit 1604 are connected to each other via the bus 1605.
[0040] The computing unit 1604 can be configured as a graphics processing unit (GPU), a neural network processor (NPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc.
[0041] The main processor 1601 can be configured as a central processing unit (CPU) to receive a service processing request of a target service and control the calculation unit 1604 to extract the quantization parameter group from the first storage unit 1602 or the second storage unit 1603, perform inverse quantization and calculation, etc.
[0042] The first storage unit 1602 or the second storage unit 1603 may be configured as a read-only memory (ROM) or a random access memory (RAM), etc. The first storage unit 1602 or the second storage unit 1603 is used to store quantization parameter groups, service data, service processing results, etc.
[0043] The following will be Figure 2 The specific operations of the business processing method shown are as follows Figure 1B The functions of the various components in the electronic device shown are further explained.
[0044] It should be noted that the sequence numbers of the operations in the following method are only used to indicate the operation for the purpose of description, and should not be regarded as indicating the order in which the operations should be performed. Unless explicitly stated, the method does not need to be performed in the order shown.
[0045] Figure 2 The flowchart of the service processing method according to the embodiment of the present disclosure is schematically shown.
[0046] like Figure 2 As shown, the method includes operations S210 to S220.
[0047] In operation S210 , in response to receiving a service processing request for a target service, service data of the target service and a quantization parameter group of a model for processing the target service are loaded from a first storage unit to a second storage unit.
[0048] In operation S220 , the following operations S221 to S223 are performed using a computing unit.
[0049] In operation S221 , a plurality of quantization parameters are extracted from the quantization parameter group acquired from the second storage unit to obtain a plurality of quantization weights.
[0050] In operation S222 , the plurality of quantized weights are respectively dequantized to obtain a plurality of dequantized weights.
[0051] In operation S223 , the service data is processed using a target model formed based on a plurality of inverse quantization weights to obtain a service processing result for the target service.
[0052] The target business includes at least one of the following: chat tasks, recommendation tasks, command control tasks, translation tasks, route planning tasks, and image recognition tasks. The business data includes at least one of the following: text, image, voice, and multimodal data.
[0053] Optionally, the model may include a large model, but is not limited thereto, and may also be a deep learning model. The quantization parameter group of the model may refer to the quantization parameter group of any network layer in the model. The network layer may include any neural network layer such as a convolutional layer, a pooling layer, a normalization layer, or may be the quantization parameter group of all network layers of the model.
[0054] In response to receiving a business processing request for the target business, business data can be determined from the data carried in the business processing request, for example, the "AAA" carried in the business processing request "Translate 'AAA' into English" is used as business data, but it is not limited to this. The business data of the target business can also be obtained by calling external tools. For example, the business processing request "Do you know the book 'BBB'?" can obtain the book "BBB" through the database as an external tool and use it as business data.
[0055] A model for processing a target service, such as a quantization parameter group of a large language model, and service data may be loaded into the first storage unit.
[0056] The main processor determines a computing unit from a plurality of candidate computing units in the electronic device for processing the target service. The service data and the quantization parameter group can be loaded from the first storage unit to a second storage unit that matches the computing unit.
[0057] The quantization parameter set can be obtained by quantizing the model weights. Quantization can refer to data type conversion. For example, model weights are converted from floating-point data types to integer data types. For example, floating-point weights are represented by FP16-bit floating-point data, while quantized weights are represented by INT8-bit integer data types. Quantization reduces the accuracy of model weights, but also reduces the amount of data required.
[0058] The quantization parameter group may include multiple quantization parameters, and the multiple quantization parameters are stored together in a unit byte. Optionally, the unit byte may include 8 bits, and the number of bits of the multiple quantization parameters is different.
[0059] The quantization parameter group is a storage form of multiple quantization weights. The computing unit can be used to extract the quantization parameter group stored in a unit byte to obtain multiple quantization parameters, and then through masking or other data conversion, the corresponding multiple quantization weights are obtained. The number of bits of the multiple quantization weights is the same.
[0060] For example, using 8 bits as a unit byte to store a quantization parameter group, the first four bits form a 4-bit quantization parameter, bits 2 to 5 form a second 4-bit quantization parameter, and bits 4 to 7 form a third 4-bit quantization parameter. Bits 2 to 5 of the second quantization parameter overlap with the other two quantization parameters, allowing 8 bits to store three quantization parameters. Data conversion can be performed on the three quantization parameters to obtain three quantization weights.
[0061] Multiple quantized weights can be dequantized separately to obtain multiple dequantized weights. Dequantization can refer to data type reversal or the restoration of quantized data. For example, quantized weights are represented by 8-bit integers, while dequantized weights are represented by 16-bit floating-point numbers. Through dequantization, the model's weight accuracy is essentially restored to the original weight accuracy.
[0062] The business data is processed using multiple inverse quantization weights of the target model, such as addition operations, subtraction operations, matrix multiplication operations, or vector multiplication operations, to obtain business processing results for the target business.
[0063] By utilizing the service processing method provided by the embodiments of the present disclosure, multiple quantization parameters can be stored in the form of unit bytes. Based on quantization, the storage amount of data can be further reduced, thereby reducing the storage resource usage of the first storage unit and the second storage unit. In addition, the bandwidth resource consumption of loading from the first storage unit to the second storage unit can be reduced, thereby improving the processing speed of the target service. It can be seen that the service processing method provided by the embodiments of the present disclosure can make specific technical associations between extraction operations, dequantization operations, etc. and the internal structure of electronic devices, thereby solving technical problems such as improving hardware execution effects by reducing data storage and data transmission.
[0064] The above Figure 2 The complete process of business processing is introduced below. Figure 2 Operation S221 shown as obtaining multiple quantization weights from the quantization parameter group is further described.
[0065] Figure 3 The flowchart of obtaining multiple quantization weights according to an embodiment of the present disclosure is schematically shown.
[0066] like Figure 3 As shown, the quantization parameter group may include INT8-bit A0-A7.
[0067] Quantization weights can be extracted from the quantization parameter group through bitwise operations. For example, these methods may include shift extraction and masking. Based on a bitwise shift parameter sequence, data matching the shifted parameters can be extracted. For example, data can be shifted from the quantization parameter group and then extracted to obtain multiple quantization parameters corresponding to the shifted parameter sequence. Each quantization parameter can be masked using the bitwise mask parameter to obtain the quantization weights.
[0068] like Figure 3 As shown, the shift parameter sequence may include a shift parameter sequence of 0, 2, and 4, where 0, 2, and 4 represent shifting by 0 bits, 2 bits, and 4 bits respectively, and then extracting the data from the current bit to the last bit as the quantization parameter. Figure 3 As shown, a first quantization parameter composed of A0-A7, a second quantization parameter composed of A2-A7, and a third quantization parameter composed of A4-A7 are obtained.
[0069] like Figure 3 As shown, the mask parameter can be used to mask the first quantization parameter, the second quantization parameter, and the third quantization parameter with different bit numbers to obtain the first quantization weight, the second quantization weight, and the third quantization weight with the same bit number.
[0070] Optionally, the mask parameters used to mask the three quantization parameters may be the same or different, as long as the quantization weight corresponding to the inverse quantization weight can be extracted.
[0071] like Figure 3 As shown, a unified mask parameter (00001111) may be used to perform an AND operation (&) to obtain first quantization weights A0-A4, second quantization weights A2-A5, and third quantization weights A4-A7.
[0072] Quantization parameter groups are used for data storage and access loading, thereby reducing resource consumption. In addition, multiple quantization parameters are parsed from the quantization parameter group using bit operations, thereby improving the subsequent restoration effect of inverse quantization weights while improving the processing accuracy of inverse quantization weights and business data.
[0073] The above describes how to obtain multiple quantization weights from a quantization parameter group in detail through an embodiment. However, it is not limited to this. Other embodiments may also be included. Figure 4 A detailed description is given on how to obtain multiple quantization weights from a quantization parameter group.
[0074] Figure 4 The figure schematically shows a flow chart of determining multiple quantization weights according to another embodiment of the present disclosure.
[0075] Like Figure 3 The differences in determining multiple quantization weights are shown in Figure 4 The illustrated process also includes first determining whether to perform an inverse quantization operation before performing the extraction operation. If so, the quantization parameter group B0-B3 obtained from the second storage unit is used as the initial quantization parameter group 410. The initial quantization parameter group 410 is inversely quantized to obtain a quantization parameter group A0-A7 420. Multiple quantization parameters 421 are extracted from the quantization parameter group 420, such as a first quantization parameter for A0-A7, a second quantization parameter for A2-A7, and a third quantization parameter for A4-A7. The multiple quantization parameters 421 are masked to obtain multiple quantization weights 430, such as a first quantization weight for A0-A3, a second quantization weight for A2-A5, and a third quantization weight for A4-A7, all having the same number of bits.
[0076] Optionally, it may be determined whether to perform an inverse quantization operation before performing an extraction operation based on the quantization strategy of the model. In the case where it is determined to perform an inverse quantization operation, the following may be performed: Figure 4 If it is determined that the dequantization operation is not to be performed and the extraction operation is to be performed directly, the following operations can be performed: Figure 3 The operation shown.
[0077] In an embodiment of the present disclosure, whether to perform a dequantization operation before performing an extraction operation can be determined according to a quantization strategy. The quantization strategy represents the way in which the weights of the quantization model are quantized.
[0078] The following will briefly explain how to quantize multiple weights of the model based on the quantization strategy to obtain a quantization parameter group.
[0079] The quantization strategy may include quantization parameters and data conversion parameters. Take the model including 64 BF16 [W1, W2, ..., W64] as an example, but it is not limited to this, and it can also be FP16 type. Based on the quantization parameters, [W1, W2, ..., W64] can be quantized to obtain 22 UNIT8bit quantized data [W1', W2', ..., W22']. Based on the data conversion parameters, the 22 UNIT8bit quantized data [W1', W2', ..., W22'] can be converted into 66 UNIT4bit quantization parameters. Among the 66 UNIT4bit quantization parameters, there is data with repeated bits, which can be combined into multiple quantization parameter groups.
[0080] There is no limitation on the quantization parameters and data conversion parameters, as long as corresponding multiple quantization parameter groups can be obtained.
[0081] According to another embodiment of the present disclosure, convolutional coding can be used as a quantization strategy to quantize model weights. For example, convolutional coding is performed based on convolutional coding parameters, and the first code space is stored in a first storage space. Multiple first code values are stored at a location corresponding to an index in the first code space. The index is obtained by removing duplicate code elements from multiple states of a single convolutional coding operation. The multiple first code values stored at the location corresponding to the index include code elements from multiple states. Based on quantization parameters, such as a quantization scaling factor, and the first code space, a second code space is stored in a second storage space. The index of the second code space is the same as the index of the first code space. The multiple second code values stored at the location corresponding to each index are obtained by multiplying the corresponding multiple first code values in the first code space by the quantization scaling factor. Multiple model weights are compared with multiple second code values in the second code space using Euclidean distance to obtain multiple second code values that match the multiple weights. The indices corresponding to the multiple second code values are stored in a third storage space to obtain quantization results, which correspond to the indices. This quantization result can be used as a quantization parameter set.
[0082] Specifically, the index in the first coding space can be obtained as follows. Convolutional coding parameters include the number of bits per state, the number of states in a convolutional coding cycle, and the number of bits transferred between adjacent states. Convolutional coding parameters may include parameters used to configure convolutional coding in channel coding. For example, convolutional coding parameters may include parameters such as L, N, and S. L may represent the number of bits required for each state in the current channel, N may represent the number of states in which convolutional coding is performed in the current channel, and S may represent the number of bits transferred between adjacent states. For example, if the convolutional coding parameters are (L=2, N=3, S=1), each state in the current channel requires L=2 bits, and the number of states in which convolutional coding is performed is N=3, for example, three 00s. The number of bits transferred between adjacent states is S=1, indicating that a state is transferred backward by one bit at a time. For example, if a state 01 is transferred backward by one bit, the next states may be 10 or 11. The number of bits in each state, the number of states in one convolutional encoding (the number of states included in a group of states), and the number of bits transferred between adjacent states can be used to obtain the deduplicated code value of each group of states, and then the index in the first coding space can be obtained.
[0083] Alternatively, taking 4-bit quantization with less loss in weight quantization as an example, each weight requires INT4bit storage after quantization. In the case of convolutional coding, when setting (L=4, N=3, S=2), only INT8bit is needed to store the quantization parameters of 3 weights, that is, the quantization parameter group, thereby achieving an equivalent 2.66-bit compression effect.
[0084] Optionally, by using different quantization strategies, such as the number of quantization times and the timing of quantization, the data size of the obtained quantization parameter group is different, and the extraction method of the quantization parameter group is also different. Adaptive adjustment can be made according to the quantization strategy, thereby improving the flexibility and accuracy of extraction and inverse quantization.
[0085] Optionally, during the process of loading the service data and the quantization parameter group from the first storage unit to the second storage unit, a policy identifier representing the quantization policy can be loaded into the second storage unit. Based on the policy identifier, it can be determined whether to perform an inverse quantization operation before performing an extraction operation. Optionally, based on the policy identifier, parsing parameters such as shift parameters and mask parameters, as well as inverse quantization parameters, that match the policy identifier can be retrieved from the second storage unit.
[0086] Therefore, based on the mapping relationship between the policy identifier and the relevant parameters of the bit operation, and the mapping relationship between the policy identifier and the inverse quantization parameter, the flexibility and accuracy of the bit operation and inverse quantization operations are improved, thereby improving the application scope of the business processing method provided by the embodiment of the present disclosure.
[0087] The above describes how to parse and obtain quantization weights. The following describes how to dequantize multiple quantization weights to obtain multiple dequantization weights.
[0088] Dequantizing the plurality of quantized weights to obtain the plurality of dequantized weights may include: dequantizing the scaling factor to obtain the target scaling factor; and dequantizing the quantized weights based on the target scaling factor and the floating-point offset to obtain the dequantized weights.
[0089] The scaling factor may be dequantized based on a dequantization parameter of the scaling factor to obtain a target scaling factor.
[0090] Taking the data with quantization weight of UNIT4bit as an example, dequantizing it to the dequantization weight of FP16bit can be performed by referring to the following formula (1).
[0091] fp16_b = (uint4_b–zp) × (uint4_scale × super_scale); Formula (1)
[0092] Among them, fp16_b represents the inverse quantization weight, uint4_b represents the quantization weight, zp represents the floating-point offset, uint4_scale represents the scaling factor, and super_scale represents the inverse quantization parameter of the scaling factor.
[0093] The target scaling factor can be directly used as an inverse quantization parameter and stored in the second storage unit. When performing the inverse quantization operation, the target scaling factor is directly loaded, and the quantization weight is inversely quantized to obtain the inverse quantization weight. However, during the inverse quantization process, it is found that because the scaling factor is data that has undergone a quantization operation, its data volume is relatively small compared to the target scaling factor, and thus the storage resources occupied are relatively small. Using this method to store data and inverse quantization can increase the speed of memory access operations and reduce the resource consumption of memory access operations.
[0094] For example, when there are redundant bits in the quantization parameter group, the scaling factor may be determined from the redundant bits. The redundant bits represent bits in the quantization parameter group that do not store the quantization parameter.
[0095] Therefore, on the basis that the scaling factor is already quantized data, by storing the scaling factor in the redundant bits of the quantization parameter group, the number of executions of the memory access operation and the consumption of storage resources are further reduced.
[0096] The following will be Figure 5 The redundant bits in the quantization parameter group are described.
[0097] Figure 5 The diagram schematically shows redundant bits of a quantization parameter group according to an embodiment of the present disclosure.
[0098] like Figure 5 As shown, taking the model including 64 BF16-bit weights [W1, W2, ..., W64] as an example, [W1, W2, ..., W64] can be quantized based on the quantization parameter to obtain 22 UNIT8-bit quantized data [W1', W2', ..., W22']. Based on the data conversion parameter, the 22 UNIT8-bit quantized data [W1', W2', ..., W22'] can be converted into 66 UNIT4-bit quantization parameters. There is bit-repeated data between multiple quantization parameters, so that 3 quantization parameters can be used as a quantization parameter group and stored in one unit byte, for example, Figure 5 The quantization parameters A0-A4, A2-A5, and A4-A7 shown in the figure form quantization parameter groups A0-A7. Of the 66 UNIT4-bit quantization parameters, only 64 are valid for dequantization into 64 BF16-bit weights [W1, W2, ..., W64]. The quantization parameter group contains bits that are irrelevant to the dequantization operation, known as redundant bits, such as A4'-A7'.
[0099] The target scaling factor may be quantized to obtain a scaling factor, and the scaling factor may be stored in redundant bits A4'-A7'.
[0100] Therefore, while quantizing the target scaling factor to reduce the amount of data storage, the quantized scaling factor is stored in the redundant bits of the quantization parameter group, thereby reducing the number of memory access operations and the resource consumption occupied by the memory access operations, and improving business processing efficiency.
[0101] For example, based on the quantization strategy of the model, it can be determined whether there are redundant bits in the quantization parameter group. If it is determined that there are redundant bits in the quantization parameter group, a scaling factor is determined from the redundant bits. If it is determined that there are no redundant bits in the quantization parameter group, the scaling factor or the target scaling factor is stored in additional storage space.
[0102] Optionally, a mapping relationship between a strategy identifier representing the quantization strategy and a redundant bit identifier can be set. When the redundant bit identifier is determined based on the strategy identifier, not only is it determined that there are redundant bits in the quantization parameter group, but the corresponding redundant bit address is determined through the redundant bit identifier, and the scaling factor is obtained based on the redundant bit address, thereby improving the reading efficiency of the scaling factor.
[0103] The above describes the dequantization of the quantization weights. The following describes the dequantization of the initial quantization parameter group to obtain the quantization parameter group.
[0104] Optionally, performing inverse quantization on the initial quantization parameter group to obtain the quantization parameter group may include: performing inverse quantization on the initial quantization parameter group based on an integer inverse quantization parameter to obtain the quantization parameter group.
[0105] Optionally, the integer inverse quantization parameter may include an integer scaling factor and an integer offset, but is not limited thereto. A rounding function may also be included. For example, an int8 value may be converted to int16 through an additional inverse quantization step to improve the numerical representation accuracy, and then the above-described quantization parameter extraction operation may be performed.
[0106] Specifically, the INT8-bit quantization parameter group can be treated as a data b_int8, and the quantization parameter group can be dequantized according to the dequantization parameter to obtain the INT16-bit intermediate quantization parameter group b_int16. More specifically, referring to the following formulas (2) and (3), the initial quantization parameter group can be dequantized according to the integer dequantization parameter.
[0107] b_fp32 = b_int8×code_scale + code_zp; formula (2)
[0108] b_int16 = round(b_fp32); Formula (3)
[0109] Among them, b_int8 represents the initial quantization parameter group, code_scale represents the integer scaling factor, code_zp represents the integer offset, round() represents the rounding function, and b_int16 represents the quantization parameter group.
[0110] By utilizing the inverse quantization operation provided by the embodiment of the present disclosure, the inverse quantization parameter type of the inverse quantization operation can be adjusted according to the inverse quantization operation type and the inverse quantization timing, thereby increasing the application scope of the inverse quantization operation and further improving the resource consumption of memory access operations during business processing.
[0111] The above describes how to obtain multiple inverse quantization weights. Figure 6A and Figure 6B , explaining how to use the target model to process business data.
[0112] Figure 6A The diagram schematically shows a diagram of performing operations on service data and multiple inverse quantization weights according to an embodiment of the present disclosure.
[0113] like Figure 6AAs shown, service data 610 may include service features of a 16×16 two-dimensional matrix type, and multiple inverse quantized weights of any network layer of the target model may include a weight matrix 620 of a 16×16 two-dimensional matrix type. The service processing result may include a processing result matrix 630 of a 16×16 two-dimensional matrix type.
[0114] like Figure 6A As shown, a quantization weight matrix 640 composed of multiple quantization weights can be subjected to an inverse quantization operation to obtain a weight matrix 620.
[0115] like Figure 6A As shown, performing a matrix multiplication operation on the business data 610 and the weight matrix 620 may include performing a vector inner product on the row vectors P1-P16 of the business data 610 and the column vectors A1-A16 of the weight matrix 620 to obtain an element B1 in the processing result matrix 630, and so on to obtain the processing result matrix 630.
[0116] The result matrix can be used as the service data of the next network layer, and matrix multiplication can be performed with multiple inverse quantization weights of the next network layer. This operation is repeated until the output of the model is obtained. The output of the model is used as the service processing result of the target service.
[0117] You can use Figure 6A The operations shown process the business data, but are not limited to this. In the case where there are multiple quantization parameter groups, multiple inverse quantization weights corresponding to the multiple quantization parameter groups can also be combined based on the extraction operation to obtain multiple inverse quantization weight groups. Based on the multiple inverse quantization weight groups, the positions of multiple business sub-data in the business data are rearranged to obtain multiple business sub-data sequences. The inverse quantization weight groups that match the business sub-sequences in the target model are used to operate on the business sub-sequences to obtain business processing results. For specific processing methods, please refer to Figure 6B .
[0118] Figure 6B The figure schematically shows a schematic diagram of operating service data and multiple inverse quantization weights according to another embodiment of the present disclosure.
[0119] like Figure 6B The operation shown is similar to Figure 6A The difference between the operation methods shown is that the quantization parameters of the multiple quantization parameter groups 650 can be extracted in parallel, for example, the quantization parameters with a shift parameter of 9 are extracted from group 1, group 2, group 3 and group 4 in the quantization parameter group (see Figure 6B9), and masking and dequantization operations are performed on each of the four quantization parameters. This yields quantization weights A0 corresponding to group 1, A4 corresponding to group 2, A8 corresponding to group 3, and A12 corresponding to group 4. This yields dequantized weight group 1, which includes A0, A4, A8, and A12. Similarly, multiple dequantized weight groups 621 are obtained. This results in a weight matrix 620'.
[0120] like Figure 6B As shown, the positions of multiple business sub-data in the business data can be rearranged based on the inverse quantization weight group to obtain multiple business sub-data sequences. For example, the business sub-data P0-P15 operated on by A0-A15 can be rearranged into the business sub-data sequence P0, P4, P8, P12, P1, P5, P9, P13, P2, P6, P10, P14, P3, P7, P11, and P15. The business sub-data sequence of the business data 610 is used as a row vector and is used as a column vector composed of multiple inverse quantization weight groups of the weight matrix 620 to perform a vector inner product, thereby obtaining an element B1 in the processing result matrix 630.
[0121] Use Figure 6B The operation method shown in the figure processes multiple quantization parameter groups in parallel, fully combining the parallel processing capability of the computing unit to improve the processing efficiency while ensuring the correctness of the business processing results by rearranging the business data.
[0122] Figure 7 A block diagram of a service processing device according to an embodiment of the present disclosure is schematically shown.
[0123] like Figure 7 As shown, the service processing device 700 may include a loading module 710 and a calculation module 720 .
[0124] The loading module 710 is used to load the business data of the target business and the quantization parameter group of the model used to process the target business from the first storage unit to the second storage unit in response to receiving a business processing request for the target business, where the quantization parameter group is obtained by quantizing the weights of the model.
[0125] The calculation module 720 includes an extraction submodule 721 , an inverse quantization submodule 722 and a calculation submodule 723 , and is configured to utilize a calculation unit to perform the following operations.
[0126] The extraction submodule 721 is configured to extract multiple quantization parameters from the quantization parameter group acquired by the second storage unit to obtain multiple quantization weights, wherein the multiple quantization parameters have different numbers of bits, and the multiple quantization weights have the same number of bits.
[0127] The dequantization submodule 722 is configured to perform dequantization on the plurality of quantization weights to obtain a plurality of dequantization weights.
[0128] The calculation submodule 723 is used to process the service data using a target model formed based on multiple inverse quantization weights to obtain a service processing result for the target service.
[0129] According to an embodiment of the present disclosure, the service processing apparatus further includes: an initial determination module and an initial inverse quantization module.
[0130] The initial determination module is configured to, when it is determined that the inverse quantization operation is to be performed before the extraction operation, use the quantization parameter group obtained from the second storage unit as the initial quantization parameter group.
[0131] The initial dequantization module is used to dequantize the initial quantization parameter group to obtain a quantization parameter group.
[0132] According to an embodiment of the present disclosure, the service processing apparatus further includes: an operation determination module.
[0133] The operation determination module is used to determine whether to perform a dequantization operation before performing an extraction operation based on a quantization strategy of the model, where the quantization strategy represents a way of quantizing the weights of the model.
[0134] According to an embodiment of the present disclosure, the extraction submodule includes: an extraction unit and a mask unit.
[0135] The extraction unit is used to extract data matching the shift parameter of the bit operation from the quantization parameter group as the quantization parameter.
[0136] The mask unit is used to mask the quantization parameter using the mask parameter of the bit operation to obtain the quantization weight.
[0137] According to an embodiment of the present disclosure, the initial inverse quantization module includes an initial inverse quantization unit.
[0138] The initial dequantization unit is used to dequantize the initial quantization parameter group based on the integer dequantization parameter to obtain a quantization parameter group.
[0139] According to an embodiment of the present disclosure, the inverse quantization submodule includes: a first inverse quantization unit and a second inverse quantization unit.
[0140] The first inverse quantization unit is configured to inverse quantize the scaling factor to obtain a target scaling factor.
[0141] The second dequantization unit is configured to dequantize the quantization weight based on the target scaling factor and the floating-point offset to obtain the dequantized weight.
[0142] According to an embodiment of the present disclosure, the service processing apparatus further includes: a scaling factor determination module.
[0143] The scaling factor determination module is configured to determine the scaling factor from redundant bits when there are redundant bits in the quantization parameter group, where the redundant bits represent bits in the quantization parameter group that do not store the quantization parameters.
[0144] According to an embodiment of the present disclosure, the service processing apparatus further includes: a redundant bit determination module.
[0145] The redundant bit determination module is used to determine whether there are redundant bits in the quantization parameter group based on the quantization strategy of the model, and the quantization strategy represents the weight of the quantization model.
[0146] According to an embodiment of the present disclosure, the quantization parameter group includes multiple quantization parameter groups.
[0147] The calculation submodule includes: a rearrangement unit and a calculation unit.
[0148] a rearrangement unit, configured to rearrange the positions of the plurality of service sub-data in the service data based on the plurality of inverse quantization weight groups to obtain the plurality of service sub-data sequences, wherein the inverse quantization weight groups are obtained by combining the plurality of inverse quantization weights corresponding to the plurality of quantization parameter groups based on an extraction operation; and
[0149] The computing unit is used to operate the business subsequence using the inverse quantization weight group in the target model that matches the business subsequence to obtain a business processing result.
[0150] According to an embodiment of the present disclosure, the service data includes at least one of the following: text, image, voice, and multimodal data. The target service includes at least one of the following: a chat task, a recommendation task, a command control task, a translation task, a route planning task, or an image recognition task.
[0151] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0152] According to an embodiment of the present disclosure, an electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described above.
[0153] According to an embodiment of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause a computer to execute the method described above.
[0154] According to an embodiment of the present disclosure, a computer program product includes a computer program, and when the computer program is executed by a processor, the computer program implements the method described above.
[0155] Figure 8 A block diagram of an electronic device suitable for implementing a business processing method according to another embodiment of the present disclosure is schematically shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0156] like Figure 8 As shown, electronic device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. RAM 803 may also store various programs and data required for the operation of device 800. Computing unit 801, ROM 802, and RAM 803 are interconnected via a bus 804. An input / output (I / O) interface 805 is also connected to bus 804.
[0157] Multiple components in electronic device 800 are connected to input / output (I / O) interface 805, including: an input unit 806, such as a keyboard, mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, optical disk, etc.; and a communication unit 809, such as a network card, modem, wireless communication transceiver, etc. The communication unit 809 allows device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0158] The computing unit 801 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 executes the various methods and processes described above, such as the business processing method. For example, in some embodiments, the business processing method may be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program may be loaded and / or installed onto the device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the business processing method described above may be performed. Alternatively, in other embodiments, the computing unit 801 may be configured to execute the business processing method via any other suitable means (e.g., via firmware).
[0159] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-a-chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0160] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0161] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fibers, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0162] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0163] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0164] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.
[0165] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not a limitation herein.
[0166] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.
Claims
1. A business processing method, comprising: In response to receiving a service processing request for a target service, loading service data of the target service and a quantization parameter group of a model for processing the target service from a first storage unit to a second storage unit, the quantization parameter group being obtained by quantizing weights of the model; Use the calculation unit to perform the following operations: Extracting a plurality of quantization parameters from the quantization parameter group acquired by the second storage unit to obtain a plurality of quantization weights, wherein the plurality of quantization parameters have different numbers of bits, and the plurality of quantization weights have the same number of bits; Dequantizing the plurality of quantized weights respectively to obtain a plurality of dequantized weights; as well as The business data is processed using a target model formed based on a plurality of the inverse quantization weights to obtain a business processing result for the target business.
2. The method according to claim 1, further comprising: In a case where it is determined that the inverse quantization operation is performed before the extraction operation is performed, the quantization parameter group obtained from the second storage unit is used as an initial quantization parameter group; Dequantizing the initial quantization parameter group to obtain the quantization parameter group.
3. The method according to claim 2, further comprising: Based on a quantization strategy of the model, it is determined whether to perform a dequantization operation before performing an extraction operation, wherein the quantization strategy characterizes a manner of quantizing weights of the model.
4. The method according to any one of claims 1 to 3, wherein Extracting a plurality of the quantization parameters from the quantization parameter group to obtain a plurality of the quantization weights includes: Extracting data matching the shift parameter of the bit operation from the quantization parameter group as the quantization parameter; and The quantization parameter is masked using the mask parameter of the bit operation to obtain the quantization weight.
5. The method according to claim 2, wherein: The dequantizing the initial quantization parameter group to obtain the quantization parameter group includes: The initial quantization parameter group is dequantized based on an integer dequantization parameter to obtain the quantization parameter group.
6. The method according to any one of claims 1 to 5, wherein The dequantizing the plurality of quantized weights to obtain a plurality of dequantized weights comprises: Dequantizing the scaling factor to obtain a target scaling factor; and The quantization weight is dequantized based on the target scaling factor and the floating-point offset to obtain the dequantized weight.
7. The method according to claim 6, further comprising: In the case that redundant bits exist in the quantization parameter group, the scaling factor is determined from the redundant bits, where the redundant bits represent bits in the quantization parameter group that do not store quantization parameters.
8. The method according to claim 7, further comprising: Based on a quantization strategy of the model, it is determined whether the redundant bits exist in the quantization parameter group, wherein the quantization strategy represents a manner of quantizing the weights of the model.
9. The method according to any one of claims 1 to 8, wherein The quantization parameter group includes a plurality of; The processing of the service data by using a target model formed based on the plurality of inverse quantization weights to obtain a service processing result for the target service includes: Rearranging the positions of the plurality of service sub-data in the service data based on a plurality of inverse quantization weight groups to obtain a plurality of service sub-data sequences, wherein the inverse quantization weight groups are obtained by combining a plurality of inverse quantization weights corresponding to the plurality of quantization parameter groups based on an extraction operation; and The service subsequence is operated by using the inverse quantization weight group in the target model that matches the service subsequence to obtain the service processing result.
10. The method according to any one of claims 1 to 9, wherein The business data includes at least one of the following: text, image, voice, and multimodal data; The target business includes at least one of the following: a chat task, a recommendation task, a command control task, a translation task, a path planning task, and an image recognition task.
11. A service processing device, comprising: a loading module, configured to, in response to receiving a service processing request for a target service, load service data of the target service and a quantization parameter group of a model for processing the target service from the first storage unit to the second storage unit, wherein the quantization parameter group is obtained by quantizing weights of the model; The computing module is configured to perform the following operations using the computing unit: an extraction submodule, configured to extract a plurality of quantization parameters from the quantization parameter group acquired by the second storage unit to obtain a plurality of quantization weights, wherein the plurality of quantization parameters have different numbers of bits, and the plurality of quantization weights have the same number of bits; a dequantization submodule, configured to dequantize the plurality of quantization weights respectively to obtain a plurality of dequantization weights; as well as The calculation submodule is used to process the business data using a target model formed based on multiple inverse quantization weights to obtain a business processing result for the target business.
12. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 10.
13. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-10.
14. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 10.