Quantization alignment method and apparatus for ai unit, and related device
By performing quantitative alignment between the sending and receiving nodes of the AI unit, the floating-point information is converted into fixed-point information using the target quantization information, the quantization complexity and accuracy loss caused by inconsistency between the training equipment and the inference equipment is solved, and the efficiency of the communication system is improved.
Patent Information
- Application Number
- PCT/CN2024/141279
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-28
- Filing Date
- 2024-12-23
- Publication Date
- 2025-07-03
AI Technical Summary
In communication networks, when the equipment training AI units and the equipment using AI units for inference are inconsistent, there are problems in the prior art that improves the quantization complexity of floating-point parameters and loses inference accuracy, resulting in a decrease in the efficiency of the communication system.
By performing quantitative alignment between the sending node and the receiving node of the AI unit, the floating point information is quantized using the target quantization information to quantize the fixed-point information, ensuring the quantitative alignment of parameters between devices, reducing the quantization complexity and reducing the loss of inference accuracy.
It realizes reducing the quantization complexity and improving the overall efficiency of the AI unit communication system, and improves the accuracy and system performance of parameter quantization.
Smart Images

Figure CN2024141279_03072025_PF_FP_ABST
Abstract
Description
Quantization alignment method, device and related equipment of AI unit
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to Chinese Patent Application No. 202311843077.0 filed in China on December 28, 2023, the entire contents of which are incorporated herein by reference. Technical Field
[0003] The present application belongs to the field of communication technology, and specifically relates to a quantization alignment method, apparatus, and related equipment for an AI unit. Background Art
[0004] In communication networks, artificial intelligence (AI) methods have been introduced to improve network communication performance.
[0005] Usually, the device that trains the AI unit and the device that uses the AI unit for inference may not be the same device. In this case, the AI unit needs to be transferred.
[0006] However, the parameters of the trained AI unit can be floating-point parameters, but when using the AI unit for inference, it is necessary to obtain the fixed-point parameters of the AI unit. In related technologies, for example, the network sends the floating-point parameters of the AI unit to the terminal, and the terminal quantizes the floating-point parameters to obtain usable fixed-point parameters of the AI unit. However, in this process, the terminal may unnecessarily quantize the floating-point parameters of the AI unit, resulting in increased quantization complexity or loss of AI unit inference accuracy due to inappropriate quantization, thereby reducing the overall efficiency of the communication system based on the AI unit. Summary of the Invention
[0007] The embodiments of the present application provide a quantization alignment method, apparatus, and related equipment for AI units. In the scenario of transmitting AI units, quantization alignment is performed between the sending node of the AI unit and the receiving node of the AI unit, which can reduce the quantization complexity and the loss of the precision of the quantized AI unit, thereby improving the overall efficiency of the communication system based on the AI unit.
[0008] In a first aspect, a quantization alignment method for an AI unit is provided, the method comprising:
[0009] The first device receives the first information or the second information from the second device;
[0010] The first device performs a first operation, where the first operation includes:
[0011] Determine a first AI unit based on the first information; or,
[0012] quantizing the second information according to target quantization information to obtain a first AI unit;
[0013] The first information includes parameter information of the first AI unit, and the first information is fixed-point number information quantized based on target quantization information; the second information includes parameter information of the first AI unit, and the second information is floating-point number information.
[0014] In a second aspect, another quantization alignment method for AI units is provided, the method comprising:
[0015] The second device performs a second operation, where the second operation includes:
[0016] quantizing the parameter information of the first AI unit according to the target quantization information to obtain first information, wherein the first information is quantized fixed-point number information; sending the first information to the first device; or
[0017] Second information is sent to the first device, where the second information includes parameter information of the first AI unit, and the second information is floating-point number information, and the second information is used to perform quantization according to target quantization information to obtain the first AI unit.
[0018] In a third aspect, a quantization alignment apparatus for an AI unit is provided, for use in a first device, the apparatus comprising:
[0019] A first receiving module, configured to receive first information or second information from a second device;
[0020] The first execution module is configured to execute a first operation, where the first operation includes:
[0021] Determine a first AI unit based on the first information; or,
[0022] quantizing the second information according to target quantization information to obtain a first AI unit;
[0023] The first information includes parameter information of the first AI unit, and the first information is fixed-point number information quantized based on target quantization information; the second information includes parameter information of the first AI unit, and the second information is floating-point number information.
[0024] In a fourth aspect, a quantization alignment apparatus for an AI unit is provided, for use in a second device, the apparatus comprising:
[0025] The second execution module is configured to execute a second operation, wherein the second operation includes:
[0026] quantizing the parameter information of the first AI unit according to the target quantization information to obtain first information, wherein the first information is quantized fixed-point number information; sending the first information to the first device; or
[0027] Second information is sent to the first device, where the second information includes parameter information of the first AI unit, and the second information is floating-point number information, and the second information is used to perform quantization according to target quantization information to obtain the first AI unit.
[0028] In a fifth aspect, a communication device is provided, which includes a processor and a memory, wherein the memory stores programs or instructions that can be run on the processor, and when the program or instructions are executed by the processor, the steps of the method described in the first aspect or the second aspect are implemented.
[0029] In a sixth aspect, a communication device is provided, including a processor and a communication interface:
[0030] Wherein, when the communication device serves as the first device, the communication interface is used to receive the first information or the second information from the second device; and the processor is used to perform a first operation, and the first operation includes:
[0031] Determine a first AI unit based on the first information; or,
[0032] quantizing the second information according to target quantization information to obtain a first AI unit;
[0033] The first information includes parameter information of the first AI unit, and the first information is fixed-point number information quantized based on target quantization information; the second information includes parameter information of the first AI unit, and the second information is floating-point number information;
[0034] or,
[0035] When the communication device serves as the second device, the communication interface and the processor are configured to perform a second operation, where the second operation includes:
[0036] quantizing the parameter information of the first AI unit according to the target quantization information to obtain first information, wherein the first information is quantized fixed-point number information; sending the first information to the first device; or
[0037] Second information is sent to the first device, where the second information includes parameter information of the first AI unit, and the second information is floating-point number information, and the second information is used to perform quantization according to target quantization information to obtain the first AI unit.
[0038] In the seventh aspect, a readable storage medium is provided, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented, or the steps of the method described in the second aspect are implemented.
[0039] In an eighth aspect, a wireless communication system is provided, comprising: a first device and a second device, wherein the first device can be used to execute the steps of the method described in the first aspect, and the second device can be used to execute the steps of the method described in the second aspect.
[0040] In a ninth aspect, a chip is provided, comprising a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is configured to run a program or instruction to implement the method described in the first aspect or the second aspect.
[0041] In a tenth aspect, a computer program / program product is provided, wherein the computer program / program product is stored in a storage medium, and the program / program product is executed by at least one processor to implement the steps of the method described in the first aspect or the second aspect.
[0042] In one embodiment of the present application, the second device can directly transmit the quantized parameter information of the first AI unit to the first device, so that the first and second devices align the parameter quantization of the first AI unit. In another embodiment of the present application, the second device transmits the parameter information of the first AI unit before quantization to the first device. In this case, the first device can quantize the parameter information of the first AI unit based on the target quantization information agreed with the second device, thereby achieving parameter quantization alignment of the first AI unit between the first and second devices. By aligning the parameter quantization of the first AI unit by the first and second devices, the quantization complexity and the loss of precision of the quantized AI unit can be reduced, thereby improving the overall efficiency of the communication system based on the AI unit. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] FIG1 is a schematic structural diagram of a wireless communication system to which an embodiment of the present application can be applied;
[0044] Figure 2 is a schematic diagram of a neural network;
[0045] Figure 3 is a schematic diagram of a neuron;
[0046] FIG4 is a flowchart of a quantization alignment method of an AI unit provided in an embodiment of the present application;
[0047] FIG5 is a flowchart of another quantization alignment method of an AI unit provided in an embodiment of the present application;
[0048] FIG6 is a schematic structural diagram of a quantization alignment device of an AI unit provided in an embodiment of the present application;
[0049] FIG7 is a schematic structural diagram of another quantization alignment device of an AI unit provided in an embodiment of the present application;
[0050] FIG8 is a schematic structural diagram of a communication device provided in an embodiment of the present application;
[0051] FIG9 is a schematic structural diagram of a terminal provided in an embodiment of the present application;
[0052] FIG10 is a schematic structural diagram of a network-side device provided in an embodiment of the present application;
[0053] FIG11 is a schematic structural diagram of another network-side device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0054] The following will be combined with the accompanying drawings in the embodiments of this application to clearly describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field are within the scope of protection of this application.
[0055] The terms "first", "second", etc. in this application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the terms used in this way are interchangeable where appropriate, so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first" and "second" are generally of the same type, and do not limit the number of objects, for example, the first object can be one or more. In addition, "or" in this application represents at least one of the connected objects. For example, "A or B" covers three options, namely, Option 1: including A but not including B; Option 2: including B but not including A; Option 3: including both A and B. The character " / " generally indicates that the objects associated before and after are in an "or" relationship.
[0056] The term "indication" in this application can be either a direct indication (or explicit indication) or an indirect indication (or implicit indication). A direct indication can be understood as the sender explicitly informing the receiver of specific information, the operation to be performed, or the requested result, etc. in the instruction sent; an indirect indication can be understood as the receiver determining the corresponding information based on the instruction sent by the sender, or making a judgment and determining the operation to be performed or the requested result, etc. based on the judgment result.
[0057] It is worth noting that the technology described in the embodiments of the present application is not limited to the Long Term Evolution (LTE) / LTE-Advanced (LTE-A) system, but can also be used in other wireless communication systems, such as Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Frequency Division Multiple Access (FDMA), Orthogonal Frequency Division Multiple Access (OFDMA), Single-carrier Frequency-Division Multiple Access (SC-FDMA) or other systems. The terms "system" and "network" in the embodiments of the present application are often used interchangeably, and the technology described can be used for the systems and radio technologies mentioned above, as well as for other systems and radio technologies. The following description describes a New Radio (NR) system for illustrative purposes, and NR terminology is used in most of the following description, but these technologies can also be applied to systems other than NR systems, such as 6th generation (6G) systems. th Generation, 6G) communication system.
[0058] FIG1 is a block diagram of a wireless communication system applicable to an embodiment of the present application. The wireless communication system includes a terminal 11 and a network-side device 12. The terminal 11 may be a mobile phone, a tablet computer (Tablet Personal Computer), a laptop computer (Laptop Computer), a notebook computer, a personal digital assistant (PDA), a handheld computer, a netbook, an ultra-mobile personal computer (UMPC), a mobile internet device (MID), an augmented reality (AR), a virtual reality (VR) device, a robot, a wearable device (Wearable Device), an aircraft (Flight Vehicle), a vehicle-mounted device (VUE), a ship-mounted device, a pedestrian user equipment (PUE), a smart home (home appliances with wireless communication capabilities, such as refrigerators, televisions, washing machines, or furniture), a game console, a personal computer (PC), an ATM, or a self-service machine, or other terminal-side devices. Wearable devices include: smart watches, smart bracelets, smart headphones, smart glasses, smart jewelry (smart bracelets, smart bracelets, smart rings, smart necklaces, smart anklets, smart anklets, etc.), smart wristbands, smart clothing, etc. Among them, the vehicle-mounted device can also be called a vehicle-mounted terminal, a vehicle-mounted controller, a vehicle-mounted module, a vehicle-mounted component, a vehicle-mounted chip or a vehicle-mounted unit, etc. It should be noted that the specific type of the terminal 11 is not limited in the embodiment of the present application. The network side device 12 may include an access network device or a core network device, wherein the access network device may also be called a radio access network (Radio Access Network, RAN) device, a radio access network function or a radio access network unit. The access network device may include a base station, a wireless local area network (WLAN) access point (AP) or a wireless fidelity (WiFi) node, etc.Among them, the base station can be referred to as Node B (NB), Evolved Node B (eNB), the next generation Node B (gNB), New Radio Node B (NR Node B), access point, Relay Base Station (RBS), Serving Base Station (SBS), Base Transceiver Station (BTS), radio base station, radio transceiver, Basic Service Set (BSS), Extended Service Set (ESS), Home Node B (HNB), Home evolved Node B (home evolved Node B), Transmission Reception Point (TRP) or other appropriate terms in the relevant field. As long as the same technical effect is achieved, the base station is not limited to specific technical vocabulary. It should be noted that in the embodiment of the present application, only the base station in the NR system is used as an example for introduction, and the specific type of the base station is not limited.
[0059] The core network equipment may include but is not limited to at least one of the following: core network node, core network function, location management function (LMF), mobility management entity (MME), access and mobility management function (AMF), session management function (SMF), user plane function (UPF), policy control function (PCF), policy and charging rules function unit (PCRF), edge application server discovery function (EASDF), unified data management (UDM), unified data repository (UDR), home subscriber server (HSS), centralized network configuration (CNC), network repository function (NRF), network exposure function (NEF), local NEF (L-NEF), binding support function (BSF), etc. Function, BSF), application function (AF), network data analysis function (NWDAF), etc. It should be noted that in the embodiment of the present application, only the core network device in the NR system is taken as an example for introduction, and the specific type of the core network device is not limited. It should be noted that in the embodiment of the present application, only the core network device in the NR system is taken as an example for introduction, and the specific type of the core network device is not limited.
[0060] Artificial intelligence is currently being widely used in various fields. AI units can be implemented in a variety of ways, such as neural networks, decision trees, support vector machines, and Bayesian classifiers. This application uses neural networks as an example, but does not limit the specific type of AI unit.
[0061] As shown in Figure 2, the neural network includes an input layer, a hidden layer, and an output layer, which can obtain the input and output information (X1~X n ) predicts the possible output result (Y). The neural network is composed of a large number of neurons, as shown in Figure 3. The parameters of the neurons include: input parameters a1~a K , weight w, bias b and activation function σ(z), and obtain the output value a with these parameters. Common activation functions include the S-shaped growth curve (Sigmoid) function, the hyperbolic tangent (tanh) function, the linear rectification function (Rectified Linear Unit, ReLU, also known as the rectified linear unit) function, etc., and the z in the above function σ(z) can be calculated by the following formula: z=a1w1+…+a k w k +a K w K +b
[0062] Here, K represents the total number of input parameters.
[0063] Neural network parameters are optimized using an optimization algorithm. An optimization algorithm is a type of algorithm that helps minimize or maximize an objective function (sometimes called a loss function). The objective function is often a mathematical combination of the AI unit parameters and the data. For example, given data X and its corresponding label Y, we construct a neural network f(.). With this modular neural network, we can obtain the predicted output f(x) based on the input x, and calculate the difference between the predicted value and the true value (f(x) - Y). This is the loss function. Our goal is to find the appropriate W and b to minimize the value of this loss function. The smaller the loss value, the closer our AI unit is to the true state.
[0064] Currently, most common optimization algorithms are based on the error backpropagation algorithm. The basic idea behind this algorithm is that the learning process consists of two steps: forward signal propagation and backward error propagation. During forward propagation, input samples are passed from the input layer, processed layer by layer through each hidden layer, and then transmitted to the output layer. If the actual output of the output layer does not match the expected output, the error backpropagation phase begins. Error backpropagation involves propagating the output error back through the hidden layers to the input layer layer by layer in some form, distributing the error to all units in each layer. This error signal is then generated for each unit in each layer, which serves as the basis for adjusting the weights of each unit. This process of adjusting the weights of each layer, through forward signal propagation and backward error propagation, is repeated over and over again. This continuous adjustment of weights is the network's learning and training process. This process continues until the error in the network output is reduced to an acceptable level, or until a pre-set number of learning cycles has been completed.
[0065] Generally speaking, the selected AI algorithms and AI units vary depending on the type of problem being solved. In related technologies, the primary approach to improving 5G network performance through AI is to enhance or replace existing algorithms or processing modules with neural network-based algorithms and AI units. In specific scenarios, neural network-based algorithms and AI units can achieve better performance than deterministic algorithms. Commonly used neural networks include deep neural networks, convolutional neural networks, and recurrent neural networks. Using existing AI tools, neural network construction, training, and verification can be achieved.
[0066] In summary, replacing non-AI processing modules in communication systems with AI or machine learning (ML) methods can effectively improve system performance. For example, in the case of Channel State Information (CSI) prediction, historical CSI can be input into an AI unit, which analyzes the time-domain variation characteristics of the channel and uses inference to derive future CSI. AI-based CSI prediction can achieve significant performance gains compared to non-prediction solutions. Furthermore, the achievable prediction accuracy varies depending on the future moment of the prediction.
[0067] However, when the device for training the AI unit (i.e., the second device) and the device for using the AI unit for reasoning (i.e., the first device) are not the same device, the AI unit needs to be transferred. Among them, the parameters of the AI unit obtained by training can be floating-point parameters, and when using the AI unit for reasoning, it is necessary to obtain the fixed-point parameters of the AI unit. In the relevant technology, taking the example of sending the AI unit from the network side to the terminal, the network side sends the floating-point parameters of the AI unit to the terminal, and the terminal quantizes the floating-point parameters to obtain the fixed-point parameters of the AI unit that can be used. However, in this process, the terminal may unnecessarily quantize the floating-point parameters of the AI unit, resulting in an increase in the complexity of quantization, or a loss of reasoning accuracy of the AI unit due to inappropriate quantization, thereby reducing the overall efficiency of the communication system based on the AI unit.
[0068] In the embodiment of the present application, the first device and the second device can quantize and align the parameter information of the AI unit to be transmitted, which can reduce the quantization complexity and the loss of the precision of the quantized AI unit, thereby improving the overall efficiency of the communication system based on the AI unit.
[0069] It should be noted that the AI unit in the embodiment of the present application can be used in the above-mentioned CSI prediction scenario, and can also be used in other scenarios, such as: AI-based CSI compression feedback, beam prediction, positioning enhancement, intelligent network selection, load balancing, energy saving, etc., which are not exhaustive here.
[0070] To facilitate understanding of the quantization alignment method of the AI unit provided in the embodiment of the present application, the following terms in the embodiment of the present application are first explained:
[0071] 1) AI unit: The AI unit in the embodiments of the present application may also be referred to as an AI model, AI structure, machine learning model, machine learning unit, neural network, etc., or the AI unit may also refer to a processing unit capable of implementing AI-related algorithms, formulas, processing procedures, capabilities, etc., or the AI unit may be a processing method, algorithm, function, module or unit for a specific data set, or the AI unit may be a processing method, algorithm, function, module or unit running on AI-related hardware such as a graphics processing unit (GPU), a natural processing unit (NPU), a tensor processing unit (TPU), an application-specific integrated circuit (ASIC), etc., and this application does not make specific restrictions on this. Optionally, the specific data set includes the input and / or output of the AI unit.
[0072] The identifier of the AI unit, such as an AI model identifier, an AI structure identifier, an AI algorithm identifier, a functional identifier (functionality ID), a physical identifier, a logical identifier, a global identifier, a local identifier, or an identifier of a specific data set associated with the AI unit, or an identifier of a specific AI-related scenario, environment, channel feature, or device, or an identifier of an AI-related function, feature, capability, or module, etc., is not specifically limited in the embodiments of the present application.
[0073] 2) First device: can be a device that sends the AI unit, for example: the first device can include a terminal or a network-side device for training the AI unit, wherein the network-side device can include at least one of an access network device and a core network device.
[0074] 3) Second device: can be a device that receives the AI unit and uses the AI unit to perform reasoning. For example, the first device can include a terminal or a network-side device, and the first device and the second device are not the same device.
[0075] The combination of the first device and the second device may include any one of the combinations shown in Table 1 below:
[0076] Table 1
[0077] It should be noted that, in the embodiments of the present application, for the sake of convenience, an example is usually given in which the first device is a terminal and the second device is a network-side device, which does not constitute a specific limitation.
[0078] 4) Quantization: refers to the process of converting floating point values into fixed point values.
[0079] Common floating-point values include single-precision floating-point (float) and double-precision floating-point (double) formats, which include decimal places. Common fixed-point values include integer (int) formats, such as unsigned integers (unsigned int), short integers (short int), long integers (long int), unsigned short integers (unsigned short int), and unsigned long integers (unsigned long int).
[0080] In the embodiment of the present application, intx represents an integer represented by x bits, such as int4 represents -16 to 15. Other common ones include int8, int16, int32, int64, etc.
[0081] 5) AI unit transfer: Taking the transfer of AI models as an example, this includes transmitting the model structure (i.e., structural information) or model parameters (i.e., parameter information). For neural network models, model parameters refer to at least some coefficients on neurons (multiplicative coefficients, additive coefficients, and activation function parameters).
[0082] The following, in conjunction with the accompanying drawings, describes in detail the quantization alignment method of the AI unit, the quantization alignment device of the AI unit, and related equipment provided in the embodiments of the present application through some embodiments and their application scenarios.
[0083] Referring to FIG. 4 , an embodiment of the present application provides a quantization alignment method for AI units, the execution subject of which may be a first device. As shown in FIG. 4 , the quantization alignment method for AI units may include the following steps:
[0084] Step 401: A first device receives first information or second information from a second device.
[0085] Step 402: The first device performs a first operation, where the first operation includes:
[0086] Determine a first AI unit based on the first information; or,
[0087] quantizing the second information according to target quantization information to obtain a first AI unit;
[0088] The first information includes parameter information of the first AI unit, and the first information is fixed-point number information quantized based on target quantization information; the second information includes parameter information of the first AI unit, and the second information is floating-point number information.
[0089] In some embodiments, the parameter information of the first AI unit may include any coefficients in the first AI unit, such as multiplicative coefficients, additive coefficients, activation function parameters, etc., which are not specifically limited here.
[0090] In some embodiments, the first information or the second information may further include structural information of the first AI unit, such as a model structure of a neural network.
[0091] In some embodiments, the first device receives the second information from the second device, which may be that the second device sends parameter information of the first AI unit before quantization to the first device, and the first device quantizes the parameter information of the first AI unit to obtain the first AI unit that can be used for reasoning.
[0092] The target quantization information used by the first device when performing quantization processing on the parameter information of the first AI unit is consistent with that of the second device.
[0093] Optionally, in order to make the first device and the second device reach a consensus on the target quantization information, at least one of the following methods may be adopted:
[0094] Method 1: The target quantitative information is agreed upon in the agreement;
[0095] Method 2: The first device reports the target quantification information to the second device;
[0096] Method three: the second device indicates the target quantization information to the first device.
[0097] Of course, at least two of the above-mentioned methods 1 to 3 may be combined to enable the first device and the second device to reach a consensus on the target quantization information.
[0098] For example, the protocol stipulates a candidate quantization information set, the first device selects a quantization information subset from the quantization information set stipulated in the protocol based on its own capability information and reports it to the second device, and the second device determines the final target quantization information from the quantization information subset reported by the first device.
[0099] For another example, the second device indicates first quantization information to the first device, and when the first device does not support the first quantization information or has better quantization information, the first device sends target quantization information determined by the first device to the second device.
[0100] For another example, the protocol stipulates a candidate quantization information set, and the second device selects target quantization information from the quantization information set stipulated in the protocol according to its own capability information, and indicates the target quantization information to the first device.
[0101] In some implementations, the first device receives the first information from the second device by quantizing the parameter information of the first AI unit by the second device and sending the quantized parameter information of the first AI unit to the first device.
[0102] The target quantization information used by the second device when performing quantization processing on the parameter information of the first AI unit is consistent with that of the first device.
[0103] The specific implementation manner in which the first device and the second device reach an agreement on the target quantization information includes at least one of the manners 1 to 3 in the previous embodiment, which will not be repeated here.
[0104] In one embodiment of the present application, the second device can directly transmit the quantized parameter information of the first AI unit to the first device, so that the first and second devices align the parameter quantization of the first AI unit. In another embodiment of the present application, the second device transmits the parameter information of the first AI unit before quantization to the first device. In this case, the first device can quantize the parameter information of the first AI unit based on the target quantization information agreed with the second device, thereby achieving parameter quantization alignment of the first AI unit between the first and second devices. By aligning the parameter quantization of the first AI unit by the first and second devices, the quantization complexity and the loss of inference accuracy of the quantized AI unit can be reduced, thereby improving the overall efficiency of the communication system based on the AI unit.
[0105] It is worth noting that, regarding the method of quantizing the parameter information of the first AI unit by the first device, the parameter information of the first AI unit transmitted by the second device to the first device is floating-point information, which can be used for a variety of quantization methods and quantization levels, and has high transmission flexibility. At the same time, by making the first device and the second device reach an agreement on the target quantization information, the second device can determine that the first AI unit it transmits is compatible with the target quantization information. For example, by determining that the target quantization information can be used to quantize at least some parameter values in the first AI unit and ensuring that the quantization loss of at least some parameter values in the quantized first AI unit is small, the second device can determine that the first AI unit is compatible with the target quantization information. In this way, the first device can use the target quantization information that is compatible with the first AI unit to quantize the parameter information of the first AI unit, which can reduce the quantization complexity of the parameter information of the first AI unit and improve the quantization accuracy.
[0106] The second device can also quantize the parameter information of the first AI unit by determining that the first AI unit it transmits is compatible with the target quantization information, thereby improving the quantization accuracy of the first AI unit's parameter information. Furthermore, by ensuring that the first and second devices agree on the target quantization information, the second device can quantize the parameter information of the first AI unit using the target quantization information compatible with the first AI unit, thereby reducing the quantization complexity of the first AI unit's parameter information and improving quantization accuracy.
[0107] In some embodiments, the target quantization information is used to indicate at least one of a quantization method and a quantization level.
[0108] Optionally, the quantification method includes at least one of the following:
[0109] Direct quantization method, which is used to quantize each parameter of the AI unit;
[0110] A group quantization method is used to quantize each group of parameters of the AI unit, wherein a group of parameters includes at least one parameter, and parameters in the same group correspond to the same quantized value;
[0111] A transform domain quantization method is used to convert floating-point parameters of the AI unit into transform domain parameters, and perform quantization and inverse transform processing on the transform domain parameters to obtain quantized fixed-point parameters of the AI unit;
[0112] The product quantization method is used to quantize a first subspace, where the first subspace is a subspace formed by at least one parameter of the AI unit.
[0113] In some implementations, direct quantization can be further divided into uniform quantization and non-uniform quantization:
[0114] 1) Uniform quantization: It means dividing the input value range into equal intervals and quantizing each value range after division.
[0115] 2) Non-uniform quantization: This refers to quantization where the quantization intervals within the dynamic range of the input are unequal. For example, the quantization intervals / levels for different input intervals are determined based on the input's probability density, probability distribution, cumulative probability distribution, etc. For example, for intervals with small input values, the quantization intervals are small; conversely, for intervals with small input values, the quantization intervals are large.
[0116] In other implementations, the above-mentioned group quantization method can also be called weight sharing quantization, that is, the parameters in the AI unit are divided into multiple sets, and the elements in each set share a quantized value.
[0117] For example: Using the K-means clustering method, the data is divided into K groups, and K objects are randomly selected from the parameters of the AI unit as the initial cluster centers. Then, the distance between each object and each seed cluster center is calculated, and each object is assigned to the cluster center closest to it. The cluster centers and the objects assigned to them represent a cluster. Each time a sample is assigned, the cluster center of the cluster is recalculated based on the existing objects in the cluster. This process will be repeated until a certain termination condition is met. The termination condition can be that no (or a minimum number) objects are reassigned to different clusters, or no (or a minimum number) cluster centers change again, or the sum of squared errors is locally minimized.
[0118] In some further embodiments, the above-mentioned transform domains may include the frequency domain corresponding to the Fourier transform, the S domain corresponding to the Laplace transform, and the Z domain corresponding to the Z transform, wherein the Z transform can also be called the Fourier transform of the discrete function * attenuation function.
[0119] For example: transform the floating-point parameters of the AI unit (such as weights, biases, convolution kernels, etc.) to the frequency domain, then quantize the floating-point parameters of the AI unit in the frequency domain, and finally perform an inverse frequency domain transform on the quantized parameters to obtain the quantized fixed-point parameters of the AI unit.
[0120] In some other implementations, product quantization can be to divide the weights of the AI unit into multiple subspaces and perform quantization operations on each subspace separately, such as performing weight sharing quantization on each subspace, to finally obtain the fixed-point parameters of the quantized AI unit.
[0121] Optionally, the quantization levels are divided based on at least one of the following:
[0122] The number of bits of a fixed-point number obtained after quantization of a floating-point number;
[0123] The difference between two adjacent fixed-point numbers after quantization, or the difference between floating-point numbers represented by the lowest bit after quantization.
[0124] In one embodiment, the more bits a fixed-point number obtained after quantization of a floating-point number has, the higher the corresponding quantization level is, and in this case, the higher the quantization accuracy is.
[0125] In another embodiment, the smaller the difference between two adjacent fixed-point numbers after quantization, or the smaller the difference between the floating-point numbers represented by the lowest bit after quantization, the higher the corresponding quantization level, and in this case, the higher the quantization accuracy.
[0126] For example, for uniform quantization, a floating-point number between -1 and 1 is quantized into 4 bits, and the lowest bit represents 2 / 16 of the floating-point number, or the difference between adjacent fixed-point numbers is 2 / 16.
[0127] In this embodiment, the quantization method may indicate the type of quantization or specific quantization parameters to be performed on the parameter information of the first AI unit; and the quantization level may indicate the degree to which the parameter information of the first AI unit is quantized.
[0128] As an optional implementation, the method further includes at least one of the following:
[0129] The first device determines the target quantitative information according to the first quantitative information from the second device;
[0130] The first device determines the target quantization information according to the second quantization information agreed upon in the protocol.
[0131] In some embodiments, when the first device determines the target quantization information based on the first quantization information from the second device, the target quantization information may be the same as the first quantization information. In this case, the first device determines the first quantization information indicated by the second device as the target quantization information.
[0132] In other embodiments, when the first device determines the target quantization information based on the first quantization information from the second device, the target quantization information may be a subset of the first quantization information, for example, the UE selects the quantization information it supports from the first quantization information indicated by the NW as the target quantization information.
[0133] In some other embodiments, when the first device determines the target quantization information based on the first quantization information from the second device, the target quantization information may be different from the first quantization information. For example, after the NW indicates the first quantization information to the UE, the UE does not support the first quantization information, or determines that there is other quantization information that is better than the first quantization information, then the UE may determine target quantization information that is different from the first quantization information.
[0134] Optionally, the method further includes:
[0135] In a case where the first quantization information is not completely identical to the target quantization information, the first device sends the target quantization information to the second device.
[0136] For example, after the UE selects the quantization information supported by the UE from the first quantization information indicated by the NW as the target quantization information, the UE reports the target quantization information to the NW.
[0137] For another example, after the UE determines target quantization information that is different from the first quantization information indicated by the NW, the UE reports the target quantization information to the NW.
[0138] It should be noted that, upon receiving the target quantization information sent by the first device, the second device may quantize the parameter information of the first AI unit according to the target quantization information to obtain the first information. Furthermore, the second device may also perform an AI unit transfer process adapted to the target quantization information. In this way, the second device may perform the quantization process or AI unit transfer process according to the quantization of the AI unit aligned with that of the first device.
[0139] In some embodiments, the first device determines the target quantitative information based on the first quantitative information from the second device, including:
[0140] The first device receives first quantized information from the second device before receiving the first information or the second information from the second device;
[0141] The first device determines the target quantization information according to the first quantization information.
[0142] It should be noted that, in one embodiment, the first device receives the first quantization information before receiving the second information, and determines the target quantization information used to quantize the second information based on the first quantization information. In another embodiment, the first device may receive the first quantization information before receiving the first information, and feed back the target quantization information to the second device, so that the second device quantizes the floating-point parameter information of the first AI unit using the target quantization information agreed with the first device. After obtaining the first information, the second device sends the first information to the first device. In this case, when the first device receives the first information, it can understand how the first information is quantized based on the target quantization information, so as to execute the operation process of the first AI unit according to the quantization method or quantization level of the first information.
[0143] In this implementation, quantization alignment may be performed before AI unit transfer.
[0144] Through the above implementation, the first device and the second device interact with each other on quantization information and reach a consensus on target quantization information.
[0145] In some embodiments, the first device and the second device may obtain consistent target quantization information through a protocol agreement. Alternatively, the protocol may agree on multiple quantization information, and the first device and the second device may interact to determine the target quantization information from the multiple quantization information agreed upon in the protocol.
[0146] As an optional implementation, the method further includes:
[0147] The first device sends first capability information to the second device, where the first capability information is used to indicate third quantization information supported by the first device, wherein the third quantization information includes the target quantization information.
[0148] In some implementations, after receiving the first capability information, the second device may determine the target quantization information accordingly.
[0149] Optionally, when the first device quantizes the parameter information of the first AI unit, the first device may further receive the target quantization information from the second device.
[0150] In this implementation, the first device sends the first capability information to the second device, so that the second device can determine the target capability information supported by the first device based on the first capability information.
[0151] In order to improve quantization accuracy, the parameter information of the first AI unit that needs to be quantized can be restricted so that the parameters that meet the restriction requirements will cause less quantization loss when quantized based on the target quantization information.
[0152] In some optional embodiments, the first AI unit includes a first parameter set, and floating-point values of the first parameter set meet at least one of the following requirements:
[0153] The maximum amplitude of the parameters in the first parameter set is less than or equal to a first threshold value;
[0154] The maximum amplitude of the parameters in the first parameter set is greater than or equal to a second threshold value;
[0155] The minimum amplitude of the parameters in the first parameter set is less than or equal to a third threshold value;
[0156] The minimum amplitude of the parameters in the first parameter set is greater than or equal to a fourth threshold value;
[0157] The average amplitude of the parameters in the first parameter set is less than or equal to a fifth threshold;
[0158] The average amplitude of the parameters in the first parameter set is greater than or equal to a sixth threshold value;
[0159] The difference between the first value and the second value is less than or equal to a seventh threshold value;
[0160] The ratio of the first value to the second value is less than or equal to an eighth threshold value;
[0161] The difference between the third value and the fourth value is less than or equal to a ninth threshold value;
[0162] The ratio of the third value to the fourth value is less than or equal to a tenth threshold value;
[0163] The difference between the fifth value and the sixth value is less than or equal to the eleventh threshold value;
[0164] The ratio of the fifth value to the sixth value is less than or equal to the twelfth threshold value;
[0165] a difference between a maximum amplitude of a parameter in the first parameter set and a minimum amplitude of a parameter in the first parameter set is less than or equal to a thirteenth threshold;
[0166] a ratio of a maximum amplitude of a parameter in the first parameter set to a minimum amplitude of a parameter in the first parameter set is less than or equal to a fourteenth threshold;
[0167] The difference between the maximum amplitude of the parameters in the first parameter set and the average amplitude of the parameters in the first parameter set is less than or equal to a fifteenth threshold;
[0168] The ratio of the maximum amplitude of the parameters in the first parameter set to the average amplitude of the parameters in the first parameter set is less than or equal to a sixteenth threshold;
[0169] a difference between an average amplitude of the parameters in the first parameter set and a minimum amplitude of the parameters in the first parameter set is less than or equal to a seventeenth threshold;
[0170] a ratio of an average amplitude of the parameters in the first parameter set to a minimum amplitude of the parameters in the first parameter set is less than or equal to an eighteenth threshold;
[0171] Among them, the first parameter set includes at least part of the parameters of the first AI unit; the first AI unit includes a first layer and a second layer; the first value is the maximum amplitude of the parameters in the first layer; the second value is the maximum amplitude of the parameters in the second layer; the third value is the minimum amplitude of the parameters in the first layer; the fourth value is the minimum amplitude of the parameters in the second layer; the fifth value is the average amplitude of the parameters in the first layer; the sixth value is the average amplitude of the parameters in the second layer; and the first value is greater than or equal to the second value, the third value is greater than or equal to the fourth value, and the fifth value is greater than or equal to the sixth value.
[0172] In some embodiments, the first device may determine whether the first AI unit meets the floating-point value requirements for the first parameter set, and if the judgment result is: the first AI unit meets the floating-point value requirements for the first parameter set, determine that the parameter information of the first AI unit can be quantized, or determine that the quantized first AI unit is usable.
[0173] In other embodiments, the second device may determine whether the first AI unit meets the floating-point value requirements for the first parameter set, and if the judgment result is that the first AI unit meets the floating-point value requirements for the first parameter set, the first AI unit is passed.
[0174] In some embodiments, the first layer is at least one layer of the first AI unit, and the second layer is at least one layer in the first AI unit that is different from the first layer.
[0175] In some embodiments, the average amplitude in this embodiment can be a value obtained by performing at least one operation such as linear averaging, geometric averaging, harmonic averaging, square averaging, weighted averaging, minimum maximization, maximum maximization, and simple variations of the amplitude, and is not specifically limited here.
[0176] In this embodiment, by limiting the maximum amplitude, minimum amplitude, average amplitude, difference in maximum amplitudes of parameters of different layers, difference in minimum amplitudes of parameters of different layers, difference in average amplitudes of parameters of different layers, difference or ratio between maximum amplitude and minimum amplitude, difference or ratio between maximum amplitude and average amplitude, and difference or ratio between average amplitude and minimum amplitude of at least some floating-point parameters in the first AI unit, the quantization range of at least some floating-point parameters in the first AI unit can be limited, thereby ensuring that the quantization accuracy of these parameters meets the requirements.
[0177] It should be noted that at least one of the above-mentioned first threshold value, the second threshold value, the third threshold value, the fourth threshold value, the fifth threshold value, the sixth threshold value, the seventh threshold value, the eighth threshold value, the ninth threshold value, the tenth threshold value, the eleventh threshold value, the twelfth threshold value, the thirteenth threshold value, the fourteenth threshold value, the fifteenth threshold value, the sixteenth threshold value, the seventeenth threshold value, the eighteenth threshold value, the nineteenth threshold value, the twentieth threshold value, the twenty-first threshold value, and the twenty-second threshold value can be agreed upon in the protocol, or indicated by the network side, or reported by the terminal, and is not specifically limited here.
[0178] In some embodiments, at least one of the first threshold value, the second threshold value, the third threshold value, the fourth threshold value, the fifth threshold value, the sixth threshold value, the seventh threshold value, the eighth threshold value, the ninth threshold value, the tenth threshold value, the eleventh threshold value, the twelfth threshold value, the thirteenth threshold value, the fourteenth threshold value, the fifteenth threshold value, the sixteenth threshold value, the seventeenth threshold value, the eighteenth threshold value, the nineteenth threshold value, the twentieth threshold value, the twenty-first threshold value, and the twenty-second threshold value can be associated with a quantization level, that is, different quantization levels can be associated with respective target threshold values.
[0179] For example, the target threshold value is related to the target difference value, and the target difference value is the difference between two adjacent fixed-point numbers after quantization or the difference between floating-point numbers represented by the lowest bit after quantization;
[0180] The target threshold value includes at least one of the first threshold value, the second threshold value, the third threshold value, the fourth threshold value, the fifth threshold value, the sixth threshold value, the seventh threshold value, the eighth threshold value, the ninth threshold value, the tenth threshold value, the eleventh threshold value, the twelfth threshold value, the thirteenth threshold value, the fourteenth threshold value, the fifteenth threshold value, the sixteenth threshold value, the seventeenth threshold value, the eighteenth threshold value, the nineteenth threshold value, the twentieth threshold value, the twenty-first threshold value, and the twenty-second threshold value.
[0181] It should be noted that the larger the difference between two adjacent fixed-point numbers after quantization or the difference between the floating-point numbers represented by the lowest bit after quantization, the lower the quantization level. At this time, the limit of the target threshold value can be appropriately relaxed. At this time, compared with the target threshold value with a high quantization level, the accuracy of the parameters that meet the target threshold value with a lower quantization level will decrease after quantization.
[0182] In this implementation, the target threshold value may be adjusted in a related manner according to different quantization levels.
[0183] It is worth mentioning that the amplitude in the embodiment of the present application can be the value of the amplitude, or it can be the value obtained after performing certain mathematical operations on the amplitude, such as the square of the amplitude or the square root of the amplitude.
[0184] Optionally, the maximum amplitude is the minimum value among the largest K1 amplitudes, where K1 is a positive integer.
[0185] In this way, the largest K1 amplitudes can be uniformly adjusted to their minimum value. For example, assuming that the three largest amplitudes in the first parameter set are A1, A2, and A3, A1 and A2 can be adjusted to A3. This can reduce the amplitude range by sacrificing a very small number of the largest amplitudes, that is, reduce the amplitude range of the quantization processing, and thus improve quantization accuracy.
[0186] Optionally, the minimum amplitude is the maximum value among the smallest K2 amplitudes, where K2 is a positive integer.
[0187] In this way, similar to the above-mentioned maximum amplitude being the minimum value among the largest K1 amplitudes, by uniformly adjusting the smallest K2 amplitudes to the maximum value therein, the amplitude range can be narrowed at the expense of a very small number of minimum amplitudes, that is, the amplitude range of the quantization processing can be narrowed, which can improve the quantization accuracy.
[0188] It should be noted that the above-mentioned maximum amplitude may include at least one of the following: the maximum amplitude of the parameters in the first parameter set, the maximum amplitude of the parameters in the first layer, and the maximum amplitude of the parameters in the second layer.
[0189] In addition, the above-mentioned minimum amplitude may include at least one of the following: the minimum amplitude of the parameters in the first parameter set, the minimum amplitude of the parameters in the first layer, and the minimum amplitude of the parameters in the second layer.
[0190] It is worth mentioning that the floating-point values of the above-mentioned first parameter set can be the floating-point parameter values of the first AI unit before quantization; or, they can also be floating-point values obtained by dequantizing the fixed-point parameter values of the quantized first AI unit.
[0191] In addition, in addition to restricting the floating-point values of the first parameter set, restrictions may also be placed on the fixed-point values of the first parameter set.
[0192] In some embodiments, the first AI unit includes a first parameter set, and the quantized fixed-point values of the first parameter set meet at least one of the following requirements:
[0193] The maximum value is less than or equal to the nineteenth threshold;
[0194] The maximum value is greater than or equal to the twentieth threshold;
[0195] The minimum value is less than or equal to the twenty-first threshold;
[0196] The minimum value is greater than or equal to the 22nd threshold;
[0197] The first parameter set includes at least part of the parameters of the first AI unit.
[0198] In this embodiment, by limiting the maximum and minimum values of at least some fixed-point parameters in the first AI unit, the quantization range of at least some parameters in the first AI unit can be limited, thereby ensuring that the quantization accuracy of these parameters meets the requirements.
[0199] It is worth mentioning that the fixed-point values of the above-mentioned first parameter set can be the fixed-point parameter values of the first AI unit after quantization; or, they can also be fixed-point values obtained by quantizing the floating-point parameter values of the first AI unit before quantization.
[0200] Optionally, the first parameter set in the embodiment of the present application includes at least one of the following:
[0201] at least some parameters of the first AI unit, or parameters in at least some structure of the first AI unit, indicated by the second device;
[0202] Parameters in at least a portion of the structure of the first AI unit reported by the first device;
[0203] Parameters of the adaptation layer of the first AI unit;
[0204] Parameters of the input layer of the first AI unit;
[0205] parameters of the hidden layer of the first AI unit;
[0206] Parameters of the output layer of the first AI unit.
[0207] In some embodiments, the model parameters (such as neuron coefficients) of the adaptive layer can be adjusted, or the model parameters and structural parameters of the adaptive layer can be adjusted.
[0208] In this embodiment, the value range of the parameters specified by the second device, or the parameters reported by the first device, or the parameters of the specified layer can be limited to improve the quantization accuracy of the parameter information of the first AI unit.
[0209] Referring to FIG. 5 , an embodiment of the present application provides a quantization alignment method for an AI unit, the execution subject of which may be a second device. As shown in FIG. 5 , the quantization alignment method for the AI unit may include the following steps:
[0210] Step 501: The second device performs a second operation, where the second operation includes:
[0211] quantizing the parameter information of the first AI unit according to the target quantization information to obtain first information, wherein the first information is quantized fixed-point number information; sending the first information to the first device; or
[0212] Second information is sent to the first device, where the second information includes parameter information of the first AI unit, and the second information is floating-point number information, and the second information is used to perform quantization according to target quantization information to obtain the first AI unit.
[0213] In some embodiments, after quantizing the parameter information of the first AI unit according to the target quantization information, the second device sends the quantized first information to the first device, so that the first device directly uses the first AI unit to perform inference based on the received first information.
[0214] In other embodiments, the second device directly sends the second information before quantization to the first device, so that the first device quantizes the second information according to the target quantization information to obtain a usable first AI unit.
[0215] It should be noted that before the second device performs the second operation, it can align the target quantization information with the first device, such as through at least one of protocol agreement, second device instruction or first device reporting, so that the first device and the second device reach an agreement on the target quantization information.
[0216] In addition, the above-mentioned first information, second information, first AI unit, parameter information, and target quantization information in the embodiments of the present application have the same meaning and function as the first information, second information, first AI unit, parameter information, and target quantization information in the first device-side method embodiment, and will not be repeated here.
[0217] The embodiments of the present application correspond to the method embodiments on the first device side. In one embodiment, the second device can quantize the parameter information of the first AI unit according to the target quantization information agreed with the first device, and directly send the quantized fixed-point parameters to the first device; in another embodiment, the second device can send floating-point parameters matching the target quantization information to the first device according to the target quantization information agreed with the first device. The methods on the first device side and the second device side are combined to jointly achieve the beneficial effect of reducing the quantization complexity and the reasoning accuracy loss of the quantized AI unit, and improving the overall efficiency of the communication system based on the AI unit.
[0218] In some embodiments, the method further comprises:
[0219] The second device receives first capability information from the first device;
[0220] The second device determines the target quantization information according to the first capability information;
[0221] The first capability information is used to indicate third quantization information supported by the first device, and the third quantization information includes the target quantization information.
[0222] In some implementations, the third quantization information is a type of quantization information. In this case, the target quantization information may be the same as the third quantization information.
[0223] In some other implementations, the third quantization information includes at least two types of quantization information. In this case, the target quantization information may be one type of quantization information in the third quantization information.
[0224] In this implementation, the first device reports the third quantization information supported by the first device to the second device, so that the second device can determine the target quantization information from the third quantization information supported by the first device.
[0225] In some embodiments, the method further comprises:
[0226] The second device determines the target quantization information according to the second quantization information agreed upon in the protocol.
[0227] Optionally, the method further includes:
[0228] The second device sends the target quantization information to the first device.
[0229] In this embodiment, when the first device quantizes the parameter information of the first AI unit, the second device can send the target quantization information to the first device after determining the target quantization information, so that the first device quantizes the parameter information of the first AI unit according to the target quantization information.
[0230] In some embodiments, the method further comprises:
[0231] The second device sends first quantization information to the first device, where the first quantization information is used to assist in determining the target quantization information.
[0232] In a possible implementation, after the second device sends the first quantization information to the first device, the first device may determine target quantization information from the first quantization information.
[0233] In another possible implementation, after the second device sends the first quantization information to the first device, the first device may not support the first quantization information, or may find other quantization information that is more suitable for the first AI unit than the first quantization information. In this case, the target quantization information determined by the first device may not be included in the first quantization information.
[0234] In this implementation, the second device may indicate the first quantization information to the first device. It should be noted that the first quantization information may be accepted or rejected by the first device.
[0235] Optionally, the second device sending the first quantization information to the first device includes:
[0236] The second device sends first quantized information to the first device before sending the first information or the second information to the first device.
[0237] In this embodiment, the quantization information may be aligned before being transmitted to the AI unit.
[0238] Optionally, when the second device sends the first quantization information or the target quantization information to the first device, the first quantization information or the target quantization information may be transmitted together with the structure information of the first AI unit. For example, the first quantization information or the target quantization information is used to be transmitted together with the model structure information before the model neuron coefficients are transmitted.
[0239] Of course, the first quantization information or the target quantization information may also be transmitted separately, for example, before transmitting the model neuron coefficients and the model structure information, the first quantization information or the target quantization information may be transmitted separately.
[0240] In some embodiments, the method further comprises:
[0241] The second device receives the target quantitative information from the first device.
[0242] For example, if the target quantization information determined by the first device is not included in the first quantization information, the first device sends the target quantization information to the second device. Alternatively, the first device can determine the target quantization information on its own and report the target quantization information to the second device.
[0243] In this implementation, the first device may report the target quantization information to the second device.
[0244] In some embodiments, the target quantization information is used to indicate at least one of a quantization method and a quantization level.
[0245] Optionally, the quantification method includes at least one of the following:
[0246] Direct quantization method, which is used to quantize each parameter of the AI unit;
[0247] A group quantization method is used to quantize each group of parameters of the AI unit, wherein a group of parameters includes at least one parameter, and the parameters in the same group correspond to the same quantization value; a transform domain quantization method is used to convert the floating-point parameters of the AI unit into transform domain parameters, and perform quantization and inverse transformation on the transform domain parameters to obtain the quantized fixed-point parameters of the AI unit;
[0248] The product quantization method is used to quantize a first subspace, where the first subspace is a subspace formed by at least one parameter of the AI unit.
[0249] Optionally, the quantization levels are divided based on at least one of the following:
[0250] The number of bits of a fixed-point number obtained after quantization of a floating-point number;
[0251] The target difference is the difference between two adjacent fixed-point numbers after quantization or the difference between floating-point numbers represented by the lowest bit after quantization.
[0252] In some embodiments, the first AI unit includes a first parameter set, wherein floating-point values of the first parameter set satisfy at least one of the following requirements:
[0253] The maximum amplitude of the parameters in the first parameter set is less than or equal to a first threshold value;
[0254] The maximum amplitude of the parameters in the first parameter set is greater than or equal to a second threshold value;
[0255] The minimum amplitude of the parameters in the first parameter set is less than or equal to a third threshold value;
[0256] The minimum amplitude of the parameters in the first parameter set is greater than or equal to a fourth threshold value;
[0257] The average amplitude of the parameters in the first parameter set is less than or equal to a fifth threshold;
[0258] The average amplitude of the parameters in the first parameter set is greater than or equal to a sixth threshold value;
[0259] The difference between the first value and the second value is less than or equal to a seventh threshold value;
[0260] The ratio of the first value to the second value is less than or equal to an eighth threshold value;
[0261] The difference between the third value and the fourth value is less than or equal to a ninth threshold value;
[0262] The ratio of the third value to the fourth value is less than or equal to a tenth threshold value;
[0263] The difference between the fifth value and the sixth value is less than or equal to the eleventh threshold value;
[0264] The ratio of the fifth value to the sixth value is less than or equal to the twelfth threshold value;
[0265] a difference between a maximum amplitude of a parameter in the first parameter set and a minimum amplitude of a parameter in the first parameter set is less than or equal to a thirteenth threshold;
[0266] a ratio of a maximum amplitude of a parameter in the first parameter set to a minimum amplitude of a parameter in the first parameter set is less than or equal to a fourteenth threshold;
[0267] The difference between the maximum amplitude of the parameters in the first parameter set and the average amplitude of the parameters in the first parameter set is less than or equal to a fifteenth threshold;
[0268] The ratio of the maximum amplitude of the parameters in the first parameter set to the average amplitude of the parameters in the first parameter set is less than or equal to a sixteenth threshold;
[0269] a difference between an average amplitude of the parameters in the first parameter set and a minimum amplitude of the parameters in the first parameter set is less than or equal to a seventeenth threshold;
[0270] a ratio of an average amplitude of the parameters in the first parameter set to a minimum amplitude of the parameters in the first parameter set is less than or equal to an eighteenth threshold;
[0271] Among them, the first parameter set includes at least part of the parameters of the first AI unit; the first AI unit includes a first layer and a second layer; the first value is the maximum amplitude of the parameters in the first layer; the second value is the maximum amplitude of the parameters in the second layer; the third value is the minimum amplitude of the parameters in the first layer; the fourth value is the minimum amplitude of the parameters in the second layer; the fifth value is the average amplitude of the parameters in the first layer; the sixth value is the average amplitude of the parameters in the second layer; and the first value is greater than or equal to the second value, the third value is greater than or equal to the fourth value, and the fifth value is greater than or equal to the sixth value.
[0272] In some embodiments, the maximum amplitude is the minimum value among the largest K1 amplitudes, where K1 is a positive integer.
[0273] In some embodiments, the minimum amplitude is the maximum value among the smallest K2 amplitudes, where K2 is a positive integer.
[0274] In some embodiments, the first AI unit includes a first parameter set, and the quantized fixed-point values of the first parameter set meet at least one of the following requirements:
[0275] The maximum value is less than or equal to the nineteenth threshold;
[0276] The maximum value is greater than or equal to the twentieth threshold;
[0277] The minimum value is less than or equal to the twenty-first threshold;
[0278] The minimum value is greater than or equal to the 22nd threshold;
[0279] The first parameter set includes at least part of the parameters of the first AI unit.
[0280] In some embodiments, the target threshold value is related to a target difference value, wherein the target difference value is a difference between two adjacent fixed-point numbers after quantization or a difference between floating-point numbers represented by the lowest bit after quantization;
[0281] The target threshold value includes at least one of the first threshold value, the second threshold value, the third threshold value, the fourth threshold value, the fifth threshold value, the sixth threshold value, the seventh threshold value, the eighth threshold value, the ninth threshold value, the tenth threshold value, the eleventh threshold value, the twelfth threshold value, the thirteenth threshold value, the fourteenth threshold value, the fifteenth threshold value, the sixteenth threshold value, the seventeenth threshold value, the eighteenth threshold value, the nineteenth threshold value, the twentieth threshold value, the twenty-first threshold value, and the twenty-second threshold value.
[0282] In some embodiments, the first parameter set includes at least one of the following:
[0283] at least some parameters of the first AI unit, or parameters in at least some structure of the first AI unit, indicated by the second device;
[0284] Parameters in at least a portion of the structure of the first AI unit reported by the first device;
[0285] Parameters of the adaptation layer of the first AI unit;
[0286] Parameters of the input layer of the first AI unit;
[0287] parameters of the hidden layer of the first AI unit;
[0288] Parameters of the output layer of the first AI unit.
[0289] In the embodiment of the present application, the steps executed by the second device correspond to the steps executed by the first device in the first device side method embodiment, and the two cooperate with each other to jointly achieve the beneficial effect of reducing the quantization complexity and the loss of inference accuracy of the quantized AI unit, and improving the overall efficiency of the communication system based on the AI unit.
[0290] The quantization alignment method for AI units provided in the embodiments of the present application can be performed by a quantization alignment device for AI units. In the embodiments of the present application, the quantization alignment method for AI units performed by the quantization alignment device for AI units is used as an example to illustrate the quantization alignment device for AI units provided in the embodiments of the present application.
[0291] 6 , an embodiment of the present application further provides a quantization alignment apparatus for an AI unit, which is applied to a first device. As shown in FIG6 , the quantization alignment apparatus 600 for an AI unit includes:
[0292] A first receiving module 601 is configured to receive first information or second information from a second device;
[0293] The first execution module 602 is configured to execute a first operation, where the first operation includes:
[0294] Determine a first AI unit based on the first information; or,
[0295] quantizing the second information according to target quantization information to obtain a first AI unit;
[0296] The first information includes parameter information of the first AI unit, and the first information is fixed-point number information quantized based on target quantization information; the second information includes parameter information of the first AI unit, and the second information is floating-point number information.
[0297] Optionally, the target quantization information is used to indicate at least one of a quantization method and a quantization level.
[0298] Optionally, the quantization alignment device 600 of the AI unit further includes at least one of the following:
[0299] a first determining module, configured to determine the target quantization information according to the first quantization information from the second device;
[0300] The second determining module is configured to determine the target quantitative information according to the second quantitative information agreed upon in the protocol.
[0301] Optionally, the quantization alignment device 600 of the AI unit further includes:
[0302] The first sending module is configured to send the target quantization information to the second device when the first quantization information is not completely the same as the target quantization information.
[0303] Optionally, the first determining module includes:
[0304] a receiving unit, configured to receive first quantized information from the second device before the first receiving module receives the first information or the second information from the second device;
[0305] A determining unit is configured to determine the target quantization information according to the first quantization information.
[0306] Optionally, the quantization alignment device 600 of the AI unit further includes:
[0307] The second sending module is configured to send first capability information to the second device, where the first capability information is used to indicate third quantization information supported by the first device, wherein the third quantization information includes the target quantization information.
[0308] Optionally, the quantification method includes at least one of the following:
[0309] Direct quantization method, which is used to quantize each parameter of the AI unit;
[0310] A group quantization method is used to quantize each group of parameters of the AI unit, wherein a group of parameters includes at least one parameter, and parameters in the same group correspond to the same quantized value;
[0311] A transform domain quantization method is used to convert floating-point parameters of the AI unit into transform domain parameters, and perform quantization and inverse transform processing on the transform domain parameters to obtain quantized fixed-point parameters of the AI unit;
[0312] The product quantization method is used to quantize a first subspace, where the first subspace is a subspace formed by at least one parameter of the AI unit.
[0313] Optionally, the quantization levels are divided based on at least one of the following:
[0314] The number of bits of a fixed-point number obtained after quantization of a floating-point number;
[0315] The difference between two adjacent fixed-point numbers after quantization, or the difference between floating-point numbers represented by the lowest bit after quantization.
[0316] Optionally, the first AI unit includes a first parameter set, and floating-point values of the first parameter set meet at least one of the following requirements:
[0317] The maximum amplitude of the parameters in the first parameter set is less than or equal to a first threshold value;
[0318] The maximum amplitude of the parameters in the first parameter set is greater than or equal to a second threshold value;
[0319] The minimum amplitude of the parameters in the first parameter set is less than or equal to a third threshold value;
[0320] The minimum amplitude of the parameters in the first parameter set is greater than or equal to a fourth threshold value;
[0321] The average amplitude of the parameters in the first parameter set is less than or equal to a fifth threshold;
[0322] The average amplitude of the parameters in the first parameter set is greater than or equal to a sixth threshold value;
[0323] The difference between the first value and the second value is less than or equal to a seventh threshold value;
[0324] The ratio of the first value to the second value is less than or equal to an eighth threshold value;
[0325] The difference between the third value and the fourth value is less than or equal to a ninth threshold value;
[0326] The ratio of the third value to the fourth value is less than or equal to a tenth threshold value;
[0327] The difference between the fifth value and the sixth value is less than or equal to the eleventh threshold value;
[0328] The ratio of the fifth value to the sixth value is less than or equal to the twelfth threshold value;
[0329] a difference between a maximum amplitude of a parameter in the first parameter set and a minimum amplitude of a parameter in the first parameter set is less than or equal to a thirteenth threshold;
[0330] a ratio of a maximum amplitude of a parameter in the first parameter set to a minimum amplitude of a parameter in the first parameter set is less than or equal to a fourteenth threshold;
[0331] The difference between the maximum amplitude of the parameters in the first parameter set and the average amplitude of the parameters in the first parameter set is less than or equal to a fifteenth threshold;
[0332] The ratio of the maximum amplitude of the parameters in the first parameter set to the average amplitude of the parameters in the first parameter set is less than or equal to a sixteenth threshold;
[0333] a difference between an average amplitude of the parameters in the first parameter set and a minimum amplitude of the parameters in the first parameter set is less than or equal to a seventeenth threshold;
[0334] a ratio of an average amplitude of the parameters in the first parameter set to a minimum amplitude of the parameters in the first parameter set is less than or equal to an eighteenth threshold;
[0335] Among them, the first parameter set includes at least part of the parameters of the first AI unit; the first AI unit includes a first layer and a second layer; the first value is the maximum amplitude of the parameters in the first layer; the second value is the maximum amplitude of the parameters in the second layer; the third value is the minimum amplitude of the parameters in the first layer; the fourth value is the minimum amplitude of the parameters in the second layer; the fifth value is the average amplitude of the parameters in the first layer; the sixth value is the average amplitude of the parameters in the second layer; and the first value is greater than or equal to the second value, the third value is greater than or equal to the fourth value, and the fifth value is greater than or equal to the sixth value.
[0336] Optionally, the maximum amplitude is the minimum value among the largest K1 amplitudes, where K1 is a positive integer.
[0337] Optionally, the minimum amplitude is the maximum value among the smallest K2 amplitudes, where K2 is a positive integer.
[0338] Optionally, the first AI unit includes a first parameter set, and quantized fixed-point values of the first parameter set meet at least one of the following requirements:
[0339] The maximum value is less than or equal to the nineteenth threshold;
[0340] The maximum value is greater than or equal to the twentieth threshold;
[0341] The minimum value is less than or equal to the twenty-first threshold;
[0342] The minimum value is greater than or equal to the 22nd threshold;
[0343] The first parameter set includes at least part of the parameters of the first AI unit.
[0344] Optionally, the target threshold value is related to a target difference value, wherein the target difference value is a difference value between two adjacent fixed-point numbers after quantization or a difference value of floating-point numbers represented by the lowest bit after quantization;
[0345] The target threshold value includes at least one of the first threshold value, the second threshold value, the third threshold value, the fourth threshold value, the fifth threshold value, the sixth threshold value, the seventh threshold value, the eighth threshold value, the ninth threshold value, the tenth threshold value, the eleventh threshold value, the twelfth threshold value, the thirteenth threshold value, the fourteenth threshold value, the fifteenth threshold value, the sixteenth threshold value, the seventeenth threshold value, the eighteenth threshold value, the nineteenth threshold value, the twentieth threshold value, the twenty-first threshold value, and the twenty-second threshold value.
[0346] Optionally, the first parameter set includes at least one of the following:
[0347] at least some parameters of the first AI unit, or parameters in at least some structure of the first AI unit, indicated by the second device;
[0348] Parameters in at least a portion of the structure of the first AI unit reported by the first device;
[0349] Parameters of the adaptation layer of the first AI unit;
[0350] Parameters of the input layer of the first AI unit;
[0351] parameters of the hidden layer of the first AI unit;
[0352] Parameters of the output layer of the first AI unit.
[0353] The quantization alignment device 600 of the AI unit provided in the embodiment of the present application can implement each process in the first device side method embodiment and achieve the same technical effect. To avoid repetition, it will not be described here.
[0354] 7 , an embodiment of the present application further provides another quantization alignment apparatus for an AI unit, which is applied to a second device. As shown in FIG7 , the quantization alignment apparatus 700 for an AI unit includes:
[0355] The second execution module 701 is configured to execute a second operation, wherein the second operation includes:
[0356] quantizing the parameter information of the first AI unit according to the target quantization information to obtain first information, wherein the first information is quantized fixed-point number information; sending the first information to the first device; or
[0357] Second information is sent to the first device, where the second information includes parameter information of the first AI unit, and the second information is floating-point number information, and the second information is used to perform quantization according to target quantization information to obtain the first AI unit.
[0358] Optionally, the quantization alignment device 700 of the AI unit further includes:
[0359] a second receiving module, configured to receive first capability information from the first device;
[0360] a third determining module, configured to determine the target quantization information according to the first capability information;
[0361] The first capability information is used to indicate third quantization information supported by the first device, and the third quantization information includes the target quantization information.
[0362] Optionally, the quantization alignment device 700 of the AI unit further includes:
[0363] The fourth determining module is configured to determine the target quantitative information according to the second quantitative information agreed upon in the protocol.
[0364] Optionally, the quantization alignment device 700 of the AI unit further includes:
[0365] A third sending module is configured to send the target quantization information to the first device.
[0366] Optionally, the quantization alignment device 700 of the AI unit further includes:
[0367] The fourth sending module is configured to send first quantization information to the first device, where the first quantization information is used to assist in determining the target quantization information.
[0368] Optionally, the fourth sending module is specifically configured to:
[0369] Before the second execution module sends the first information or the second information to the first device, the second execution module sends first quantized information to the first device.
[0370] Optionally, the quantization alignment device 700 of the AI unit further includes:
[0371] The third receiving module is configured to receive the target quantization information from the first device.
[0372] Optionally, the target quantization information is used to indicate at least one of a quantization method and a quantization level.
[0373] Optionally, the quantification method includes at least one of the following:
[0374] Direct quantization method, which is used to quantize each parameter of the AI unit;
[0375] A group quantization method is used to quantize each group of parameters of the AI unit, wherein a group of parameters includes at least one parameter, and parameters in the same group correspond to the same quantized value;
[0376] A transform domain quantization method is used to convert floating-point parameters of the AI unit into transform domain parameters, and perform quantization and inverse transform processing on the transform domain parameters to obtain quantized fixed-point parameters of the AI unit;
[0377] The product quantization method is used to quantize a first subspace, where the first subspace is a subspace formed by at least one parameter of the AI unit.
[0378] Optionally, the quantization levels are divided based on at least one of the following:
[0379] The number of bits of a fixed-point number obtained after quantization of a floating-point number;
[0380] The difference between two adjacent fixed-point numbers after quantization, or the difference between floating-point numbers represented by the lowest bit after quantization.
[0381] Optionally, the first AI unit includes a first parameter set, and floating-point values of the first parameter set meet at least one of the following requirements:
[0382] The maximum amplitude of the parameters in the first parameter set is less than or equal to a first threshold value;
[0383] The maximum amplitude of the parameters in the first parameter set is greater than or equal to a second threshold value;
[0384] The minimum amplitude of the parameters in the first parameter set is less than or equal to a third threshold value;
[0385] The minimum amplitude of the parameters in the first parameter set is greater than or equal to a fourth threshold value;
[0386] The average amplitude of the parameters in the first parameter set is less than or equal to a fifth threshold;
[0387] The average amplitude of the parameters in the first parameter set is greater than or equal to a sixth threshold value;
[0388] The difference between the first value and the second value is less than or equal to a seventh threshold value;
[0389] The ratio of the first value to the second value is less than or equal to an eighth threshold value;
[0390] The difference between the third value and the fourth value is less than or equal to a ninth threshold value;
[0391] The ratio of the third value to the fourth value is less than or equal to a tenth threshold value;
[0392] The difference between the fifth value and the sixth value is less than or equal to the eleventh threshold value;
[0393] The ratio of the fifth value to the sixth value is less than or equal to the twelfth threshold value;
[0394] a difference between a maximum amplitude of a parameter in the first parameter set and a minimum amplitude of a parameter in the first parameter set is less than or equal to a thirteenth threshold;
[0395] a ratio of a maximum amplitude of a parameter in the first parameter set to a minimum amplitude of a parameter in the first parameter set is less than or equal to a fourteenth threshold;
[0396] The difference between the maximum amplitude of the parameters in the first parameter set and the average amplitude of the parameters in the first parameter set is less than or equal to a fifteenth threshold;
[0397] The ratio of the maximum amplitude of the parameters in the first parameter set to the average amplitude of the parameters in the first parameter set is less than or equal to a sixteenth threshold;
[0398] a difference between an average amplitude of the parameters in the first parameter set and a minimum amplitude of the parameters in the first parameter set is less than or equal to a seventeenth threshold;
[0399] a ratio of an average amplitude of the parameters in the first parameter set to a minimum amplitude of the parameters in the first parameter set is less than or equal to an eighteenth threshold;
[0400] Among them, the first parameter set includes at least part of the parameters of the first AI unit; the first AI unit includes a first layer and a second layer; the first value is the maximum amplitude of the parameters in the first layer; the second value is the maximum amplitude of the parameters in the second layer; the third value is the minimum amplitude of the parameters in the first layer; the fourth value is the minimum amplitude of the parameters in the second layer; the fifth value is the average amplitude of the parameters in the first layer; the sixth value is the average amplitude of the parameters in the second layer; and the first value is greater than or equal to the second value, the third value is greater than or equal to the fourth value, and the fifth value is greater than or equal to the sixth value.
[0401] Optionally, the maximum amplitude is the minimum value among the largest K1 amplitudes, where K1 is a positive integer.
[0402] Optionally, the minimum amplitude is the maximum value among the smallest K2 amplitudes, where K2 is a positive integer.
[0403] Optionally, the first AI unit includes a first parameter set, and quantized fixed-point values of the first parameter set meet at least one of the following requirements:
[0404] The maximum value is less than or equal to the nineteenth threshold;
[0405] The maximum value is greater than or equal to the twentieth threshold;
[0406] The minimum value is less than or equal to the twenty-first threshold;
[0407] The minimum value is greater than or equal to the 22nd threshold;
[0408] The first parameter set includes at least part of the parameters of the first AI unit.
[0409] Optionally, the target threshold value is related to a target difference value, wherein the target difference value is a difference value between two adjacent fixed-point numbers after quantization or a difference value of floating-point numbers represented by the lowest bit after quantization;
[0410] The target threshold value includes at least one of the first threshold value, the second threshold value, the third threshold value, the fourth threshold value, the fifth threshold value, the sixth threshold value, the seventh threshold value, the eighth threshold value, the ninth threshold value, the tenth threshold value, the eleventh threshold value, the twelfth threshold value, the thirteenth threshold value, the fourteenth threshold value, the fifteenth threshold value, the sixteenth threshold value, the seventeenth threshold value, the eighteenth threshold value, the nineteenth threshold value, the twentieth threshold value, the twenty-first threshold value, and the twenty-second threshold value.
[0411] Optionally, the first parameter set includes at least one of the following:
[0412] at least some parameters of the first AI unit, or parameters in at least some structure of the first AI unit, indicated by the second device;
[0413] Parameters in at least a portion of the structure of the first AI unit reported by the first device;
[0414] Parameters of the adaptation layer of the first AI unit;
[0415] Parameters of the input layer of the first AI unit;
[0416] parameters of the hidden layer of the first AI unit;
[0417] Parameters of the output layer of the first AI unit.
[0418] The quantization alignment device 700 of the AI unit provided in the embodiment of the present application can implement each process in the second device side method embodiment and achieve the same technical effect. To avoid repetition, it will not be described here.
[0419] The quantization alignment device of the AI unit in the embodiment of the present application can be an electronic device, such as an electronic device with an operating system, or a component in an electronic device, such as an integrated circuit or chip. The electronic device can be a terminal or a network-side device. For example, the terminal can include but is not limited to the type of terminal 11 listed above, and the network-side device can include but is not limited to the type of access network device or core network device listed above, etc., which are not specifically limited in the embodiment of the present application.
[0420] Optionally, as shown in Figure 8, an embodiment of the present application also provides a communication device 800, including a processor 801 and a memory 802, and the memory 802 stores programs or instructions that can be run on the processor 801. For example: when the communication device 800 acts as a first device, the program or instruction is executed by the processor 801 to implement the various steps of the aforementioned first device side method embodiment and can achieve the same technical effect; when the communication device 800 acts as a second device, the program or instruction is executed by the processor 801 to implement the various steps of the aforementioned second device side method embodiment and can achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0421] An embodiment of the present application also provides a communication device, including a processor and a communication interface.
[0422] In a case where the communication device is a first device, the communication interface is used to receive the first information or the second information from the second device; and the processor is used to perform a first operation, where the first operation includes:
[0423] Determine a first AI unit based on the first information; or,
[0424] quantizing the second information according to target quantization information to obtain a first AI unit;
[0425] The first information includes parameter information of the first AI unit, and the first information is fixed-point number information quantized based on target quantization information; the second information includes parameter information of the first AI unit, and the second information is floating-point number information.
[0426] In a case where the communication device is a second device, the communication interface and the processor are configured to perform a second operation, where the second operation includes:
[0427] quantizing the parameter information of the first AI unit according to the target quantization information to obtain first information, wherein the first information is quantized fixed-point number information; sending the first information to the first device; or
[0428] Second information is sent to the first device, where the second information includes parameter information of the first AI unit, and the second information is floating-point number information, and the second information is used to perform quantization according to target quantization information to obtain the first AI unit.
[0429] This communication device embodiment corresponds to the aforementioned quantization alignment method embodiment of the AI unit on the first device and the second device side. The various implementation processes and implementation methods of the aforementioned method embodiments are applicable to this communication device embodiment and can achieve the same technical effect.
[0430] In some implementations, FIG9 is a schematic diagram of the hardware structure of a terminal implementing an embodiment of the present application.
[0431] The terminal 900 includes but is not limited to: a radio frequency unit 901, a network module 902, an audio output unit 903, an input unit 904, a sensor 905, a display unit 906, a user input unit 907, an interface unit 908, a memory 909 and at least some of the components of the processor 910.
[0432] Those skilled in the art will appreciate that the terminal 900 may also include a power supply (such as a battery) to power various components. The power supply may be logically connected to the processor 910 via a power management system, thereby enabling the power management system to manage charging, discharging, and power consumption. The terminal structure shown in FIG9 does not limit the terminal. The terminal may include more or fewer components than shown, or may combine certain components, or have different component arrangements, which will not be described in detail here.
[0433] It should be understood that in an embodiment of the present application, the input unit 904 may include a graphics processing unit (GPU) 9041 and a microphone 9042, and the graphics processor 9041 processes the image data of a static picture or video obtained by an image capture device (such as a camera) in a video capture mode or an image capture mode. The display unit 906 may include a display panel 9061, and the display panel 9061 may be configured in the form of a liquid crystal display, an organic light emitting diode, etc. The user input unit 907 includes a touch panel 9071 and at least one of other input devices 9072. The touch panel 9071 is also called a touch screen. The touch panel 9071 may include two parts: a touch detection device and a touch controller. Other input devices 9072 may include, but are not limited to, a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, and an operating stick, which will not be repeated here.
[0434] In the embodiment of the present application, after receiving downlink data from a network-side device, the RF unit 901 may transmit the data to the processor 910 for processing. Furthermore, the RF unit 901 may send uplink data to the network-side device. Typically, the RF unit 901 includes, but is not limited to, an antenna, an amplifier, a transceiver, a coupler, a low-noise amplifier, a duplexer, and the like.
[0435] The memory 909 can be used to store software programs or instructions and various data. The memory 909 may mainly include a first storage area for storing programs or instructions and a second storage area for storing data, wherein the first storage area may store an operating system, applications or instructions required for at least one function (such as a sound playback function, an image playback function, etc.). In addition, the memory 909 may include a volatile memory or a non-volatile memory. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct RAM bus random access memory (DRRAM). The memory 909 in the embodiment of the present application includes but is not limited to these and any other suitable types of memory.
[0436] Processor 910 may include one or more processing units. Optionally, processor 910 integrates an application processor and a modem processor. The application processor primarily handles operations related to the operating system, user interface, and application programs, while the modem processor primarily processes wireless communication signals, such as a baseband processor. It is understood that the modem processor may not be integrated into processor 910.
[0437] In one implementation, the terminal 900 serves as the first device.
[0438] The radio frequency unit 901 is configured to receive the first information or the second information from the second device;
[0439] The processor 910 is configured to perform a first operation, where the first operation includes:
[0440] Determine a first AI unit based on the first information; or,
[0441] quantizing the second information according to target quantization information to obtain a first AI unit;
[0442] The first information includes parameter information of the first AI unit, and the first information is fixed-point number information quantized based on target quantization information; the second information includes parameter information of the first AI unit, and the second information is floating-point number information.
[0443] Optionally, the target quantization information is used to indicate at least one of a quantization method and a quantization level.
[0444] Optionally, the processor 910 is further configured to perform at least one of the following:
[0445] determining the target quantitative information according to the first quantitative information from the second device;
[0446] The target quantitative information is determined according to the second quantitative information agreed upon in the protocol.
[0447] Optionally, the radio frequency unit 901 is further configured to, when the first quantization information is not completely identical to the target quantization information, send the target quantization information by the first device to the second device.
[0448] Optionally, the determining, by the processor 910, the target quantization information according to the first quantization information from the second device includes:
[0449] Before receiving the first information or the second information from the second device through the radio frequency unit 901, controlling the radio frequency unit 901 to receive the first quantized information from the second device;
[0450] The target quantization information is determined according to the first quantization information.
[0451] Optionally, the radio frequency unit 901 is further configured to send first capability information to the second device, where the first capability information is used to indicate third quantization information supported by the first device, wherein the third quantization information includes the target quantization information.
[0452] Optionally, the quantification method includes at least one of the following:
[0453] Direct quantization method, which is used to quantize each parameter of the AI unit;
[0454] A group quantization method is used to quantize each group of parameters of the AI unit, wherein a group of parameters includes at least one parameter, and parameters in the same group correspond to the same quantized value;
[0455] A transform domain quantization method is used to convert floating-point parameters of the AI unit into transform domain parameters, and perform quantization and inverse transform processing on the transform domain parameters to obtain quantized fixed-point parameters of the AI unit;
[0456] The product quantization method is used to quantize a first subspace, where the first subspace is a subspace formed by at least one parameter of the AI unit.
[0457] Optionally, the quantization levels are divided based on at least one of the following:
[0458] The number of bits of a fixed-point number obtained after quantization of a floating-point number;
[0459] The difference between two adjacent fixed-point numbers after quantization, or the difference between floating-point numbers represented by the lowest bit after quantization.
[0460] Optionally, the first AI unit includes a first parameter set, and floating-point values of the first parameter set meet at least one of the following requirements:
[0461] The maximum amplitude of the parameters in the first parameter set is less than or equal to a first threshold value;
[0462] The maximum amplitude of the parameters in the first parameter set is greater than or equal to a second threshold value;
[0463] The minimum amplitude of the parameters in the first parameter set is less than or equal to a third threshold value;
[0464] The minimum amplitude of the parameters in the first parameter set is greater than or equal to a fourth threshold value;
[0465] The average amplitude of the parameters in the first parameter set is less than or equal to a fifth threshold;
[0466] The average amplitude of the parameters in the first parameter set is greater than or equal to a sixth threshold value;
[0467] The difference between the first value and the second value is less than or equal to a seventh threshold value;
[0468] The ratio of the first value to the second value is less than or equal to an eighth threshold value;
[0469] The difference between the third value and the fourth value is less than or equal to a ninth threshold value;
[0470] The ratio of the third value to the fourth value is less than or equal to a tenth threshold value;
[0471] The difference between the fifth value and the sixth value is less than or equal to the eleventh threshold value;
[0472] The ratio of the fifth value to the sixth value is less than or equal to the twelfth threshold value;
[0473] a difference between a maximum amplitude of a parameter in the first parameter set and a minimum amplitude of a parameter in the first parameter set is less than or equal to a thirteenth threshold;
[0474] a ratio of a maximum amplitude of a parameter in the first parameter set to a minimum amplitude of a parameter in the first parameter set is less than or equal to a fourteenth threshold;
[0475] The difference between the maximum amplitude of the parameters in the first parameter set and the average amplitude of the parameters in the first parameter set is less than or equal to a fifteenth threshold;
[0476] The ratio of the maximum amplitude of the parameters in the first parameter set to the average amplitude of the parameters in the first parameter set is less than or equal to a sixteenth threshold;
[0477] a difference between an average amplitude of the parameters in the first parameter set and a minimum amplitude of the parameters in the first parameter set is less than or equal to a seventeenth threshold;
[0478] a ratio of an average amplitude of the parameters in the first parameter set to a minimum amplitude of the parameters in the first parameter set is less than or equal to an eighteenth threshold;
[0479] Among them, the first parameter set includes at least part of the parameters of the first AI unit; the first AI unit includes a first layer and a second layer; the first value is the maximum amplitude of the parameters in the first layer; the second value is the maximum amplitude of the parameters in the second layer; the third value is the minimum amplitude of the parameters in the first layer; the fourth value is the minimum amplitude of the parameters in the second layer; the fifth value is the average amplitude of the parameters in the first layer; the sixth value is the average amplitude of the parameters in the second layer; and the first value is greater than or equal to the second value, the third value is greater than or equal to the fourth value, and the fifth value is greater than or equal to the sixth value.
[0480] Optionally, the maximum amplitude is the minimum value among the largest K1 amplitudes, where K1 is a positive integer.
[0481] Optionally, the minimum amplitude is the maximum value among the smallest K2 amplitudes, where K2 is a positive integer.
[0482] Optionally, the first AI unit includes a first parameter set, and quantized fixed-point values of the first parameter set meet at least one of the following requirements:
[0483] The maximum value is less than or equal to the nineteenth threshold;
[0484] The maximum value is greater than or equal to the twentieth threshold;
[0485] The minimum value is less than or equal to the twenty-first threshold;
[0486] The minimum value is greater than or equal to the 22nd threshold;
[0487] The first parameter set includes at least part of the parameters of the first AI unit.
[0488] Optionally, the target threshold value is related to a target difference value, wherein the target difference value is a difference value between two adjacent fixed-point numbers after quantization or a difference value of floating-point numbers represented by the lowest bit after quantization;
[0489] The target threshold value includes at least one of the first threshold value, the second threshold value, the third threshold value, the fourth threshold value, the fifth threshold value, the sixth threshold value, the seventh threshold value, the eighth threshold value, the ninth threshold value, the tenth threshold value, the eleventh threshold value, the twelfth threshold value, the thirteenth threshold value, the fourteenth threshold value, the fifteenth threshold value, the sixteenth threshold value, the seventeenth threshold value, the eighteenth threshold value, the nineteenth threshold value, the twentieth threshold value, the twenty-first threshold value, and the twenty-second threshold value.
[0490] Optionally, the first parameter set includes at least one of the following:
[0491] at least some parameters of the first AI unit, or parameters in at least some structure of the first AI unit, indicated by the second device;
[0492] Parameters in at least a portion of the structure of the first AI unit reported by the first device;
[0493] Parameters of the adaptation layer of the first AI unit;
[0494] Parameters of the input layer of the first AI unit;
[0495] parameters of the hidden layer of the first AI unit;
[0496] Parameters of the output layer of the first AI unit.
[0497] It can be understood that the implementation process of each implementation method mentioned in this embodiment can refer to the relevant description of the aforementioned first device side method embodiment, and achieve the same or corresponding technical effects. To avoid repetition, it will not be repeated here.
[0498] In another embodiment, the terminal 900 serves as the second device.
[0499] The processor 910 is configured to perform a second operation, where the second operation includes:
[0500] quantize the parameter information of the first AI unit according to the target quantization information to obtain first information, wherein the first information is quantized fixed-point number information; control the radio frequency unit 901 to send the first information to the first device; or,
[0501] Control the RF unit 901 to send second information to the first device, where the second information includes parameter information of the first AI unit, and the second information is floating-point number information, and the second information is used to quantize according to the target quantization information to obtain the first AI unit.
[0502] Optionally, the radio frequency unit 901 is further configured to receive first capability information from the first device;
[0503] The processor 910 is further configured to determine the target quantization information according to the first capability information;
[0504] The first capability information is used to indicate third quantization information supported by the first device, and the third quantization information includes the target quantization information.
[0505] Optionally, the processor 910 is further configured to determine the target quantization information according to second quantization information agreed upon in the protocol.
[0506] Optionally, the radio frequency unit 901 is further configured to send the target quantization information to the first device.
[0507] Optionally, the radio frequency unit 901 is further configured to send first quantization information to the first device, where the first quantization information is used to assist in determining the target quantization information.
[0508] Optionally, the sending of the first quantization information to the first device performed by the radio frequency unit 901 includes:
[0509] Before sending the first information or the second information to the first device, first quantized information is sent to the first device.
[0510] Optionally, the radio frequency unit 901 is further configured to receive the target quantization information from the first device.
[0511] Optionally, the target quantization information is used to indicate at least one of a quantization method and a quantization level.
[0512] Optionally, the quantification method includes at least one of the following:
[0513] Direct quantization method, which is used to quantize each parameter of the AI unit;
[0514] A group quantization method is used to quantize each group of parameters of the AI unit, wherein a group of parameters includes at least one parameter, and the parameters in the same group correspond to the same quantization value; a transform domain quantization method is used to convert the floating-point parameters of the AI unit into transform domain parameters, and perform quantization and inverse transformation on the transform domain parameters to obtain the quantized fixed-point parameters of the AI unit;
[0515] The product quantization method is used to quantize a first subspace, where the first subspace is a subspace formed by at least one parameter of the AI unit.
[0516] Optionally, the quantization levels are divided based on at least one of the following:
[0517] The number of bits of a fixed-point number obtained after quantization of a floating-point number;
[0518] The target difference is the difference between two adjacent fixed-point numbers after quantization or the difference between floating-point numbers represented by the lowest bit after quantization.
[0519] Optionally, the first AI unit includes a first parameter set, and floating-point values of the first parameter set meet at least one of the following requirements:
[0520] The maximum amplitude of the parameters in the first parameter set is less than or equal to a first threshold value;
[0521] The maximum amplitude of the parameters in the first parameter set is greater than or equal to a second threshold value;
[0522] The minimum amplitude of the parameters in the first parameter set is less than or equal to a third threshold value;
[0523] The minimum amplitude of the parameters in the first parameter set is greater than or equal to a fourth threshold value;
[0524] The average amplitude of the parameters in the first parameter set is less than or equal to a fifth threshold;
[0525] The average amplitude of the parameters in the first parameter set is greater than or equal to a sixth threshold value;
[0526] The difference between the first value and the second value is less than or equal to a seventh threshold value;
[0527] The ratio of the first value to the second value is less than or equal to an eighth threshold value;
[0528] The difference between the third value and the fourth value is less than or equal to a ninth threshold value;
[0529] The ratio of the third value to the fourth value is less than or equal to a tenth threshold value;
[0530] The difference between the fifth value and the sixth value is less than or equal to the eleventh threshold value;
[0531] The ratio of the fifth value to the sixth value is less than or equal to the twelfth threshold value;
[0532] a difference between a maximum amplitude of a parameter in the first parameter set and a minimum amplitude of a parameter in the first parameter set is less than or equal to a thirteenth threshold;
[0533] a ratio of a maximum amplitude of a parameter in the first parameter set to a minimum amplitude of a parameter in the first parameter set is less than or equal to a fourteenth threshold;
[0534] The difference between the maximum amplitude of the parameters in the first parameter set and the average amplitude of the parameters in the first parameter set is less than or equal to a fifteenth threshold;
[0535] The ratio of the maximum amplitude of the parameters in the first parameter set to the average amplitude of the parameters in the first parameter set is less than or equal to a sixteenth threshold;
[0536] a difference between an average amplitude of the parameters in the first parameter set and a minimum amplitude of the parameters in the first parameter set is less than or equal to a seventeenth threshold;
[0537] a ratio of an average amplitude of the parameters in the first parameter set to a minimum amplitude of the parameters in the first parameter set is less than or equal to an eighteenth threshold;
[0538] Among them, the first parameter set includes at least part of the parameters of the first AI unit; the first AI unit includes a first layer and a second layer; the first value is the maximum amplitude of the parameters in the first layer; the second value is the maximum amplitude of the parameters in the second layer; the third value is the minimum amplitude of the parameters in the first layer; the fourth value is the minimum amplitude of the parameters in the second layer; the fifth value is the average amplitude of the parameters in the first layer; the sixth value is the average amplitude of the parameters in the second layer; and the first value is greater than or equal to the second value, the third value is greater than or equal to the fourth value, and the fifth value is greater than or equal to the sixth value.
[0539] Optionally, the maximum amplitude is the minimum value among the largest K1 amplitudes, where K1 is a positive integer.
[0540] Optionally, the minimum amplitude is the maximum value among the smallest K2 amplitudes, where K2 is a positive integer.
[0541] Optionally, the first AI unit includes a first parameter set, and quantized fixed-point values of the first parameter set meet at least one of the following requirements:
[0542] The maximum value is less than or equal to the nineteenth threshold;
[0543] The maximum value is greater than or equal to the twentieth threshold;
[0544] The minimum value is less than or equal to the twenty-first threshold;
[0545] The minimum value is greater than or equal to the 22nd threshold;
[0546] The first parameter set includes at least part of the parameters of the first AI unit.
[0547] Optionally, the target threshold value is related to a target difference value, wherein the target difference value is a difference value between two adjacent fixed-point numbers after quantization or a difference value of floating-point numbers represented by the lowest bit after quantization;
[0548] The target threshold value includes at least one of the first threshold value, the second threshold value, the third threshold value, the fourth threshold value, the fifth threshold value, the sixth threshold value, the seventh threshold value, the eighth threshold value, the ninth threshold value, the tenth threshold value, the eleventh threshold value, the twelfth threshold value, the thirteenth threshold value, the fourteenth threshold value, the fifteenth threshold value, the sixteenth threshold value, the seventeenth threshold value, the eighteenth threshold value, the nineteenth threshold value, the twentieth threshold value, the twenty-first threshold value, and the twenty-second threshold value.
[0549] Optionally, the first parameter set includes at least one of the following:
[0550] at least some parameters of the first AI unit, or parameters in at least some structure of the first AI unit, indicated by the second device;
[0551] Parameters in at least a portion of the structure of the first AI unit reported by the first device;
[0552] Parameters of the adaptation layer of the first AI unit;
[0553] Parameters of the input layer of the first AI unit;
[0554] parameters of the hidden layer of the first AI unit;
[0555] Parameters of the output layer of the first AI unit.
[0556] It can be understood that the implementation process of each implementation method mentioned in this embodiment can refer to the relevant description of the aforementioned second device side method embodiment, and achieve the same or corresponding technical effects. To avoid repetition, it will not be repeated here.
[0557] The present application also provides a network-side device, including a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is configured to execute a program or instruction to implement the steps of the aforementioned method embodiment on the first device side or the second device side. This network-side device embodiment corresponds to the aforementioned method embodiment on the first device side or the second device side, and each implementation process and implementation method of the aforementioned method embodiment can be applied to this network-side device embodiment and can achieve the same technical effects.
[0558] In one embodiment, as shown in Figure 10, the network-side device 1000 includes an antenna 1001, a radio frequency device 1002, a baseband device 1003, a processor 1004, and a memory 1005. Antenna 1001 is connected to radio frequency device 1002. In the uplink direction, radio frequency device 1002 receives information via antenna 1001 and sends the received information to baseband device 1003 for processing. In the downlink direction, baseband device 1003 processes the information to be transmitted and sends it to radio frequency device 1002. Radio frequency device 1002 processes the received information and then sends it through antenna 1001.
[0559] The method executed by the network-side device in the above embodiment may be implemented in the baseband device 1003 , which includes a baseband processor.
[0560] The baseband device 1003 may include, for example, at least one baseband board, on which multiple chips are arranged, as shown in Figure 11, one of the chips is, for example, a baseband processor, which is connected to the memory 1005 through a bus interface to call the program in the memory 1005 and execute the network device operations shown in the above method embodiment.
[0561] The network side device may further include a network interface 1006, which is, for example, a Common Public Radio Interface (CPRI).
[0562] Specifically, the network side device 1000 of the embodiment of the present application also includes: instructions or programs stored in the memory 1005 and executable on the processor 1004. The processor 1004 calls the instructions or programs in the memory 1005 to execute the method executed by each module shown in at least one of Figures 6 and 7, and achieves the same technical effect. To avoid repetition, it will not be described here.
[0563] In another embodiment, the present application also provides a network-side device. As shown in FIG11 , the network-side device 1100 includes a processor 1101, a network interface 1102, and a memory 1103. The network interface 1102 is, for example, a Common Public Radio Interface (CPRI).
[0564] Specifically, the network side device 1100 of the embodiment of the present application also includes: instructions or programs stored in the memory 1103 and can be run on the processor 1101. The processor 1101 calls the instructions or programs in the memory 1103 to execute the method executed by each module as shown in at least one of Figures 6 and 7, and achieves the same technical effect. To avoid repetition, it will not be repeated here.
[0565] An embodiment of the present application also provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by the processor, the various processes of the aforementioned quantization alignment method embodiment of the AI unit on the first device side or the second device side are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.
[0566] The processor is the processor in the terminal described in the above embodiment. The readable storage medium includes a computer-readable storage medium, such as a computer read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk. In some examples, the readable storage medium may be a non-transitory readable storage medium.
[0567] An embodiment of the present application further provides a chip, which includes a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the various processes of the aforementioned quantization alignment method embodiment of the AI unit on the first device side or the second device side, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0568] It should be understood that the chip mentioned in the embodiments of the present application can also be called a system-level chip, a system chip, a chip system or a system-on-chip chip, etc.
[0569] An embodiment of the present application further provides a computer program / program product, which is stored in a storage medium. The computer program / program product is executed by at least one processor to implement the various processes of the aforementioned quantization alignment method embodiment of the AI unit on the first device side or the second device side, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0570] An embodiment of the present application also provides a communication system, including: a first device and a second device, wherein the first device can be used to execute the steps of the embodiment of the quantization alignment method of the AI unit on the first device side, and the second device can be used to execute the steps of the embodiment of the quantization alignment method of the AI unit on the second device side.
[0571] It should be noted that, in this article, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the presence of other identical elements in the process, method, article or device comprising the element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the opposite order according to the functions involved. For example, the described method may be performed in an order different from that described, and various steps may also be added, omitted or combined. In addition, the features described with reference to certain examples may be combined in other examples.
[0572] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of a computer software product plus a necessary general-purpose hardware platform, or of course, by hardware. The computer software product is stored in a storage medium (such as ROM, RAM, magnetic disk, optical disk, etc.) and includes a number of instructions for enabling a terminal or network-side device to execute the methods described in each embodiment of the present application.
[0573] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of this application, ordinary technicians in this field can also make many forms of implementation methods without departing from the purpose of this application and the scope of protection of the claims. These implementation methods are all within the protection of this application.
Claims
1. A quantization alignment method for an AI unit, comprising: The first device receives the first information or the second information from the second device; The first device performs a first operation, and the first operation includes: Determining a first AI unit according to the first information; or Quantizing the second information according to the target quantization information to obtain a first AI unit; Wherein, the first information includes parameter information of the first AI unit, and the first information is fixed-point number information quantized based on the target quantization information; the second information includes parameter information of the first AI unit, and the second information is floating-point number information.
2. The method according to claim 1, wherein, The target quantization information is used to indicate at least one of a quantization method and a quantization level.
3. The method according to claim 1 or 2, further comprising at least one of the following: The first device determines the target quantization information according to the first quantization information from the second device; The first device determines the target quantization information according to the second quantization information agreed upon by the protocol.
4. The method according to claim 3, further comprising: In the case where the first quantization information is not completely the same as the target quantization information, the first device sends the target quantization information to the second device.
5. The method according to claim 3, wherein The first device determines the target quantization information according to the first quantization information from the second device, including: Before receiving the first information or the second information from the second device, the first device receives the first quantization information from the second device; The first device determines the target quantization information according to the first quantization information.
6. The method according to any one of claims 1 to 5, further comprising: The first device sends first capability information to the second device, and the first capability information is used to indicate third quantization information supported by the first device, wherein the third quantization information includes the target quantization information.
7. The method according to claim 2, wherein The quantization method includes at least one of the following: Direct quantization method, which is used to quantize each parameter of the AI unit; Group quantization method, which is used to quantize each group of parameters of the AI unit, wherein a group of parameters includes at least one parameter, and the parameters in the same group correspond to the same quantization value; Transform domain quantization method, which is used to convert the floating-point parameters of the AI unit into transform domain parameters, and perform quantization processing and inverse transform processing on the transform domain parameters to obtain the quantized fixed-point parameters of the AI unit; Product quantization method, which is used to quantize a first subspace, and the first subspace is a subspace composed of at least one parameter of the AI unit.
8. The method according to claim 2 or 7, wherein, The quantization level is divided based on at least one of the following: The number of bits of the fixed-point number obtained after quantizing a floating-point number; The difference between two adjacent fixed-point numbers after quantization, or the floating-point number difference represented by the lowest bit after quantization.
9. A quantization alignment method for an AI unit, comprising: The second device performs a second operation, wherein the second operation includes: Quantize the parameter information of the first AI unit according to the target quantization information to obtain first information, where the first information is quantized fixed-point number information; send the first information to the first device; or, Send second information to the first device, where the second information includes the parameter information of the first AI unit, and the second information is floating-point number information, and the second information is used to be quantized according to the target quantization information to obtain the first AI unit.
10. The method according to claim 9 further includes: The second device receives first capability information from the first device; The second device determines the target quantization information according to the first capability information; Wherein, the first capability information is used to indicate third quantization information supported by the first device, and the third quantization information includes the target quantization information.
11. The method according to claim 9 further includes: The second device determines the target quantization information according to second quantization information agreed upon by the protocol.
12. The method according to claim 10 or 11 further includes: The second device sends the target quantization information to the first device.
13. The method according to claim 9 further includes: The second device sends first quantization information to the first device, and the first quantization information is used to assist in determining the target quantization information.
14. The method according to claim 13, wherein, The second device sending first quantization information to the first device includes: Before the second device sends the first information or the second information to the first device, the second device sends the first quantization information to the first device.
15. The method according to claim 9 or 13 further includes: The second device receives the target quantization information from the first device.
16. The method according to any one of claims 9 to 15, wherein The target quantization information is used to indicate at least one of a quantization method and a quantization level.
17. The method according to claim 16, wherein, The quantization method includes at least one of the following: Direct quantization method, which is used to quantize each parameter of the AI unit; Group quantization method, which is used to quantize each group of parameters of the AI unit, where a group of parameters includes at least one parameter, and the parameters in the same group correspond to the same quantization value; Transform domain quantization method, which is used to convert the floating-point parameters of the AI unit into transform domain parameters, and perform quantization processing and inverse transform processing on the transform domain parameters to obtain the quantized fixed-point parameters of the AI unit; Product quantization method, which is used to quantize the first subspace, and the first subspace is a subspace composed of at least one parameter of the AI unit.
18. The method according to claim 16 or 17, wherein The quantization level is divided based on at least one of the following: The number of bits of the fixed-point number obtained after quantizing a floating-point number; Target difference, where the target difference is the difference between two adjacent fixed-point numbers after quantization or the floating-point number difference represented by the lowest bit after quantization.
19. The method according to any one of claims 1 to 18, wherein, The first AI unit includes a first parameter set, and the floating-point values of the first parameter set satisfy at least one of the following requirements: The maximum amplitude of the parameters in the first parameter set is less than or equal to the first threshold; The maximum amplitude of the parameters in the first parameter set is greater than or equal to the second threshold; The minimum amplitude of the parameters in the first parameter set is less than or equal to the third threshold value; The minimum amplitude of the parameters in the first parameter set is greater than or equal to the fourth threshold value; The average amplitude of the parameters in the first parameter set is less than or equal to the fifth threshold value; The average amplitude of the parameters in the first parameter set is greater than or equal to the sixth threshold value; The difference between the first value and the second value is less than or equal to the seventh threshold value; The ratio of the first value to the second value is less than or equal to the eighth threshold value; The difference between the third value and the fourth value is less than or equal to the ninth threshold value; The ratio of the third value to the fourth value is less than or equal to the tenth threshold value; The difference between the fifth value and the sixth value is less than or equal to the eleventh threshold value; The ratio of the fifth value to the sixth value is less than or equal to the twelfth threshold value; The difference between the maximum amplitude and the minimum amplitude of the parameters in the first parameter set is less than or equal to the thirteenth threshold value; The ratio of the maximum amplitude to the minimum amplitude of the parameters in the first parameter set is less than or equal to the fourteenth threshold value; The difference between the maximum amplitude and the average amplitude of the parameters in the first parameter set is less than or equal to the fifteenth threshold value; The ratio of the maximum amplitude to the average amplitude of the parameters in the first parameter set is less than or equal to the sixteenth threshold value; The difference between the average amplitude and the minimum amplitude of the parameters in the first parameter set is less than or equal to the seventeenth threshold value; The ratio of the average amplitude to the minimum amplitude of the parameters in the first parameter set is less than or equal to the eighteenth threshold value; Wherein, the first parameter set includes at least some parameters of the first AI unit; the first AI unit includes a first layer and a second layer; the first value is the maximum amplitude of the parameters in the first layer; the second value is the maximum amplitude of the parameters in the second layer; the third value is the minimum amplitude of the parameters in the first layer; the fourth value is the minimum amplitude of the parameters in the second layer; the fifth value is the average amplitude of the parameters in the first layer; the sixth value is the average amplitude of the parameters in the second layer; and the first value is greater than or equal to the second value, the third value is greater than or equal to the fourth value, and the fifth value is greater than or equal to the sixth value.
20. The method according to claim 19, wherein, The maximum amplitude is the minimum value among the largest K1 amplitudes, where K1 is a positive integer.
21. The method according to claim 19, wherein, The minimum amplitude is the maximum value among the smallest K2 amplitudes, where K2 is a positive integer.
22. The method according to any one of claims 1 to 18, wherein The first AI unit includes a first parameter set, and the fixed-point values after quantization of the first parameter set satisfy at least one of the following requirements: The maximum value is less than or equal to the nineteenth threshold; The maximum value is greater than or equal to the twentieth threshold; The minimum value is less than or equal to the twenty-first threshold; The minimum value is greater than or equal to the twenty-second threshold; Wherein, the first parameter set includes at least some parameters of the first AI unit.
23. The method according to claim 19 or 22, wherein The target threshold value is related to the target difference, and the target difference is the difference between two adjacent fixed-point numbers after quantization or the floating-point number difference represented by the lowest bit after quantization; Among them, the target threshold value includes at least one of the first threshold value, the second threshold value, the third threshold value, the fourth threshold value, the fifth threshold value, the sixth threshold value, the seventh threshold value, the eighth threshold value, the ninth threshold value, the tenth threshold value, the eleventh threshold value, the twelfth threshold value, the thirteenth threshold value, the fourteenth threshold value, the fifteenth threshold value, the sixteenth threshold value, the seventeenth threshold value, the eighteenth threshold value, the nineteenth threshold value, the twentieth threshold value, the twenty - first threshold value, and the twenty - second threshold value.
24. The method according to claim 19 or 22, wherein The first parameter set includes at least one of the following: At least some parameters of the first AI unit indicated by the second device, or parameters in at least some structures of the first AI unit; Parameters in at least some structures of the first AI unit reported by the first device; Parameters of the adaptive layer of the first AI unit; Parameters of the input layer of the first AI unit; Parameters of the hidden layer of the first AI unit; Parameters of the output layer of the first AI unit.
25. A quantization alignment device for an AI unit, used for a first device, the device includes: A first receiving module, configured to receive the first information or the second information from a second device; A first execution module, configured to execute a first operation, and the first operation includes: Determining a first AI unit according to the first information; or, Quantizing the second information according to target quantization information to obtain a first AI unit; Among them, the first information includes parameter information of the first AI unit, and the first information is fixed - point number information quantized based on the target quantization information; the second information includes parameter information of the first AI unit, and the second information is floating - point number information.
26. The device according to claim 25, further includes at least one of the following: A first determining module, configured to determine the target quantization information according to the first quantization information from the second device; A second determining module, configured to determine the target quantization information according to the second quantization information agreed upon by the protocol.
27. The device according to claim 26, further includes: A first sending module, configured to send the target quantization information to the second device when the first quantization information is not exactly the same as the target quantization information.
28. The apparatus according to claim 26, wherein, The first determining module includes: A receiving unit, configured to receive the first quantization information from the second device before the first receiving module receives the first information or the second information from the second device; A determining unit, configured to determine the target quantization information according to the first quantization information.
29. The device according to any one of claims 25 to 28, further includes: A second sending module, configured to send first capability information to the second device, and the first capability information is used to indicate the third quantization information supported by the first device, where the third quantization information includes the target quantization information.
30. A quantization alignment device for an AI unit, used for a second device, the device includes: A second execution module, configured to execute a second operation, where the second operation includes: Quantifying the parameter information of the first AI unit according to target quantization information to obtain first information, where the first information is fixed-point number information after quantization; sending the first information to a first device; or, Sending second information to the first device, where the second information includes the parameter information of the first AI unit, and the second information is floating-point number information, and the second information is used to be quantized according to the target quantization information to obtain the first AI unit.
31. The apparatus according to claim 30, further comprising: A second receiving module, configured to receive first capability information from the first device; A third determining module, configured to determine the target quantization information according to the first capability information; where the first capability information is used to indicate third quantization information supported by the first device, and the third quantization information includes the target quantization information.
32. The apparatus according to claim 30, further comprising: A fourth determining module, configured to determine the target quantization information according to second quantization information agreed upon by a protocol.
33. The apparatus according to claim 31 or 32, further comprising: A third sending module, configured to send the target quantization information to the first device.
34. The apparatus according to claim 30, further comprising: A fourth sending module, configured to send first quantization information to the first device, where the first quantization information is used to assist in determining the target quantization information.
35. The apparatus according to claim 34, wherein, The fourth sending module is specifically configured to: Before the second execution module sends the first information or the second information to the first device, send the first quantization information to the first device.
36. The apparatus according to claim 30 or 34, further comprising: A third receiving module, configured to receive the target quantization information from the first device.
37. A communication device, comprising a processor and a memory, where the memory stores a program or instruction that can run on the processor, and when the program or instruction is executed by the processor, the steps of the quantization alignment method of the AI unit according to any one of claims 1 to 24 are implemented.
38. A readable storage medium, where a program or instruction is stored on the readable storage medium, and when the program or instruction is executed by a processor, the steps of the quantization alignment method of the AI unit according to any one of claims 1 to 24 are implemented.
Citation Information
Patent Citations
Fixed point method and apparatus, and computer device
CN108596328A
Bit width fixed-point method and device in neural network, terminal and storage medium
CN110929838A
Convolution operation device and method and related product
CN115469827A
Dynamic precision management for integer deep learning primitives
US20210110508A1