Quantization alignment method and device of AI unit and related equipment

By performing quantitative alignment between the sending node and the receiving node of the AI ​​unit, the problem of unnecessary quantization in the transmission of AI unit parameters is solved, and the inference accuracy and communication system efficiency are improved.

CN120238399APending Publication Date: 2025-07-01VIVO MOBILE COMM CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311843077.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-28
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

In a communication network, when the parameters of the AI ​​unit are transmitted from the training device to the inference device, unnecessary quantization may be performed, resulting in an increase in quantization complexity and loss of inference accuracy, reducing the efficiency of the overall communication system.

Method used

By quantizing the alignment between the sending node and the receiving node of the AI ​​unit, we ensure that the quantization of the AI ​​unit parameters at both ends is consistent, reducing the quantization complexity and improving the inference accuracy.

Benefits of technology

It realizes reducing the quantization complexity and improving the inference accuracy of AI units, and improving the overall efficiency of the communication system based on AI units.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120238399A_ABST
    Figure CN120238399A_ABST
Patent Text Reader

Abstract

The invention discloses a quantization alignment method and device of an AI unit and related equipment, and belongs to the technical field of communication, and the quantization alignment method of the AI unit in the embodiment of the invention comprises the steps that first equipment receives first information or second information from second equipment; the first device executes a first operation, and the first operation comprises: determining a first AI unit according to the first information; or, the second information is quantized according to target quantization information, and a first AI unit is obtained; wherein the first information comprises parameter information of the first AI unit, and the first information is fixed point number information quantized based on target quantization information; the second information comprises parameter information of the first AI unit, and the second information is floating point number information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of communication technologies, and particularly relates to a quantization alignment method, apparatus, and related equipment for an AI unit. Background Art

[0002] In a communication network, methods of artificial intelligence (AI) are introduced to improve network communication performance.

[0003] Generally, the device for training an AI unit and the device for using the AI unit for inference may not be the same device. At this time, it is necessary to transfer the AI unit.

[0004] However, the parameters of the trained AI unit can be floating-point parameters, while fixed-point parameters of the AI unit are required for inference using the AI unit. In related technologies, taking the network sending the AI unit to the terminal as an example, the network sends the floating-point parameters of the AI unit to the terminal, and the terminal quantizes the floating-point parameters to obtain the fixed-point parameters of the AI unit that can be used. However, in this process, the terminal may perform unnecessary quantization on the floating-point parameters of the AI unit, resulting in an increase in quantization complexity, or a loss of inference accuracy of the AI unit due to inappropriate quantization, thereby reducing the overall efficiency of the communication system based on the AI unit. Summary of the Invention

[0005] Embodiments of this application provide a quantization alignment method, apparatus, and related equipment for an AI unit. In the scenario of transferring an AI unit, quantization alignment is performed between the sending node and the receiving node of the AI unit, which can reduce quantization complexity and loss of inference accuracy of the quantized AI unit, thereby improving the overall efficiency of the communication system based on the AI unit.

[0006] In a first aspect, a quantization alignment method for an AI unit is provided. The method includes:

[0007] A first device receives first information or second information from a second device;

[0008] The first device performs a first operation, and the first operation includes:

[0009] Determine a first AI unit according to the first information; or,

[0010] Quantize the second information according to target quantization information to obtain a first AI unit;

[0011] Among them, the first information includes parameter information of the first AI unit, and the first information is fixed-point number information quantized based on target quantization information; the second information includes parameter information of the first AI unit, and the second information is floating-point number information.

[0012] In a second aspect, another quantization alignment method for an AI unit is provided. The method includes:

[0013] A second device performs a second operation, where the second operation includes:

[0014] Quantize the parameter information of the first AI unit according to the target quantization information to obtain first information, where the first information is fixed-point number information after quantization; send the first information to a first device; or,

[0015] Send second information to the first device, where the second information includes parameter information of the first AI unit, and the second information is floating-point number information, and the second information is used to be quantized according to the target quantization information to obtain the first AI unit.

[0016] In a third aspect, a quantization alignment device for an AI unit is provided for a first device. The device includes:

[0017] A first receiving module, configured to receive the first information or the second information from a second device;

[0018] A first execution module, configured to perform a first operation, where the first operation includes:

[0019] Determine a first AI unit according to the first information; or,

[0020] Quantize the second information according to the target quantization information to obtain a first AI unit;

[0021] Among them, the first information includes parameter information of the first AI unit, and the first information is fixed-point number information quantized based on target quantization information; the second information includes parameter information of the first AI unit, and the second information is floating-point number information.

[0022] In a fourth aspect, a quantization alignment device for an AI unit is provided for a second device. The device includes:

[0023] A second execution module, configured to perform a second operation, where the second operation includes:

[0024] Quantize the parameter information of the first AI unit according to the target quantization information to obtain first information, where the first information is fixed-point number information after quantization; send the first information to a first device; or,

[0025] Send second information to a first device, where the second information includes parameter information of a first AI unit, and the second information is floating-point number information, and the second information is used to be quantized according to target quantization information to obtain the first AI unit.

[0026] In a fifth aspect, a communication device is provided, which includes a processor and a memory. The memory stores a program or instructions that can be run on the processor. When the program or instructions are executed by the processor, the steps of the method described in the first aspect or the second aspect are implemented.

[0027] In a sixth aspect, a communication device is provided, including a processor and a communication interface:

[0028] Wherein, when the communication device is used as the first device, the communication interface is used to receive first information or second information from a second device; the processor is used to perform a first operation, and the first operation includes:

[0029] Determine a first AI unit according to the first information; or,

[0030] Quantize the second information according to target quantization information to obtain a first AI unit;

[0031] Wherein, the first information includes parameter information of the first AI unit, and the first information is fixed-point number information quantized based on target quantization information; the second information includes parameter information of the first AI unit, and the second information is floating-point number information;

[0032] Or,

[0033] When the communication device is used as the second device, the communication interface and the processor are used to perform a second operation, where the second operation includes:

[0034] Quantize the parameter information of the first AI unit according to target quantization information to obtain first information, where the first information is quantized fixed-point number information; send the first information to the first device; or,

[0035] Send second information to the first device, where the second information includes parameter information of the first AI unit, and the second information is floating-point number information, and the second information is used to be quantized according to target quantization information to obtain the first AI unit.

[0036] In a seventh aspect, a readable storage medium is provided. A program or instructions are stored on the readable storage medium. When the program or instructions are executed by a processor, the steps of the method described in the first aspect are implemented, or the steps of the method described in the second aspect are implemented.

[0037] In an eighth aspect, a wireless communication system is provided, including: a first device and a second device. The first device can be used to execute the steps of the method described in the first aspect, and the second device can be used to execute the steps of the method described in the second aspect.

[0038] In a ninth aspect, a chip is provided. The chip includes a processor and a communication interface. The communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the method described in the first aspect or the second aspect.

[0039] In a tenth aspect, a computer program / program product is provided. The computer program / program product is stored in a storage medium, and the program / program product is executed by at least one processor to implement the steps of the method described in the first aspect or the second aspect.

[0040] In an implementation manner of the embodiments of the present application, the second device can directly transmit the parameter information of the quantized first AI unit to the first device, so that the first device and the second device align the parameter quantization of the first AI unit; in another implementation manner of the embodiments of the present application, the second device transmits the parameter information of the first AI unit before quantization to the first device. At this time, the first device can quantize the parameter information of the first AI unit based on the target quantization information agreed upon with the second device, so as to achieve the parameter quantization alignment of the first AI unit between the first device and the second device. By aligning the parameter quantization of the first AI unit between the first device and the second device, the quantization complexity and the loss of the inference accuracy of the quantized AI unit can be reduced, thereby improving the overall efficiency of the communication system based on the AI unit. Description of the Drawings

[0041] Figure 1 is a schematic structural diagram of a wireless communication system to which the embodiments of the present application can be applied;

[0042] Figure 2 is a schematic diagram of a neural network;

[0043] Figure 3 is a schematic diagram of a neuron;

[0044] Figure 4 is a flowchart of a method for quantizing and aligning an AI unit provided by the embodiments of the present application;

[0045] Figure 5 is a flowchart of another method for quantizing and aligning an AI unit provided by the embodiments of the present application;

[0046] Figure 6 is a schematic structural diagram of a device for quantizing and aligning an AI unit provided by the embodiments of the present application;

[0047] Figure 7 It is a schematic structural diagram of another quantization alignment device of the AI unit provided by an embodiment of the present application;

[0048] Figure 8 It is a schematic structural diagram of a communication device provided by an embodiment of the present application;

[0049] Figure 9 It is a schematic structural diagram of a terminal provided by an embodiment of the present application;

[0050] Figure 10 It is a schematic structural diagram of a network-side device provided by an embodiment of the present application;

[0051] Figure 11 It is a schematic structural diagram of another network-side device provided by an embodiment of the present application. Detailed implementation manners

[0052] Next, the technical solutions in the embodiments of the present application will be clearly described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art belong to the scope protected by the present application.

[0053] The terms "first", "second", etc. in the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first" and "second" are usually of the same type, and the number of objects is not limited. For example, the first object can be one or multiple. In addition, "or" in the present application means at least one of the connected objects. For example, "A or B" covers three scenarios, namely, Scenario 1: including A and not including B; Scenario 2: including B and not including A; Scenario 3: including both A and B. The character " / " generally indicates an "or" relationship between the associated objects before and after.

[0054] The term "indication" in the present application can be either a direct indication (or an explicit indication) or an indirect indication (or an implicit indication). Among them, a direct indication can be understood as that the sender clearly informs the receiver of specific information, operations to be performed, or request results, etc. in the sent indication; an indirect indication can be understood as that the receiver determines the corresponding information according to the indication sent by the sender, or makes a judgment and determines the operations to be performed or request results, etc. according to the judgment result.

[0055] It should be noted that the technology described in the embodiments of this application is not limited to Long Term Evolution (LTE) / LTE-Advanced (LTE-A) systems, and can also be used in other wireless communication systems, such as Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Frequency Division Multiple Access (FDMA), Orthogonal Frequency Division Multiple Access (OFDMA), Single-carrier Frequency-Division Multiple Access (SC-FDMA), or other systems. The terms "system" and "network" in the embodiments of this application are often used interchangeably, and the described technology can be used in the above-mentioned systems and radio technologies, as well as in other systems and radio technologies. The following description describes the New Radio (NR) system for example purposes, and the NR term is used in most of the following descriptions, but these technologies can also be applied to systems other than the NR system, such as the 6th th Generation (6G) communication system.

[0056] Figure 1Block diagram of a wireless communication system to which embodiments of the present application can be applied. The wireless communication system includes a terminal 11 and a network-side device 12. Among them, the terminal 11 can be a mobile phone, a tablet personal computer, a laptop computer, a notebook computer, a personal digital assistant (PDA), a handheld computer, a netbook, an ultra-mobile personal computer (UMPC), a mobile internet device (MID), an augmented reality (AR), a virtual reality (VR) device, a robot, a wearable device, a flight vehicle, a vehicle user equipment (VUE), a shipborne device, a pedestrian user equipment (PUE), a smart home (home devices with wireless communication functions, such as refrigerators, TVs, washing machines or furniture, etc.), a game console, a personal computer (PC), a teller machine or a self-service machine, etc. Wearable devices include: smart watches, smart bracelets, smart earphones, smart glasses, smart jewelry (smart bracelets, smart bracelets, smart rings, smart necklaces, smart anklets, smart ankle chains, etc.), smart wristbands, smart clothing, etc. Among them, the vehicle user equipment can also be called a vehicle terminal, a vehicle controller, a vehicle module, a vehicle component, a vehicle chip or a vehicle unit, etc. It should be noted that the specific type of the terminal 11 is not limited in the embodiments of the present application. The network-side device 12 can include an access network device or a core network device. Among them, the access network device can also be called a radio access network (RAN) device, a radio access network function or a radio access network unit. The access network device can include a base station, a wireless local area network (WLAN) access point (AP) or a wireless fidelity (WiFi) node, etc.Among them, the base station may be referred to as Node B (NB), Evolved Node B (eNB), the next generation Node B (gNB), New Radio Node B (NR Node B), access point, Relay Base Station (RBS), Serving Base Station (SBS), Base Transceiver Station (BTS), radio base station, radio transceiver, Basic Service Set (BSS), Extended Service Set (ESS), home Node B (HNB), home evolved Node B, Transmission Reception Point (TRP), or some other suitable term in the art. As long as the same technical effect is achieved, the base station is not limited to specific technical terms. It should be noted that in the embodiments of this application, only the base station in the NR system is taken as an example for introduction, and the specific type of the base station is not limited.

[0057] The core network device may include, but is not limited to, at least one of the following: core network node, core network function, Location Management Function (LMF), Mobility Management Entity (MME), Access and Mobility Management Function (AMF), Session Management Function (SMF), User Plane Function (UPF), Policy Control Function (PCF), Policy and Charging Rules Function (PCRF), Edge Application Server Discovery Function (EASDF), Unified Data Management (UDM), Unified Data Repository (UDR), Home Subscriber Server (HSS), Centralized network configuration (CNC), Network Repository Function (NRF), Network Exposure Function (NEF), Local NEF (L-NEF), Binding Support Function (BSF), Application Function (AF), Network Data Analytics Function (NWDAF), etc. It should be noted that in the embodiments of this application, only the core network devices in the NR system are taken as examples for introduction, and the specific types of core network devices are not limited. It should be noted that in the embodiments of this application, only the core network devices in the NR system are taken as examples for introduction, and the specific types of core network devices are not limited.

[0058] Artificial intelligence has currently been widely applied in various fields. There are multiple implementation methods for the AI unit, such as neural network, decision tree, support vector machine, Bayesian classifier, etc. In this application, the neural network is taken as an example for illustration, but the specific type of the AI unit is not limited.

[0059] Such asFigure 2 As shown, the neural network includes an input layer, a hidden layer, and an output layer, which can predict possible output results (Y) based on the input information (X1 to X n ) obtained by the input layer. The neural network is composed of a large number of neurons. As Figure 3 shown, the parameters of a neuron include: input parameters a1 to a K , weights w, biases b, and activation function σ(z), and the output value a is obtained from these parameters. Among them, common activation functions include the sigmoid function, the hyperbolic tangent (tanh) function, the rectified linear unit (ReLU, also known as the rectified linear unit) function, etc. And z in the above function σ(z) can be calculated by the following formula:

[0060] z = a1w1 + … + a k w k + a K w K + b

[0061] where K represents the total number of input parameters.

[0062] The parameters of the neural network are optimized by an optimization algorithm. An optimization algorithm is a type of algorithm that can help us minimize or maximize an objective function (sometimes also called a loss function). And the objective function is often a mathematical combination of AI unit parameters and data. For example, given data X and its corresponding label Y, we construct a neural network f(.). After having the neural network model, based on the input x, we can obtain the predicted output f(x), and we can calculate the difference between the predicted value and the true value (f(x) - Y), which is the loss function. Our goal is to find the appropriate W and b to minimize the value of the above loss function. The smaller the loss value, the closer our AI unit is to the real situation.

[0063] Currently, the common optimization algorithms are basically based on the backpropagation algorithm. The basic idea of the backpropagation algorithm is that the learning process consists of two processes: the forward propagation of signals and the backpropagation of errors. During forward propagation, the input samples are fed into the input layer and processed layer by layer through each hidden layer, and then transmitted to the output layer. If the actual output of the output layer does not match the expected output, it enters the stage of error backpropagation. Error backpropagation is to transmit the output error back layer by layer through the hidden layer in a certain form and allocate the error to all units in each layer, so as to obtain the error signals of each layer of units. This error signal is used as the basis for correcting the weights of each unit. This process of adjusting the weights of each layer during the forward propagation of signals and the backpropagation of errors is carried out cyclically. The process of continuously adjusting the weights is also the learning and training process of the network. This process continues until the error of the network output is reduced to an acceptable level or until a pre-set number of learning times is reached.

[0064] Generally speaking, according to the different types of problems to be solved, the selected AI algorithms and the adopted AI units are also different. In the related technologies, the main method to improve the 5G network performance by means of AI is to enhance or replace the existing algorithms or processing modules through algorithms and AI units based on neural networks. In specific scenarios, algorithms and AI units based on neural networks can achieve better performance than those based on deterministic algorithms. Commonly used neural networks include deep neural networks, convolutional neural networks, and recurrent neural networks, etc. With the help of existing AI tools, the construction, training, and verification of neural networks can be realized.

[0065] In summary, replacing non-AI processing modules in the communication system by means of AI or machine learning (ML) methods can effectively improve the system performance. For example, taking the prediction of channel state information (CSI) as an example, historical CSI can be input into the AI unit, and the AI unit is used to analyze the time-domain change characteristics of the channel, and future CSI is inferred based on this. The CSI prediction based on AI will have a very large performance gain compared with the scheme without prediction. At the same time, the different future moments for prediction will result in different achievable prediction accuracies.

[0066] However, when the device for training the AI unit (i.e., the second device) and the device for performing inference using the AI unit (i.e., the first device) are not the same device, it is necessary to transfer the AI unit. Among them, the parameters of the AI unit obtained by training can be floating-point parameters, while fixed-point parameters of the AI unit need to be obtained when using the AI unit for inference. In the related art, taking the network sending the AI unit to the terminal as an example, the network sends the floating-point parameters of the AI unit to the terminal, and the terminal quantifies the floating-point parameters to obtain the fixed-point parameters of the AI unit that can be used. However, in this process, the terminal may perform unnecessary quantization on the floating-point parameters of the AI unit, resulting in an increase in quantization complexity, or a loss of inference accuracy of the AI unit due to inappropriate quantization, thereby reducing the overall efficiency of the communication system based on the AI unit.

[0067] In the embodiments of the present application, it is possible to achieve quantization alignment of the parameter information of the AI unit to be transferred between the first device and the second device, which can reduce the quantization complexity and the loss of inference accuracy of the quantized AI unit, thereby improving the overall efficiency of the communication system based on the AI unit.

[0068] It should be noted that the AI unit in the embodiments of the present application can be used in the above CSI prediction scenario, or in other scenarios, such as: AI-based CSI compression feedback, beam prediction, positioning enhancement, intelligent network selection, load balancing, energy saving, etc., which are not enumerated here.

[0069] To facilitate the understanding of the quantization alignment method of the AI unit provided in the embodiments of the present application, the following terms in the embodiments of the present application are first explained:

[0070] 1) AI unit: The AI unit in the embodiments of the present application can also be referred to as an AI model, an AI structure, a machine learning model, a machine learning unit, a neural network, etc. Or the AI unit can also refer to a processing unit that can implement AI-related algorithms, formulas, processing flows, capabilities, etc. Or the AI unit can be a processing method, algorithm, function, module or unit for a specific data set. Or the AI unit can be a processing method, algorithm, function, module or unit running on AI-related hardware such as a Graphic Processing Unit (GPU), a Natural Processing Unit (NPU), a Tensor Processing Unit (TPU), an Application-Specific Integrated Circuit (ASIC), etc. The present application does not make specific limitations on this. Optionally, the specific data set includes the input and / or output of the AI unit.

[0071] The identification of the AI unit, such as the AI model identification, AI structure identification, AI algorithm identification, functionality ID, physical identification, logical identification, global identification, local identification, or the identification of a specific dataset associated with the AI unit, or the identification of a specific scenario, environment, channel characteristic, device related to the AI, or the identification of a function, feature, capability, or module related to the AI, etc., is not specifically limited in the embodiments of this application.

[0072] 2) First device: It can be a device that sends the AI unit. For example, the first device may include a terminal or a network-side device that trains the AI unit. Among them, the network-side device may include at least one of an access network device and a core network device.

[0073] 3) Second device: It can be a device that receives the AI unit and uses the AI unit for inference. For example, the first device may include a terminal or a network-side device, and the first device and the second device are not the same device.

[0074] Among them, the combination of the above first device and second device may include any one of the combinations shown in Table 1 below:

[0075] Table 1

[0076]

[0077] It should be noted that in the embodiments of this application, for the convenience of description, usually the first device is a terminal and the second device is a network-side device as an example for illustration, which does not constitute a specific limitation here.

[0078] 4) Quantization: It refers to the process of converting a floating-point value into a fixed-point value.

[0079] Among them, common floating-point values include floating-point numbers in formats such as single-precision floating-point type (abbreviation: float) and double-precision floating-point type (abbreviation: double), which contain decimal places. Common fixed-point values include fixed-point numbers in formats such as integer (int), such as unsigned integer (unsigned int), and also short integer (short int), long integer (long int), unsigned short integer (unsigned short int), unsigned long integer (unsigned long int), etc.

[0080] In the embodiments of this application, intx represents an integer represented by x bits (bit). For example, int4 represents -16 to 15. Commonly seen are also int8, int16, int32, int64, etc.

[0081] 5) AI unit transfer: Taking the transfer of an AI model as an example, it includes transmitting the model structure (i.e., structure information) or model parameters (i.e., parameter information). Among them, for a neural network model, the model parameters at least refer to at least some coefficients on the neurons (multiplicative coefficients, additive coefficients, parameters of the activation function).

[0082] Next, in conjunction with the accompanying drawings, through some embodiments and their application scenarios, the quantization alignment method of the AI unit provided by the embodiments of the present application, the quantization alignment device of the AI unit, and related devices will be described in detail.

[0083] Please refer to Figure 4 , a quantization alignment method of an AI unit provided by an embodiment of the present application, the execution subject of which can be a first device, as Figure 4 shown, the quantization alignment method of the AI unit may include the following steps:

[0084] Step 401, the first device receives the first information or the second information from the second device.

[0085] Step 402, the first device performs a first operation, and the first operation includes:

[0086] Determine the first AI unit according to the first information; or,

[0087] Quantize the second information according to the target quantization information to obtain the first AI unit;

[0088] Among them, the first information includes the parameter information of the first AI unit, and the first information is fixed-point number information quantized based on the target quantization information; the second information includes the parameter information of the first AI unit, and the second information is floating-point number information.

[0089] In some embodiments, the above parameter information of the first AI unit may include any coefficients in the first AI unit, such as multiplicative coefficients, additive coefficients, parameters of the activation function, etc., which are not specifically limited herein.

[0090] In some embodiments, the above first information or second information may further include the structure information of the first AI unit, such as the model structure of a neural network.

[0091] In some embodiments, the first device receives the second information from the second device, which may be that the second device sends the parameter information of the first AI unit before quantization to the first device, and the first device quantizes the parameter information of the first AI unit to obtain the first AI unit that can be used for inference.

[0092] Among them, the target quantization information used by the first device for quantizing the parameter information of the first AI unit is consistent with the second device.

[0093] Optionally, in order to make the first device and the second device agree on the target quantization information, at least one of the following methods can be adopted:

[0094] Method 1: Agree on the target quantization information through a protocol;

[0095] Method 2: The first device reports the target quantization information to the second device;

[0096] Method 3: The second device indicates the target quantization information to the first device.

[0097] Of course, at least two of the above Methods 1 to 3 can also be combined to make the first device and the second device agree on the target quantization information.

[0098] For example: Agree on a set of candidate quantization information through a protocol. The first device selects a subset of quantization information from the set of quantization information agreed upon by the protocol according to its own capability information and reports it to the second device. The second device determines the final target quantization information from the subset of quantization information reported by the first device.

[0099] Another example: The second device indicates the first quantization information to the first device. When the first device does not support the first quantization information or there is better quantization information, the first device sends the target quantization information determined by the first device to the second device.

[0100] Another example: Agree on a set of candidate quantization information through a protocol. The second device selects the target quantization information from the set of quantization information agreed upon by the protocol according to its own capability information and indicates the target quantization information to the first device.

[0101] In some embodiments, the first device receiving the first information from the second device may be that the second device quantizes the parameter information of the first AI unit and sends the quantized parameter information of the first AI unit to the first device.

[0102] Among them, the target quantization information used by the second device for quantizing the parameter information of the first AI unit is consistent with the first device.

[0103] Among them, the specific implementation manners for the first device and the second device to agree on the target quantization information include at least one of Methods 1 to 3 in the previous embodiment, which will not be elaborated here.

[0104] In one implementation manner of the embodiment of the present application, the second device can directly transmit the parameter information of the first AI unit after quantization to the first device, so that the first device and the second device align the quantization of the parameters of the first AI unit; in another implementation manner of the embodiment of the present application, the second device transmits the parameter information of the first AI unit before quantization to the first device. At this time, the first device can quantize the parameter information of the first AI unit based on the target quantization information agreed upon with the second device, so as to achieve the alignment of the quantization of the parameters of the first AI unit between the first device and the second device. By aligning the quantization of the parameters of the first AI unit between the first device and the second device, the quantization complexity and the loss of the inference accuracy of the AI unit after quantization can be reduced, thereby improving the overall efficiency of the communication system based on the AI unit.

[0105] It is worth noting that for the method of quantizing the parameter information of the first AI unit by the first device, the parameter information of the first AI unit transmitted by the second device to the first device is floating-point information, which can be used for various quantization methods and quantization levels, and the transmission flexibility is very high. At the same time, by making the first device and the second device agree on the target quantization information, the second device can determine that the first AI unit it transmits is adapted to the target quantization information. For example, by judging that the target quantization information can be used for the quantization of at least some parameter values in the first AI unit and ensuring that the quantization loss of at least some parameter values in the first AI unit after quantization is small, it is determined that the first AI unit is adapted to the target quantization information. In this way, the first device can use the target quantization information adapted to the first AI unit to quantize the parameter information of the first AI unit, which can reduce the quantization complexity of the parameter information of the first AI unit and improve the quantization accuracy.

[0106] For the method of quantizing the parameter information of the first AI unit by the second device, it is also possible to make the second device determine that the first AI unit it transmits is adapted to the target quantization information, so as to improve the quantization accuracy of the parameter information of the first AI unit. At the same time, by making the first device and the second device agree on the target quantization information, the second device can also use the target quantization information adapted to the first AI unit to quantize the parameter information of the first AI unit, which can reduce the quantization complexity of the parameter information of the first AI unit and improve the quantization accuracy.

[0107] In some implementation manners, the target quantization information is used to indicate at least one of a quantization method and a quantization level.

[0108] Optionally, the quantization method includes at least one of the following:

[0109] The direct quantization method, which is used to quantize each parameter of the AI unit;

[0110] Group quantization method, which is used to quantize each group of parameters of the AI unit, where a group of parameters includes at least one parameter, and the parameters within the same group correspond to the same quantization value;

[0111] Transform domain quantization method, which is used to convert the floating-point parameters of the AI unit into transform domain parameters, and perform quantization processing and inverse transform processing on the transform domain parameters to obtain the quantized fixed-point parameters of the AI unit;

[0112] Product quantization method, which is used to quantize the first subspace, and the first subspace is a subspace composed of at least one parameter of the AI unit.

[0113] In some embodiments, direct quantization can be further divided into uniform quantization and non-uniform quantization:

[0114] 1) Uniform quantization: It means dividing the input value range at equal intervals and quantizing each divided value range separately.

[0115] 2) Non-uniform quantization: It means quantization with unequal quantization intervals within the dynamic range of the input. For example: According to the probability density, probability distribution, cumulative probability distribution, etc. of the input, determine the quantization intervals / quantization levels of different intervals of the input. For example, for the interval with a small input value, its quantization interval is also small; on the contrary, the quantization interval is large.

[0116] In some other embodiments, the above-mentioned group quantization method can also be called weight sharing quantization, that is, dividing the parameters in the AI unit into multiple sets, and the elements in each set share a quantization value.

[0117] For example: Using the K-means clustering partition method, divide the data into K groups, and randomly select K objects from the parameters of the AI unit as the initial clustering centers, then calculate the distance between each object and each seed clustering center, and assign each object to the clustering center closest to it. The clustering centers and the objects assigned to them represent a cluster. Each time a sample is assigned, the clustering center of the cluster will be recalculated according to the existing objects in the cluster. This process will be repeated continuously until a certain termination condition is met. The termination condition can be that no (or the minimum number) of objects are reassigned to different clusters, or, no (or the minimum number) of clustering centers change anymore, or, the sum of squared errors is locally minimized.

[0118] In some other embodiments, the above-mentioned transform domain can include the frequency domain corresponding to the Fourier transform, the S domain corresponding to the Laplace transform, and the Z domain corresponding to the Z transform, where the Z transform can also be called the Fourier transform of the discrete function * attenuation function.

[0119] For example: transform the floating-point parameters (such as weights, biases, convolution kernels, etc.) of the AI unit to the frequency domain, then quantize the floating-point parameters of the AI unit in the frequency domain, and finally perform an inverse frequency-domain transformation on the quantized parameters to obtain the fixed-point parameters of the quantized AI unit.

[0120] In some other embodiments, Product Quantization may be to divide the weights of the AI unit into multiple subspaces and perform quantization operations on each subspace respectively, such as performing weight-sharing quantization method on each subspace, etc., and finally obtaining the fixed-point parameters of the quantized AI unit.

[0121] Optionally, the quantization levels are divided based on at least one of the following:

[0122] The number of bits of the fixed-point number obtained after quantizing a floating-point number;

[0123] The difference between two adjacent fixed-point numbers after quantization, or the floating-point number difference represented by the least significant bit after quantization.

[0124] In one embodiment, the more bits of the fixed-point number obtained after quantizing a floating-point number, the higher the corresponding quantization level, and at this time, the quantization accuracy is also higher.

[0125] In another embodiment, the smaller the difference between two adjacent fixed-point numbers after quantization, or the smaller the floating-point number difference represented by the least significant bit after quantization, the higher the corresponding quantization level, and at this time, the quantization accuracy is also higher.

[0126] For example: for uniform quantization, quantize the floating-point number between -1 and 1 to 4 bits, then 1 bit of the least significant bit represents 2 / 16 of the floating-point number, or the difference between adjacent fixed-point numbers is 2 / 16.

[0127] In this embodiment, the quantization method can indicate what kind of quantization or specific quantization parameters are performed on the parameter information of the first AI unit; the quantization level can indicate the degree of quantization of the parameter information of the first AI unit.

[0128] As an optional embodiment, the method further includes at least one of the following:

[0129] The first device determines the target quantization information according to the first quantization information from the second device;

[0130] The first device determines the target quantization information according to the second quantization information agreed upon by the protocol.

[0131] In some embodiments, when the first device determines the target quantization information based on the first quantization information from the second device, the target quantization information may be the same as the first quantization information. In this case, the first device determines the first quantization information indicated by the second device as the target quantization information.

[0132] In other embodiments, when the first device determines the target quantization information based on the first quantization information from the second device, the target quantization information may be a subset of the first quantization information. For example, the UE selects the quantization information it supports from the first quantization information indicated by the NW as the target quantization information.

[0133] In still other embodiments, when the first device determines the target quantization information based on the first quantization information from the second device, the target quantization information may be different from the first quantization information. For example, after the NW indicates the first quantization information to the UE, if the UE does not support the first quantization information or determines that there is other quantization information better than the first quantization information, the UE may determine target quantization information different from the first quantization information.

[0134] Optionally, the method further includes:

[0135] When the first quantization information is not exactly the same as the target quantization information, the first device sends the target quantization information to the second device.

[0136] For example: After the UE selects the quantization information it supports from the first quantization information indicated by the NW as the target quantization information, the UE reports the target quantization information to the NW.

[0137] For another example: After the UE determines target quantization information different from the first quantization information indicated by the NW, the UE reports the target quantization information to the NW.

[0138] It should be noted that after receiving the target quantization information sent by the first device, the second device may quantize the parameter information of the first AI unit according to the target quantization information to obtain the first information. In addition, it may also perform an AI unit transfer process adapted to the target quantization information. In this way, the second device can perform the quantization process or the AI unit transfer process according to the quantization of the AI unit aligned with the first device.

[0139] In some embodiments, the first device determines the target quantization information based on the first quantization information from the second device, including:

[0140] Before receiving the first information or the second information from the second device, the first device receives the first quantization information from the second device;

[0141] The first device determines the target quantization information according to the first quantization information.

[0142] It should be noted that, in one implementation, before receiving the second information, the first device receives the first quantization information and determines the target quantization information for quantifying the second information accordingly. In another implementation, the first device may receive the first quantization information before receiving the first information and feedback the target quantization information to the second device, so that the second device uses the target quantization information agreed with the first device to quantify the floating-point parameter information of the first AI unit. After obtaining the first information, the second device then sends the first information to the first device. At this time, when the first device receives the first information, it can know how the first information is quantified based on the target quantization information, so as to execute the operation process of the first AI unit according to the quantization method or quantization level of the first information.

[0143] In this implementation, quantization alignment can be performed before the transfer of the AI unit.

[0144] Through the above implementation, the first device and the second device interact with each other on the quantization information and reach an agreement on the target quantization information.

[0145] In addition, in some implementations, the first device and the second device can obtain the consistent target quantization information through protocol agreement. Or, the protocol stipulates multiple quantization information, and through the interaction between the first device and the second device, the target quantization information is determined from the multiple quantization information stipulated by the protocol.

[0146] As an alternative implementation, the method further includes:

[0147] The first device sends first capability information to the second device, and the first capability information is used to indicate the third quantization information supported by the first device, where the third quantization information includes the target quantization information.

[0148] In some implementations, after receiving the first capability information, the second device can determine the target quantization information accordingly.

[0149] Optionally, in the case where the first device quantifies the parameter information of the first AI unit, the first device may also receive the target quantization information from the second device.

[0150] In this implementation, the first device sending the first capability information to the second device enables the second device to determine the target capability information supported by the first device accordingly.

[0151] To improve the quantization accuracy, the parameter information of the first AI unit to be quantized can be restricted so that the quantization loss caused by the parameters that meet the restriction requirements during quantization processing based on the target quantization information is small.

[0152] In some alternative embodiments, the first AI unit includes a first parameter set, and the floating-point values of the first parameter set satisfy at least one of the following requirements:

[0153] The maximum amplitude of the parameters in the first parameter set is less than or equal to a first threshold value;

[0154] The maximum amplitude of the parameters in the first parameter set is greater than or equal to a second threshold value;

[0155] The minimum amplitude of the parameters in the first parameter set is less than or equal to a third threshold value;

[0156] The minimum amplitude of the parameters in the first parameter set is greater than or equal to a fourth threshold value;

[0157] The average amplitude of the parameters in the first parameter set is less than or equal to a fifth threshold value;

[0158] The average amplitude of the parameters in the first parameter set is greater than or equal to a sixth threshold value;

[0159] The difference between the first value and the second value is less than or equal to a seventh threshold value;

[0160] The ratio of the first value to the second value is less than or equal to an eighth threshold value;

[0161] The difference between the third value and the fourth value is less than or equal to a ninth threshold value;

[0162] The ratio of the third value to the fourth value is less than or equal to a tenth threshold value;

[0163] The difference between the fifth value and the sixth value is less than or equal to an eleventh threshold value;

[0164] The ratio of the fifth value to the sixth value is less than or equal to a twelfth threshold value;

[0165] The difference between the maximum amplitude of the parameters in the first parameter set and the minimum amplitude of the parameters in the first parameter set is less than or equal to a thirteenth threshold value;

[0166] The ratio of the maximum amplitude of the parameters in the first parameter set to the minimum amplitude of the parameters in the first parameter set is less than or equal to a fourteenth threshold value;

[0167] The difference between the maximum amplitude of the parameters in the first parameter set and the average amplitude of the parameters in the first parameter set is less than or equal to a fifteenth threshold value;

[0168] The ratio of the maximum amplitude of the parameters in the first parameter set to the average amplitude of the parameters in the first parameter set is less than or equal to the sixteenth threshold value;

[0169] The difference between the average amplitude of the parameters in the first parameter set and the minimum amplitude of the parameters in the first parameter set is less than or equal to the seventeenth threshold value;

[0170] The ratio of the average amplitude of the parameters in the first parameter set to the minimum amplitude of the parameters in the first parameter set is less than or equal to the eighteenth threshold value;

[0171] Wherein, the first parameter set includes at least some parameters of the first AI unit; the first AI unit includes a first layer and a second layer; the first value is the maximum amplitude of the parameters in the first layer; the second value is the maximum amplitude of the parameters in the second layer; the third value is the minimum amplitude of the parameters in the first layer; the fourth value is the minimum amplitude of the parameters in the second layer; the fifth value is the average amplitude of the parameters in the first layer; the sixth value is the average amplitude of the parameters in the second layer; and the first value is greater than or equal to the second value, the third value is greater than or equal to the fourth value, and the fifth value is greater than or equal to the sixth value.

[0172] In some embodiments, the first device may determine whether the first AI unit meets the requirements for the floating-point values of the first parameter set, and when the determination result is that the first AI unit meets the requirements for the floating-point values of the first parameter set, it is determined that the parameter information of the first AI unit can be quantized, or it is determined that the quantized first AI unit is available.

[0173] In other embodiments, the second device may determine whether the first AI unit meets the requirements for the floating-point values of the first parameter set, and when the determination result is that the first AI unit meets the requirements for the floating-point values of the first parameter set, the first AI unit is transmitted.

[0174] In some embodiments, the first layer is at least one layer of the first AI unit, and the second layer is at least one other layer of the first AI unit different from the first layer.

[0175] In some embodiments, the average amplitude in this embodiment may be a value obtained by performing at least one operation such as linear averaging, geometric averaging, harmonic averaging, quadratic averaging, weighted averaging, minimum value maximization, maximum value minimization, etc. on the amplitude and its simple variations, and no specific limitation is made here.

[0176] In this embodiment, by limiting the maximum amplitude, minimum amplitude, average amplitude of at least some of the floating-point parameters in the first AI unit, the differences in the maximum amplitudes of the parameters in different layers, the differences in the minimum amplitudes of the parameters in different layers, the differences in the average amplitudes of the parameters in different layers, the difference or ratio between the maximum amplitude and the minimum amplitude, the difference or ratio between the maximum amplitude and the average amplitude, or the difference or ratio between the average amplitude and the minimum amplitude, the quantization range of at least some of the floating-point parameters in the first AI unit can be restricted, and thus the quantization accuracy of these parameters can be ensured to meet the requirements.

[0177] It should be noted that at least one of the above first threshold, second threshold, third threshold, fourth threshold, fifth threshold, sixth threshold, seventh threshold, eighth threshold, ninth threshold, tenth threshold, eleventh threshold, twelfth threshold, thirteenth threshold, fourteenth threshold, fifteenth threshold, sixteenth threshold, seventeenth threshold, eighteenth threshold, nineteenth threshold, twentieth threshold, twenty-first threshold, and twenty-second threshold can be agreed upon in the protocol, or indicated by the network side, or reported by the terminal, and no specific limitation is made here.

[0178] In some embodiments, at least one of the above first threshold, second threshold, third threshold, fourth threshold, fifth threshold, sixth threshold, seventh threshold, eighth threshold, ninth threshold, tenth threshold, eleventh threshold, twelfth threshold, thirteenth threshold, fourteenth threshold, fifteenth threshold, sixteenth threshold, seventeenth threshold, eighteenth threshold, nineteenth threshold, twentieth threshold, twenty-first threshold, and twenty-second threshold can be associated with the quantization level, that is, different quantization levels can be associated with their respective target thresholds.

[0179] For example: the target threshold is related to the target difference, and the target difference is the difference between two adjacent fixed-point numbers after quantization or the floating-point number difference represented by the lowest bit after quantization;

[0180] Among them, the target threshold value includes at least one of the first threshold value, the second threshold value, the third threshold value, the fourth threshold value, the fifth threshold value, the sixth threshold value, the seventh threshold value, the eighth threshold value, the ninth threshold value, the tenth threshold value, the eleventh threshold value, the twelfth threshold value, the thirteenth threshold value, the fourteenth threshold value, the fifteenth threshold value, the sixteenth threshold value, the seventeenth threshold value, the eighteenth threshold value, the nineteenth threshold value, the twentieth threshold value, the twenty - first threshold value, and the twenty - second threshold value.

[0181] It should be noted that the larger the difference between two adjacent fixed - point numbers after quantization or the difference between floating - point numbers represented by the lowest bit after quantization, the lower the quantization level. At this time, the limitation of the target threshold value can be appropriately relaxed. At this time, compared with the target threshold value of a higher quantization level, the accuracy of the parameters that meet the target threshold value of a lower quantization level will decrease after quantization.

[0182] In this embodiment, the target threshold value can be adjusted correlatively according to different quantization levels.

[0183] It is worth noting that the amplitude value in the embodiments of this application can be the value of the amplitude, or it may be the value obtained after performing certain mathematical operations on the amplitude. For example: the square of the amplitude or the square root of the amplitude, etc.

[0184] Optionally, the maximum amplitude is the minimum value among the largest K1 amplitudes, where K1 is a positive integer.

[0185] In this way, the largest K1 amplitudes can be uniformly adjusted to the minimum value among them. For example: assuming that the largest 3 amplitudes in the first parameter set are A1, A2, and A3 in descending order, then A1 and A2 can be adjusted to A3. In this way, by sacrificing a very small number of the largest amplitudes, the amplitude range can be reduced, that is, the amplitude range of the quantization process can be reduced, and the quantization accuracy can be improved.

[0186] Optionally, the minimum amplitude is the maximum value among the smallest K2 amplitudes, where K2 is a positive integer.

[0187] In this way, similar to the above where the maximum amplitude is the minimum value among the largest K1 amplitudes, by uniformly adjusting the smallest K2 amplitudes to the maximum value among them, by sacrificing a very small number of the smallest amplitudes, the amplitude range can be reduced, that is, the amplitude range of the quantization process can be reduced, and the quantization accuracy can be improved.

[0188] It should be noted that the above - mentioned maximum amplitude may include at least one of the following: the maximum amplitude of the parameters in the first parameter set, the maximum amplitude of the parameters in the first layer, and the maximum amplitude of the parameters in the second layer.

[0189] In addition, the above minimum amplitude may include at least one of the following: the minimum amplitude of the parameters in the first parameter set, the minimum amplitude of the parameters in the first layer, and the minimum amplitude of the parameters in the second layer.

[0190] It is worth noting that the floating-point value of the first parameter set may be the floating-point parameter value of the first AI unit before quantization; or, it may also be a floating-point value obtained by dequantizing the fixed-point parameter value of the first AI unit after quantization.

[0191] In addition to restricting the floating-point value of the first parameter set, the fixed-point value of the first parameter set may also be restricted.

[0192] In some embodiments, the first AI unit includes a first parameter set, and the fixed-point value of the first parameter set after quantization satisfies at least one of the following requirements:

[0193] The maximum value is less than or equal to the nineteenth threshold;

[0194] The maximum value is greater than or equal to the twentieth threshold;

[0195] The minimum value is less than or equal to the twenty-first threshold;

[0196] The minimum value is greater than or equal to the twenty-second threshold;

[0197] Wherein, the first parameter set includes at least some parameters of the first AI unit.

[0198] In this embodiment, by limiting the maximum and minimum values of at least some of the fixed-point parameters in the first AI unit, the quantization range of at least some of the parameters in the first AI unit can be restricted, and thus the quantization accuracy of these parameters can be ensured to meet the requirements.

[0199] It is worth noting that the fixed-point value of the first parameter set may be the fixed-point parameter value of the first AI unit after quantization; or, it may also be a fixed-point value obtained by quantizing the floating-point parameter value of the first AI unit before quantization.

[0200] Optionally, the first parameter set in the embodiments of the present application includes at least one of the following:

[0201] At least some parameters of the first AI unit indicated by the second device, or parameters in at least some structures of the first AI unit;

[0202] Parameters in at least some structures of the first AI unit reported by the first device;

[0203] Parameters of the adaptive layer of the first AI unit;

[0204] The parameters of the input layer of the first AI unit;

[0205] The parameters of the hidden layer of the first AI unit;

[0206] The parameters of the output layer of the first AI unit.

[0207] In some embodiments, the model parameters (such as neuron coefficients) of the above-mentioned adaptive layer can be adjusted, or the model parameters and structural parameters of the adaptive layer can be adjusted.

[0208] In this embodiment, value range limits can be set for the parameters specified by the second device, or the parameters reported by the first device, or the parameters of the specified layer, so as to improve the quantization accuracy of the parameter information of the first AI unit.

[0209] Please refer to Figure 5 , a quantization alignment method for an AI unit provided by an embodiment of the present application, the execution subject of which can be a second device, such as Figure 5 shown, the quantization alignment method of the AI unit may include the following steps:

[0210] Step 501, the second device performs a second operation, where the second operation includes:

[0211] Quantize the parameter information of the first AI unit according to the target quantization information to obtain first information, where the first information is quantized fixed-point number information; send the first information to the first device; or,

[0212] Send second information to the first device, where the second information includes the parameter information of the first AI unit, and the second information is floating-point number information, and the second information is used to be quantized according to the target quantization information to obtain the first AI unit.

[0213] In some embodiments, after the second device quantizes the parameter information of the first AI unit according to the target quantization information, it sends the quantized first information to the first device, so that the first device directly uses the first AI unit for inference according to the received first information.

[0214] In other embodiments, the second device directly sends the second information before quantization to the first device, so that the first device quantizes the second information according to the target quantization information to obtain a usable first AI unit.

[0215] It should be noted that before the second device performs the second operation, the target quantization information can be aligned with the first device, such as through at least one of the ways of protocol agreement, second device indication or first device reporting, so that the first device and the second device reach an agreement on the target quantization information.

[0216] In addition, the first information, the second information, the first AI unit, the parameter information, and the target quantization information in the embodiments of the present application have the same meanings and functions as the first information, the second information, the first AI unit, the parameter information, and the target quantization information in the embodiments of the first device side method, and will not be elaborated here.

[0217] The embodiments of the present application correspond to the embodiments of the first device side method. In one implementation, the second device can quantize the parameter information of the first AI unit according to the target quantization information agreed with the first device and directly send the quantized fixed-point parameters to the first device. In another implementation, the second device can send floating-point parameters matching the target quantization information to the first device according to the target quantization information agreed with the first device. The methods on the first device side and the second device side are combined to jointly achieve the beneficial effects of reducing the quantization complexity and the accuracy loss of the quantized AI unit, and improving the overall efficiency of the communication system based on the AI unit.

[0218] In some embodiments, the method further includes:

[0219] The second device receives first capability information from the first device;

[0220] The second device determines the target quantization information according to the first capability information;

[0221] Wherein, the first capability information is used to indicate third quantization information supported by the first device, and the third quantization information includes the target quantization information.

[0222] In some embodiments, the third quantization information is a kind of quantization information. In this case, the target quantization information may be the same as the third quantization information.

[0223] In other embodiments, the third quantization information includes at least two kinds of quantization information. In this case, the target quantization information may be one kind of quantization information in the third quantization information.

[0224] In this embodiment, the first device reports the third quantization information it supports to the second device, so that the second device can determine the target quantization information from the third quantization information supported by the first device's capabilities.

[0225] In some embodiments, the method further includes:

[0226] The second device determines the target quantization information according to the second quantization information agreed by the protocol.

[0227] Optionally, the method further includes:

[0228] The second device sends the target quantization information to the first device.

[0229] In this embodiment, when the first device performs quantization processing on the parameter information of the first AI unit, after determining the target quantization information, the second device may send the target quantization information to the first device, so that the first device performs quantization processing on the parameter information of the first AI unit according to the target quantization information.

[0230] In some embodiments, the method further includes:

[0231] The second device sends first quantization information to the first device, and the first quantization information is used to assist in determining the target quantization information.

[0232] In a possible embodiment, after the second device sends the first quantization information to the first device, the first device may determine the target quantization information from the first quantization information.

[0233] In another possible embodiment, after the second device sends the first quantization information to the first device, if the first device does not support the first quantization information or finds other quantization information that is more suitable for the first AI unit, the target quantization information determined by the first device may not be included in the first quantization information.

[0234] In this embodiment, the second device may indicate the first quantization information to the first device. It should be noted that the first quantization information may be accepted or rejected by the first device.

[0235] Optionally, the second device sending the first quantization information to the first device includes:

[0236] Before sending the first information or the second information to the first device, the second device sends the first quantization information to the first device.

[0237] In this embodiment, quantization information alignment may be performed before transmitting the AI unit.

[0238] Optionally, when the second device sends the first quantization information or the target quantization information to the first device, the first quantization information or the target quantization information may be transmitted together with the structure information of the first AI unit. For example: the first quantization information or the target quantization information is used to be transmitted together with the model structure information before transmitting the model neuron coefficients.

[0239] Of course, the above first quantization information or the target quantization information may also be transmitted separately. For example: before transmitting the model neuron coefficients and the model structure information, the above first quantization information or the target quantization information is transmitted separately.

[0240] In some embodiments, the method further includes:

[0241] The second device receives the target quantization information from the first device.

[0242] For example: If the target quantization information determined by the first device is not included in the first quantization information, the first device sends the target quantization information to the second device. Alternatively, the first device can determine the target quantization information by itself and report the target quantization information to the second device.

[0243] In this embodiment, the first device can report the target quantization information to the second device.

[0244] In some embodiments, the target quantization information is used to indicate at least one of a quantization method and a quantization level.

[0245] Optionally, the quantization method includes at least one of the following:

[0246] Direct quantization method, which is used to quantize each parameter of the AI unit;

[0247] Group quantization method, which is used to quantize each group of parameters of the AI unit, where a group of parameters includes at least one parameter, and the parameters in the same group correspond to the same quantization value; Transform domain quantization method, which is used to convert the floating-point parameters of the AI unit into transform domain parameters, and perform quantization processing and inverse transform processing on the transform domain parameters to obtain the quantized fixed-point parameters of the AI unit;

[0248] Product quantization method, which is used to quantize the first subspace, and the first subspace is a subspace composed of at least one parameter of the AI unit.

[0249] Optionally, the quantization level is divided based on at least one of the following:

[0250] The number of bits of the fixed-point number obtained after quantizing a floating-point number;

[0251] Target difference, where the target difference is the difference between two adjacent fixed-point numbers after quantization or the floating-point number difference represented by the lowest bit after quantization.

[0252] In some embodiments, the first AI unit includes a first parameter set, and the floating-point values of the first parameter set satisfy at least one of the following requirements:

[0253] The maximum amplitude of the parameters in the first parameter set is less than or equal to the first threshold;

[0254] The maximum amplitude of the parameters in the first parameter set is greater than or equal to the second threshold value;

[0255] The minimum amplitude of the parameters in the first parameter set is less than or equal to the third threshold value;

[0256] The minimum amplitude of the parameters in the first parameter set is greater than or equal to the fourth threshold value;

[0257] The average amplitude of the parameters in the first parameter set is less than or equal to the fifth threshold value;

[0258] The average amplitude of the parameters in the first parameter set is greater than or equal to the sixth threshold value;

[0259] The difference between the first value and the second value is less than or equal to the seventh threshold value;

[0260] The ratio of the first value to the second value is less than or equal to the eighth threshold value;

[0261] The difference between the third value and the fourth value is less than or equal to the ninth threshold value;

[0262] The ratio of the third value to the fourth value is less than or equal to the tenth threshold value;

[0263] The difference between the fifth value and the sixth value is less than or equal to the eleventh threshold value;

[0264] The ratio of the fifth value to the sixth value is less than or equal to the twelfth threshold value;

[0265] The difference between the maximum amplitude of the parameters in the first parameter set and the minimum amplitude of the parameters in the first parameter set is less than or equal to the thirteenth threshold value;

[0266] The ratio of the maximum amplitude of the parameters in the first parameter set to the minimum amplitude of the parameters in the first parameter set is less than or equal to the fourteenth threshold value;

[0267] The difference between the maximum amplitude of the parameters in the first parameter set and the average amplitude of the parameters in the first parameter set is less than or equal to the fifteenth threshold value;

[0268] The ratio of the maximum amplitude of the parameters in the first parameter set to the average amplitude of the parameters in the first parameter set is less than or equal to the sixteenth threshold value;

[0269] The difference between the average amplitude of the parameters in the first parameter set and the minimum amplitude of the parameters in the first parameter set is less than or equal to the seventeenth threshold value;

[0270] The ratio of the average amplitude of the parameters in the first parameter set to the minimum amplitude of the parameters in the first parameter set is less than or equal to the eighteenth threshold value;

[0271] Among them, the first parameter set includes at least some parameters of the first AI unit; the first AI unit includes a first layer and a second layer; the first value is the maximum amplitude of the parameters in the first layer; the second value is the maximum amplitude of the parameters in the second layer; the third value is the minimum amplitude of the parameters in the first layer; the fourth value is the minimum amplitude of the parameters in the second layer; the fifth value is the average amplitude of the parameters in the first layer; the sixth value is the average amplitude of the parameters in the second layer; and the first value is greater than or equal to the second value, the third value is greater than or equal to the fourth value, and the fifth value is greater than or equal to the sixth value.

[0272] In some embodiments, the maximum amplitude is the minimum value among the largest K1 amplitudes, where K1 is a positive integer.

[0273] In some embodiments, the minimum amplitude is the maximum value among the smallest K2 amplitudes, where K2 is a positive integer.

[0274] In some embodiments, the first AI unit includes a first parameter set, and the fixed-point values after quantization of the first parameter set satisfy at least one of the following requirements:

[0275] The maximum value is less than or equal to the nineteenth threshold;

[0276] The maximum value is greater than or equal to the twentieth threshold;

[0277] The minimum value is less than or equal to the twenty-first threshold;

[0278] The minimum value is greater than or equal to the twenty-second threshold;

[0279] Among them, the first parameter set includes at least some parameters of the first AI unit.

[0280] In some embodiments, the target threshold value is related to the target difference, and the target difference is the difference between two adjacent fixed-point numbers after quantization or the floating-point number difference represented by the lowest bit after quantization;

[0281] Among them, the target threshold value includes at least one of the first threshold value, the second threshold value, the third threshold value, the fourth threshold value, the fifth threshold value, the sixth threshold value, the seventh threshold value, the eighth threshold value, the ninth threshold value, the tenth threshold value, the eleventh threshold value, the twelfth threshold value, the thirteenth threshold value, the fourteenth threshold value, the fifteenth threshold value, the sixteenth threshold value, the seventeenth threshold value, the eighteenth threshold value, the nineteenth threshold value, the twentieth threshold value, the twenty-first threshold value, and the twenty-second threshold value.

[0282] In some embodiments, the first parameter set includes at least one of the following:

[0283] At least some of the parameters of the first AI unit indicated by the second device, or the parameters in at least some of the structures of the first AI unit;

[0284] The parameters in at least some of the structures of the first AI unit reported by the first device;

[0285] The parameters of the adaptive layer of the first AI unit;

[0286] The parameters of the input layer of the first AI unit;

[0287] The parameters of the hidden layer of the first AI unit;

[0288] The parameters of the output layer of the first AI unit.

[0289] In the embodiments of the present application, the steps executed by the second device correspond to the steps executed by the first device in the method embodiments on the first device side, and the two cooperate with each other to jointly achieve the beneficial effects of reducing the quantization complexity and the loss of inference accuracy of the quantized AI unit, and improving the overall efficiency of the communication system based on the AI unit.

[0290] For the quantization alignment method of the AI unit provided in the embodiments of the present application, the execution subject may be a quantization alignment device of the AI unit. In the embodiments of the present application, taking the quantization alignment device of the AI unit executing the quantization alignment method of the AI unit as an example, the quantization alignment device of the AI unit provided in the embodiments of the present application is described.

[0291] Referring to Figure 6 , the embodiments of the present application further provide a quantization alignment device of an AI unit, which is applied to a first device. As Figure 6 shown, the quantization alignment device 600 of the AI unit includes:

[0292] A first receiving module 601, configured to receive the first information or the second information from the second device;

[0293] A first execution module 602, configured to execute a first operation, and the first operation includes:

[0294] Determine a first AI unit according to the first information; or,

[0295] Quantize the second information according to the target quantization information to obtain a first AI unit;

[0296] Among them, the first information includes parameter information of the first AI unit, and the first information is fixed-point number information quantized based on target quantization information; the second information includes parameter information of the first AI unit, and the second information is floating-point number information.

[0297] Optionally, the target quantization information is used to indicate at least one of a quantization method and a quantization level.

[0298] Optionally, the quantization alignment device 600 of the AI unit further includes at least one of the following:

[0299] The first determination module is configured to determine the target quantization information according to the first quantization information from the second device;

[0300] The second determination module is configured to determine the target quantization information according to the second quantization information agreed upon by the protocol.

[0301] Optionally, the quantization alignment device 600 of the AI unit further includes:

[0302] The first sending module is configured to send the target quantization information to the second device when the first quantization information is not exactly the same as the target quantization information.

[0303] Optionally, the first determination module includes:

[0304] The receiving unit is configured to receive the first quantization information from the second device before the first receiving module receives the first information or the second information from the second device;

[0305] The determination unit is configured to determine the target quantization information according to the first quantization information.

[0306] Optionally, the quantization alignment device 600 of the AI unit further includes:

[0307] The second sending module is configured to send first capability information to the second device, where the first capability information is used to indicate third quantization information supported by the first device, and the third quantization information includes the target quantization information.

[0308] Optionally, the quantization method includes at least one of the following:

[0309] The direct quantization method, which is used to quantize each parameter of the AI unit;

[0310] The grouped quantization method, which is used to quantize each group of parameters of the AI unit, where a group of parameters includes at least one parameter, and the parameters in the same group correspond to the same quantization value;

[0311] The transform domain quantization method is used to convert the floating-point parameters of the AI unit into transform domain parameters, and perform quantization processing and inverse transform processing on the transform domain parameters to obtain the fixed-point parameters after quantization of the AI unit;

[0312] The product quantization method is used to quantize the first subspace, and the first subspace is a subspace composed of at least one parameter of the AI unit.

[0313] Optionally, the quantization levels are divided based on at least one of the following:

[0314] The number of bits of the fixed-point number obtained after quantizing a floating-point number;

[0315] The difference between two adjacent fixed-point numbers after quantization, or the difference of the floating-point number represented by the lowest bit after quantization.

[0316] Optionally, the first AI unit includes a first parameter set, and the floating-point values of the first parameter set satisfy at least one of the following requirements:

[0317] The maximum amplitude of the parameters in the first parameter set is less than or equal to the first threshold value;

[0318] The maximum amplitude of the parameters in the first parameter set is greater than or equal to the second threshold value;

[0319] The minimum amplitude of the parameters in the first parameter set is less than or equal to the third threshold value;

[0320] The minimum amplitude of the parameters in the first parameter set is greater than or equal to the fourth threshold value;

[0321] The average amplitude of the parameters in the first parameter set is less than or equal to the fifth threshold value;

[0322] The average amplitude of the parameters in the first parameter set is greater than or equal to the sixth threshold value;

[0323] The difference between the first value and the second value is less than or equal to the seventh threshold value;

[0324] The ratio of the first value to the second value is less than or equal to the eighth threshold value;

[0325] The difference between the third value and the fourth value is less than or equal to the ninth threshold value;

[0326] The ratio of the third value to the fourth value is less than or equal to the tenth threshold value;

[0327] The difference between the fifth value and the sixth value is less than or equal to the eleventh threshold value;

[0328] The ratio of the fifth value to the sixth value is less than or equal to the twelfth threshold value;

[0329] The difference between the maximum amplitude of the parameters in the first parameter set and the minimum amplitude of the parameters in the first parameter set is less than or equal to the thirteenth threshold value;

[0330] The ratio of the maximum amplitude of the parameters in the first parameter set to the minimum amplitude of the parameters in the first parameter set is less than or equal to the fourteenth threshold value;

[0331] The difference between the maximum amplitude of the parameters in the first parameter set and the average amplitude of the parameters in the first parameter set is less than or equal to the fifteenth threshold value;

[0332] The ratio of the maximum amplitude of the parameters in the first parameter set to the average amplitude of the parameters in the first parameter set is less than or equal to the sixteenth threshold value;

[0333] The difference between the average amplitude of the parameters in the first parameter set and the minimum amplitude of the parameters in the first parameter set is less than or equal to the seventeenth threshold value;

[0334] The ratio of the average amplitude of the parameters in the first parameter set to the minimum amplitude of the parameters in the first parameter set is less than or equal to the eighteenth threshold value;

[0335] Wherein, the first parameter set includes at least some parameters of the first AI unit; the first AI unit includes a first layer and a second layer; the first value is the maximum amplitude of the parameters in the first layer; the second value is the maximum amplitude of the parameters in the second layer; the third value is the minimum amplitude of the parameters in the first layer; the fourth value is the minimum amplitude of the parameters in the second layer; the fifth value is the average amplitude of the parameters in the first layer; the sixth value is the average amplitude of the parameters in the second layer; and the first value is greater than or equal to the second value, the third value is greater than or equal to the fourth value, and the fifth value is greater than or equal to the sixth value.

[0336] Optionally, the maximum amplitude is the minimum value among the largest K1 amplitudes, where K1 is a positive integer.

[0337] Optionally, the minimum amplitude is the maximum value among the smallest K2 amplitudes, where K2 is a positive integer.

[0338] Optionally, the first AI unit includes a first parameter set, and the fixed-point values after quantization of the first parameter set satisfy at least one of the following requirements:

[0339] The maximum value is less than or equal to the nineteenth threshold;

[0340] The maximum value is greater than or equal to the twentieth threshold;

[0341] The minimum value is less than or equal to the twenty-first threshold;

[0342] The minimum value is greater than or equal to the twenty-second threshold;

[0343] Wherein, the first parameter set includes at least some parameters of the first AI unit.

[0344] Optionally, the target threshold value is related to a target difference, and the target difference is the difference between two adjacent fixed-point numbers after quantization or the difference of the floating-point number represented by the least significant bit after quantization;

[0345] Wherein, the target threshold value includes at least one of the first threshold value, the second threshold value, the third threshold value, the fourth threshold value, the fifth threshold value, the sixth threshold value, the seventh threshold value, the eighth threshold value, the ninth threshold value, the tenth threshold value, the eleventh threshold value, the twelfth threshold value, the thirteenth threshold value, the fourteenth threshold value, the fifteenth threshold value, the sixteenth threshold value, the seventeenth threshold value, the eighteenth threshold value, the nineteenth threshold value, the twentieth threshold value, the twenty-first threshold value, and the twenty-second threshold value.

[0346] Optionally, the first parameter set includes at least one of the following:

[0347] At least some parameters of the first AI unit indicated by the second device, or parameters in at least some structures of the first AI unit;

[0348] Parameters in at least some structures of the first AI unit reported by the first device;

[0349] Parameters of the adaptive layer of the first AI unit;

[0350] Parameters of the input layer of the first AI unit;

[0351] Parameters of the hidden layer of the first AI unit;

[0352] Parameters of the output layer of the first AI unit.

[0353] The quantization alignment device 600 of the AI unit provided in the embodiments of the present application can implement each process in the method embodiments on the first device side and achieve the same technical effects. To avoid repetition, it will not be described in detail here.

[0354] Referring to Figure 7 , the embodiments of the present application also provide another quantization alignment device of an AI unit, which is applied to a second device. As Figure 7 shown, the quantization alignment device 700 of the AI unit includes:

[0355] The second execution module 701 is used to execute a second operation, where the second operation includes:

[0356] Quantize the parameter information of the first AI unit according to the target quantization information to obtain first information, where the first information is quantized fixed-point number information; send the first information to the first device; or,

[0357] Send second information to the first device, where the second information includes the parameter information of the first AI unit, and the second information is floating-point number information, and the second information is used to be quantized according to the target quantization information to obtain the first AI unit.

[0358] Optionally, the quantization alignment device 700 of the AI unit further includes:

[0359] A second receiving module, configured to receive first capability information from the first device;

[0360] A third determination module, configured to determine the target quantization information according to the first capability information;

[0361] Wherein, the first capability information is used to indicate third quantization information supported by the first device, and the third quantization information includes the target quantization information.

[0362] Optionally, the quantization alignment device 700 of the AI unit further includes:

[0363] A fourth determination module, configured to determine the target quantization information according to second quantization information agreed upon by the protocol.

[0364] Optionally, the quantization alignment device 700 of the AI unit further includes:

[0365] A third sending module, configured to send the target quantization information to the first device.

[0366] Optionally, the quantization alignment device 700 of the AI unit further includes:

[0367] A fourth sending module, configured to send first quantization information to the first device, and the first quantization information is used to assist in determining the target quantization information.

[0368] Optionally, the fourth sending module is specifically configured to:

[0369] Before the second execution module sends the first information or the second information to the first device, send the first quantization information to the first device.

[0370] Optionally, the quantization alignment device 700 of the AI unit further includes:

[0371] A third receiving module, configured to receive the target quantization information from the first device.

[0372] Optionally, the target quantization information is used to indicate at least one of a quantization method and a quantization level.

[0373] Optionally, the quantization method includes at least one of the following:

[0374] Direct quantization method, which is used to quantize each parameter of the AI unit;

[0375] Group quantization method, which is used to quantize each group of parameters of the AI unit, where a group of parameters includes at least one parameter, and the parameters within the same group correspond to the same quantization value;

[0376] Transform domain quantization method, which is used to convert the floating-point parameters of the AI unit into transform domain parameters, and perform quantization processing and inverse transform processing on the transform domain parameters to obtain the fixed-point parameters after quantization of the AI unit;

[0377] Product quantization method, which is used to quantize the first subspace, and the first subspace is a subspace composed of at least one parameter of the AI unit.

[0378] Optionally, the quantization level is divided based on at least one of the following:

[0379] The number of bits of the fixed-point number obtained after quantizing a floating-point number;

[0380] The difference between two adjacent fixed-point numbers after quantization, or the floating-point number difference represented by the lowest bit after quantization.

[0381] Optionally, the first AI unit includes a first parameter set, and the floating-point values of the first parameter set satisfy at least one of the following requirements:

[0382] The maximum amplitude of the parameters in the first parameter set is less than or equal to a first threshold;

[0383] The maximum amplitude of the parameters in the first parameter set is greater than or equal to a second threshold;

[0384] The minimum amplitude of the parameters in the first parameter set is less than or equal to a third threshold;

[0385] The minimum amplitude of the parameters in the first parameter set is greater than or equal to a fourth threshold;

[0386] The average amplitude of the parameters in the first parameter set is less than or equal to a fifth threshold;

[0387] The average amplitude of the parameters in the first parameter set is greater than or equal to the sixth threshold value;

[0388] The difference between the first value and the second value is less than or equal to the seventh threshold value;

[0389] The ratio of the first value to the second value is less than or equal to the eighth threshold value;

[0390] The difference between the third value and the fourth value is less than or equal to the ninth threshold value;

[0391] The ratio of the third value to the fourth value is less than or equal to the tenth threshold value;

[0392] The difference between the fifth value and the sixth value is less than or equal to the eleventh threshold value;

[0393] The ratio of the fifth value to the sixth value is less than or equal to the twelfth threshold value;

[0394] The difference between the maximum amplitude and the minimum amplitude of the parameters in the first parameter set is less than or equal to the thirteenth threshold value;

[0395] The ratio of the maximum amplitude to the minimum amplitude of the parameters in the first parameter set is less than or equal to the fourteenth threshold value;

[0396] The difference between the maximum amplitude and the average amplitude of the parameters in the first parameter set is less than or equal to the fifteenth threshold value;

[0397] The ratio of the maximum amplitude to the average amplitude of the parameters in the first parameter set is less than or equal to the sixteenth threshold value;

[0398] The difference between the average amplitude and the minimum amplitude of the parameters in the first parameter set is less than or equal to the seventeenth threshold value;

[0399] The ratio of the average amplitude to the minimum amplitude of the parameters in the first parameter set is less than or equal to the eighteenth threshold value;

[0400] Wherein, the first parameter set includes at least some parameters of the first AI unit; the first AI unit includes a first layer and a second layer; the first value is the maximum amplitude of the parameters in the first layer; the second value is the maximum amplitude of the parameters in the second layer; the third value is the minimum amplitude of the parameters in the first layer; the fourth value is the minimum amplitude of the parameters in the second layer; the fifth value is the average amplitude of the parameters in the first layer; the sixth value is the average amplitude of the parameters in the second layer; and the first value is greater than or equal to the second value, the third value is greater than or equal to the fourth value, and the fifth value is greater than or equal to the sixth value.

[0401] Optionally, the maximum amplitude is the minimum value among the largest K1 amplitudes, where K1 is a positive integer.

[0402] Optionally, the minimum amplitude is the maximum value among the smallest K2 amplitudes, where K2 is a positive integer.

[0403] Optionally, the first AI unit includes a first parameter set, and the fixed-point values after quantization of the first parameter set satisfy at least one of the following requirements:

[0404] The maximum value is less than or equal to the nineteenth threshold;

[0405] The maximum value is greater than or equal to the twentieth threshold;

[0406] The minimum value is less than or equal to the twenty-first threshold;

[0407] The minimum value is greater than or equal to the twenty-second threshold;

[0408] Wherein, the first parameter set includes at least part of the parameters of the first AI unit.

[0409] Optionally, the target threshold value is related to the target difference, and the target difference is the difference between two adjacent fixed-point numbers after quantization or the floating-point number difference represented by the lowest bit after quantization;

[0410] Wherein, the target threshold value includes at least one of the first threshold value, the second threshold value, the third threshold value, the fourth threshold value, the fifth threshold value, the sixth threshold value, the seventh threshold value, the eighth threshold value, the ninth threshold value, the tenth threshold value, the eleventh threshold value, the twelfth threshold value, the thirteenth threshold value, the fourteenth threshold value, the fifteenth threshold value, the sixteenth threshold value, the seventeenth threshold value, the eighteenth threshold value, the nineteenth threshold value, the twentieth threshold value, the twenty-first threshold value, and the twenty-second threshold value.

[0411] Optionally, the first parameter set includes at least one of the following:

[0412] At least part of the parameters of the first AI unit indicated by the second device, or the parameters in at least part of the structure of the first AI unit;

[0413] The parameters in at least part of the structure of the first AI unit reported by the first device;

[0414] The parameters of the adaptive layer of the first AI unit;

[0415] The parameters of the input layer of the first AI unit;

[0416] The parameters of the hidden layer of the first AI unit;

[0417] The parameters of the output layer of the first AI unit.

[0418] The quantization alignment device 700 of the AI unit provided by the embodiments of the present application can implement each process in the method embodiments on the second device side and achieve the same technical effects. To avoid repetition, details are not described here again.

[0419] The quantization alignment device of the AI unit in the embodiments of the present application can be an electronic device, such as an electronic device with an operating system, or a component in an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or a network-side device. Exemplarily, the terminal can include, but is not limited to, the types of the terminal 11 listed above, and the network-side device includes, but is not limited to, the types of the access network device or the core network device listed above, etc. The embodiments of the present application do not make specific limitations.

[0420] Optionally, as Figure 8 shown, the embodiments of the present application further provide a communication device 800, including a processor 801 and a memory 802. A program or instruction that can run on the processor 801 is stored on the memory 802. For example: when the communication device 800 is used as the first device, when the program or instruction is executed by the processor 801, each step of the method embodiment on the first device side described above is implemented, and the same technical effects can be achieved; when the communication device 800 is used as the second device, when the program or instruction is executed by the processor 801, each step of the method embodiment on the second device side described above is implemented, and the same technical effects can be achieved. To avoid repetition, details are not described here again.

[0421] The embodiments of the present application further provide a communication device, including a processor and a communication interface.

[0422] When the communication device is the first device, the communication interface is used to receive the first information or the second information from the second device; the processor is used to perform a first operation, and the first operation includes:

[0423] Determining a first AI unit according to the first information; or,

[0424] Quantizing the second information according to the target quantization information to obtain a first AI unit;

[0425] Wherein, the first information includes parameter information of the first AI unit, and the first information is fixed-point number information quantized based on the target quantization information; the second information includes parameter information of the first AI unit, and the second information is floating-point number information.

[0426] When the communication device is the second device, the communication interface and the processor are used to perform a second operation, where the second operation includes:

[0427] Quantize the parameter information of the first AI unit according to the target quantization information to obtain first information, where the first information is quantized fixed-point number information; send the first information to the first device; or,

[0428] Send second information to the first device, where the second information includes the parameter information of the first AI unit, and the second information is floating-point number information, and the second information is used to be quantized according to the target quantization information to obtain the first AI unit.

[0429] This embodiment of the communication device corresponds to the foregoing embodiment of the quantization alignment method of the AI unit on the first device and the second device side. Each implementation process and implementation manner of the foregoing method embodiment can be applied to this embodiment of the communication device, and the same technical effect can be achieved.

[0430] In some embodiments, Figure 9 It is a schematic diagram of the hardware structure of a terminal for implementing an embodiment of the present application.

[0431] The terminal 900 includes but is not limited to at least some components such as a radio frequency unit 901, a network module 902, an audio output unit 903, an input unit 904, a sensor 905, a display unit 906, a user input unit 907, an interface unit 908, a memory 909, and a processor 910.

[0432] Those skilled in the art can understand that the terminal 900 may further include a power supply (such as a battery) for supplying power to each component. The power supply can be logically connected to the processor 910 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system. Figure 9 The terminal structure shown does not limit the terminal. The terminal may include more or fewer components than shown, or combine some components, or have different component arrangements, which will not be elaborated here.

[0433] It should be understood that in the embodiments of the present application, the input unit 904 may include a Graphics Processing Unit (GPU) 9041 and a microphone 9042. The graphics processor 9041 processes the image data of static pictures or videos obtained by an image capture device (such as a camera) in the video capture mode or the image capture mode. The display unit 906 may include a display panel 9061, and the display panel 9061 may be configured in the form of, for example, a liquid crystal display, an organic light emitting diode, etc. The user input unit 907 includes at least one of a touch panel 9071 and other input devices 9072. The touch panel 9071 is also referred to as a touch screen. The touch panel 9071 may include two parts: a touch detection device and a touch controller. The other input devices 9072 may include, but are not limited to, a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, and a joystick, which will not be elaborated herein.

[0434] In the embodiments of the present application, after receiving downlink data from a network-side device, the radio frequency unit 901 may transmit it to the processor 910 for processing; in addition, the radio frequency unit 901 may send uplink data to the network-side device. Generally, the radio frequency unit 901 includes, but is not limited to, an antenna, an amplifier, a transceiver, a coupler, a low noise amplifier, a duplexer, etc.

[0435] The memory 909 can be used to store software programs or instructions as well as various data. The memory 909 may mainly include a first storage area for storing programs or instructions and a second storage area for storing data. Among them, the first storage area may store an operating system, application programs or instructions required for at least one function (such as a sound playback function, an image playback function, etc.). In addition, the memory 909 may include volatile memory or non-volatile memory. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), a static random access memory (SRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a double data rate synchronous dynamic random access memory (DDR SDRAM), an enhanced synchronous dynamic random access memory (ESDRAM), a synchronous link dynamic random access memory (SLDRAM), and a direct rambus random access memory (DRRAM). The memory 909 in the embodiments of the present application includes but is not limited to these and any other suitable types of memory.

[0436] The processor 910 may include one or more processing units; optionally, the processor 910 integrates an application processor and a modem processor. Among them, the application processor mainly processes operations related to the operating system, user interface, and application programs, etc., and the modem processor mainly processes wireless communication signals, such as a baseband processor. It can be understood that the above modem processor may not be integrated into the processor 910 either.

[0437] In one implementation, the terminal 900 serves as the first device.

[0438] The radio frequency unit 901 is used to receive the first information or the second information from the second device;

[0439] The processor 910 is used to execute a first operation, and the first operation includes:

[0440] Determine a first AI unit according to the first information; or,

[0441] Quantize the second information according to the target quantization information to obtain a first AI unit;

[0442] Among them, the first information includes parameter information of the first AI unit, and the first information is fixed-point number information quantized based on the target quantization information; the second information includes parameter information of the first AI unit, and the second information is floating-point number information.

[0443] Optionally, the target quantization information is used to indicate at least one of a quantization method and a quantization level.

[0444] Optionally, the processor 910 is further configured to perform at least one of the following:

[0445] Determine the target quantization information according to the first quantization information from the second device;

[0446] Determine the target quantization information according to the second quantization information agreed upon by the protocol.

[0447] Optionally, the radio frequency unit 901 is further configured to, when the first quantization information is not completely the same as the target quantization information, the first device sends the target quantization information to the second device.

[0448] Optionally, the determining the target quantization information according to the first quantization information from the second device performed by the processor 910 includes:

[0449] Before receiving the first information or the second information from the second device through the radio frequency unit 901, control the radio frequency unit 901 to receive the first quantization information from the second device;

[0450] Determine the target quantization information according to the first quantization information.

[0451] Optionally, the radio frequency unit 901 is further configured to send first capability information to the second device, where the first capability information is used to indicate third quantization information supported by the first device, and the third quantization information includes the target quantization information.

[0452] Optionally, the quantization method includes at least one of the following:

[0453] Direct quantization method, which is used to quantize each parameter of the AI unit;

[0454] Group quantization method, which is used to quantize each group of parameters of the AI unit, where a group of parameters includes at least one parameter, and the parameters in the same group correspond to the same quantization value;

[0455] The transform domain quantization method is used to convert the floating-point parameters of the AI unit into transform domain parameters, and perform quantization processing and inverse transform processing on the transform domain parameters to obtain the fixed-point parameters after quantization of the AI unit;

[0456] The product quantization method is used to quantize the first subspace, where the first subspace is a subspace composed of at least one parameter of the AI unit.

[0457] Optionally, the quantization levels are divided based on at least one of the following:

[0458] The number of bits of the fixed-point number obtained after quantizing a floating-point number;

[0459] The difference between two adjacent fixed-point numbers after quantization, or the difference of the floating-point numbers represented by the lowest bit after quantization.

[0460] Optionally, the first AI unit includes a first parameter set, and the floating-point values of the first parameter set satisfy at least one of the following requirements:

[0461] The maximum amplitude of the parameters in the first parameter set is less than or equal to the first threshold value;

[0462] The maximum amplitude of the parameters in the first parameter set is greater than or equal to the second threshold value;

[0463] The minimum amplitude of the parameters in the first parameter set is less than or equal to the third threshold value;

[0464] The minimum amplitude of the parameters in the first parameter set is greater than or equal to the fourth threshold value;

[0465] The average amplitude of the parameters in the first parameter set is less than or equal to the fifth threshold value;

[0466] The average amplitude of the parameters in the first parameter set is greater than or equal to the sixth threshold value;

[0467] The difference between the first value and the second value is less than or equal to the seventh threshold value;

[0468] The ratio of the first value to the second value is less than or equal to the eighth threshold value;

[0469] The difference between the third value and the fourth value is less than or equal to the ninth threshold value;

[0470] The ratio of the third value to the fourth value is less than or equal to the tenth threshold value;

[0471] The difference between the fifth value and the sixth value is less than or equal to the eleventh threshold value;

[0472] The ratio of the fifth value to the sixth value is less than or equal to the twelfth threshold value;

[0473] The difference between the maximum amplitude and the minimum amplitude of the parameters in the first parameter set is less than or equal to the thirteenth threshold value;

[0474] The ratio of the maximum amplitude to the minimum amplitude of the parameters in the first parameter set is less than or equal to the fourteenth threshold value;

[0475] The difference between the maximum amplitude and the average amplitude of the parameters in the first parameter set is less than or equal to the fifteenth threshold value;

[0476] The ratio of the maximum amplitude to the average amplitude of the parameters in the first parameter set is less than or equal to the sixteenth threshold value;

[0477] The difference between the average amplitude and the minimum amplitude of the parameters in the first parameter set is less than or equal to the seventeenth threshold value;

[0478] The ratio of the average amplitude to the minimum amplitude of the parameters in the first parameter set is less than or equal to the eighteenth threshold value;

[0479] Wherein, the first parameter set includes at least some parameters of the first AI unit; the first AI unit includes a first layer and a second layer; the first value is the maximum amplitude of the parameters in the first layer; the second value is the maximum amplitude of the parameters in the second layer; the third value is the minimum amplitude of the parameters in the first layer; the fourth value is the minimum amplitude of the parameters in the second layer; the fifth value is the average amplitude of the parameters in the first layer; the sixth value is the average amplitude of the parameters in the second layer; and the first value is greater than or equal to the second value, the third value is greater than or equal to the fourth value, and the fifth value is greater than or equal to the sixth value.

[0480] Optionally, the maximum amplitude is the minimum value among the largest K1 amplitudes, where K1 is a positive integer.

[0481] Optionally, the minimum amplitude is the maximum value among the smallest K2 amplitudes, where K2 is a positive integer.

[0482] Optionally, the first AI unit includes a first parameter set, and the fixed-point values after quantization of the first parameter set satisfy at least one of the following requirements:

[0483] The maximum value is less than or equal to the nineteenth threshold;

[0484] The maximum value is greater than or equal to the twentieth threshold;

[0485] The minimum value is less than or equal to the twenty-first threshold;

[0486] The minimum value is greater than or equal to the twenty-second threshold;

[0487] Wherein, the first parameter set includes at least part of the parameters of the first AI unit.

[0488] Optionally, the target threshold value is related to a target difference, where the target difference is the difference between two adjacent fixed-point numbers after quantization or the floating-point number difference represented by the least significant bit after quantization;

[0489] Wherein, the target threshold value includes at least one of the first threshold value, the second threshold value, the third threshold value, the fourth threshold value, the fifth threshold value, the sixth threshold value, the seventh threshold value, the eighth threshold value, the ninth threshold value, the tenth threshold value, the eleventh threshold value, the twelfth threshold value, the thirteenth threshold value, the fourteenth threshold value, the fifteenth threshold value, the sixteenth threshold value, the seventeenth threshold value, the eighteenth threshold value, the nineteenth threshold value, the twentieth threshold value, the twenty-first threshold value, and the twenty-second threshold value.

[0490] Optionally, the first parameter set includes at least one of the following:

[0491] At least part of the parameters of the first AI unit indicated by the second device, or the parameters in at least part of the structure of the first AI unit;

[0492] The parameters in at least part of the structure of the first AI unit reported by the first device;

[0493] The parameters of the adaptive layer of the first AI unit;

[0494] The parameters of the input layer of the first AI unit;

[0495] The parameters of the hidden layer of the first AI unit;

[0496] The parameters of the output layer of the first AI unit.

[0497] It can be understood that the implementation processes of the implementation manners mentioned in this embodiment can refer to the relevant descriptions of the foregoing method embodiment on the first device side, and achieve the same or corresponding technical effects. To avoid repetition, they will not be elaborated here.

[0498] In another implementation manner, the above terminal 900 serves as the second device.

[0499] The processor 910 is configured to perform a second operation, where the second operation includes:

[0500] Quantize the parameter information of the first AI unit according to the target quantization information to obtain the first information, where the first information is quantized fixed-point number information; control the radio frequency unit 901 to send the first information to the first device; or,

[0501] Control the radio frequency unit 901 to send second information to the first device, where the second information includes the parameter information of the first AI unit, and the second information is floating-point number information, and the second information is used to be quantized according to the target quantization information to obtain the first AI unit.

[0502] Optionally, the radio frequency unit 901 is further configured to receive first capability information from the first device;

[0503] The processor 910 is further configured to determine the target quantization information according to the first capability information;

[0504] Wherein, the first capability information is used to indicate third quantization information supported by the first device, and the third quantization information includes the target quantization information.

[0505] Optionally, the processor 910 is further configured to determine the target quantization information according to second quantization information agreed upon by the protocol.

[0506] Optionally, the radio frequency unit 901 is further configured to send the target quantization information to the first device.

[0507] Optionally, the radio frequency unit 901 is further configured to send first quantization information to the first device, and the first quantization information is used to assist in determining the target quantization information.

[0508] Optionally, the sending of the first quantization information to the first device performed by the radio frequency unit 901 includes:

[0509] Before sending the first information or the second information to the first device, send the first quantization information to the first device.

[0510] Optionally, the radio frequency unit 901 is further configured to receive the target quantization information from the first device.

[0511] Optionally, the target quantization information is used to indicate at least one of a quantization method and a quantization level.

[0512] Optionally, the quantization method includes at least one of the following:

[0513] Direct quantization method, which is used to quantize each parameter of the AI unit;

[0514] Group quantization method, where the group quantization method is used to quantize each group of parameters of the AI unit. Among them, a group of parameters includes at least one parameter, and the parameters within the same group correspond to the same quantization value; Transform domain quantization method, where the transform domain quantization method is used to convert the floating-point parameters of the AI unit into transform domain parameters, and perform quantization processing and inverse transform processing on the transform domain parameters to obtain the quantized fixed-point parameters of the AI unit;

[0515] Product quantization method, where the product quantization method is used to quantize the first subspace, and the first subspace is a subspace composed of at least one parameter of the AI unit.

[0516] Optionally, the quantization levels are divided based on at least one of the following:

[0517] The number of bits of the fixed-point number obtained after quantizing a floating-point number;

[0518] Target difference, where the target difference is the difference between two adjacent quantized fixed-point numbers or the floating-point number difference represented by the lowest bit after quantization.

[0519] Optionally, the first AI unit includes a first parameter set, and the floating-point values of the first parameter set satisfy at least one of the following requirements:

[0520] The maximum amplitude of the parameters in the first parameter set is less than or equal to the first threshold;

[0521] The maximum amplitude of the parameters in the first parameter set is greater than or equal to the second threshold;

[0522] The minimum amplitude of the parameters in the first parameter set is less than or equal to the third threshold;

[0523] The minimum amplitude of the parameters in the first parameter set is greater than or equal to the fourth threshold;

[0524] The average amplitude of the parameters in the first parameter set is less than or equal to the fifth threshold;

[0525] The average amplitude of the parameters in the first parameter set is greater than or equal to the sixth threshold;

[0526] The difference between the first value and the second value is less than or equal to the seventh threshold;

[0527] The ratio of the first value to the second value is less than or equal to the eighth threshold;

[0528] The difference between the third value and the fourth value is less than or equal to the ninth threshold;

[0529] The ratio of the third value to the fourth value is less than or equal to the tenth threshold;

[0530] The difference between the fifth value and the sixth value is less than or equal to the eleventh threshold value;

[0531] The ratio of the fifth value to the sixth value is less than or equal to the twelfth threshold value;

[0532] The difference between the maximum amplitude of the parameters in the first parameter set and the minimum amplitude of the parameters in the first parameter set is less than or equal to the thirteenth threshold value;

[0533] The ratio of the maximum amplitude of the parameters in the first parameter set to the minimum amplitude of the parameters in the first parameter set is less than or equal to the fourteenth threshold value;

[0534] The difference between the maximum amplitude of the parameters in the first parameter set and the average amplitude of the parameters in the first parameter set is less than or equal to the fifteenth threshold value;

[0535] The ratio of the maximum amplitude of the parameters in the first parameter set to the average amplitude of the parameters in the first parameter set is less than or equal to the sixteenth threshold value;

[0536] The difference between the average amplitude of the parameters in the first parameter set and the minimum amplitude of the parameters in the first parameter set is less than or equal to the seventeenth threshold value;

[0537] The ratio of the average amplitude of the parameters in the first parameter set to the minimum amplitude of the parameters in the first parameter set is less than or equal to the eighteenth threshold value;

[0538] Wherein, the first parameter set includes at least some parameters of the first AI unit; the first AI unit includes a first layer and a second layer; the first value is the maximum amplitude of the parameters in the first layer; the second value is the maximum amplitude of the parameters in the second layer; the third value is the minimum amplitude of the parameters in the first layer; the fourth value is the minimum amplitude of the parameters in the second layer; the fifth value is the average amplitude of the parameters in the first layer; the sixth value is the average amplitude of the parameters in the second layer; and the first value is greater than or equal to the second value, the third value is greater than or equal to the fourth value, and the fifth value is greater than or equal to the sixth value.

[0539] Optionally, the maximum amplitude is the minimum value among the largest K1 amplitudes, and K1 is a positive integer.

[0540] Optionally, the minimum amplitude is the maximum value among the smallest K2 amplitudes, and K2 is a positive integer.

[0541] Optionally, the first AI unit includes a first parameter set, and the fixed-point values after quantization of the first parameter set satisfy at least one of the following requirements:

[0542] The maximum value is less than or equal to the nineteenth threshold;

[0543] The maximum value is greater than or equal to the twentieth threshold;

[0544] The minimum value is less than or equal to the twenty - first threshold;

[0545] The minimum value is greater than or equal to the twenty - second threshold;

[0546] Wherein, the first parameter set includes at least part of the parameters of the first AI unit.

[0547] Optionally, the target threshold value is related to a target difference, and the target difference is the difference between two adjacent fixed - point numbers after quantization or the floating - point number difference represented by the least significant bit after quantization;

[0548] Wherein, the target threshold value includes at least one of the first threshold value, the second threshold value, the third threshold value, the fourth threshold value, the fifth threshold value, the sixth threshold value, the seventh threshold value, the eighth threshold value, the ninth threshold value, the tenth threshold value, the eleventh threshold value, the twelfth threshold value, the thirteenth threshold value, the fourteenth threshold value, the fifteenth threshold value, the sixteenth threshold value, the seventeenth threshold value, the eighteenth threshold value, the nineteenth threshold value, the twentieth threshold value, the twenty - first threshold value, and the twenty - second threshold value.

[0549] Optionally, the first parameter set includes at least one of the following:

[0550] At least part of the parameters of the first AI unit indicated by the second device, or the parameters in at least part of the structure of the first AI unit;

[0551] The parameters in at least part of the structure of the first AI unit reported by the first device;

[0552] The parameters of the adaptive layer of the first AI unit;

[0553] The parameters of the input layer of the first AI unit;

[0554] The parameters of the hidden layer of the first AI unit;

[0555] The parameters of the output layer of the first AI unit.

[0556] It can be understood that the implementation processes of the various implementation manners mentioned in this embodiment can refer to the relevant descriptions of the method embodiment on the second device side, and achieve the same or corresponding technical effects. To avoid repetition, they are not described herein again.

[0557] The embodiment of the present application further provides a network-side device, including a processor and a communication interface. The communication interface is coupled to the processor, and the processor is configured to run programs or instructions to implement the steps of the method embodiments of the foregoing first device side or second device side. This network-side device embodiment corresponds to the method embodiments of the foregoing first device side or second device side. Each implementation process and implementation manner of the above method embodiments can be applied to this network-side device embodiment, and the same technical effects can be achieved.

[0558] In one implementation, as Figure 10 shown, the network-side device 1000 includes: an antenna 1001, a radio frequency device 1002, a baseband device 1003, a processor 1004, and a memory 1005. The antenna 1001 is connected to the radio frequency device 1002. In the uplink direction, the radio frequency device 1002 receives information through the antenna 1001 and sends the received information to the baseband device 1003 for processing. In the downlink direction, the baseband device 1003 processes the information to be sent and sends it to the radio frequency device 1002. After processing the received information, the radio frequency device 1002 sends it out through the antenna 1001.

[0559] The method executed by the network-side device in the above embodiments can be implemented in the baseband device 1003, and the baseband device 1003 includes a baseband processor.

[0560] The baseband device 1003 may include, for example, at least one baseband board, and a plurality of chips are provided on the baseband board. As Figure 11 shown, one of the chips is, for example, a baseband processor, which is connected to the memory 1005 through a bus interface to call the programs in the memory 1005 and execute the operations of the network device shown in the above method embodiments.

[0561] The network-side device may further include a network interface 1006, and this interface is, for example, a Common Public Radio Interface (CPRI).

[0562] Specifically, the network-side device 1000 in the embodiment of the present application further includes: instructions or programs stored on the memory 1005 and executable on the processor 1004. The processor 1004 calls the instructions or programs in the memory 1005 to execute Figure 6 and Figure 7 the methods executed by at least one of the modules shown in, and achieve the same technical effects. To avoid repetition, it will not be elaborated here.

[0563] In another implementation, the embodiment of the present application further provides a network-side device. As Figure 11As shown in the figure, the network-side device 1100 includes: a processor 1101, a network interface 1102, and a memory 1103. Among them, the network interface 1102 is, for example, a Common Public Radio Interface (CPRI).

[0564] Specifically, the network-side device 1100 in the embodiment of the present application further includes: instructions or programs stored on the memory 1103 and executable on the processor 1101. The processor 1101 calls the instructions or programs in the memory 1103 to execute the methods performed by the respective modules shown in 7 and achieve the same technical effects. To avoid repetition, it will not be elaborated here.

[0565] The embodiment of the present application further provides a readable storage medium. Programs or instructions are stored on the readable storage medium. When the programs or instructions are executed by a processor, they implement the respective processes of the quantization alignment method embodiment of the AI unit on the first device or the second device side, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.

[0566] Among them, the processor is the processor in the terminal described in the above embodiment. The readable storage medium includes computer-readable storage media, such as computer read-only memory ROM, random access memory RAM, magnetic disks, or optical discs, etc. In some examples, the readable storage medium may be a non-transitory readable storage medium.

[0567] The embodiment of the present application further provides a chip. The chip includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the respective processes of the quantization alignment method embodiment of the AI unit on the first device or the second device side, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.

[0568] It should be understood that the chip mentioned in the embodiment of the present application may also be referred to as a system-on-chip, system chip, chip system, or system-on-chip, etc.

[0569] The embodiment of the present application further provides a computer program / program product. The computer program / program product is stored in a storage medium. The computer program / program product is executed by at least one processor to implement the respective processes of the quantization alignment method embodiment of the AI unit on the first device or the second device side, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.

[0570] The embodiments of the present application further provide a communication system, including: a first device and a second device. The first device can be used to execute the steps of the embodiment of the quantization alignment method of the AI unit on the first device side, and the second device can be used to execute the steps of the embodiment of the quantization alignment method of the AI unit on the second device side.

[0571] It should be noted that in this document, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device including a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including that element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in a reverse order according to the functions involved. For example, the described methods may be performed in an order different from that described, and various steps may be added, omitted or combined. Additionally, the features described with reference to certain examples may be combined in other examples.

[0572] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of a computer software product plus a necessary general hardware platform, and of course, they can also be implemented by hardware. The computer software product is stored in a storage medium (such as ROM, RAM, magnetic disk, optical disk, etc.) and includes several instructions for causing a terminal or a network-side device to execute the methods described in various embodiments of the present application.

[0573] The above describes the embodiments of the present application with reference to the drawings, but the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms of embodiments without departing from the purpose of the present application and the scope protected by the claims. These embodiments are all within the protection scope of the present application.

Claims

1. A quantization alignment method for an AI unit, characterized in that, Including: The first device receives the first information or the second information from the second device; The first device performs a first operation, and the first operation includes: Determining a first AI unit according to the first information; or Quantifying the second information according to target quantization information to obtain a first AI unit; Wherein, the first information includes parameter information of the first AI unit, and the first information is fixed-point number information quantized based on the target quantization information; the second information includes parameter information of the first AI unit, and the second information is floating-point number information.

2. The method according to claim 1, characterized in that, The target quantization information is used to indicate at least one of a quantization method and a quantization level.

3. The method according to claim 1 or 2, characterized in that, The method further includes at least one of the following: The first device determines the target quantization information according to first quantization information from the second device; The first device determines the target quantization information according to second quantization information agreed upon by the protocol.

4. The method according to claim 3, characterized in that, The method further includes: In the case where the first quantization information is not exactly the same as the target quantization information, the first device sends the target quantization information to the second device.

5. The method according to claim 3, wherein The first device determines the target quantization information according to first quantization information from the second device, including: Before receiving the first information or the second information from the second device, the first device receives first quantization information from the second device; The first device determines the target quantization information according to the first quantization information.

6. The method according to any one of claims 1 to 5, characterized in that, The method further includes: The first device sends first capability information to the second device, and the first capability information is used to indicate third quantization information supported by the first device, wherein the third quantization information includes the target quantization information.

7. The method according to claim 2, characterized in that, The quantization method includes at least one of the following: Direct quantization method, which is used to quantize each parameter of the AI unit; Group quantization method, which is used to quantize each group of parameters of the AI unit, wherein a group of parameters includes at least one parameter, and the parameters in the same group correspond to the same quantization value; Transform domain quantization method, which is used to convert the floating-point parameters of the AI unit into transform domain parameters, and perform quantization processing and inverse transform processing on the transform domain parameters to obtain the fixed-point parameters of the quantized AI unit; Product quantization method, which is used to quantize a first subspace, and the first subspace is a subspace composed of at least one parameter of the AI unit.

8. The method according to claim 2 or 7, characterized in that, The quantization level is divided based on at least one of the following: The number of bits of the fixed-point number obtained after quantizing a floating-point number; The difference between two adjacent fixed-point numbers after quantization, or the floating-point number difference represented by the lowest bit after quantization.

9. A quantization alignment method for an AI unit, characterized in that, Including: The second device performs a second operation, wherein the second operation includes: Quantifying the parameter information of the first AI unit according to the target quantization information to obtain first information, wherein the first information is quantized fixed-point number information; sending the first information to the first device; or Send second information to the first device, where the second information includes parameter information of a first AI unit, and the second information is floating-point information, and the second information is used to be quantized according to target quantization information to obtain the first AI unit.

10. The method according to claim 9, characterized in that, The method further includes: The second device receives first capability information from the first device; The second device determines the target quantization information according to the first capability information; Wherein, the first capability information is used to indicate third quantization information supported by the first device, and the third quantization information includes the target quantization information.

11. The method according to claim 9, characterized in that, The method further includes: The second device determines the target quantization information according to second quantization information agreed upon by the protocol.

12. The method according to claim 10 or 11, characterized in that, The method further includes: The second device sends the target quantization information to the first device.

13. The method according to claim 9, wherein The method further includes: The second device sends first quantization information to the first device, and the first quantization information is used to assist in determining the target quantization information.

14. The method according to claim 13, wherein The second device sending first quantization information to the first device includes: Before the second device sends the first information or the second information to the first device, the second device sends the first quantization information to the first device.

15. The method according to claim 9 or 13, characterized in that The method further includes: The second device receives the target quantization information from the first device.

16. The method according to any one of claims 9 to 15, characterized in that, The target quantization information is used to indicate at least one of a quantization method and a quantization level.

17. The method according to claim 16, wherein The quantization method includes at least one of the following: Direct quantization method, which is used to quantize each parameter of the AI unit; Group quantization method, which is used to quantize each group of parameters of the AI unit, where a group of parameters includes at least one parameter, and the parameters within the same group correspond to the same quantization value; Transform domain quantization method, which is used to convert the floating-point parameters of the AI unit into transform domain parameters, and perform quantization processing and inverse transform processing on the transform domain parameters to obtain the fixed-point parameters of the quantized AI unit; Product quantization method, which is used to quantize a first subspace, and the first subspace is a subspace composed of at least one parameter of the AI unit.

18. The method according to claim 16 or 17, characterized in that, The quantization level is divided based on at least one of the following: The number of bits of the fixed-point number obtained after quantizing a floating-point number; Target difference, where the target difference is the difference between two adjacent fixed-point numbers after quantization or the floating-point number difference represented by the lowest bit after quantization.

19. The method according to any one of claims 1 to 18, characterized in that, The first AI unit includes a first parameter set, and the floating-point values of the first parameter set satisfy at least one of the following requirements: The maximum amplitude of the parameters in the first parameter set is less than or equal to a first threshold value; The maximum amplitude of the parameters in the first parameter set is greater than or equal to a second threshold value; The minimum amplitude of the parameters in the first parameter set is less than or equal to a third threshold value; The minimum amplitude of the parameters in the first parameter set is greater than or equal to a fourth threshold value; The average amplitude of the parameters in the first parameter set is less than or equal to a fifth threshold value; The average amplitude of the parameters in the first parameter set is greater than or equal to a sixth threshold value; The difference between the first value and the second value is less than or equal to a seventh threshold value; The ratio of the first value to the second value is less than or equal to the eighth threshold value; The difference between the third value and the fourth value is less than or equal to the ninth threshold value; The ratio of the third value to the fourth value is less than or equal to the tenth threshold value; The difference between the fifth value and the sixth value is less than or equal to the eleventh threshold value; The ratio of the fifth value to the sixth value is less than or equal to the twelfth threshold value; The difference between the maximum amplitude of the parameters in the first parameter set and the minimum amplitude of the parameters in the first parameter set is less than or equal to the thirteenth threshold value; The ratio of the maximum amplitude of the parameters in the first parameter set to the minimum amplitude of the parameters in the first parameter set is less than or equal to the fourteenth threshold value; The difference between the maximum amplitude of the parameters in the first parameter set and the average amplitude of the parameters in the first parameter set is less than or equal to the fifteenth threshold value; The ratio of the maximum amplitude of the parameters in the first parameter set to the average amplitude of the parameters in the first parameter set is less than or equal to the sixteenth threshold value; The difference between the average amplitude of the parameters in the first parameter set and the minimum amplitude of the parameters in the first parameter set is less than or equal to the seventeenth threshold value; The ratio of the average amplitude of the parameters in the first parameter set to the minimum amplitude of the parameters in the first parameter set is less than or equal to the eighteenth threshold value; Wherein, the first parameter set includes at least some parameters of the first AI unit; the first AI unit includes a first layer and a second layer; the first value is the maximum amplitude of the parameters in the first layer; the second value is the maximum amplitude of the parameters in the second layer; the third value is the minimum amplitude of the parameters in the first layer; the fourth value is the minimum amplitude of the parameters in the second layer; the fifth value is the average amplitude of the parameters in the first layer; the sixth value is the average amplitude of the parameters in the second layer; and the first value is greater than or equal to the second value, the third value is greater than or equal to the fourth value, and the fifth value is greater than or equal to the sixth value.

20. The method according to claim 19, wherein The maximum amplitude is the minimum value among the largest K1 amplitudes, and K1 is a positive integer.

21. The method according to claim 19, wherein The minimum amplitude is the maximum value among the smallest K2 amplitudes, and K2 is a positive integer.

22. The method according to any one of claims 1 to 18, characterized in that, The first AI unit includes a first parameter set, and the fixed-point values after quantization of the first parameter set meet at least one of the following requirements: The maximum value is less than or equal to the nineteenth threshold; The maximum value is greater than or equal to the twentieth threshold; The minimum value is less than or equal to the twenty-first threshold; The minimum value is greater than or equal to the twenty-second threshold; Wherein, the first parameter set includes at least some parameters of the first AI unit.

23. The method according to claim 19 or 22, characterized in that, The target threshold value is related to the target difference, and the target difference is the difference between two adjacent fixed-point numbers after quantization or the floating-point number difference represented by the lowest bit after quantization; Among them, the target threshold value includes at least one of the first threshold value, the second threshold value, the third threshold value, the fourth threshold value, the fifth threshold value, the sixth threshold value, the seventh threshold value, the eighth threshold value, the ninth threshold value, the tenth threshold value, the eleventh threshold value, the twelfth threshold value, the thirteenth threshold value, the fourteenth threshold value, the fifteenth threshold value, the sixteenth threshold value, the seventeenth threshold value, the eighteenth threshold value, the nineteenth threshold value, the twentieth threshold value, the twenty-first threshold value, and the twenty-second threshold value.

24. The method according to claim 19 or 22, characterized in that, The first parameter set includes at least one of the following: At least some parameters of the first AI unit indicated by the second device, or parameters in at least some structures of the first AI unit; Parameters in at least some structures of the first AI unit reported by the first device; Parameters of the adaptive layer of the first AI unit; Parameters of the input layer of the first AI unit; Parameters of the hidden layer of the first AI unit; Parameters of the output layer of the first AI unit.

25. A quantization alignment device for an AI unit, characterized in that, For the first device, the apparatus includes: A first receiving module, configured to receive the first information or the second information from the second device; A first execution module, configured to execute a first operation, where the first operation includes: Determining a first AI unit according to the first information; or, Quantifying the second information according to the target quantization information to obtain a first AI unit; Among them, the first information includes parameter information of the first AI unit, and the first information is fixed-point number information quantized based on the target quantization information; the second information includes parameter information of the first AI unit, and the second information is floating-point number information.

26. The device according to claim 25, characterized in that, The apparatus further includes at least one of the following: A first determining module, configured to determine the target quantization information according to the first quantization information from the second device; A second determining module, configured to determine the target quantization information according to the second quantization information agreed upon by the protocol.

27. The device according to claim 26, characterized in that, The apparatus further includes: A first sending module, configured to send the target quantization information to the second device when the first quantization information is not exactly the same as the target quantization information.

28. The device according to claim 26, wherein The first determining module includes: A receiving unit, configured to receive the first quantization information from the second device before the first receiving module receives the first information or the second information from the second device; A determining unit, configured to determine the target quantization information according to the first quantization information.

29. The device according to any one of claims 25 to 28, characterized in that The apparatus further includes: A second sending module, configured to send first capability information to the second device, where the first capability information is used to indicate the third quantization information supported by the first device, and the third quantization information includes the target quantization information.

30. A quantization alignment device for an AI unit, characterized in that, For the second device, the apparatus includes: A second execution module, configured to execute a second operation, where the second operation includes: Quantize the parameter information of the first AI unit according to the target quantization information to obtain first information, where the first information is quantized fixed-point number information; send the first information to the first device; or, Send second information to the first device, where the second information includes the parameter information of the first AI unit, and the second information is floating-point number information, and the second information is used to be quantized according to the target quantization information to obtain the first AI unit.

31. The device according to claim 30, characterized in that, The device further includes: A second receiving module, configured to receive first capability information from the first device; A third determining module, configured to determine the target quantization information according to the first capability information; Wherein the first capability information is used to indicate third quantization information supported by the first device, and the third quantization information includes the target quantization information.

32. The device according to claim 30, characterized in that, The device further includes: A fourth determining module, configured to determine the target quantization information according to second quantization information agreed upon by the protocol.

33. The device according to claim 31 or 32, characterized in that, The device further includes: A third sending module, configured to send the target quantization information to the first device.

34. The device according to claim 30, characterized in that, The device further includes: A fourth sending module, configured to send first quantization information to the first device, and the first quantization information is used to assist in determining the target quantization information.

35. The device according to claim 34, characterized in that, The fourth sending module is specifically configured to: Before the second execution module sends the first information or the second information to the first device, send the first quantization information to the first device.

36. The device according to claim 30 or 34, characterized in that, The device further includes: A third receiving module, configured to receive the target quantization information from the first device.

37. A communication device, characterized in that, It includes a processor and a memory, and the memory stores a program or instruction that can run on the processor. When the program or instruction is executed by the processor, the steps of the quantization alignment method of the AI unit as described in any one of claims 1 to 24 are implemented.

38. A readable storage medium, characterized in that, A program or instruction is stored on the readable storage medium, and when the program or instruction is executed by a processor, the steps of the quantization alignment method of the AI unit as described in any one of claims 1 to 24 are implemented.