Quantization parameter determination method and related apparatus

By dynamically adjusting the quantization parameters, the problem of the quantization scheme being unable to adapt to dynamically changing computing power and channel conditions is solved, and the effect of reducing communication overhead while ensuring the accuracy of AI tasks is achieved.

WO2025214068A1 Publication Date: 2025-10-16HUAWEI TECH CO LTD
5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/082367
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-10
Filing Date
2025-03-13
Publication Date
2025-10-16

AI Technical Summary

Technical Problem

Existing quantization schemes cannot dynamically adapt to the dynamically changing computing power and channel states of terminal devices, resulting in large communication overhead and affecting the processing accuracy and efficiency of AI tasks.

Method used

Through data interaction between the first device and the second device, the quantization parameters are dynamically adjusted, and the quantization parameters of the quantizer are optimized based on the channel state information and computing power information to adapt to the dynamically changing computing power and channel state, thereby reducing communication overhead.

Benefits of technology

While ensuring the accuracy of AI task processing, communication overhead is reduced and the adaptability and communication efficiency of the quantizer are improved.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

The present application discloses a quantization parameter determination method and a related apparatus. The quantization parameter determination method comprises: receiving first data and second data sent by a second device, wherein the first data is obtained by a first quantizer quantizing third data on the basis of quantization parameters configured in the m-th round, and the second data comprises at least one of channel state information and computing power information; determining fourth data on the basis of the first data and the second data, the fourth data being the gradient of an input layer of a first neural network model; and sending the fourth data to the second device, the fourth data being used by the second device to determine at least one first quantization parameter of the first quantizer in the (m+1)-th round. According to embodiments of the present application, quantization parameters of a quantizer can be dynamically adjusted, which is conducive to reducing communication overhead while ensuring the processing accuracy of AI tasks, or maximizes the processing accuracy of AI tasks when the given communication overhead is met.
Need to check novelty before this filing date? Find Prior Art