Quantized ML Configuration Transfer for Wireless DNN Deployment
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Communicating machine-learning configurations for deep neural networks in wireless communication systems consumes large amounts of air interface resources and power, straining network capacity and device battery life, especially when dealing with complex architectures like multi-layered DNNs.
Innovation Solution
Implementing quantized machine-learning configurations using selected quantization formats such as vector or scalar quantization to reduce data transmission, where the base station and user equipment generate and transmit quantized representations of ML configuration information, allowing for efficient use of air interface resources and reduced power consumption.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If machine-learning configuration information is transmitted in full precision format, then the accuracy of the deep neural network is improved, but the consumption of air interface resources and power increases significantly
Solution Approach 1:
The patent applies quantization to change the precision parameter of ML configuration information from full precision (e.g., 32-bit floating point) to lower precision representations (e.g., 8-bit integers or even 1-bit signs). This parameter change reduces the amount of data to be transmitted while maintaining sufficient accuracy for the DNN to function effectively. The base station and UE negotiate and apply appropriate quantization levels based on channel conditions and DNN requirements.
Solution Approach 2:
The patent extracts only the most critical components of the ML configuration information for transmission. Instead of transmitting complete high-precision parameters, the system identifies and transmits only the essential configuration elements that have the greatest impact on DNN performance, leaving out redundant or less significant details that can be reconstructed or inferred at the receiving end.
2Measurement precision
If machine-learning configuration information is transmitted in full precision format, then the accuracy of the deep neural network is improved, but the network resource consumption increases
Solution Approach 1:
The patent changes the data representation parameter from full precision floating-point format to quantized formats with fewer bits. This transformation reduces the volume of configuration data that needs to be transmitted over the air interface. For example, 32-bit floating-point numbers are converted to 8-bit or 16-bit quantized values, achieving significant compression while preserving the essential information needed for DNN operation.
Solution Approach 2:
The patent transmits only a partial representation of the complete ML configuration information. Rather than sending all configuration parameters in full precision, the system transmits a quantized subset that contains sufficient information for the DNN to operate effectively, accepting some loss in precision in exchange for dramatically reduced data transmission requirements.
3Productivity
If complex deep neural network architectures are used, then the processing capability is improved, but the machine-learning configuration size increases
Solution Approach 1:
The patent applies quantization parameters to reduce the precision of DNN configuration data. Complex DNN architectures with many layers and nodes generate large configuration files, but by changing the parameter precision from 32-bit to 8-bit or lower, the overall configuration size is reduced proportionally while the DNN maintains its complex structure and processing capabilities.
Solution Approach 2:
The patent creates quantized copies of the original high-precision DNN configuration. Instead of transmitting the full-precision configuration data, the system transmits quantized versions that serve as sufficient copies for the receiving device to reconstruct and deploy the DNN. These quantized copies occupy significantly less space but preserve the essential architectural and parameter information needed for complex processing tasks.
Data Source
AI summary
Aspects describe communicating quantized machine-learning, ML, configuration information over a wireless network. A base station selects (605) a quantization configuration for quantizing ML configuration information for a deep neural network, DNN, where the quantization configuration indicates one or more quantization formats associated with quantizing the ML configuration information. The base station transmits (610) an indication of the quantization configuration to a user equipment, UE and transfers (615), over the wireless network and with the UE, quantized ML configuration information using the quantization configuration.


