Quantized ML Configuration Transfer for Wireless DNN Deployment

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Communicating machine-learning configurations for deep neural networks in wireless communication systems consumes large amounts of air interface resources and power, straining network capacity and device battery life, especially when dealing with complex architectures like multi-layered DNNs.

Innovation Solution

Implementing quantized machine-learning configurations using selected quantization formats such as vector or scalar quantization to reduce data transmission, where the base station and user equipment generate and transmit quantized representations of ML configuration information, allowing for efficient use of air interface resources and reduced power consumption.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If machine-learning configuration information is transmitted in full precision format, then the accuracy of the deep neural network is improved, but the consumption of air interface resources and power increases significantly

Engineering Contradiction:
ImproveML configuration precisionVSAvoidpower consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent applies quantization to change the precision parameter of ML configuration information from full precision (e.g., 32-bit floating point) to lower precision representations (e.g., 8-bit integers or even 1-bit signs). This parameter change reduces the amount of data to be transmitted while maintaining sufficient accuracy for the DNN to function effectively. The base station and UE negotiate and apply appropriate quantization levels based on channel conditions and DNN requirements.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent extracts only the most critical components of the ML configuration information for transmission. Instead of transmitting complete high-precision parameters, the system identifies and transmits only the essential configuration elements that have the greatest impact on DNN performance, leaving out redundant or less significant details that can be reconstructed or inferred at the receiving end.

Inventive Principle:
Principle #2Taking out (Extraction)

2Measurement precision

If machine-learning configuration information is transmitted in full precision format, then the accuracy of the deep neural network is improved, but the network resource consumption increases

Engineering Contradiction:
ImproveML configuration precisionVSAvoiddata transmission volume
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent changes the data representation parameter from full precision floating-point format to quantized formats with fewer bits. This transformation reduces the volume of configuration data that needs to be transmitted over the air interface. For example, 32-bit floating-point numbers are converted to 8-bit or 16-bit quantized values, achieving significant compression while preserving the essential information needed for DNN operation.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent transmits only a partial representation of the complete ML configuration information. Rather than sending all configuration parameters in full precision, the system transmits a quantized subset that contains sufficient information for the DNN to operate effectively, accepting some loss in precision in exchange for dramatically reduced data transmission requirements.

Inventive Principle:
Principle #16Partial or excessive action

3Productivity

If complex deep neural network architectures are used, then the processing capability is improved, but the machine-learning configuration size increases

Engineering Contradiction:
Improveprocessing capabilityVSAvoidconfiguration data size
Core Design Contradiction:
ProductivityVSQuantity of substance

Solution Approach 1:

The patent applies quantization parameters to reduce the precision of DNN configuration data. Complex DNN architectures with many layers and nodes generate large configuration files, but by changing the parameter precision from 32-bit to 8-bit or lower, the overall configuration size is reduced proportionally while the DNN maintains its complex structure and processing capabilities.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent creates quantized copies of the original high-precision DNN configuration. Instead of transmitting the full-precision configuration data, the system transmits quantized versions that serve as sufficient copies for the receiving device to reconstruct and deploy the DNN. These quantized copies occupy significantly less space but preserve the essential architectural and parameter information needed for complex processing tasks.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS20240419953A1Quantized machine-learning configuration information
Publication Date: 2024.12.19 GOOGLE LLC
  • US20240419953A1 patent drawing
  • US20240419953A1 patent drawing
  • US20240419953A1 patent drawing

AI summary

Aspects describe communicating quantized machine-learning, ML, configuration information over a wireless network. A base station selects (605) a quantization configuration for quantizing ML configuration information for a deep neural network, DNN, where the quantization configuration indicates one or more quantization formats associated with quantizing the ML configuration information. The base station transmits (610) an indication of the quantization configuration to a user equipment, UE and transfers (615), over the wireless network and with the UE, quantized ML configuration information using the quantization configuration.