CoLor Component Address Translation for DNN Memory Footprint

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Deep Neural Networks (DNNs) face significant challenges in memory and bandwidth efficiency due to the substantial increase in memory footprint and bandwidth requirements caused by the lowering process in convolutional layers, which results in prohibitive computing resources and time for processing high-resolution images or videos.

Innovation Solution

The introduction of a convolutional lowering (CoLor) component that implements address translation logic, transparent to both processor and memory sub-systems, maps locations in a lowered matrix to equivalent locations in a non-lowered matrix, reducing memory footprint by K^2 times and improving bandwidth by merging multiple requests to the same memory location, while also streamlining zero padding in convolutional layers.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If convolutional layers are implemented using lowering process to enable GEMM operations, then computational efficiency is improved, but memory footprint increases by K^2 times

Engineering Contradiction:
Improvecomputational efficiencyVSAvoidmemory footprint
Core Design Contradiction:
ProductivityVSQuantity of substance

Solution Approach 1:

The patent introduces a CoLor component as an intermediary layer between the processor and memory subsystem. This component implements address translation logic that maps lowered matrix indices to non-lowered matrix indices, allowing the system to use efficient GEMM operations while storing data in the compact non-lowered format. The intermediary translates addresses on-the-fly, eliminating the need to physically replicate data K^2 times in memory.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent changes the address parameter representation from lowered matrix indices to non-lowered matrix indices through the CoLor component's address translation logic. By transforming how addresses are represented and interpreted, the system maintains computational efficiency while reducing memory footprint. The parameter change occurs in the address space rather than the physical data layout.

Inventive Principle:
Principle #35Parameter changes

2Speed

If lowering process is used to enable GEMM operations, then processing speed is improved, but bandwidth requirements increase significantly

Engineering Contradiction:
Improveprocessing speedVSAvoidbandwidth requirements
Core Design Contradiction:
SpeedVSQuantity of substance

Solution Approach 1:

The CoLor component acts as an intermediary that intercepts and translates memory address requests between the processor and memory subsystem. By translating lowered indices to non-lowered indices at the intermediary level, the system reduces redundant data transfers and bandwidth consumption while maintaining fast GEMM processing speed.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

Instead of physically copying data K^2 times in memory as traditional lowering requires, the patent uses virtual copying through address translation. The CoLor component creates the effect of replicated data by translating addresses, allowing multiple logical copies without physical duplication, thereby reducing bandwidth requirements.

Inventive Principle:
Principle #26Copying

3Quantity of substance

If address translation is implemented to reduce memory footprint, then memory efficiency is improved, but system complexity increases

Engineering Contradiction:
Improvememory efficiencyVSAvoidsystem complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The patent encapsulates the address translation complexity within a dedicated CoLor intermediary component, keeping the rest of the system simple. The component implements the complex address mapping logic internally while presenting a simple interface to both the processor and memory subsystem, effectively hiding the complexity from the rest of the system.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent segments the memory management functionality by introducing a separate CoLor component dedicated to address translation. This segmentation isolates the complexity of address mapping into a manageable module, allowing the processor and memory subsystem to remain simple while achieving improved memory efficiency through the specialized translation component.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS10565285B2Processor and memory transparent convolutional lowering and auto zero padding for deep neural network implementations
Publication Date: 2020.02.18 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US10565285B2 patent drawing
  • US10565285B2 patent drawing
  • US10565285B2 patent drawing

AI summary

A convolutional lowering component (CoLor component) between processor and memory units (or within a memory hierarchy) maps location in a lowered matrix to an equivalent location in a non-lowered matrix and provides auto zero padding in computational heavy convolutional layers. An identification component identifies processing components that execute computations in deep neural networks (DNNs) in which convolutions are realized as general matrix to matrix multiplications (GEMM) operations, and identifies a subset of the processing components that store deep neural network (DNN) features in a non-lowered form component that determines output for successively larger neural networks of a set. An address translation component translates address requests, generated by the subset of processing components to a memory subsystem, from a lowered index form to a non-lowered index form.