Shared L1 Cache Crossbar for Read-Write Data Transmission

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current computing architectures that use dedicated data crossbars for reads and writes in L1 cache memory are inefficient and costly due to the significant on-chip die space required, as they separate read and write operations.

Innovation Solution

A method where data requests are processed by determining instruction types and addresses to identify subsets for simultaneous processing, allowing read and write data to be transmitted on the same crossbar unit, thereby reducing the need for separate crossbar units and optimizing on-chip space.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If separate data crossbars are used for reads and writes in L1 cache, then data transmission reliability is improved, but on-chip die space consumption increases significantly

Engineering Contradiction:
Improvedata transmission reliabilityVSAvoidon-chip die space
Core Design Contradiction:
ReliabilityVSArea of stationary object

Solution Approach 1:

The patent merges separate read and write data crossbars into a single shared crossbar unit. The crossbar is configured to handle both read operations (transmitting data from L1 cache to clients) and write operations (transmitting data from clients to L1 cache) using the same physical infrastructure, thereby reducing on-chip die space while maintaining functional separation through logical control mechanisms.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The shared crossbar unit is designed with universal functionality to perform multiple roles: it can transmit read data, transmit write data, and switch between these functions dynamically. This multi-functional design eliminates the need for dedicated separate crossbars while preserving the reliability of data transmission through proper arbitration and control logic.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Productivity

If dedicated crossbars are implemented for read and write operations, then data transmission efficiency is maintained, but manufacturing cost increases

Engineering Contradiction:
Improvedata transmission efficiencyVSAvoidmanufacturing cost
Core Design Contradiction:
ProductivityVSEase of manufacture

Solution Approach 1:

By combining read and write crossbar functions into a single shared unit, the patent reduces the total number of crossbar components that need to be manufactured and integrated. This consolidation directly lowers manufacturing complexity and cost while the internal architecture maintains efficient data transmission through dedicated transmission paths within the shared crossbar.

Inventive Principle:
Principle #5Merging (Combining)

3Reliability

If separate crossbars are used for reads and writes, then operational reliability is improved, but device complexity increases

Engineering Contradiction:
Improveoperational reliabilityVSAvoidcrossbar architecture complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent reduces device complexity by merging separate crossbar architectures into a single shared crossbar unit. While the physical structure is simplified, operational reliability is maintained through logical separation mechanisms including request arbitration, operation tagging, and control logic that ensures read and write operations are handled correctly without interference.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS9286256B2Sharing data crossbar for reads and writes in a data cache
Publication Date: 2016.03.15 NVIDIA CORP
  • US9286256B2 patent drawing
  • US9286256B2 patent drawing
  • US9286256B2 patent drawing

AI summary

The invention sets forth an L1 cache architecture that includes a crossbar unit configured to transmit data associated with both read data requests and write data requests. Data associated with read data requests is retrieved from a cache memory and transmitted to the client subsystems. Similarly, data associated with write data requests is transmitted from the client subsystems to the cache memory. To allow for the transmission of both read and write data on the crossbar unit, an arbiter is configured to schedule the crossbar unit transmissions as well and arbitrate between data requests received from the client subsystems.