Mask Load Store Logic for Vector Data Exceptions

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current processor architectures face inefficiencies in performing mask load and store operations due to their complexity and the need for multiple instructions, leading to increased processing cycles and power consumption, especially in handling misaligned data and page faults.

Innovation Solution

The introduction of multiple flavors of mask load and store instructions that enable conditional SIMD packed data operations, allowing for efficient loading and storing of packed data based on mask values, and the use of speculative full-width loads to optimize performance while handling exceptions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If mask load and store operations are implemented using traditional block load methods, then memory operations can be performed in blocks, but it becomes challenging to support mask operations at reasonable performance and causes unnecessary power consumption

Engineering Contradiction:
Improvemask operation performanceVSAvoidpower consumption
Core Design Contradiction:
ProductivityVSUse of energy by moving object

Solution Approach 1:

The patent divides the mask load/store operation into two independent phases: (1) a full block load/store operation that transfers all data elements in the memory block, and (2) a separate mask application phase that selectively enables or disables individual data elements using mask bits. This segmentation allows the memory subsystem to operate efficiently on full blocks while the processor applies masks only to the necessary elements, avoiding the need to transfer and process entire blocks when only some elements are needed.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs the memory block load/store operation in advance without waiting for mask evaluation, allowing the memory subsystem to operate at full speed. The mask application is then applied afterward to selectively enable or disable data elements. This preliminary action eliminates the need to wait for mask evaluation before initiating memory operations, thereby improving overall performance while avoiding unnecessary power consumption from conditional processing.

Inventive Principle:
Principle #10Preliminary action

2Adaptability or versatility

If traditional block loads are used for SIMD operations, then memory can be accessed in blocks, but mask operations cannot be efficiently supported as loads are done without reference to mask

Engineering Contradiction:
Improvemask operation supportVSAvoidoperation throughput
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The patent separates the memory access operation from the mask application operation. The block load/store executes first to transfer all data elements to the processor, and then the mask application phase selectively enables or disables individual elements based on mask bits. This segmentation allows the memory subsystem to operate efficiently on full blocks while the processor applies masks only to the necessary elements.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a mask application phase as an intermediary step between the block load operation and the actual data usage. This intermediary phase evaluates the mask bits and selectively enables or disables data elements after the block load is complete, allowing the system to maintain high memory throughput while still supporting conditional mask operations.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Adaptability or versatility

If mask loads are performed on data spanning multiple pages, then comprehensive data access is enabled, but page faults occur when pages are not present

Engineering Contradiction:
Improvemisaligned data handlingVSAvoidpage fault handling
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent performs the full block load operation before evaluating mask bits or handling page faults. The memory subsystem attempts to load the entire block including any data spanning multiple pages, and only after this preliminary load attempt does the system evaluate which elements should actually be processed based on the mask. This approach allows the system to handle misaligned data access while concentrating page fault handling at a specific point in the operation.

Inventive Principle:
Principle #10Preliminary action

4Adaptability or versatility

If multiple instructions are used to perform mathematical operations on operands, then complex operations can be executed, but throughput is diminished and clock cycles increase

Engineering Contradiction:
Improvecomplex operation capabilityVSAvoidinstruction throughput
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The patent combines the mask evaluation and the data selection into a single unified operation. Rather than using separate instructions to evaluate masks and then selectively process data elements, the system performs both functions simultaneously during the mask application phase, where mask bits directly control which data elements are enabled or disabled in a single processing step.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS10120684B2Instructions and logic to perform mask load and store operations as sequential or one-at-a-time operations after exceptions and for un-cacheable type memory
Publication Date: 2018.11.06 TAHOE RES LTD
  • US10120684B2 patent drawing
  • US10120684B2 patent drawing
  • US10120684B2 patent drawing

AI summary

Logic is provided to receive and execute a mask move instruction to transfer unmasked data elements of a vector data element including a plurality of packed data elements from a source location to a destination location, subject to mask information for the instruction. The logic is to execute a speculative full width operation, and if an exception occurs is to perform operations sequentially or one at a time. Other embodiments are described and claimed.