Cache Memory Partitioning for Latency-Critical Application QoS

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In cloud computing environments, applications sensitive to cache and memory bandwidth interference experience performance degradation due to bottlenecks in cache space and memory bandwidth, which affects quality of service (QoS) for latency-critical services.

Innovation Solution

A method for dynamically partitioning cache memory and memory channels, combined with prioritization of memory controller requests, to allocate resources effectively and reduce interference, ensuring sufficient cache space and memory bandwidth for latency-critical applications.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If multiple applications share cache memory and memory channels, then resource utilization increases, but performance degradation occurs due to interference and bottlenecks

Engineering Contradiction:
Improveresource utilizationVSAvoidapplication performance
Core Design Contradiction:
Quantity of substanceVSProductivity

Solution Approach 1:

The patent divides cache memory into multiple partitions and memory channels into separate groups, assigning specific partitions to specific applications. This segmentation prevents interference between applications while maintaining high resource utilization, as each application has dedicated cache space and memory bandwidth without excluding the ability to share the overall system resources.

Inventive Principle:
Principle #1Segmentation

2Reliability

If cache memory is dynamically partitioned to reduce interference, then quality of service for latency-critical applications improves, but system complexity increases

Engineering Contradiction:
Improvequality of serviceVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent implements dynamic partitioning where cache memory allocations and memory channel assignments can be adjusted in real-time based on application requirements and system conditions. This dynamic approach allows the system to adapt to changing workloads while maintaining QoS guarantees, with the complexity managed through automated control logic rather than static rigid configurations.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system incorporates monitoring and control mechanisms that track cache usage patterns, memory bandwidth utilization, and application performance metrics. This feedback information is used to dynamically adjust partition sizes and channel assignments, ensuring QoS requirements are met while automatically managing the complexity of resource coordination without requiring manual intervention.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS10318425B2Coordination of cache and memory reservation
Publication Date: 2019.06.11 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US10318425B2 patent drawing
  • US10318425B2 patent drawing
  • US10318425B2 patent drawing

AI summary

A method for coordinating cache and memory reservation in a computerized system includes identifying at least one running application, recognizing the at least one application as a latency-critical application, monitoring information associated with a current cache access rate and a required memory bandwidth of the at least one application, allocating a cache partition, a size of the cache partition corresponds to the cache access rate and the required memory bandwidth of the at least one application, defining a threshold value including a number of cache misses per time unit, determining a reduction of cache misses per time unit, in response to the reduction of cache misses per time unit being above the threshold value, retaining the cache partition, assigning a priority of scheduling memory request including a medium priority level, and assigning a memory channel to the at least one application to avoid memory channel contention.