Cache Memory Partitioning for Latency-Critical Application QoS
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In cloud computing environments, applications sensitive to cache and memory bandwidth interference experience performance degradation due to bottlenecks in cache space and memory bandwidth, which affects quality of service (QoS) for latency-critical services.
Innovation Solution
A method for dynamically partitioning cache memory and memory channels, combined with prioritization of memory controller requests, to allocate resources effectively and reduce interference, ensuring sufficient cache space and memory bandwidth for latency-critical applications.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If multiple applications share cache memory and memory channels, then resource utilization increases, but performance degradation occurs due to interference and bottlenecks
Solution Approach 1:
The patent divides cache memory into multiple partitions and memory channels into separate groups, assigning specific partitions to specific applications. This segmentation prevents interference between applications while maintaining high resource utilization, as each application has dedicated cache space and memory bandwidth without excluding the ability to share the overall system resources.
2Reliability
If cache memory is dynamically partitioned to reduce interference, then quality of service for latency-critical applications improves, but system complexity increases
Solution Approach 1:
The patent implements dynamic partitioning where cache memory allocations and memory channel assignments can be adjusted in real-time based on application requirements and system conditions. This dynamic approach allows the system to adapt to changing workloads while maintaining QoS guarantees, with the complexity managed through automated control logic rather than static rigid configurations.
Solution Approach 2:
The system incorporates monitoring and control mechanisms that track cache usage patterns, memory bandwidth utilization, and application performance metrics. This feedback information is used to dynamically adjust partition sizes and channel assignments, ensuring QoS requirements are met while automatically managing the complexity of resource coordination without requiring manual intervention.
Data Source
AI summary
A method for coordinating cache and memory reservation in a computerized system includes identifying at least one running application, recognizing the at least one application as a latency-critical application, monitoring information associated with a current cache access rate and a required memory bandwidth of the at least one application, allocating a cache partition, a size of the cache partition corresponds to the cache access rate and the required memory bandwidth of the at least one application, defining a threshold value including a number of cache misses per time unit, determining a reduction of cache misses per time unit, in response to the reduction of cache misses per time unit being above the threshold value, retaining the cache partition, assigning a priority of scheduling memory request including a medium priority level, and assigning a memory channel to the at least one application to avoid memory channel contention.


