Pooled Memory Architecture for Cross-Server Error Logging
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Data centers face resource waste and performance degradation due to uneven storage occupancy across servers, and existing solutions for pooled memory either require additional memory, increase costs, or suffer from network latency and blast radius issues.
Innovation Solution
A data processing system with a pooled memory architecture that allows servers in the same rack to share memory spaces using an interconnection circuit, enabling memory error detection and handling across machines, thus achieving a low-cost, high-performance shared memory pool with reduced latency and blast radius.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If servers use isolated memory spaces, then memory reliability is maintained, but memory resource utilization deteriorates due to uneven storage occupancy
Solution Approach 1:
The patent merges memory spaces from multiple servers into a pooled memory architecture where memory resources are shared across servers. The memory pool manager combines available memory from participating servers into a unified resource pool that can be dynamically allocated to any server needing memory, thereby improving overall memory utilization while maintaining reliability through centralized management and error logging mechanisms.
2Quantity of substance
If additional memory is added to balance storage occupancy, then storage capacity is improved, but system cost deteriorates
Solution Approach 1:
Instead of adding memory to individual servers, the patent combines existing memory resources across multiple servers into a shared pool. This allows the system to achieve balanced storage capacity utilization without purchasing additional memory hardware, thereby avoiding increased system costs while still providing adequate storage capacity to all servers.
Solution Approach 2:
The pooled memory architecture makes existing memory resources universal and multi-functional, allowing the same physical memory to serve multiple servers dynamically. This eliminates the need for each server to have dedicated memory capacity, reducing overall system cost while maintaining adequate storage capacity for all workloads.
3Quantity of substance
If memory is pooled across servers, then memory utilization is improved, but network latency increases
Solution Approach 1:
The patent introduces a memory pool manager as an intermediary layer that sits between servers and the pooled memory resources. This manager handles memory allocation, error logging, and resource coordination locally, reducing the need for frequent network communications and thereby minimizing network latency while maintaining high memory utilization through efficient resource sharing.
4Reliability
If error logs are shared across servers, then memory reliability is improved, but system complexity increases
Solution Approach 1:
The patent merges error logging functionality into the centralized memory pool manager, combining error detection, logging, and management operations into a single coordinated system. This approach improves memory reliability across the pooled memory while managing system complexity through unified error handling rather than distributed independent logging mechanisms.
Data Source
AI summary
A data processing system includes a first server and a second server. The first server includes a first processor group, a first memory space and a first interface circuit. The second server includes a second processor group, a second memory space and a second interface circuit. The first memory space and the second memory space are allocated to the first processor group. The first processor group is configured to perform memory error detection to generate an error log corresponding to a memory error. When the memory error occurs in the second memory space, the first interface circuit is configured to send the error log to the second interface circuit, and the second processor group is configured to log the memory error according to the error log received by the second interface circuit. The data processing system is capable of realizing memory reliability architecture supporting operations across different servers.


