Cross-Architecture Kernel Data Recovery for OS Failure Collection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Storage systems crash due to software or hardware defects, leading to unresponsive shutdowns during kernel dump processing, which prevents data collection and classification, and requires extensive recovery time, affecting user experience and system availability.
Innovation Solution
A method for data processing that synchronizes data between operating systems with different architectures, allowing instant data collection during system failures by using a second operating system to continue data collection when the first operating system fails, utilizing a data collector to transfer data between a first and second operating system.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If data collection is performed during kernel dump processing, then data can be collected, but the operating system becomes unresponsive and shutdown is prevented
Solution Approach 1:
The patent introduces a second operating system as an intermediary to perform data collection when the first operating system is unresponsive. The data collector in the second operating system can still access and collect data from the storage system even when the first operating system is frozen during kernel dump processing, thus ensuring data collection continuity without affecting system responsiveness.
Solution Approach 2:
The patent creates a copy of the operating system (second operating system) that can independently perform data collection functions. This copy includes a data collector that mirrors the functionality of the first operating system's data collector, enabling redundant data collection capabilities that activate when the primary system becomes unresponsive.
2Reliability
If the operating system remains unresponsive during shutdown, then data collection can be performed, but system recovery time increases
Solution Approach 1:
The second operating system is pre-configured with data collection capabilities before the first operating system fails. When failure occurs, the second operating system can immediately begin data collection without waiting for the first system to become responsive, performing preliminary actions that reduce overall recovery time while ensuring data collection completion.
Solution Approach 2:
The patent enables continuous data collection by transitioning from the first operating system to the second operating system. The data collection process does not interrupt when the first system fails; instead, it seamlessly continues on the second system, maintaining useful action continuity and reducing the perceived recovery time from the perspective of data collection operations.
3Reliability
If a single operating system performs all data collection, then system resources are conserved, but data collection fails during system failures
Solution Approach 1:
The patent assigns different roles to different operating systems based on their operational state. The first operating system handles normal data collection operations, while the second operating system is specifically activated for data collection during failure conditions. This local quality differentiation ensures data collection success under various conditions without requiring both systems to maintain full functionality simultaneously.
Solution Approach 2:
The system changes the operational parameter of data collection availability by switching between operating systems. When the first operating system is healthy, it handles data collection with normal resource allocation. When it fails, the parameter changes to activate the second operating system's data collection capabilities, adjusting the system state to maintain data collection success rate despite increased architectural complexity.
Data Source
AI summary
Techniques perform data processing. Such techniques involve obtaining, by a first operating system, first data of a storage system from a data collector. Such techniques further involve synchronizing the first data from the first operating system to a second operating system, wherein an architecture of the second operating system is different from that of the first operating system. Such techniques further involve, in response to that the first operating system has a failure, obtaining, by the second operating system, second data of the storage system from the data collector. Such techniques further involve, in response to that the first operating system is recovered, synchronizing the second data from the second operating system to the first operating system. Accordingly, such techniques can minimize processing time, ensure high availability of data collection, save computing and memory resources of a storage system, and help improve user experience.


