Hadoop Distributed Data Processing for Memory Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for analyzing data between a memory controller and a DRAM device are inefficient due to the short data collection period of logic analyzers and the time-consuming processing of large-capacity data using existing software tools, which results in fragmentary analysis and prolonged processing times.
Innovation Solution
A data processing system utilizing the High-Availability Distributed Object-Oriented Platform (HADOOP) framework for parallel processing of large-capacity data, where data is collected and split into block-based files, stored in a distributed file system, and processed using a Map-Reduce engine, enabling distributed and parallel processing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If existing software tools are used to process large-capacity data, then data processing can be performed, but the processing time becomes excessively long
Solution Approach 1:
The patent divides large-capacity data into multiple block-based files and distributes them across a Hadoop distributed file system. The data processing is segmented into map and reduce tasks that can be executed in parallel across multiple nodes, thereby significantly reducing the overall processing time compared to sequential processing with existing tools.
Solution Approach 2:
The patent combines multiple processing nodes into a distributed system that works collectively on the same data set. By merging computational resources across multiple machines, the system achieves higher throughput and faster processing of large-capacity data than single-machine existing tools.
2Measurement precision
If a logic analyzer with short data collection section is used, then data can be collected and analyzed, but the analysis becomes fragmentary
Solution Approach 1:
The patent segments the data collection process by using multiple data collecting devices that can operate simultaneously over extended periods. Each device collects data for a predetermined period, and the results are aggregated to form a complete picture, avoiding the fragmentary analysis caused by single-device short-duration collection.
Solution Approach 2:
The patent enables continuous data collection over extended periods by distributing the collection task across multiple devices that can run concurrently. This continuous collection ensures no data is missed and provides comprehensive analysis coverage, unlike short-duration single-device collection.
3Quantity of substance
If data is collected for extended periods with multiple devices, then comprehensive data is obtained, but data storage and management complexity increases
Solution Approach 1:
The patent introduces a Hadoop distributed file system as an intermediary layer between data collecting devices and processing devices. This intermediary automatically manages the storage, organization, and distribution of large volumes of data across the network, simplifying data management complexity while enabling extended data collection capacity.
Solution Approach 2:
The Hadoop ecosystem provides universal functionality that handles multiple tasks including data storage, processing, and management in a unified framework. This multi-functional platform reduces overall system complexity by consolidating what would otherwise require separate specialized systems for each function.
Data Source
AI summary
A data processing system includes: a memory device suitable for performing an operation corresponding to a command and outputting a memory data; a data collecting device suitable for collecting big data by integrating the command and the memory data at a predetermined cycle or at every predetermined time, splitting the collected big data based on a predetermined unit, and transferring the split big data; and a data processing device suitable for storing the split big data received from the data collecting device in block-based files in a High-Availability Distributed Object-Oriented Platform (HADOOP) distributed file system (HDFS), classifying the block-based files based on a particular memory command, and processing the block-based files.


