Data Extraction Apparatus Using Clustering for Railroad Maintenance
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional data extraction systems struggle with efficient data analysis when the extraction condition is indefinite, leading to missed data that could lead to new knowledge and inefficient processing of excessive data volumes.
Innovation Solution
A data extraction apparatus that includes a parameter analysis unit for morphological analysis of learning text information, a grouping settings display unit to finalize search-target data and clustering conditions, and clustering units to perform clustering based on these conditions, ensuring efficient data extraction and analysis.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If the extraction condition is set to be definite, then data analysis efficiency is improved, but data that leads to new knowledge is missed
Solution Approach 1:
The patent applies dynamics by making the data extraction condition adaptable rather than fixed. The system dynamically adjusts the extraction condition based on clustering results, transitioning from a definite initial condition to a modified condition that balances efficiency and completeness. This is achieved through the determination unit that modifies the extraction condition according to clustering outcomes, allowing the system to adapt between extracting only necessary data and extracting additional data that may lead to new knowledge.
2Loss of information
If the extraction condition is set to be indefinite, then data that leads to new knowledge is preserved, but data analysis efficiency deteriorates due to increased data volume
Solution Approach 1:
The patent applies partial or excessive action by initially extracting more data than strictly necessary (excessive action) to ensure no potential new knowledge is missed. The system then uses clustering to identify and focus on the most relevant portions of this expanded dataset, effectively performing partial action on the subset of data that is most likely to yield new insights while maintaining the option to explore additional data if needed.
3Reliability
If the extraction condition is indefinite, then completeness of data is improved, but processing time increases due to excessive data volume
Solution Approach 1:
The patent applies segmentation by dividing the data extraction and processing task into distinct stages: initial extraction with a first condition, clustering analysis, determination of a second extraction condition based on clustering results, and subsequent extraction. This segmentation allows the system to process data in manageable portions rather than handling all possible data at once, reducing overall processing time while maintaining completeness through the multi-stage approach.
4Productivity
If the extraction condition is definite, then processing efficiency is improved, but adaptability to different analysis needs deteriorates
Solution Approach 1:
The patent applies universality by creating a system that can serve multiple functions: it can extract data efficiently when the extraction condition is definite, and it can also adapt to extract more comprehensive data when clustering results indicate potential for new knowledge. The determination unit acts as a universal component that decides between these different modes of operation based on the specific analysis needs, making the system versatile rather than specialized for a single approach.
Data Source
AI summary
A data extraction apparatus includes a parameter analysis unit that performs analysis of learning text information, extracts words that serve as machine learning parameters, and classifies the words into types of parameters; a grouping settings display unit that finalizes search-target data and clustering conditions based on the parameters; at least one clustering training data extraction unit that extracts training data from a database based on the search-target data and the clustering conditions; at least one clustering unit that performs clustering based on the clustering condition on the training data; an applicable-clustering determination unit that performs analysis of search text information and identifies search-target data serving as a narrowing-down condition and which clustering unit is to be operated; and a search range specification unit that causes the clustering unit to operate and extracts a narrowed range of search-target data from the database based on an operation result.


