Automatic Category Tree Creation for Unstructured Data Stocks
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data processing systems struggle to provide an effective overview of the contents of data stocks, particularly non-structured data stocks, which are difficult to navigate and understand.
Innovation Solution
A method for automatically creating a category tree by filtering out stop words, calculating significance values, sorting and reducing lists of words, detecting co-occurrences, and iteratively building levels of the category tree, utilizing an index and database to enhance search efficiency and visualization.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If mechanical processing of contents as data is used, then data can be stored and retrieved, but the overview and understanding of non-structured data stocks becomes difficult
Solution Approach 1:
The patent segments the data stock into hierarchical categories through automatic category tree creation. The system divides unstructured data into organized groups based on co-occurrence analysis, creating a multi-level structure that enables both efficient processing and easy user navigation through the segmented categories.
Solution Approach 2:
The patent introduces an intermediary category tree structure between the raw data and the user interface. This intermediary representation transforms unstructured data into an organized hierarchical form that maintains processing efficiency while significantly improving user understanding and navigation capabilities.
2Reliability
If comprehensive indexing of all words is performed, then search coverage is improved, but processing time and computational resources increase
Solution Approach 1:
The patent extracts only the most significant words and co-occurrence patterns from the data stock, rather than indexing all words. By identifying and retaining only the key terms that define category structures, the system achieves comprehensive search coverage for meaningful queries while dramatically reducing processing time and computational resource requirements.
Solution Approach 2:
The patent changes the parameter of word selection from including all words to including only significant words based on co-occurrence frequency and relevance metrics. This parameter change maintains search reliability for important terms while reducing the overall data volume that requires processing and storage.
3Manufacturing precision
If manual organization of data categories is performed, then accuracy of category structure is improved, but labor requirements and cost increase
Solution Approach 1:
The patent enables the system to automatically organize data categories through self-service mechanisms. The category tree is generated autonomously by analyzing co-occurrence patterns in the data, eliminating the need for manual intervention while maintaining high accuracy in the category structure through algorithmic identification of meaningful relationships.
Solution Approach 2:
The patent replaces the mechanical manual process of category organization with an automated computational system. The manual labor of analyzing and categorizing data is substituted by algorithmic processes that detect co-occurrences and generate hierarchical structures automatically, reducing system complexity in terms of human operation while maintaining precision.
4Loss of information
If detailed analysis of all data contents is performed, then completeness of information is improved, but system complexity and processing load increase
Solution Approach 1:
The patent applies partial action by analyzing only the necessary co-occurrence patterns required to build the category tree, rather than performing exhaustive analysis of all data contents. This approach maintains information completeness for categorization purposes while avoiding unnecessary processing that would increase system complexity and processing load.
Data Source
AI summary
Methods for the automatic creation of a category tree with respect to the contents of a data stock, wherein a taxonomy of the data stock will be created on the base of co-occurrences. Another object of the present invention is furthermore a data processing system comprising data which represent information in at least one data stock which is accessible via at least one data source, which is designed and/or adapted to at least partially carry out a method according to the invention. Another object of the present invention is furthermore a data processing device for the electronic processing of data, comprising a control and/or computer unit, an input unit and an output unit, which is designed and/or adapted to at least partially carry out a method according to the invention, preferably using at least a part of a data processing system according to the invention.


