Automatic Category Tree Creation for Unstructured Data Stocks

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data processing systems struggle to provide an effective overview of the contents of data stocks, particularly non-structured data stocks, which are difficult to navigate and understand.

Innovation Solution

A method for automatically creating a category tree by filtering out stop words, calculating significance values, sorting and reducing lists of words, detecting co-occurrences, and iteratively building levels of the category tree, utilizing an index and database to enhance search efficiency and visualization.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If mechanical processing of contents as data is used, then data can be stored and retrieved, but the overview and understanding of non-structured data stocks becomes difficult

Engineering Contradiction:
Improvedata processing efficiencyVSAvoiduser understanding of data contents
Core Design Contradiction:
ProductivityVSEase of operation

Solution Approach 1:

The patent segments the data stock into hierarchical categories through automatic category tree creation. The system divides unstructured data into organized groups based on co-occurrence analysis, creating a multi-level structure that enables both efficient processing and easy user navigation through the segmented categories.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary category tree structure between the raw data and the user interface. This intermediary representation transforms unstructured data into an organized hierarchical form that maintains processing efficiency while significantly improving user understanding and navigation capabilities.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If comprehensive indexing of all words is performed, then search coverage is improved, but processing time and computational resources increase

Engineering Contradiction:
Improvesearch coverageVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent extracts only the most significant words and co-occurrence patterns from the data stock, rather than indexing all words. By identifying and retaining only the key terms that define category structures, the system achieves comprehensive search coverage for meaningful queries while dramatically reducing processing time and computational resource requirements.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent changes the parameter of word selection from including all words to including only significant words based on co-occurrence frequency and relevance metrics. This parameter change maintains search reliability for important terms while reducing the overall data volume that requires processing and storage.

Inventive Principle:
Principle #35Parameter changes

3Manufacturing precision

If manual organization of data categories is performed, then accuracy of category structure is improved, but labor requirements and cost increase

Engineering Contradiction:
Improvecategory structure accuracyVSAvoidsystem complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent enables the system to automatically organize data categories through self-service mechanisms. The category tree is generated autonomously by analyzing co-occurrence patterns in the data, eliminating the need for manual intervention while maintaining high accuracy in the category structure through algorithmic identification of meaningful relationships.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical manual process of category organization with an automated computational system. The manual labor of analyzing and categorizing data is substituted by algorithmic processes that detect co-occurrences and generate hierarchical structures automatically, reducing system complexity in terms of human operation while maintaining precision.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

4Loss of information

If detailed analysis of all data contents is performed, then completeness of information is improved, but system complexity and processing load increase

Engineering Contradiction:
Improveinformation completenessVSAvoidsystem complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent applies partial action by analyzing only the necessary co-occurrence patterns required to build the category tree, rather than performing exhaustive analysis of all data contents. This approach maintains information completeness for categorization purposes while avoiding unnecessary processing that would increase system complexity and processing load.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS8745069B2Creation of a category tree with respect to the contents of a data stock
Publication Date: 2014.06.03 KNECON IQSER HLDG GMBH
  • US8745069B2 patent drawing
  • US8745069B2 patent drawing
  • US8745069B2 patent drawing

AI summary

Methods for the automatic creation of a category tree with respect to the contents of a data stock, wherein a taxonomy of the data stock will be created on the base of co-occurrences. Another object of the present invention is furthermore a data processing system comprising data which represent information in at least one data stock which is accessible via at least one data source, which is designed and/or adapted to at least partially carry out a method according to the invention. Another object of the present invention is furthermore a data processing device for the electronic processing of data, comprising a control and/or computer unit, an input unit and an output unit, which is designed and/or adapted to at least partially carry out a method according to the invention, preferably using at least a part of a data processing system according to the invention.