Categorized Document Base Multi-Dimensional Classification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current statistical natural language processing methods do not adequately support access to information in text documents, particularly in automatically categorizing and retrieving documents across multiple classification systems, limiting the ability to provide comprehensive information for decision-making.
Innovation Solution
The method involves using pre-existing classification schemes to automatically assign documents to categories through Information Retrieval techniques, generating numerical scores based on document composition, and searching documents according to various criteria to create a categorized document base that can be analyzed and presented in an array format, allowing for multiple dimensions of classification and search results.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If automated classification using search terms and thesauruses is implemented, then document categorization efficiency is improved, but the depth and quality of classification information is reduced
Solution Approach 1:
The patent segments the classification process into multiple independent analysis dimensions. Instead of assigning a single category label, the system generates separate classification results across multiple taxonomic systems (e.g., subject classification, document type classification, source classification), allowing each dimension to be processed independently while preserving comprehensive classification information.
Solution Approach 2:
The patent transitions from traditional single-dimension classification to multi-dimensional classification by organizing documents across multiple independent taxonomic systems simultaneously. Each taxonomic system represents a different dimension of classification, and the system provides tools to navigate and analyze documents along any of these dimensions independently or in combination.
2Adaptability or versatility
If multiple classification systems are applied to documents, then information access comprehensiveness is improved, but system complexity increases
Solution Approach 1:
The patent creates a universal classification framework that can accommodate multiple independent taxonomic systems within a single platform. The system is designed to import, store, and process documents according to different classification schemes (subject, type, source, etc.) using a common infrastructure, allowing the same system to serve multiple classification purposes without requiring separate systems for each taxonomic approach.
Solution Approach 2:
The patent implements a nested structure where multiple classification systems are organized hierarchically within a unified framework. Individual taxonomic systems can be nested within the overall classification platform, with each system maintaining its own structure while being accessible through the universal interface. This allows complex multi-system classification to be managed through layered organization.
3Speed
If automated assessment using IR techniques is used, then classification speed is improved, but measurement precision of document-category matching is reduced
Solution Approach 1:
The patent introduces classification experts as an intermediary layer between automated IR techniques and final classification decisions. The system uses automated IR methods to generate initial classification candidates and scores, then applies expert judgment to review and refine these assignments, particularly for complex or ambiguous cases. This intermediary step bridges the gap between fast automated processing and high-precision expert classification.
Data Source
AI summary
A method of managing information comprises generating a categorized document base. Generating the document base comprises providing a pre-existing classification of things other than documents, providing a source collection of documents, and automatically assessing the documents using Information Retrieval techniques to assign at least some of the documents to one or more taxa of the classification. For each taxon in the classification one or more numerical scores are assigned, based at least in part on a composition, makeup or constitution of the documents assigned to the taxon of the categorized document base.


