Electronic Document Clustering Using Multi-Attribute Similarity
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing classification systems for electronic documents often misclassify documents, lack sufficient granularity, and are burdensome to manage, leading to difficulties in researchers finding relevant documents.
Innovation Solution
A method that compares electronic documents to form pairs, calculates similarity values based on attributes like citations, text, authors, and institutions, and applies clustering algorithms to create hierarchical clusters, allowing for reclassification and more accurate organization.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional classification systems are used to organize electronic documents, then documents can be categorized using existing classification codes, but documents are often misclassified and researchers cannot find relevant documents
Solution Approach 1:
The patent replaces manual classification systems with automated text mining and natural language processing techniques. The system automatically extracts keywords, analyzes document content, and generates classification tags without human intervention, thereby improving classification accuracy and preventing document misclassification
Solution Approach 2:
The patent introduces multiple parameters for document classification including keyword frequency, term weight, and relevance scoring. By changing from simple classification codes to multi-parameter analysis, the system achieves more precise document categorization and improves the ability to retrieve relevant documents
2Adaptability or versatility
If classification systems provide broad categorization, then documents can be organized into general groups, but the systems lack sufficient granularity to be beneficial to researchers
Solution Approach 1:
The patent segments documents into multiple hierarchical classification levels and assigns multiple tags to each document. This segmentation allows researchers to navigate from broad categories to specific topics, providing the needed granularity without creating a single overly complex classification structure
Solution Approach 2:
The patent creates a multi-functional classification system that serves both broad overview needs and detailed research needs simultaneously. The same system provides both high-level categorization for general navigation and fine-grained classification for specialized research, eliminating the need for separate classification systems
3Measurement precision
If manual creation and management of multiple sub-levels in hierarchical classification system is implemented, then sufficient granularity can be achieved, but it becomes too burdensome
Solution Approach 1:
The patent implements self-service classification where the system automatically analyzes document content, extracts relevant keywords, and generates classification tags without human intervention. This eliminates the burden of manual classification management while maintaining high granularity through automated text mining and natural language processing
Solution Approach 2:
The patent replaces manual classification operations with automated computational processes including text mining, keyword extraction, and natural language processing. This substitution eliminates the need for human classifiers to manually create and manage sub-levels, making the system easy to operate while achieving sufficient granularity
Data Source
AI summary
Methods of organizing documents by reclassification and clustering are disclosed. In one embodiment, a method of clustering electronic documents of a document corpus includes comparing, by a computer, each individual electronic document in the document corpus with each other electronic document in the document corpus, thereby forming document pairs. The electronic documents of the document pairs are compared by calculating a similarity value with respect to the electronic documents of a document pair, associating the similarity value with both electronic documents of the document pair, and applying a clustering algorithm to the document corpus using the similarity values to create a plurality of hierarchical clusters. The similarity value is based on a plurality of attributes of the electronic documents in the document corpus. The plurality of attributes includes a citation attribute, a text-based attribute and one or more of an author-attribute, a publication-attribute, an institution-attribute, a downloads-attribute, and a clustering-results-attribute.


