Multi-document clustering via co-occurrence density analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Users face an overwhelming amount of information that needs to be filtered, processed, and analyzed, with existing sentiment filtering approaches failing to provide context-aware results, leading to difficulties in retrieving relevant content and clustering documents that discuss the same concepts or interests, especially on constrained devices.
Innovation Solution
A method involving a microprocessor-based process to receive and analyze content, extract themes, determine co-occurrence densities, select seed terms, and create cohesive clusters by removing items with low saliency, allowing for the presentation of salient text and automatic summarization for improved content processing and filtering.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional sentiment filtering approaches are used to analyze content, then processing speed is maintained, but context-awareness and accuracy of sentiment analysis deteriorate
Solution Approach 1:
The patent segments the document collection into multiple clusters based on thematic coherence. Each cluster represents a distinct topic or subject area, allowing sentiment analysis to be performed contextually within each segment rather than uniformly across all documents. This segmentation enables more accurate sentiment detection by considering the specific context of each cluster while maintaining manageable processing complexity.
Solution Approach 2:
The patent changes the parameter of analysis from simple keyword matching to co-occurrence density calculations. By computing how frequently themes co-occur within documents and clusters, the system transforms the analysis approach to capture contextual relationships. This parameter change improves sentiment analysis accuracy by considering thematic context while the automated calculation process keeps processing complexity manageable.
2Ease of operation
If all retrieved content is presented to users, then completeness of information is maintained, but ease of accessing relevant information deteriorates
Solution Approach 1:
The patent extracts and presents only the most salient information from each document cluster. By identifying and highlighting key themes and representative documents within each cluster, the system enables users to access relevant information efficiently without being overwhelmed by the complete document collection. This extraction approach maintains information completeness at the cluster level while improving individual document accessibility.
Solution Approach 2:
The patent applies partial action by presenting a curated subset of documents from each cluster rather than the entire collection. Users receive a manageable portion of information that captures the essential content of each theme, making information accessible without presenting all possible documents. This partial presentation balances accessibility with information completeness.
3Manufacturing precision
If simple keyword-based clustering is used, then processing speed is maintained, but manufacturing precision of document clustering deteriorates
Solution Approach 1:
The patent changes the clustering parameter from simple keyword matching to co-occurrence density calculations. By computing how frequently themes co-occur together within documents, the system achieves more accurate document clustering that reflects actual thematic relationships. The automated computation of co-occurrence densities maintains processing speed while significantly improving clustering precision compared to traditional keyword-based approaches.
Solution Approach 2:
The patent substitutes manual or simple keyword-based clustering mechanisms with an automated co-occurrence analysis system. The microprocessor-based computation of theme co-occurrence densities replaces simpler mechanical clustering methods, achieving higher precision in document grouping while maintaining productivity through automated processing.
Data Source
AI summary
Individuals receive overwhelming barrage of information which must be filtered, processed, analyzed, reviewed, consolidated and distributed or acted upon. However, prior art tools for automatically processing content, such as for example returning search results from an Internet or database search for example are ineffective. Prior art search techniques merely provide large numbers of “hits” with at most removal of multiple occurrences of identical items. However, it would be beneficial to present searches as a series of multi-document clusters wherein occurrences of commonly themed content are clustered allowing the user to rapidly see the number of different themes and review a selected theme. Further, it would be beneficial, in repeated searches, for new clusters to be identified automatically as well as new items of content associated with existing clusters to be associated to these clusters.


