Multi-document clustering via co-occurrence density analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Users face an overwhelming amount of information that needs to be filtered, processed, and analyzed, with existing sentiment filtering approaches failing to provide context-aware results, leading to difficulties in retrieving relevant content and clustering documents that discuss the same concepts or interests, especially on constrained devices.

Innovation Solution

A method involving a microprocessor-based process to receive and analyze content, extract themes, determine co-occurrence densities, select seed terms, and create cohesive clusters by removing items with low saliency, allowing for the presentation of salient text and automatic summarization for improved content processing and filtering.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional sentiment filtering approaches are used to analyze content, then processing speed is maintained, but context-awareness and accuracy of sentiment analysis deteriorate

Engineering Contradiction:
Improvesentiment analysis accuracyVSAvoidprocessing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the document collection into multiple clusters based on thematic coherence. Each cluster represents a distinct topic or subject area, allowing sentiment analysis to be performed contextually within each segment rather than uniformly across all documents. This segmentation enables more accurate sentiment detection by considering the specific context of each cluster while maintaining manageable processing complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent changes the parameter of analysis from simple keyword matching to co-occurrence density calculations. By computing how frequently themes co-occur within documents and clusters, the system transforms the analysis approach to capture contextual relationships. This parameter change improves sentiment analysis accuracy by considering thematic context while the automated calculation process keeps processing complexity manageable.

Inventive Principle:
Principle #35Parameter changes

2Ease of operation

If all retrieved content is presented to users, then completeness of information is maintained, but ease of accessing relevant information deteriorates

Engineering Contradiction:
Improveinformation accessibilityVSAvoidinformation completeness
Core Design Contradiction:
Ease of operationVSLoss of information

Solution Approach 1:

The patent extracts and presents only the most salient information from each document cluster. By identifying and highlighting key themes and representative documents within each cluster, the system enables users to access relevant information efficiently without being overwhelmed by the complete document collection. This extraction approach maintains information completeness at the cluster level while improving individual document accessibility.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent applies partial action by presenting a curated subset of documents from each cluster rather than the entire collection. Users receive a manageable portion of information that captures the essential content of each theme, making information accessible without presenting all possible documents. This partial presentation balances accessibility with information completeness.

Inventive Principle:
Principle #16Partial or excessive action

3Manufacturing precision

If simple keyword-based clustering is used, then processing speed is maintained, but manufacturing precision of document clustering deteriorates

Engineering Contradiction:
Improvedocument clustering accuracyVSAvoidprocessing speed
Core Design Contradiction:
Manufacturing precisionVSProductivity

Solution Approach 1:

The patent changes the clustering parameter from simple keyword matching to co-occurrence density calculations. By computing how frequently themes co-occur together within documents, the system achieves more accurate document clustering that reflects actual thematic relationships. The automated computation of co-occurrence densities maintains processing speed while significantly improving clustering precision compared to traditional keyword-based approaches.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent substitutes manual or simple keyword-based clustering mechanisms with an automated co-occurrence analysis system. The microprocessor-based computation of theme co-occurrence densities replaces simpler mechanical clustering methods, achieving higher precision in document grouping while maintaining productivity through automated processing.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS9600470B2Method and system relating to re-labelling multi-document clusters
Publication Date: 2017.03.21 WHYZ TECH
  • US9600470B2 patent drawing
  • US9600470B2 patent drawing
  • US9600470B2 patent drawing

AI summary

Individuals receive overwhelming barrage of information which must be filtered, processed, analyzed, reviewed, consolidated and distributed or acted upon. However, prior art tools for automatically processing content, such as for example returning search results from an Internet or database search for example are ineffective. Prior art search techniques merely provide large numbers of “hits” with at most removal of multiple occurrences of identical items. However, it would be beneficial to present searches as a series of multi-document clusters wherein occurrences of commonly themed content are clustered allowing the user to rapidly see the number of different themes and review a selected theme. Further, it would be beneficial, in repeated searches, for new clusters to be identified automatically as well as new items of content associated with existing clusters to be associated to these clusters.