Facet Analysis for Positive Negative Classification in Document Search

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing text mining and similar document ranking technologies require users to manually identify and differentiate between similar and dissimilar documents, often resulting in the inclusion of noisy facets that do not reflect the current search context, making it time-consuming and inefficient.

Innovation Solution

A computer-implemented method that utilizes facet analysis processing to classify facets into positive and negative sets based on their relevance to the search context, highlighting these facets differently in the user interface to facilitate quicker identification of relevant documents.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual facet identification is used in existing text mining systems, then users can identify document similarities and differences, but the process becomes time-consuming and inefficient with noisy facets included

Engineering Contradiction:
Improvefacet identification accuracyVSAvoidtime for evaluating document similarities
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system automatically performs facet identification, classification, and highlighting without requiring manual user intervention. The facet analysis processing autonomously generates candidate facets, classifies them into positive and negative sets based on document subset commonality, and highlights relevant facets in the user interface, eliminating the time-consuming manual process while maintaining high accuracy

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the manual mechanical process of facet identification with automated computational processing. The system uses facet analysis processing to automatically generate, classify, and evaluate facets, substituting human manual evaluation with algorithmic processing that is both faster and more consistent, thereby reducing time loss while preserving measurement precision

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Loss of information

If all candidate facets are displayed without classification, then comprehensive information is provided, but users struggle to identify relevant documents due to noise

Engineering Contradiction:
Improveinformation completenessVSAvoidease of identifying relevant documents
Core Design Contradiction:
Loss of informationVSEase of operation

Solution Approach 1:

The patent segments the set of candidate facets into two distinct subsets: positive facets that are common across the document subset and indicate relevance, and negative facets that are not common and represent noise. This segmentation is visually represented through differential highlighting in the user interface, allowing users to quickly identify relevant documents while preserving access to comprehensive facet information

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system applies different visual qualities (highlighting) to different facets based on their relevance. Positive facets that are common across the document subset are highlighted to draw user attention, while negative facets are not highlighted. This local differentiation of quality helps users easily identify relevant documents without losing access to the complete facet information

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS11361030B2Positive/negative facet identification in similar documents to search context
Publication Date: 2022.06.14 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11361030B2 patent drawing
  • US11361030B2 patent drawing
  • US11361030B2 patent drawing

AI summary

Facet-based search processing is provided which includes receiving a query search context for querying documents of a document set, and retrieving, by similar document search processing, a document subset from the document set. The document subset includes documents of the set most similar to a search document of the query search context. Facet analysis processing is used to generate M candidate facets most-related to the query search context, and facets of the M candidate facets associated with documents of the subset are identified, and classified into a positive facet set and a negative facet set based, at least in part, on extent of facet commonality across the documents. A listing is of the documents in the document subset is provided, with the listing highlighting facets of the positive facet set.