Cascade Claim Detection for Unstructured Text Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current text analysis technologies face challenges in automatically detecting claims relevant to a topic, particularly in distinguishing between context-dependent claims and other text types, and in pinpointing exact claim boundaries within large volumes of unstructured content.

Innovation Solution

A method and system utilizing a hardware processor to receive a topic under consideration and relevant content, detect claim boundaries, and output a list of detected claims, employing a cascade of components for sentence and section detection, with features that characterize claims and assess relevance, and applying filters to refine claim detection and scoring.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If text analysis is performed on large volumes of unstructured content to detect claims, then the quantity of analyzed text increases, but the precision of claim detection decreases

Engineering Contradiction:
Improvevolume of content analyzedVSAvoidprecision of claim detection
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The patent segments the claim detection process into multiple cascade stages: (1) section detection to identify potential claim sections, (2) sentence detection to find claim sentences within sections, and (3) claim detection to pinpoint exact claim boundaries. Each stage applies filters and scoring mechanisms to progressively refine results, maintaining precision while handling large volumes of unstructured content.

Inventive Principle:
Principle #1Segmentation

2Extent of automation

If automated claim detection is applied to distinguish context-dependent claims from other text types, then the extent of automation increases, but the difficulty of detecting and measuring claims increases

Engineering Contradiction:
Improveautomation of claim detectionVSAvoiddifficulty of detecting claim boundaries
Core Design Contradiction:
Extent of automationVSDifficulty of detecting and measuring

Solution Approach 1:

The patent introduces intermediary components between raw text input and final claim detection: (1) section detectors that identify potential claim sections using features like section titles and headings, (2) sentence detectors that locate claim sentences within sections using linguistic features, and (3) filters that apply scoring mechanisms to refine detections. These intermediaries break down the complex task into manageable sub-tasks, making automated detection feasible.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Manufacturing precision

If claim boundaries are pinpointed within large volumes of unstructured content, then the manufacturing precision of claim extraction improves, but the loss of time increases

Engineering Contradiction:
Improveprecision of claim extractionVSAvoidtime for claim detection
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The patent applies preliminary actions by pre-processing content into structured sections with detected headings and titles before performing claim detection. The cascade architecture pre-identifies potential claim sections and sentences using linguistic features and patterns, so that when exact claim boundaries are detected, the search space is already significantly reduced, saving time while maintaining precision.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11113471B2Automatic detection of claims with respect to a topic
Publication Date: 2021.09.07 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11113471B2 patent drawing
  • US11113471B2 patent drawing

AI summary

A method comprising using at least one hardware processor for: receiving a topic under consideration (TUC) and content relevant to the TUC; detecting one or more claims relevant to the TUC in the content, based on detection of boundaries of the claims in the content; and outputting a list of said detected one or more claims.