Semantic Article Clustering for Newspaper Digitization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The complexity of newspaper page layouts and poor quality of scanned documents hinder effective digitalization, particularly in identifying and clustering article components across multiple pages, due to varying formats, low print quality, and mixed text and image structures.

Innovation Solution

A method and system that analyze text and image objects within an electronic document to determine their closeness, clustering them when the closeness exceeds a threshold, and storing these clusters in an indexed database for improved searchability and analytics, using engines for text and image analysis, and object recognition techniques.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If traditional layout-based rules are used to detect article elements, then the detection process can be simple, but it fails to handle varying newspaper formats and complex multi-column layouts

Engineering Contradiction:
Improvehandling of varying newspaper formatsVSAvoiddetection process complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent replaces traditional mechanical layout-based rules with a semantic analysis system that uses topic modeling and similarity computation. Instead of relying on fixed positional rules, the system analyzes the semantic content of text blocks and images to determine article boundaries, enabling it to adapt to various newspaper formats and complex layouts.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system changes the parameter space from spatial layout coordinates to semantic topic similarity metrics. By computing topic distributions and measuring similarity between text blocks and images based on their semantic content rather than their physical positions, the system achieves format independence while handling complex layouts effectively.

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If text content analysis is used to determine article boundaries, then articles can be identified across different layouts, but the process becomes computationally intensive and time-consuming

Engineering Contradiction:
Improvearticle identification across layoutsVSAvoidprocessing time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent applies partial action by using topic modeling to extract only the most salient semantic features from text blocks rather than analyzing the entire text content. This selective approach maintains the ability to identify articles across different layouts while significantly reducing computational intensity and processing time.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The system introduces topic models as an intermediary representation between raw text content and article boundary detection. Instead of directly comparing full text contents, the system first transforms text into compact topic distributions, which serve as efficient intermediaries for similarity computation and article identification.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If semantic analysis is performed on all text blocks to cluster articles, then accurate article extraction can be achieved, but the computational resources required increase significantly

Engineering Contradiction:
Improvearticle extraction accuracyVSAvoidcomputational resource consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent segments the semantic analysis process into two stages: first, topic modeling is performed on individual text blocks to extract compact topic representations; second, similarity computation is performed only on these segmented topic representations rather than on the full text content. This segmentation maintains extraction accuracy while reducing computational resource consumption.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Instead of performing expensive semantic analysis operations on the original text blocks, the system creates compact topic distribution copies that capture the essential semantic information. These copied representations are then used for similarity computation and article clustering, preserving accuracy while minimizing resource usage.

Inventive Principle:
Principle #26Copying

4Productivity

If images are excluded from article detection, then processing speed increases, but articles with images or complex visual layouts cannot be properly identified

Engineering Contradiction:
Improveprocessing speedVSAvoidarticle identification reliability
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent merges image analysis with text analysis by computing topic distributions for both modality types and measuring their similarity. Images are processed to extract semantic content in the same topic space as text, allowing the system to identify articles containing images or complex visual layouts while maintaining processing efficiency through unified semantic representation.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS10572528B2System and method for automatic detection and clustering of articles using multimedia information
Publication Date: 2020.02.25 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US10572528B2 patent drawing
  • US10572528B2 patent drawing
  • US10572528B2 patent drawing

AI summary

The disclosure provides methods and systems that automatically detect and cluster related articles in a publication for archival, search, and other purposes. Text and images are recognized and scored in order to cluster related content into coherent and searchable articles.