NLP Content Segmentation for Redundant Data Elimination

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The abundance of electronic information makes it difficult for users to distinguish between new and previously consumed content, leading to inefficiencies in processing and storage, as well as a waste of computing resources due to repetitive content presentation.

Innovation Solution

A method utilizing natural language processing (NLP) to identify previously consumed content within machine-encoded files, modifying the electronic presentment structure to differentiate new from old content, thereby reducing redundant processing and storage.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of time

If all electronic content is presented to users without distinction, then users receive complete information, but users waste time processing previously consumed content

Engineering Contradiction:
Improveuser timeVSAvoidinformation completeness
Core Design Contradiction:
Loss of timeVSLoss of information

Solution Approach 1:

The system segments electronic content into previously consumed portions and new portions using natural language processing to identify matching content. This segmentation allows the system to selectively present only new information to users while maintaining access to complete content when needed, thereby reducing time waste without complete information loss.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary comparison of incoming electronic content with previously consumed content using NLP techniques before presentation. By pre-identifying matching content segments, the system prepares the content in advance for selective presentation, eliminating the need for users to manually review entire documents and saving their time.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If redundant content is processed and stored, then complete content archives are maintained, but computing resources are wasted

Engineering Contradiction:
Improvecontent archive completenessVSAvoidcomputing resources
Core Design Contradiction:
ReliabilityVSLoss of energy

Solution Approach 1:

The system extracts and identifies redundant content segments through NLP comparison with previously consumed content. By taking out only the unique new portions for processing and storage, the system maintains reliable content archives while eliminating waste of computing resources on duplicate material.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system changes the parameter of content representation by using NLP-derived features and classifiers to identify redundancy. Instead of storing and processing entire content files, the system uses transformed parametric representations (NLP elements, classifiers) to detect duplicates, reducing computational energy while maintaining archive reliability.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If natural language processing is performed on all content, then accurate content identification is achieved, but processing time increases

Engineering Contradiction:
Improvecontent matching accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system applies NLP processing partially - only to the extent needed for segment classification and redundancy detection. By performing NLP selectively on new content to identify matching segments rather than processing entire documents uniformly, the system achieves accurate content identification while minimizing processing time.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The system segments content processing into distinct NLP stages (element extraction, classifier application, matching comparison) and applies them only where needed. This segmented approach maintains measurement precision for content matching while reducing overall processing time by avoiding unnecessary NLP operations on already-identified redundant segments.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11437038B2Recognition and restructuring of previously presented materials
Publication Date: 2022.09.06 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11437038B2 patent drawing
  • US11437038B2 patent drawing
  • US11437038B2 patent drawing

AI summary

Information recognition and restructuring includes analyzing electronic media content embedded in an electronic presentment structure presented to a user, and based on the analyzing, detecting portions of the electronic media content previously consumed by the user. The method includes modifying the electronic presentment structure, based on the detecting, to distinguish the electronic media content previously consumed by the user from other portions of the electronic media content.