Text Extraction Module for Contextual Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for evaluating digital content consumed by users are inefficient and unreliable, relying on manual analysis and inconsistent data, which hinders the ability of publishers and advertisers to deliver targeted content effectively, as they struggle to systematically understand and categorize the interests of website visitors from vast amounts of semi-structured or unstructured content.

Innovation Solution

A contextual analysis engine that extracts, analyzes, and organizes digital content using a text extraction module to separate relevant content from irrelevant, and a text analytics module to generate structured categorization of topics, leveraging semantic, statistical, and natural language analytics, with an input/output interface managing workflows and caching results to provide actionable insights.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If manual analysis methods are used to evaluate digital content, then implementation simplicity is maintained, but measurement precision and reliability of user interest evaluation deteriorate

Engineering Contradiction:
Improveimplementation simplicityVSAvoiduser interest evaluation accuracy
Core Design Contradiction:
Ease of manufactureVSMeasurement precision

Solution Approach 1:

The patent replaces manual analysis methods with automated text extraction and analytics systems. The text extraction module automatically extracts content from digital sources, and the text analytics module applies natural language processing, statistical analysis, and semantic analysis to evaluate user interests, eliminating the need for manual content assessment while significantly improving measurement precision.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables self-service evaluation by implementing automated workflows where the text extraction module and text analytics module work together to independently extract, analyze, and categorize digital content without human intervention. The workflow management coordinates these modules to automatically generate user interest profiles from consumed content.

Inventive Principle:
Principle #25Self-service

2Productivity

If automated text extraction and analytics are implemented, then analysis productivity and measurement precision improve, but device complexity increases

Engineering Contradiction:
Improvecontent analysis throughputVSAvoidsystem architecture complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent divides the content analysis system into distinct functional modules: a text extraction module that handles content acquisition from various digital sources, a text analytics module that performs analysis using multiple techniques, and a workflow management component that coordinates operations. This segmentation allows each module to be optimized independently while maintaining high overall productivity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The text analytics module is designed with multi-functionality, incorporating natural language processing, statistical analysis, and semantic analysis capabilities within a single integrated system. This universal approach allows the system to handle diverse content types and analysis requirements without requiring separate specialized systems, thereby managing complexity while maintaining high productivity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Measurement precision

If multiple analytics techniques are integrated, then measurement precision and reliability improve, but device complexity increases

Engineering Contradiction:
Improveuser interest categorization accuracyVSAvoidanalytics processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent merges multiple analytics techniques including natural language processing, statistical analysis, and semantic analysis into a unified text analytics module. These techniques are integrated to work together synergistically, where each technique contributes to the overall user interest evaluation, improving measurement precision through combined analysis while managing complexity through unified architecture.

Inventive Principle:
Principle #5Merging (Combining)

4Device complexity

If manual content evaluation is used, then device complexity is minimized, but loss of time in understanding user interests increases

Engineering Contradiction:
Improvesystem structure simplicityVSAvoiduser interest analysis time
Core Design Contradiction:
Device complexityVSLoss of time

Solution Approach 1:

The patent replaces time-consuming manual content evaluation with automated text extraction and analytics processing. The system automatically extracts text from digital content sources and applies multiple analytics techniques to rapidly generate user interest profiles, reducing analysis time from manual processes to automated real-time or near-real-time processing.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS10235681B2Text extraction module for contextual analysis engine
Publication Date: 2019.03.19 ADOBE INC
  • US10235681B2 patent drawing
  • US10235681B2 patent drawing
  • US10235681B2 patent drawing

AI summary

A contextual analysis engine systematically extracts, analyzes and organizes digital content stored in an electronic file such as a webpage. Content can be extracted using a text extraction module which is capable of separating the content which is to be analyzed from less meaningful content such as format specifications and programming scripts. The resulting unstructured corpus of plain text can then be passed to a text analytics module capable of generating a structured categorization of topics included within the content. This structured categorization can be organized based on a content topic ontology which may have been previously defined or which may be developed in real-time. The systems disclosed herein optionally include an input/output interface capable of managing workflows of the text extraction module and the text analytics module, administering a cache of previously generated results, and interfacing with other applications that leverage the disclosed contextual analysis services.