Automated Content Analysis System Using NLP Entity Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing research systems are inefficient and time-consuming due to the need for manual analysis and updating of large volumes of unstructured data from multiple sources, making it difficult to identify and classify relevant content for analysis.

Innovation Solution

A system that filters and processes data feeds using natural language processing and risk mining to extract entities and activities, generating an enhanced output with graphical annotations and GUI controls to facilitate analysis and database updates.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual analysis and processing of data feeds is performed, then analysts can identify and classify relevant content, but the process becomes extremely time-consuming and inefficient

Engineering Contradiction:
Improveaccuracy of content analysisVSAvoidtime required for manual analysis
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system enables self-service automated content analysis through AI-powered entity recognition, classification, and tagging that processes data feeds without requiring manual analyst intervention for routine tasks, while still allowing human oversight when needed

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

Manual mechanical analysis processes are replaced with automated computational systems that use natural language processing, machine learning models, and entity recognition algorithms to perform content analysis, entity extraction, and classification tasks that were previously done manually

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If all data from multiple sources is presented to analysts for review, then comprehensive analysis is possible, but the volume of information makes it difficult to find relevant content

Engineering Contradiction:
Improvecompleteness of data reviewVSAvoiddifficulty of finding relevant information
Core Design Contradiction:
ReliabilityVSDifficulty of detecting and measuring

Solution Approach 1:

The system extracts and isolates relevant information from large volumes of unstructured data by identifying and pulling out key entities, activities, and risk indicators, separating them from irrelevant content through automated entity recognition and classification

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system applies different processing and filtering criteria to different portions of the data based on their relevance and importance, enhancing the presentation of high-value information while reducing or filtering out less critical content

Inventive Principle:
Principle #3Local quality

3Reliability

If analysts manually update databases with new information, then database accuracy is maintained, but the update process is time-consuming and inefficient

Engineering Contradiction:
Improveaccuracy of database informationVSAvoidspeed of database updates
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system performs self-service automated database updates by automatically ingesting processed content, extracting entities and attributes, and updating database records without requiring manual data entry, while maintaining data quality through automated validation

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system performs preliminary data processing, entity extraction, and validation actions before database updates are executed, preparing the data in advance and reducing the time required for actual database insertion and updates

Inventive Principle:
Principle #10Preliminary action

4Adaptability or versatility

If unstructured content from various formats is processed, then diverse data sources can be analyzed, but the lack of standardized format increases processing time

Engineering Contradiction:
Improveability to process multiple content formatsVSAvoidcomplexity of processing system
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system implements universal processing capabilities that can handle multiple content formats (news articles, blogs, data feeds, streams) through a single integrated platform that automatically adapts to different input types using format-detection and transformation algorithms

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11609959B2System and methods for generating an enhanced output of relevant content to facilitate content analysis
Publication Date: 2023.03.21 REFINITIV US ORGANIZATION LLC
  • US11609959B2 patent drawing
  • US11609959B2 patent drawing
  • US11609959B2 patent drawing

AI summary

The present disclosure relates to methods and systems for ingesting content from data feeds and to generate an enhanced output of relevant content for presentation to a user to facilitate analysis of the relevant content. Content is received from data feeds, and filtered to identify relevant content with respect to a particular context. The relevant content is then processed, e.g., using natural language processes, to extract entities involved, and to also identify particular activities detailed in the relevant content. Activity-mining is applied to the identified relevant data to classify and assigned activity tags to the extracted entity. Based on the extracted and identified information, an enhanced output is generated for presentation to facilitate research operations. The enhanced output may include overlaid graphical annotations, indicators, and graphical controls over the relevant articles to provide a means for updating a database based on the relevant content.