Topic Popularity Analysis via Gram Correlation in Big Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current information retrieval systems face challenges in analyzing the popularity of user-defined topics within big data due to the complexity of correlations between structured and unstructured data sources, leading to difficulties in identifying relevant trends amidst vast amounts of information.

Innovation Solution

A system and method that analyze the popularity of user-defined topics by identifying correlations between 'grams' in user-identified topical anchor documents and raw documents, utilizing a processor and program modules to determine rarity, importance, relevancy, and popularity values, and display the most relevant information.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If traditional information retrieval systems analyze big data by creating indexes and retrieving documents based on query terms, then document retrieval is achieved, but the ability to identify relevant trends and relationships between topics deteriorates due to the sheer volume of data

Engineering Contradiction:
Improvevolume of data analyzedVSAvoidaccuracy of trend identification
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The patent segments the analysis process into distinct components: (1) identifying anchor documents that represent specific topics, (2) extracting grams (word sequences) from these anchor documents, (3) comparing these grams against grams in raw documents, and (4) calculating popularity values based on correlation metrics. This segmentation allows the system to manage big data by breaking it down into manageable analytical units while maintaining precision in trend identification.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary mechanism - the gram-based correlation analysis - that bridges anchor documents and raw documents. Instead of directly comparing entire documents or using traditional keyword matching, the system uses grams as intermediaries to measure the relationship between topical anchor documents and raw documents, enabling precise trend identification even in large data volumes.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Loss of information

If more data is collected and analyzed to provide better insight into topic popularity, then the volume of information increases, but the complexity of the algorithm required to process this data increases with little or no enlargement of insights

Engineering Contradiction:
Improveinsight qualityVSAvoidalgorithm complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent extracts only the essential elements needed for analysis - grams from anchor documents that represent key topical concepts. Rather than processing entire documents or all data points, the system extracts and compares only the relevant gram sequences, significantly reducing algorithmic complexity while maintaining insight quality. The extraction focus is on grammatical units that capture topic essence rather than exhaustive data processing.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent changes the analytical parameter from traditional document-level or keyword-level analysis to gram-level analysis. By measuring topic popularity through gram correlations rather than document frequencies or keyword counts, the system achieves better insights with simpler algorithms. The parameter transformation allows the system to capture nuanced topic relationships without requiring complex processing of entire document corpora.

Inventive Principle:
Principle #35Parameter changes

3Productivity

If traditional search engines present retrieved documents in ranked order without grouping or hierarchy, then document retrieval is efficient, but the user cannot easily identify relationships between topics or understand trends

Engineering Contradiction:
Improveinformation retrieval efficiencyVSAvoidtopic relationship information
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

The patent adds a new dimension to information presentation by calculating and displaying popularity values that represent the relationship between anchor documents and raw documents. Instead of only showing documents in ranked order, the system introduces a popularity metric dimension that groups and hierarchizes documents based on their correlation with topical anchors, enabling users to see both retrieval efficiency and topic relationships simultaneously.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent implements feedback by using the calculated popularity values to inform the presentation and grouping of results. The system feeds back the correlation analysis results into the display mechanism, organizing documents by their popularity scores and relationships to anchor topics. This feedback loop allows users to see not just retrieved documents but also the underlying topic relationships and trends.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS10067964B2System and method for analyzing popularity of one or more user defined topics among the big data
Publication Date: 2018.09.04 HALLER JR JOHN L
  • US10067964B2 patent drawing
  • US10067964B2 patent drawing
  • US10067964B2 patent drawing

AI summary

A method to analyze popularity of user defined topics by identifying correlations between grams contained in user identified anchor documents and the grams contained in raw documents is provided. The method includes following steps: (a) a user input data that includes (i) user identified topics for user identified subject matter, (ii) user identified topical anchor documents, and (iii) a plurality of user identified raw documents internet source with respective source addresses; (b) the raw document sources is accessed using the source addresses to retrieve and store data in a database; (c) grams and gram document dictionaries together with gram values for each topical anchor document and raw document are identified and stored; and (d) the grams in each of the topical anchor documents against the grams in all the raw documents are analyzed to determine a relative popularity of the topical anchor documents.