Topic Popularity Analysis via Gram Correlation in Big Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current information retrieval systems face challenges in analyzing the popularity of user-defined topics within big data due to the complexity of correlations between structured and unstructured data formats, leading to difficulties in identifying relevant relationships and trends amidst vast amounts of information.
Innovation Solution
A system and method that analyze the popularity of user-defined topics by identifying correlations between 'grams' in user-identified topical anchor documents and raw documents, utilizing a processor and program modules to determine rarity, importance, relevancy, and popularity values, and display the most relevant information to users.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If traditional information retrieval systems analyze big data by creating indexes that relate documents to individual words, then document retrieval is enabled, but the ability to identify relevant relationships and trends between topics is lost due to the sheer volume of data
Solution Approach 1:
The patent segments the analysis process into multiple stages: first creating traditional word-based indexes for efficient retrieval, then performing secondary analysis on retrieved documents to identify gramm patterns, correlations, and trends. This segmentation allows the system to handle large volumes of data while preserving relationship information through multi-layered processing.
Solution Approach 2:
The patent introduces an intermediary analysis layer that operates between traditional document retrieval and final results. This intermediary process analyzes gramm patterns, topic correlations, and relationships in retrieved documents, acting as a mediator that transforms raw document data into meaningful relationship insights without requiring the entire big data set to be processed at once.
2Loss of information
If more data is analyzed to provide better insight into topic popularity, then the volume of information increases, but the complexity of algorithms required to process and analyze this data increases with little or no enlargement of insights
Solution Approach 1:
The patent performs preliminary actions by pre-processing and indexing gramm patterns from documents before actual topic analysis is needed. Grammatic structures, topic associations, and relationship patterns are pre-computed and stored, allowing the system to answer topic popularity queries efficiently without performing complex algorithms on the entire data set at query time.
Solution Approach 2:
The patent applies local quality by focusing computational resources on specific local patterns and relationships relevant to user-defined topics rather than uniformly analyzing all data. The system identifies and analyzes only the gramm patterns and correlations that are locally relevant to each topic, reducing overall algorithmic complexity while maintaining insight quality.
3Quantity of substance
If unstructured data is processed using greater computer power, memory space, and processor time, then more data can be handled, but better or more accurate analysis is not necessarily provided
Solution Approach 1:
The patent changes key parameters by transforming unstructured text data into structured gramm patterns and relationship metrics. Instead of directly analyzing raw unstructured data with increasing computational power, the system transforms the data into standardized gramm representations with defined attributes, enabling accurate analysis through parameter-based processing rather than brute-force computation.
Data Source
AI summary
A method to analyze popularity of user defined topics by identifying correlations between grams contained in user identified anchor documents and the grams contained in raw documents includes the following steps: (a) a user input data that includes (i) user identified topics for user identified subject matter, (ii) user identified topical anchor documents, and (iii) a plurality of user identified raw documents internet source with respective source addresses; (b) the raw document sources is accessed using the source addresses to retrieve and store data in a database; (c) grams and gram document dictionaries together with gram values for each topical anchor document and raw document are identified and stored; and (d) the grams in each of the topical anchor documents against the grams in all the raw documents are analyzed to determine a relative popularity of the topical anchor documents.


