Relevancy Index Table for Big Data Analytics
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Big data systems face challenges in analyzing and quantifying data relevancy due to the large size and complexity of data sets, leading to low performance in analytics and predictive analytics, as traditional systems struggle to identify relevant tables and fields within these systems.
Innovation Solution
A computerized method that uses in-memory databases and machine learning algorithms to monitor interactions with database tables, generating a relevancy index table that updates and scores fields based on interaction frequency and recency, with customizable relevancy rules to identify and prioritize relevant data for analytics.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional data processing applications are used to analyze big data sets, then the system structure is simple, but the analysis performance is inadequate and the system cannot effectively identify relevant data
Solution Approach 1:
The patent introduces a relevancy index table as an intermediary data structure between the monitored data sources and the analytics system. This table stores pre-computed relevancy scores for fields and tables based on monitored interactions, allowing the analytics system to quickly identify relevant data without complex real-time analysis, thus improving reliability while managing system complexity
Solution Approach 2:
The system performs preliminary actions by continuously monitoring interactions with data sources and pre-computing relevancy scores before analytics operations are needed. The relevancy index table is updated in advance with field and table relevancy information, enabling fast retrieval and effective identification of relevant data when analytics operations occur
2Reliability
If all data in big data sets is analyzed, then comprehensive analysis coverage is achieved, but the processing time and computational resources increase significantly
Solution Approach 1:
The patent extracts only the relevant portions of data from the big data sets by using the relevancy index table to identify fields and tables with high relevancy scores. Instead of analyzing all data, the system extracts and processes only those data elements that have been determined to be relevant based on monitored interactions, significantly reducing processing time while maintaining analytics accuracy
Solution Approach 2:
The system applies local quality by assigning different relevancy scores to different fields and tables based on their specific interaction patterns. Rather than treating all data uniformly, the system identifies and focuses on locally relevant data regions (specific fields and tables) that have higher relevancy scores, enabling efficient targeted analysis
3Reliability
If the relevancy index table is continuously updated with interaction data, then data relevancy information remains current, but the update overhead increases system complexity
Solution Approach 1:
The relevancy index table serves multiple functions: it stores field relevancy information, table relevancy information, and serves as the basis for generating cleanup rules. This multi-functionality reduces the need for separate data structures and mechanisms, managing complexity while maintaining current relevancy information through continuous updates
Data Source
AI summary
The present disclosure involves analyzing data relevancy of particular fields within one or more databases in a big data system. In one example method, an interaction with at least one of a plurality of monitored data sources is identified, wherein the identified interactions is associated with a particular field of a database table of one of the monitored data sources. A set of data associated with the interaction is determined which includes an identification of each field associated with the identified interaction and a count of a number of interactions associated with each particular field. A relevancy index table is updated to include the determined set of data, wherein each identified field is associated with a row in the index table. At least one relevancy rule is identified for the relevancy index table and is executed to generate a relevancy score for at least one of the fields.


