Data Aggregation via Temporary Extraction and Normalization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data mining and reporting methods often result in data duplication leading to lag between production and duplicated databases, interfering with production database use and providing outdated reports, and may compromise security by making report outputs accessible to unauthorized users.
Innovation Solution
A method and system for simultaneously collecting, aggregating, normalizing, and storing data from multiple locations within a source database, eliminating the need for duplication and ensuring current data availability while maintaining security by determining data locations, acquiring access configuration logs, and releasing data back into the source database.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If data is duplicated for data mining, then data can be mined without interfering with production database use, but data lag increases between 8-24 hours
Solution Approach 1:
The patent extracts data from the production database temporarily, processes it for mining purposes, and then returns it to the production database. This extraction approach allows data mining operations to proceed without creating permanent duplicates, eliminating the 8-24 hour lag associated with traditional duplication methods while maintaining production database integrity.
Solution Approach 2:
The system performs preliminary data extraction and processing before the actual mining operation. By preparing data in advance and returning it to the production database before mining completes, the system eliminates the need for long-term duplication and reduces data lag to minimal levels while allowing concurrent production operations.
2Measurement precision
If production database is queried for reporting, then current and accurate data is obtained, but production database use is interfered with for several hours
Solution Approach 1:
The patent segments the reporting function from the production database by creating a separate reporting database. This segmentation allows reporting operations to query historical data independently without blocking production database access. The system maintains accurate, current reports by periodically synchronizing data while preserving full production database availability during reporting operations.
Solution Approach 2:
The system introduces an intermediary reporting database that mediates between production database queries and reporting needs. This intermediary layer captures current data states and makes them available for reporting without requiring concurrent access to the production database, thereby maintaining both report accuracy and production availability simultaneously.
3Productivity
If report output is stored outside database, then repeated queries are avoided, but report outputs representing multiple non-current states become available
Solution Approach 1:
The patent implements dynamic report generation where reports are created on-demand based on current database states rather than storing static historical outputs. This dynamic approach ensures reports always reflect the most current data while avoiding the need to repeat queries. The system generates reports dynamically from the production database, eliminating both the need for repeated queries and the problem of non-current report states.
4Productivity
If report output is stored outside database, then query repetition is avoided, but security is compromised by allowing unauthorized access
Solution Approach 1:
The patent merges the reporting function with the production database by implementing reports within the database structure itself. This integration ensures that report access is governed by the same security permissions as production database access, eliminating the security vulnerability where unauthorized users could access stored report outputs. The unified approach maintains both efficient report access and proper security controls.
Data Source
AI summary
The invention provides a method, system, and program product for managing data for data aggregation, including data mining and reporting. Locations of a plurality of data to be collected are determined within a source database. Data are simultaneously collected from the plurality of locations and aggregated. The aggregated data are normalized by adding an encryption key and the normalized data are stored. Data at each of the plurality of locations are then released in the source database.


