Query Reconstruction for Privacy-Compliant Big Data Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing big data analysis platforms face limitations due to network-separated environments that hinder interactive data analysis and compliance with rapidly changing personal information protection regulations, leading to inefficiencies and risks of infringement.
Innovation Solution
A big data analysis device that decomposes user queries into sub-queries, identifies those referring to personal information, reconstructs them to be independent of personal information, and generates a reconstruction query to ensure compliance with personal information protection laws, using a query reconstruction module and natural language interpretation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If a network-separated environment is implemented to protect personal information, then personal information protection is improved, but data analysis functionality and user interaction are worsened
Solution Approach 1:
The patent introduces a query reconstruction module as an intermediary between the user and the personal information database. This module automatically identifies personal information in queries, reconstructs them to remove direct references to personal information, and executes the reconstructed queries. This mediator enables data analysis functionality while maintaining network separation and personal information protection.
Solution Approach 2:
The patent segments the data analysis process into distinct components: original query input, personal information identification, query reconstruction, and result generation. By dividing the query processing into separate stages, the system can protect personal information in the network-separated environment while still providing comprehensive data analysis capabilities through the reconstructed queries.
2Reliability
If manual compliance verification with personal information protection laws is performed, then legal compliance is improved, but time consumption and operational complexity are worsened
Solution Approach 1:
The query reconstruction module operates autonomously to identify personal information in user queries and automatically reconstruct the queries to comply with personal information protection laws. The system performs self-verification of compliance without requiring manual legal review, thereby ensuring legal compliance while eliminating time-consuming manual verification processes.
Solution Approach 2:
The system performs preliminary reconstruction of queries before execution, proactively identifying and removing personal information references in advance. This preliminary action ensures that all queries are pre-compliant with personal information protection laws, eliminating the need for subsequent manual compliance verification and reducing time loss.
3Productivity
If direct access to personal information for analysis is allowed, then data analysis efficiency is improved, but personal information infringement risk is worsened
Solution Approach 1:
The patent converts the constraint of personal information protection into a benefit by automatically reconstructuring queries to remove personal information references. This process maintains data analysis efficiency by preserving the analytical intent while eliminating infringement risk, effectively converting the harmful factor (potential infringement) into a beneficial outcome (protected analysis).
Solution Approach 2:
The query reconstruction module serves as an intermediary that processes queries to eliminate direct access to personal information while preserving analytical capabilities. It reconstructs queries using aggregated or anonymized data representations, enabling efficient data analysis without creating personal information infringement risks.
Data Source
AI summary
A method of operating a big data analysis device is provided. The method includes receiving a user query, decomposing the received user query into multiple subqueries, identifying first subqueries configured to refer to personal information among the multiple subqueries, reconfiguring the first subqueries into second subqueries independent of the personal information based on a component accessible to the personal information, and generating a reconfigured query by merging the second subqueries with subqueries other than the first subqueries configured to refer to the personal information among the multiple subqueries.


