Machine Learning Dataset Similarity for Access Control
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional computing systems face challenges in reducing storage costs and managing data access effectively due to overlapping and similar derivative datasets, leading to increased costs and security issues, as they struggle to determine accurate access rights and schema matches.
Innovation Solution
The use of data lineage information and machine learning to merge datasets, determine similarity scores, and modify access rights, allowing for the identification of related datasets and optimizing storage by merging or restricting access based on similarity thresholds.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional rules-based approaches are used to manage datasets, then data access control is implemented, but storage costs increase due to inability to identify and merge similar derivative datasets
Solution Approach 1:
The patent introduces machine learning models as an intermediary between data access control and storage management. The ML models analyze dataset similarities and generate similarity scores, enabling the system to identify derivative datasets that should be merged. This intermediary analysis layer resolves the contradiction by providing intelligent insights that rules-based systems cannot detect, allowing storage optimization without compromising access control reliability.
Solution Approach 2:
The system dynamically changes the parameter of dataset similarity assessment by using machine learning-generated similarity scores instead of fixed rules. This parameter change enables the system to adaptively identify which datasets should be merged based on their actual content similarity rather than rigid schema matching, thereby reducing storage costs while maintaining appropriate access control.
2Loss of energy
If machine learning models are used to determine dataset similarity, then storage costs are reduced through merging, but system complexity increases
Solution Approach 1:
The patent segments the data management system into distinct functional components: data access control mechanisms, machine learning similarity assessment models, and dataset merging operations. This segmentation allows each component to be independently optimized and managed, reducing the perceived complexity while enabling the system to achieve storage cost reductions through intelligent dataset identification and consolidation.
3Reliability
If access rights are modified based on dataset similarity, then security is enhanced by preventing inadvertent access, but ease of operation decreases
Solution Approach 1:
The system implements self-service by automatically modifying access rights based on machine learning-determined dataset similarities. Instead of requiring manual security policy updates, the system autonomously identifies when derivative datasets should inherit or restrict access based on their similarity to source datasets, enhancing security while reducing the operational burden on users.
4Ease of operation
If conventional data access control systems are used, then access rights are managed, but they fail to prevent access to derivative datasets when access to original datasets is restricted
Solution Approach 1:
The patent implements feedback by using machine learning models to continuously assess dataset similarities and provide insights back to the access control system. This feedback loop enables the system to dynamically adjust access rights for derivative datasets based on their relationship to restricted source datasets, ensuring that security policies are effectively enforced across the entire data ecosystem rather than in isolation.
Data Source
AI summary
In certain embodiments, machine learning and lineage data may be used to manage data. In some embodiments, a computing system may use lineage data to identify two datasets that may be related. The computing system may determine that a user has access to a derivative dataset but does not have access to an original dataset that was used to create the derivative dataset. In response, the computing system may use a machine learning model to generate a similarity score indicating a level of similarity between the original dataset and the derivative dataset. If the similarity score satisfies a threshold score, the computing system may modify access rights of the user so that the user is unable to access a portion of the data in the derivative dataset.


