Hierarchical Vector Store for Privacy-Protective Knowledge Sharing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Organizations face challenges in sharing knowledge across or within their entities while protecting sensitive and proprietary data, as existing methods often require direct access to sensitive information, which may be restricted by regulatory or privacy concerns.
Innovation Solution
A hierarchical vector store with fine-grained access control rules and a machine learning model that generates responses based on contextual knowledge, ensuring data privacy by excluding sensitive information and employing masking techniques to prevent information leakage.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If direct access to sensitive data is provided for knowledge sharing, then knowledge sharing effectiveness is improved, but data privacy and security are compromised
Solution Approach 1:
The patent introduces a synthetic data generation system as an intermediary between the sensitive data and the knowledge sharing process. This system generates artificial datasets that preserve statistical properties and analytical value while containing no actual sensitive information, thus enabling effective knowledge sharing without direct access to proprietary data
Solution Approach 2:
The patent creates synthetic copies of sensitive datasets that replicate the structural and statistical characteristics without containing actual sensitive values. These synthetic copies serve as substitutes for the original data in analytical processes, maintaining knowledge sharing effectiveness while eliminating privacy risks associated with direct data access
2Loss of information
If sensitive data is shared across organizations, then collaborative insights are improved, but compliance with access control rules deteriorates
Solution Approach 1:
The synthetic data generation system acts as a mediator that enables collaborative insights to be derived without violating access control policies. Organizations can share synthetic datasets or models trained on synthetic data, allowing collaborative analysis while maintaining compliance with each organization's data protection rules
Solution Approach 2:
The system transforms the parameters of data sharing by changing from sharing actual sensitive values to sharing synthetic generated data with controlled statistical properties. This parameter transformation maintains the analytical value needed for collaborative insights while ensuring compliance with access control requirements
3Object-affected harmful factors
If proprietary data is protected from sharing, then data security is improved, but knowledge derivation capability deteriorates
Solution Approach 1:
The patent creates synthetic copies that preserve the analytical and statistical properties of proprietary data while containing no actual sensitive information. These copies enable knowledge derivation and analytical processing to proceed with the same effectiveness as if using original data, while maintaining enhanced data security
Solution Approach 2:
The system extracts the essential statistical properties and patterns from proprietary data to create synthetic representations. By separating the analytical value from the sensitive information, the system enables knowledge derivation capability to be maintained while proprietary data remains protected and unseen
Data Source
AI summary
Disclosed herein are various approaches for sharing knowledge within and between organizations while protecting sensitive data. A machine learning model may be trained using training prompts querying a vector store to prevent unauthorized user disclosure of data derived from the vector store. A prompt may be received and a response to the prompt may be generated using the machine learning model based at least in part on the vector store.


