Unstructured Data Valuation via Domain-Aware Tokenization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data lake architectures lack methods to programmatically calculate business value of unstructured data, provide contextual business value, calculate data value without rescan, observe fluctuation in business value, and offer data valuation reporting and prioritization, leading to inefficiencies in data management and decision-making.
Innovation Solution
A data valuation framework that generates domain-aware tokens through tokenization, metadata extraction, and domain mapping, allowing for contextual business value calculation and dynamic re-evaluation of data sets without full rescan, and provides a dashboard for reporting and prioritization.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Volume of stationary object
If data lake architectures store enormous amounts of unstructured content, then data storage capacity is improved, but the ability to calculate business value of the content deteriorates
Solution Approach 1:
The patent segments unstructured data into discrete tokens and organizes them in an inverted index structure. This segmentation allows the system to process and evaluate individual data elements independently, enabling business value calculation without requiring analysis of the entire data lake, thus resolving the contradiction between storage capacity and value assessment capability.
Solution Approach 2:
The patent introduces domain-aware tokens as an intermediary layer between raw unstructured data and business value assessment. These tokens capture semantic meaning and enable the valuation system to interpret unstructured content without requiring full content analysis, allowing business value information to be derived from stored unstructured data.
2Measurement precision
If full data rescan is performed to recalculate business value, then valuation accuracy is improved, but processing time and computational resources worsen
Solution Approach 1:
The patent performs preliminary tokenization and indexing of unstructured data during the data ingestion phase. By pre-processing data into domain-aware tokens and organizing them in an inverted index, the system eliminates the need for full rescans when recalculating business value, maintaining valuation accuracy while significantly reducing processing time.
Solution Approach 2:
The patent implements a dynamic valuation system that can recalculate business value by selectively processing only the data elements that have changed or are relevant to the current context. This dynamic approach allows the system to maintain accurate valuations without requiring complete reprocessing of the entire data lake.
3Adaptability or versatility
If contextual business value calculation is implemented, then data valuation relevance is improved, but system complexity worsens
Solution Approach 1:
The patent applies local quality by computing business value in the context of specific data sets and their relevant contexts rather than applying a uniform valuation approach to all data. This allows the system to provide context-relevant valuations while managing complexity through localized computation rather than system-wide complexity.
Data Source
AI summary
A set of domain aware tokens generated for a given unstructured data set are obtained. A value is computed for the given unstructured data set as a function of the set of domain aware tokens and a given context of interest. The value represents a valuation of the unstructured data set for the given context of interest. A placement of the unstructured data set within a data storage environment may be determined based on the computed value.


