Automated Data Set Valuation via Semantic Matching
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional information processing systems are limited in their ability to support similarity measures and fail to automate the valuation of data sets, particularly in determining suitability for specific purposes or roles in analytic processes.
Innovation Solution
The implementation of an information processing system that includes a data set discovery engine and a data set valuation engine, which generates similarity measures and valuation measures based on semantic matching, allowing for automated assignment of data set values by examining data sets using a semantic hierarchy and combining valuation measures using weights.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If conventional document-oriented data processing tools are used, then basic data processing is supported, but similarity measures are limited and automated valuation is not provided
Solution Approach 1:
The system implements a unified data processing platform that supports multiple similarity measures (frequency-based, semantic, and custom measures) through a common architecture. The graph database framework enables diverse valuation algorithms to operate on the same data structure, providing multi-functionality without requiring separate specialized tools for each measure type.
Solution Approach 2:
The patent introduces intermediate processing layers including a graph database as mediator between raw data and valuation results. This intermediary structure enables complex similarity computations and automated valuation by transforming data into graph representations that can be processed by various algorithms, thereby extending capabilities beyond conventional direct processing tools.
2Extent of automation
If conventional data processing tools are used, then processing operations are simple, but automated valuation of data sets is not provided
Solution Approach 1:
The system performs preliminary actions by pre-computing and storing similarity measures and data characteristics in graph database structures before valuation is needed. This advance preparation enables automated valuation to occur quickly when requested, as the foundational computational work has already been completed and stored for reuse.
Solution Approach 2:
The automated valuation system operates autonomously by automatically computing similarity measures, applying valuation algorithms, and generating valuation results without manual intervention. The system self-manages the entire valuation pipeline from data ingestion through result generation, eliminating the need for manual valuation processes.
3Measurement precision
If frequency-based similarity measures are used, then computation is efficient, but suitability for specific purposes cannot be determined
Solution Approach 1:
The patent segments the valuation process into distinct computational stages: frequency-based similarity computation for efficient initial filtering, followed by semantic analysis for precision suitability assessment. This segmentation allows the system to use efficient methods where appropriate while applying more precise but computationally intensive methods only where needed, balancing both criteria.
Solution Approach 2:
The system applies a tiered approach where frequency-based measures provide partial valuation for quick assessment, and semantic analysis provides excessive (more thorough) action for cases requiring high precision. This partial/excessive action strategy ensures efficiency for routine cases while maintaining the capability for precise suitability assessment when required.
Data Source
AI summary
An apparatus comprises a processing platform implementing a data set discovery engine and a data set valuation engine. The data set discovery engine is configured to generate data set similarity measures each relating a corresponding one of a plurality of data sets to one or more other ones of the plurality of data sets. The data set valuation engine is coupled to the data set discovery engine and configured to generate valuation measures for respective ones of at least a subset of the plurality of data sets based at least in part on respective ones of the data set similarity measures generated by the data set discovery engine. For example, the data set valuation engine may generate the valuation measure for a given data set as a function of valuation measures previously generated for respective other data sets determined to exhibit at least a threshold similarity to the given data set.


