Automated Data Set Valuation via Semantic Matching

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional information processing systems are limited in their ability to support similarity measures and fail to automate the valuation of data sets, particularly in determining suitability for specific purposes or roles in analytic processes.

Innovation Solution

The implementation of an information processing system that includes a data set discovery engine and a data set valuation engine, which generates similarity measures and valuation measures based on semantic matching, allowing for automated assignment of data set values by examining data sets using a semantic hierarchy and combining valuation measures using weights.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If conventional document-oriented data processing tools are used, then basic data processing is supported, but similarity measures are limited and automated valuation is not provided

Engineering Contradiction:
Improvesimilarity measure supportVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system implements a unified data processing platform that supports multiple similarity measures (frequency-based, semantic, and custom measures) through a common architecture. The graph database framework enables diverse valuation algorithms to operate on the same data structure, providing multi-functionality without requiring separate specialized tools for each measure type.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent introduces intermediate processing layers including a graph database as mediator between raw data and valuation results. This intermediary structure enables complex similarity computations and automated valuation by transforming data into graph representations that can be processed by various algorithms, thereby extending capabilities beyond conventional direct processing tools.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Extent of automation

If conventional data processing tools are used, then processing operations are simple, but automated valuation of data sets is not provided

Engineering Contradiction:
Improveautomated valuationVSAvoidsystem complexity
Core Design Contradiction:
Extent of automationVSDevice complexity

Solution Approach 1:

The system performs preliminary actions by pre-computing and storing similarity measures and data characteristics in graph database structures before valuation is needed. This advance preparation enables automated valuation to occur quickly when requested, as the foundational computational work has already been completed and stored for reuse.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The automated valuation system operates autonomously by automatically computing similarity measures, applying valuation algorithms, and generating valuation results without manual intervention. The system self-manages the entire valuation pipeline from data ingestion through result generation, eliminating the need for manual valuation processes.

Inventive Principle:
Principle #25Self-service

3Measurement precision

If frequency-based similarity measures are used, then computation is efficient, but suitability for specific purposes cannot be determined

Engineering Contradiction:
Improvesuitability assessmentVSAvoidprocessing efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent segments the valuation process into distinct computational stages: frequency-based similarity computation for efficient initial filtering, followed by semantic analysis for precision suitability assessment. This segmentation allows the system to use efficient methods where appropriate while applying more precise but computationally intensive methods only where needed, balancing both criteria.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system applies a tiered approach where frequency-based measures provide partial valuation for quick assessment, and semantic analysis provides excessive (more thorough) action for cases requiring high precision. This partial/excessive action strategy ensures efficiency for routine cases while maintaining the capability for precise suitability assessment when required.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS11550834B1Automated assignment of data set value via semantic matching
Publication Date: 2023.01.10 EMC IP HLDG CO LLC
  • US11550834B1 patent drawing
  • US11550834B1 patent drawing
  • US11550834B1 patent drawing

AI summary

An apparatus comprises a processing platform implementing a data set discovery engine and a data set valuation engine. The data set discovery engine is configured to generate data set similarity measures each relating a corresponding one of a plurality of data sets to one or more other ones of the plurality of data sets. The data set valuation engine is coupled to the data set discovery engine and configured to generate valuation measures for respective ones of at least a subset of the plurality of data sets based at least in part on respective ones of the data set similarity measures generated by the data set discovery engine. For example, the data set valuation engine may generate the valuation measure for a given data set as a function of valuation measures previously generated for respective other data sets determined to exhibit at least a threshold similarity to the given data set.