Dataset Ranking via Field Suitability and Attribute Scoring
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing dataset evaluation methods fail to effectively match user-specific data requirements, leading to inefficiencies in identifying and ranking datasets based on their relevance and suitability for specific use cases.
Innovation Solution
A computer-implemented method that identifies target data fields and attributes from user-provided documents, generates metadata sets for datasets, and assesses their suitability using a comparison scoring system to rank datasets based on their relevance and content alignment with user needs.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional dataset evaluation methods are used, then the evaluation process is simple, but the accuracy of matching user-specific data requirements is poor
Solution Approach 1:
The patent segments the dataset evaluation process into multiple independent components: metadata extraction from process documents, metadata extraction from data use documents, candidate dataset identification, and scoring/ranking. Each component handles a specific aspect of the evaluation, improving matching accuracy while keeping individual components manageable in complexity
Solution Approach 2:
The patent introduces multiple parameters for dataset evaluation including field suitability values, metadata attribute scores, and compared attribute scores. These parameters transform the evaluation from a simple binary assessment to a multi-dimensional scoring system that accurately reflects user requirements while providing structured complexity
2Measurement precision
If comprehensive metadata assessment is performed for all datasets, then the ranking accuracy improves, but the processing time increases
Solution Approach 1:
The patent performs preliminary actions by extracting and storing metadata from process documents and data use documents before the actual dataset evaluation. Candidate datasets are pre-identified based on initial criteria, so that when comprehensive metadata assessment is performed, it is only on a reduced set of promising candidates, maintaining accuracy while reducing overall processing time
Solution Approach 2:
The patent applies partial assessment action by performing comprehensive metadata evaluation only on candidate datasets that meet initial suitability thresholds, rather than assessing all datasets in the repository. This selective approach maintains ranking accuracy for relevant datasets while significantly reducing processing time for the overall system
3Reliability
If multiple data attributes are evaluated, then the dataset selection quality improves, but the computational complexity increases
Solution Approach 1:
The patent segments the multiple data attributes into distinct metadata categories extracted from different document types. Process documents provide one set of attributes while data use documents provide another set. This segmentation allows the system to evaluate multiple attributes through separate, specialized extraction processes rather than a single complex evaluation
Solution Approach 2:
The patent introduces metadata as an intermediary layer between the raw datasets and the evaluation process. Instead of directly comparing datasets against user requirements, the system uses extracted metadata attributes as intermediaries that bridge the two, simplifying the computational complexity while maintaining selection quality
Data Source
AI summary
Ranking a group of datasets using a computer includes determining a set of target data fields from a set of process documents that indicate user data field preferences. A set of target dataset attributes from a set of data use documents indicate user data scope preferences. A plurality of metadata sets for an associated plurality of datasets the computer determines having a field suitability value exceeding a predetermined suitability threshold value. The FSV represents a degree of similarity between a set of fields associated with said dataset and the set of target data fields. The computer assesses metadata sets with regard to the target attributes and generates a compared attribute score for each candidate dataset. A degree of likelihood is indicated that an associated dataset will have content exhibiting said target dataset attributes. The computer candidate datasets is based on the compared attribute score.


