Semantic Filtering Coefficient for Dataset Relevance Ranking
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data matching systems often include irrelevant and unrelated datasets in search results, making it difficult for users to find truly relevant and related datasets for data enrichment and analytical insights.
Innovation Solution
A system utilizing a semantic filtering module that calculates a semantic correlation coefficient by combining weighted domain and geography coefficients with correlation coefficients to determine the relevance of reference datasets to a user dataset, thereby filtering out less relevant results.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If data matching systems include more datasets in search results, then the quantity of available data increases, but the relevance and quality of matching results deteriorates due to inclusion of irrelevant datasets
Solution Approach 1:
The patent applies parameter changes by introducing a semantic filtering coefficient that combines multiple parameters (correlation coefficient, domain weight, geography weight) to evaluate dataset relevance. This multi-parameter approach transforms the single-dimensional matching into a comprehensive evaluation system that maintains precision while handling large quantities of datasets.
Solution Approach 2:
The semantic filtering coefficient acts as an intermediary between the correlation coefficient and the final relevance determination. It mediates the evaluation process by integrating domain and geography weights with correlation data, providing a nuanced assessment that improves relevance accuracy without limiting the quantity of datasets considered.
2Measurement precision
If semantic filtering calculation is performed for all reference datasets, then the relevance accuracy improves, but the computational complexity and processing time increases
Solution Approach 1:
The patent segments the semantic filtering process into distinct components: correlation coefficient calculation, domain weight determination, geography weight determination, and final coefficient computation. This segmentation allows for modular implementation and optimization of each component independently, reducing overall computational complexity while maintaining accuracy.
Solution Approach 2:
The system applies partial action by calculating the semantic filtering coefficient for datasets based on their likelihood of relevance. Rather than uniformly processing all reference datasets with equal computational resources, the system can prioritize datasets that show higher correlation or belong to more relevant domains, reducing unnecessary computational effort on clearly irrelevant datasets.
3Loss of information
If the semantic filtering module calculates correlation coefficients for all datasets, then the completeness of matching results improves, but the processing time and computational resources increase
Solution Approach 1:
The patent applies preliminary action by pre-determining domain weights and geography weights before performing the actual semantic filtering. These pre-calculated weights are stored and reused during the filtering process, eliminating the need to recalculate them for each dataset comparison. This preliminary preparation significantly reduces processing time while maintaining the completeness of the matching results.
Data Source
AI summary
A computer-implemented method for finding related datasets includes, for each reference dataset from multiple reference datasets, determining domains and geographies for a user dataset and the reference dataset, obtaining a weighted domain coefficient and a weighted geography coefficient using the determined domains and geographies for the user dataset and the reference dataset, calculating a correlation coefficient between the user dataset and the reference dataset and calculating a semantic filtering coefficient for the user dataset and the reference dataset using the calculated correlation coefficient, the weighted domain coefficient and the weighted geography coefficient.


