Semantic Data Linking for Cross-Repository Relationship Discovery
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data classification and search methodologies fail to uncover relevant data due to the use of various terminologies and synonyms, leading to data democratization issues and the inability to discover latent relationships between data sets.
Innovation Solution
Assigning tags to data items that include synonyms and using controlled terms, ranking data items based on similarity, and displaying them at distances related to shared term frequency to uncover new relationships.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If keyword or Boolean search tools are used to search large data repositories, then search operations can be performed, but relevant data cannot be found due to the size of the repository and concealed information
Solution Approach 1:
The patent introduces an intermediary system that acts as a mediator between the large data repository and the search function. This intermediary automatically generates and maintains indexes, metadata, and summaries of data items, enabling efficient search without requiring users to navigate the entire repository. The intermediary translates complex repository structures into searchable formats, resolving the contradiction between repository size and data discoverability.
Solution Approach 2:
The system performs preliminary actions by automatically indexing and categorizing data items before they are searched. Data are pre-processed to extract key features, create metadata profiles, and establish relationships with other data items. This preliminary organization enables rapid retrieval and discovery of relevant data without requiring full repository scanning during search operations.
2Loss of information
If data are recorded in searchable format with known data types, then search can be performed, but latent relationships between data sets cannot be discovered
Solution Approach 1:
The patent replaces manual search operations with an automated intelligent system that uses machine learning and semantic analysis. Instead of requiring users to manually construct complex search queries, the system automatically analyzes data relationships, identifies patterns, and discovers latent connections between data sets. This substitution transforms the search process from a manual mechanical operation to an automated intelligent process, enabling relationship discovery without increasing user burden.
Solution Approach 2:
The system implements feedback mechanisms where search results and user interactions are continuously analyzed to improve relationship discovery. The system learns from search patterns, user preferences, and data correlations to refine its understanding of data relationships. This feedback loop enables the system to automatically identify and present latent relationships that would otherwise remain hidden, while maintaining ease of operation through adaptive result presentation.
3Adaptability or versatility
If various terminologies and synonyms are used in data classification, then data can be categorized, but data democratization is hindered and uniform searching becomes difficult
Solution Approach 1:
The patent applies parameter changes by transforming various terminologies and synonyms into a unified semantic framework. The system maintains the original diverse terminology in the data while introducing standardized semantic parameters and controlled vocabularies for indexing and search. This dual-parameter approach allows the system to accommodate terminology flexibility in the source data while enforcing search uniformity through standardized semantic mapping and normalization processes.
Data Source
AI summary
A method and system for identifying connections between data items in an updatable data repository includes dynamically receiving a plurality of data items, each data item including one or more controlled terms, displaying the dynamically received data items on a display, and identifying one or more connections between data items from different updatable sources based on the display. A method and system for identifying connections between data items in an updatable data repository also includes dynamically receiving a plurality of data items, each data item including one or more controlled terms, defining degrees of similarity between the data items, ranking the data items with respect to one another based on the degree of similarity therebetween, and identifying one or more connections between data items from different updatable sources based on the ranking. The data items are also updated contemporaneously when the sources thereof are updated.


