Semantic Fingerprinting for Third-Party Knowledge Graph Onboarding

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional techniques for onboarding third-party entity data with general-purpose knowledge graphs are inefficient, requiring significant computational and human resources, and often involve trial and error due to format differences and domain-specific data that may not represent widely-known information.

Innovation Solution

The described techniques improve the onboarding process by analyzing third-party entity data to identify semantic fingerprints, which group similar entities and apply rules to determine success or failure, providing failure statistics and suggested actions to streamline data integration and reduce errors.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional techniques are used to onboard third-party entity data with knowledge graphs, then data integration can be achieved, but the process requires significant computational and human resources and involves trial and error

Engineering Contradiction:
Improveonboarding success rateVSAvoidonboarding efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system performs preliminary analysis of third-party entity data to identify semantic fingerprints and generate failure statistics before actual onboarding. This preliminary action includes analyzing data formats, identifying potential mismatches with knowledge graph schemas, and predicting failure points, thereby reducing trial-and-error iterations during the actual onboarding process

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system implements a feedback mechanism where failure statistics from preliminary analysis are provided back to the third-party data provider. This feedback includes specific information about format mismatches, missing entities, and relationship issues, enabling the provider to correct data before re-submission, thereby improving onboarding success rate and reducing iterative corrections

Inventive Principle:
Principle #23Feedback

2Manufacturing precision

If manual analysis and correction of third-party entity data is performed, then data quality can be improved, but significant human resources are required

Engineering Contradiction:
Improvedata integration accuracyVSAvoidhuman resources required
Core Design Contradiction:
Manufacturing precisionVSQuantity of substance

Solution Approach 1:

The system enables self-service automated analysis where the onboarding platform automatically identifies semantic fingerprints in third-party data, compares them against knowledge graph schemas, and generates detailed failure statistics without human intervention. This automated self-analysis maintains high data integration accuracy while eliminating the need for manual data quality assessment

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system replaces manual human analysis with automated computational mechanisms that use semantic fingerprinting algorithms and statistical analysis. These mechanical/automated systems perform data quality assessment, format validation, and mismatch identification with higher precision and without requiring human resources

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Adaptability or versatility

If third-party entity data with format differences and domain-specific information is onboarded, then comprehensive data coverage is achieved, but errors and iterations increase

Engineering Contradiction:
Improvedata format compatibilityVSAvoidonboarding error rate
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The system dynamically adjusts analysis parameters and semantic fingerprinting criteria based on the specific format and domain characteristics of third-party entity data. By changing parameters such as matching thresholds, semantic similarity weights, and validation strictness according to data type, the system maintains high adaptability across different formats while reducing error rates through optimized analysis configurations

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20240086735A1Onboarding of entity data
Publication Date: 2024.03.14 GOOGLE LLC
  • US20240086735A1 patent drawing
  • US20240086735A1 patent drawing
  • US20240086735A1 patent drawing

AI summary

Techniques are described herein to improve onboarding of third party entity data with existing knowledge graphs. In various implementations, a computing system associated with an existing knowledge graph may receive a request from a third party to onboard, with the existing knowledge graph, a plurality of entities. Each entity may have associated identifier(s) and relationship(s) with other entities of the plurality of entities. First third party entity data that describes the plurality of entities and associated identifiers/relationships may be received from the third party. The first third entity party data may be analyzed to identify semantic fingerprint(s) matching respective subsets of the entities. Results related to the analyzing may be determined. The results may include a statistic representing success or failure of applying rule(s) to a respective subset of entities that match a given semantic fingerprint. Remedial action(s) may be triggered based on the failure statistic.