Trait-Based Framework for Linking Text Documents and Structured Records

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current solutions for linking information from various data sources are limited by the need for accurate pre-categorization of documents and structured records, and are often dependent on good taxonomy and classification, which is difficult to achieve in practice, especially with the vast and varied forms of data available online.

Innovation Solution

A framework that utilizes instance-based 'traits' - sets of characteristics serving as proxies for object identity - to map and join information from text documents and structured records, allowing for the association of records with documents without assuming the structure of the data sources, and using a scoring function to determine the relevance of text documents to structured records based on shared traits.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If current solutions use pre-categorization and taxonomy-based approaches to link information from various data sources, then matching accuracy can be improved, but the complexity of achieving good taxonomy and classification increases significantly

Engineering Contradiction:
Improvematching accuracyVSAvoidtaxonomy construction complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent inverts the traditional approach by not starting with taxonomy and classification, but rather with computing instance-based traits directly from the data. Instead of categorizing data first and then matching, the system computes traits from records and documents independently, then uses these traits for matching. This inversion eliminates the need for complex pre-categorization while maintaining matching accuracy.

Inventive Principle:
Principle #13The other way round (Inversion)

Solution Approach 2:

The patent introduces traits as an intermediary representation that bridges structured records and unstructured documents. Traits serve as a common language that both data sources can be mapped to independently, without requiring complex taxonomy. The scoring function then acts as another intermediary to evaluate the quality of matches based on shared traits, resolving the contradiction between accuracy and complexity.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If traditional methods require accurate classification of documents and structured records according to taxonomy, then information linking can be achieved, but the difficulty of achieving good taxonomy and classification in practice increases

Engineering Contradiction:
Improveinformation linking reliabilityVSAvoidtaxonomy construction difficulty
Core Design Contradiction:
ReliabilityVSDifficulty of detecting and measuring

Solution Approach 1:

The patent enables the system to self-organize by computing traits directly from the data without human intervention for taxonomy construction. The trait computation process automatically identifies distinguishing characteristics from the data itself, and the scoring function automatically evaluates matches. This self-service approach eliminates the manual taxonomy construction difficulty while maintaining reliable information linking.

Inventive Principle:
Principle #25Self-service

3Adaptability or versatility

If the framework uses instance-based traits instead of schema-based database keys, then flexibility in handling various data sources is improved, but the complexity of computing and mapping traits increases

Engineering Contradiction:
Improvedata source flexibilityVSAvoidtrait computation complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent creates a universal trait-based framework that can handle both structured records and unstructured documents through the same mechanism. The trait computation process is designed to work with different data types and formats, and the scoring function universally evaluates matches across all data sources. This universality provides flexibility while the automated computation keeps complexity manageable.

Inventive Principle:
Principle #6Universality (Multi-functionality)

4Ease of operation

If the system processes text documents as bags of words without categorization, then ease of processing is improved, but the ability to accurately identify objects being discussed decreases

Engineering Contradiction:
Improvetext processing simplicityVSAvoidobject identification accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent uses traits as an intermediary that bridges simple bag-of-words representation and accurate object identification. The trait computation process extracts meaningful characteristics from the bag of words, and the scoring function uses these traits to accurately identify objects being discussed. This intermediary approach maintains processing simplicity while achieving accurate object identification.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS8996539B2Composing text and structured databases
Publication Date: 2015.03.31 MICROSOFT TECHNOLOGY LICENSING LLC
  • US8996539B2 patent drawing
  • US8996539B2 patent drawing
  • US8996539B2 patent drawing

AI summary

A framework is provided for composing texts about objects with structured information about these objects, and thus disclosed are methodologies for linking information from at least two data sources—one comprising a plurality of documents comprising text pertaining to at least one object, and one comprising a plurality of structured records comprising at least one characteristic of the at least one object, each characteristic comprising one property name and an associated property value corresponding to the property name for the at least one object—by determining one or more instance-based traits for each object in both data sources and associating at least one record with at least one document that refers to each object, each trait comprising one or more characteristics that identifiably distinguish each object from all other objects.