Data Set Alignment via Dynamic Term Weighting

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for analyzing large data sets, such as keyword searches and structured tagging systems, are limited by their reliance on predefined categories and lack of adaptability to changing areas of interest, leading to inconsistent and less relevant results when comparing or combining different data sets.

Innovation Solution

A content analyzer subsystem that employs co-clustering techniques and weighting algorithms to extract and prioritize features from data sets, allowing for the comparison and alignment of unstructured data across different datasets by generating semantic representations and specificity measures, thereby enabling dynamic and accurate analysis.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If keyword searches and structured tagging systems are used to analyze large data sets, then the analysis process is simple and fast, but the results are inconsistent and less relevant when comparing or combining different data sets due to reliance on predefined categories

Engineering Contradiction:
Improvealignment accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent introduces an intermediary alignment process that maps terms from different data sets to a common reference framework. This intermediary layer enables accurate comparison and combination of data sets without requiring complex preprocessing or predefined categorization schemes, thereby improving alignment accuracy while maintaining system simplicity.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system dynamically adjusts weighting parameters for different terms based on their specificity and relevance to the analysis task. By changing these parameters adaptively rather than using fixed predefined categories, the system achieves higher alignment accuracy across diverse data sets without increasing structural complexity.

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If predefined categories are used for classifying records, then the classification process is straightforward, but the system lacks adaptability to changing areas of interest

Engineering Contradiction:
Improveadaptability to changing interestsVSAvoidoperation simplicity
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

The patent implements dynamic term weighting that automatically adapts to changing areas of interest by adjusting the importance of different terms based on their specificity measures. This dynamic approach allows the system to remain versatile and adaptable without requiring manual recategorization or complex operational changes, thus maintaining ease of operation while improving adaptability.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system performs self-adjustment by automatically computing specificity measures and updating term weights based on the data being analyzed. This self-service capability enables the system to adapt to changing interests autonomously without requiring external intervention or complex operational procedures, thereby maintaining simplicity while enhancing versatility.

Inventive Principle:
Principle #25Self-service

3Adaptability or versatility

If traditional tagging methods are used, then the implementation is simple, but the system cannot effectively handle diverse data types and changing interests

Engineering Contradiction:
Improvehandling diverse data typesVSAvoidinformation relevance
Core Design Contradiction:
Adaptability or versatilityVSLoss of information

Solution Approach 1:

The patent employs parameter changes by dynamically adjusting term weights based on specificity measures computed from the actual data. This allows the system to effectively handle diverse data types by adapting to their unique characteristics while preserving information relevance, overcoming the limitations of static traditional tagging methods.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The system performs preliminary computation of specificity measures for all terms before the actual analysis. This preliminary action prepares the system to handle diverse data types effectively by pre-establishing relevance metrics, thereby preventing information loss during subsequent comparison and alignment operations without adding operational complexity.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10366108B2Distributional alignment of sets
Publication Date: 2019.07.30 SRI INTERNATIONAL
  • US10366108B2 patent drawing
  • US10366108B2 patent drawing
  • US10366108B2 patent drawing

AI summary

Technology for classifying a data set includes extracting one or more features from items of the data set, computing a specificity measure for the extracted features, and measuring the similarity of the extracted features to a set of characteristic features associated with the property of one or more reference models.