Federated Genealogical Search Ranking with ML Weighting

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Genealogical search systems face challenges in scoring results from diverse data sources due to variations in data type, relevance, and inexact matching, making it difficult to return the best possible results while maintaining variety.

Innovation Solution

A search system performs a federated search across multiple databases, using a machine learning model to rank records based on user queries, with query expansion and weighting to account for variations in data relevance and type, trained with historical search data and optimized using algorithms like Nelder-Mead and simulated annealing.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a search system scores records from diverse databases using traditional methods, then the system structure remains simple, but the accuracy of ranking results deteriorates due to variations in data type, relevance, and record size across different sources

Engineering Contradiction:
Improveranking accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent introduces a machine learning model as an intermediary component between the diverse databases and the search results. This model learns optimal weighting strategies from historical data and acts as a mediator to harmonize the varying data types, formats, and relevance criteria across different databases, thereby improving ranking accuracy without requiring complex manual scoring rules

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system implements feedback mechanisms by using historical user interactions (clicks, views, saves) to train and continuously improve the machine learning model. The model learns from past search behaviors and adjusts its weighting strategy accordingly, creating a closed-loop system where ranking accuracy improves over time based on actual user preferences

Inventive Principle:
Principle #23Feedback

2Adaptability or versatility

If the system searches across multiple diverse databases with different data types and formats, then the variety of results increases, but the difficulty of equitably scoring and ranking results worsens

Engineering Contradiction:
Improveresult varietyVSAvoidscoring difficulty
Core Design Contradiction:
Adaptability or versatilityVSDifficulty of detecting and measuring

Solution Approach 1:

The machine learning model dynamically changes the weighting parameters for different databases based on the query characteristics and historical performance. Instead of using fixed scoring rules, the system adjusts parameters like database weights, relevance thresholds, and matching criteria adaptively, allowing equitable scoring across diverse data types while maintaining result variety

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent segments the scoring process into multiple independent components: query expansion, database-specific filtering, relevance scoring, and final ranking. Each component handles specific aspects of the diverse data independently, making the overall scoring process more manageable and equitable across different database types

Inventive Principle:
Principle #1Segmentation

3Measurement precision

If the system uses exact matching for queries, then the precision of matches improves, but the quantity of relevant records decreases due to missing data, misspellings, and variations in record specifications

Engineering Contradiction:
Improvematch precisionVSAvoidnumber of records
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The system performs query expansion as a preliminary action before searching the databases. It generates multiple variations of the query including phonetic spellings, synonym expansions, and fuzzy matching patterns. This preliminary expansion ensures that the subsequent exact matching can find more records while maintaining precision through the expanded search space

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11500884B2Search and ranking of records across different databases
Publication Date: 2022.11.15 ANCESTRY COM OPERATIONS INC
  • US11500884B2 patent drawing
  • US11500884B2 patent drawing
  • US11500884B2 patent drawing

AI summary

A search system performs a federated search across multiple databases and generates a ranked combined list of found genealogical records. The system receives a user query with one or more specified characteristics. The system may determine expanded characteristics derived from the specified characteristics. The system searches the various databases with the characteristics retrieving records according to the characteristics. The system combines the retrieved records and ranks them using a machine learning model. The machine learning model is configured to assign a weight to the records returned from each of the genealogical databases based on the characteristics specified in the user query. The machine learning model may be trained by any combination of one or more of: a Nelder-Mead method, a coordinate ascent method, and a simulated annealing method. The ranked combined results are provided in response to the user query.