Discovery Avatars Using Vector Models for Relevant Data Search

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional data search techniques, such as keyword searching and Boolean operators, are inadequate for efficiently organizing and discovering relevant data in large repositories due to mismatches in keyword definitions and lack of intelligence, leading to incomplete or overly inclusive search results.

Innovation Solution

Development of computer-based discovery avatars that learn from human analysts' interactions to provide intelligent data organization and retrieval by tokenizing data, extracting features, and using mathematical models to cluster and score data elements based on relevance to a topic, allowing for iterative improvement and deployment across various data sources.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional keyword searching and Boolean operators are used, then the search process is simple and fast, but the search results are incomplete or overly inclusive due to keyword mismatches and lack of intelligence

Engineering Contradiction:
Improvesearch accuracyVSAvoidsearch system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent transforms the search approach by changing parameters from simple keyword matching to multi-dimensional data characterization using quantitative vectors. Each data element is represented by multiple features (textual, structural, contextual) that are converted into numerical vectors, enabling more precise comparison and matching beyond exact keyword matches.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces an intermediary mathematical model that acts as a mediator between the search query and data elements. This model uses quantitative vector representation and similarity calculations to bridge the gap between simple keyword search and intelligent relevance assessment, resolving the contradiction between simplicity and accuracy.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If keyword matching is used to find relevant data, then the search covers all documents containing the keyword, but the results are too voluminous for human review in acceptable time

Engineering Contradiction:
Improverelevance identificationVSAvoidreview time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies local quality by evaluating different features of data elements differently. Instead of treating all data equally, the system assigns weights to different quantitative features (textual content, structure, context) and calculates localized relevance scores for each data element based on its specific characteristics, enabling efficient filtering of voluminous results.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent uses partial action by not requiring exact keyword matches but rather partial feature matching through quantitative vector comparison. The system calculates similarity scores based on overlapping features and presents results in order of relevance, allowing reviewers to focus on the most promising candidates first rather than reviewing all matching documents.

Inventive Principle:
Principle #16Partial or excessive action

3Measurement precision

If keyword search is used, then the search process is straightforward, but keyword matches lack intelligence and combine documents simply on the basis of sharing a word with substantively different meanings

Engineering Contradiction:
Improvesemantic understandingVSAvoidanalysis complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the analysis process into multiple independent feature extraction steps. Each data element is broken down into distinct quantitative features (textual features, structural features, contextual features) that are analyzed separately and then combined. This segmentation enables intelligent semantic understanding by considering multiple aspects of the data rather than relying on a single keyword match.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates a composite representation of data elements by combining multiple quantitative features into a unified vector profile. This composite approach integrates textual content, structural properties, and contextual information to create a holistic view of each data element's relevance, enabling intelligent differentiation between documents that share keywords but have different meanings.

Inventive Principle:
Principle #40Composite materials

Data Source

PatentUS12361314B2Creation, use and training of computer-based discovery avatars
Publication Date: 2025.07.15 GCP IP HLDG I LLC
  • US12361314B2 patent drawing
  • US12361314B2 patent drawing
  • US12361314B2 patent drawing

AI summary

In embodiments of the present invention improved capabilities are described for developing, training, validating and deploying discovery avatars embodying mathematical models that may be used for document and data discovery and deployed within large data repositories. For example, an avatar may be constructed by machine learning processes, including by processing information related to what types of information analysts find useful in large data sets. Once constructed, an avatar may be deployed as an aid to human intuition in a wide range of analytical processes, such as related to national security, enterprise management (e.g., programs related to sales, marketing, product, promotions, placement, pricing and the like), dispute resolution (including litigation), forensic analysis, criminal, administrative, civil and private investigations, scientific investigations, research and development, and a wide range of others.