Domain-Specific Relevance Service for Theme Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing search engines face challenges in accurately identifying and providing relevant information to users due to the abundance of data available, as they rely on pre-processing and manual categorization methods that are inefficient in determining themes and content relevance across vast domains.

Innovation Solution

The development of a Domain-Specific Relevance Determination (DSRD) service that automatically analyzes documents to determine relevant themes and content items by using techniques like TF-IDF analysis and machine learning, enabling the generation of relevance information that can be used to improve search results and user feedback loops.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If automated pre-processing and manual categorization are used to index documents, then search engines can identify documents matching search terms, but the accuracy of determining relevant themes and content remains insufficient due to the abundance of data

Engineering Contradiction:
Improveaccuracy of determining relevant themes and contentVSAvoidabundance of available information
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent replaces manual categorization and traditional automated indexing mechanisms with a machine learning-based system. The system automatically determines theme relevance and content relevance by analyzing document features, user interactions, and feedback data through trained models, eliminating the need for manual categorization while improving accuracy in identifying relevant information among abundant data

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent implements feedback loops where user interactions with search results (clicks, dwell time, refinements) are collected and used to retrain and improve the machine learning models. This continuous feedback mechanism enhances the system's ability to accurately determine theme and content relevance over time, addressing the challenge of filtering relevant information from abundant data

Inventive Principle:
Principle #23Feedback

2Adaptability or versatility

If traditional search indexing methods are used, then documents can be mapped to search terms, but the system cannot effectively determine themes across vast domains

Engineering Contradiction:
Improveability to determine themes across domainsVSAvoidcomplexity of theme determination system
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent segments the theme determination process into distinct components: document feature extraction, theme identification models, content relevance analysis, and user interaction processing. This segmentation allows the system to handle complex multi-domain theme determination through modular machine learning components, making the overall system manageable despite its versatility across vast domains

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates a universal machine learning-based theme determination system that can handle multiple domains and document types through a single integrated platform. The system uses generalizable models trained on diverse data that can adapt to different domains (technology, health, finance, etc.) without requiring separate specialized systems for each domain

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Productivity

If automated programs crawl the web to create indexes, then document identification is improved, but the process is inefficient in determining content relevance

Engineering Contradiction:
Improveefficiency of determining content relevanceVSAvoidtime for pre-processing and indexing
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent performs preliminary extraction and analysis of document features during the indexing phase, preparing data structures that enable rapid relevance determination during search operations. By pre-processing documents to extract key features, metadata, and thematic elements before search queries are submitted, the system reduces the time required for content relevance determination when actual search requests are processed

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces traditional mechanical web crawling and manual relevance assessment with automated machine learning models that efficiently evaluate content relevance. The system uses trained models to rapidly analyze document features and determine relevance to search queries and identified themes, significantly improving productivity compared to traditional automated indexing methods

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS8706664B2Determining relevant information for domains of interest
Publication Date: 2014.04.22 VERITONE INC
  • US8706664B2 patent drawing
  • US8706664B2 patent drawing
  • US8706664B2 patent drawing

AI summary

Techniques are described for determining and using relevant information related to domains of interest. In at least some situations, the techniques include automatically analyzing documents, terms and other information related to a domain of interest in order to automatically determine information about relevant themes within the domain and/or about which documents have contents that are relevant to such themes. Such automatically determined information related to a domain may then be used in various ways, including to assist users in specifying themes of interest and/or in obtaining documents and/or document fragments with contents that are relevant to specified themes. In addition, information about how the automatically determined information is used by users may be tracked and used as feedback for learning improved determinations of relevant themes and relevant documents within the domain, such as by using automated machine learning techniques.