Generative Model Text Labeling for Search Accuracy
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current search technologies are limited by the need for users to provide specific keywords and rely on predefined taxonomies, which can lead to irrelevant results if the user's vocabulary is limited, and often fail to capture documents using different terminology, resulting in inefficient and incomplete search outcomes.
Innovation Solution
A system that allows users to specify labels in natural language without prior training, using a generative model to classify candidate text into requested classes, producing semantically rich examples and keywords, reducing the need for manual input and processing power.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If users provide specific keywords for search, then search results can be obtained, but the results may be irrelevant if the user's vocabulary is limited or does not match the terminology in the documents
Solution Approach 1:
The patent introduces an intermediary labeling system that automatically generates labels for documents based on their content. This intermediary layer translates document content into standardized labels that can be used for search, eliminating the need for users to know specific terminology. The system acts as a mediator between the document corpus and the user query, resolving the vocabulary mismatch problem.
Solution Approach 2:
The patent changes the parameter of search from keyword-based matching to label-based classification. Instead of requiring users to provide exact keywords, the system transforms documents into labeled categories and performs classification based on these labels. This parameter change allows users to search using natural language while the system handles the terminology translation automatically.
2Extent of automation
If predefined taxonomies are used for document classification, then documents can be organized, but the system requires great processing power and cost for manual labeling
Solution Approach 1:
The patent implements a self-service labeling system where the documentation system itself generates the labels automatically without requiring external manual intervention. The system uses the content of the documents to generate appropriate labels through automated analysis, making the labeling process self-sufficient and eliminating the need for expensive manual labeling operations.
Solution Approach 2:
The patent performs preliminary labeling actions during the document ingestion and processing phase, rather than requiring separate manual labeling operations. By generating labels automatically as part of the initial document processing workflow, the system prepares the data for future search operations without requiring additional processing power at query time.
3Reliability
If manual labeling systems are developed, then text data can be classified, but the systems are not generally available to search processes and require extensive resources
Solution Approach 1:
The patent creates a universal labeling system that can be applied to any document corpus regardless of domain or terminology. The system generates labels that are both specific to the document content and general enough to be used across different search contexts. This multi-functional labeling approach allows the same system to serve various search needs without requiring domain-specific customization.
Data Source
AI summary
The technology described herein determines whether a candidate text is in a requested class by using a generative model that may not be trained on the requested class. The present technology may use of a model trained primarily in an unsupervised mode, without requiring a large number of manual user-input examples of a label class. The may produce a semantically rich positive example of label text from a candidate text and label. Likewise, the technology may produce from the candidate text and the label a semantically rich negative example of label text. The labeling service makes use of a generative model to produce a generative result, which estimates the likelihood that the label properly applies to the candidate text. In another aspect, the technology is directed toward a method for obtaining a semantically rich example that is similar to a candidate text.


