Generative Model Text Labeling for Search Accuracy

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current search technologies are limited by the need for users to provide specific keywords and rely on predefined taxonomies, which can lead to irrelevant results if the user's vocabulary is limited, and often fail to capture documents using different terminology, resulting in inefficient and incomplete search outcomes.

Innovation Solution

A system that allows users to specify labels in natural language without prior training, using a generative model to classify candidate text into requested classes, producing semantically rich examples and keywords, reducing the need for manual input and processing power.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If users provide specific keywords for search, then search results can be obtained, but the results may be irrelevant if the user's vocabulary is limited or does not match the terminology in the documents

Engineering Contradiction:
Improvesearch accuracyVSAvoiduser effort
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The patent introduces an intermediary labeling system that automatically generates labels for documents based on their content. This intermediary layer translates document content into standardized labels that can be used for search, eliminating the need for users to know specific terminology. The system acts as a mediator between the document corpus and the user query, resolving the vocabulary mismatch problem.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent changes the parameter of search from keyword-based matching to label-based classification. Instead of requiring users to provide exact keywords, the system transforms documents into labeled categories and performs classification based on these labels. This parameter change allows users to search using natural language while the system handles the terminology translation automatically.

Inventive Principle:
Principle #35Parameter changes

2Extent of automation

If predefined taxonomies are used for document classification, then documents can be organized, but the system requires great processing power and cost for manual labeling

Engineering Contradiction:
Improveautomatic classificationVSAvoidprocessing power
Core Design Contradiction:
Extent of automationVSUse of energy by stationary object

Solution Approach 1:

The patent implements a self-service labeling system where the documentation system itself generates the labels automatically without requiring external manual intervention. The system uses the content of the documents to generate appropriate labels through automated analysis, making the labeling process self-sufficient and eliminating the need for expensive manual labeling operations.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent performs preliminary labeling actions during the document ingestion and processing phase, rather than requiring separate manual labeling operations. By generating labels automatically as part of the initial document processing workflow, the system prepares the data for future search operations without requiring additional processing power at query time.

Inventive Principle:
Principle #10Preliminary action

3Reliability

If manual labeling systems are developed, then text data can be classified, but the systems are not generally available to search processes and require extensive resources

Engineering Contradiction:
Improveclassification accuracyVSAvoidsystem availability
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent creates a universal labeling system that can be applied to any document corpus regardless of domain or terminology. The system generates labels that are both specific to the document content and general enough to be used across different search contexts. This multi-functional labeling approach allows the same system to serve various search needs without requiring domain-specific customization.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20240370484A1Automatic labeling of text data
Publication Date: 2024.11.07 MICROSOFT TECHNOLOGY LICENSING LLC
  • US20240370484A1 patent drawing
  • US20240370484A1 patent drawing
  • US20240370484A1 patent drawing

AI summary

The technology described herein determines whether a candidate text is in a requested class by using a generative model that may not be trained on the requested class. The present technology may use of a model trained primarily in an unsupervised mode, without requiring a large number of manual user-input examples of a label class. The may produce a semantically rich positive example of label text from a candidate text and label. Likewise, the technology may produce from the candidate text and the label a semantically rich negative example of label text. The labeling service makes use of a generative model to produce a generative result, which estimates the likelihood that the label properly applies to the candidate text. In another aspect, the technology is directed toward a method for obtaining a semantically rich example that is similar to a candidate text.