Generative Model Text Labeling via Semantic Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current search technologies rely on predefined keywords and taxonomies, limiting relevance due to vocabulary differences between user queries and document terminology, and require extensive manual labeling and processing power, making them inefficient and inaccessible to users.

Innovation Solution

A system that classifies candidate text into user-defined labels without prior training data, using a generative model to produce semantically rich examples and keywords, reducing the need for manual input and processing power, and allowing for context-aware keyword extraction.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual labeling systems are used to classify text data, then classification accuracy can be achieved, but processing power and cost increase significantly

Engineering Contradiction:
Improveclassification accuracyVSAvoidprocessing power
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The system enables users to perform their own text classification by providing them with extracted keywords and examples directly, eliminating the need for external manual labeling services while maintaining high classification accuracy

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system extracts key semantic elements (keywords and examples) from the text data that capture the essential classification characteristics, allowing accurate classification without processing the entire text corpus manually

Inventive Principle:
Principle #2Taking out (Extraction)

2Stability of the object's composition

If predefined taxonomies and keywords are used for search, then search structure is maintained, but relevance decreases due to vocabulary differences between user queries and document terminology

Engineering Contradiction:
Improvesearch structureVSAvoidsearch relevance
Core Design Contradiction:
Stability of the object's compositionVSMeasurement precision

Solution Approach 1:

The system dynamically adapts the search parameters by extracting context-relevant keywords and examples that bridge the vocabulary gap between user queries and document terminology, maintaining structural integrity while improving relevance matching

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The extracted keywords and examples act as intermediaries that translate user queries into the terminology used in the document corpus, enabling accurate matching without requiring users to know the predefined taxonomy

Inventive Principle:
Principle #24Intermediary (Mediator)

3Reliability

If extensive manual training data is collected for model training, then classification performance improves, but time and resource requirements increase

Engineering Contradiction:
Improveclassification performanceVSAvoidtraining time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system extracts a small set of highly representative keywords and examples from the training data that capture the essential classification patterns, achieving good performance without requiring extensive manual training data collection

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system performs preliminary extraction of semantic elements (keywords and examples) that pre-process the training data into a condensed form that is ready for immediate use in classification, reducing the time needed for data preparation and model training

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12197486B2Automatic labeling of text data
Publication Date: 2025.01.14 MICROSOFT TECHNOLOGY LICENSING LLC
  • US12197486B2 patent drawing
  • US12197486B2 patent drawing
  • US12197486B2 patent drawing

AI summary

The technology described herein determines whether a candidate text is in a requested class by using a generative model that may not be trained on the requested class. The present technology may use of a model trained primarily in an unsupervised mode, without requiring a large number of manual user-input examples of a label class. The may produce a semantically rich positive example of label text from a candidate text and label. Likewise, the technology may produce from the candidate text and the label a semantically rich negative example of label text. The labeling service makes use of a generative model to produce a generative result, which estimates the likelihood that the label properly applies to the candidate text. In another aspect, the technology is directed toward a method for obtaining a semantically rich example that is similar to a candidate text.