Unsupervised Annotation Generation for NLU Models

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for generating annotated data for natural language understanding are costly, time-consuming, and often hindered by privacy concerns, making it difficult to train effective machine learning models.

Innovation Solution

The system automatically populates target templates with vocabulary words to generate annotated training data, allowing for unsupervised training of machine learning models without human annotation, thus maintaining data security.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual annotation by human readers is used to generate training data, then the quality and accuracy of annotated data is improved, but the cost and time consumption increase significantly

Engineering Contradiction:
Improveannotation accuracyVSAvoidannotation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent uses automated algorithms to generate synthetic annotated training data by copying and transforming existing unannotated data, replacing the need for manual human annotation. This produces sufficient training data without the time and cost overhead of human annotators while maintaining adequate quality for model training.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system performs self-annotation by automatically processing unannotated data through computational algorithms that generate annotations without human intervention. The machine learning model trains on this self-generated annotated data, eliminating dependency on external human annotators and significantly reducing annotation time and cost.

Inventive Principle:
Principle #25Self-service

2Measurement precision

If manual annotation by human readers is used to generate training data, then the quality and accuracy of annotated data is improved, but the cost increases significantly

Engineering Contradiction:
Improveannotation accuracyVSAvoidcost effectiveness
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent generates training data by copying and transforming existing unannotated datasets through automated processes, eliminating the need to pay human annotators. This approach maintains sufficient data quality for training while dramatically reducing the cost burden associated with manual annotation services.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system autonomously generates its own training data through automated annotation algorithms, making the process cost-effective by eliminating human labor costs. The self-service approach allows the organization to produce unlimited training data without recurring annotation expenses.

Inventive Principle:
Principle #25Self-service

3Quantity of substance

If unannotated data containing personal or confidential information is used for training, then data availability is improved, but data security and privacy concerns arise

Engineering Contradiction:
Improvetraining data availabilityVSAvoidprivacy risk
Core Design Contradiction:
Quantity of substanceVSObject-affected harmful factors

Solution Approach 1:

The patent creates synthetic copies of unannotated data through automated processing, allowing the original confidential data to remain secure while generating derived training data. The copying process transforms the data into annotated format without requiring human reviewers to see the sensitive information, thus maintaining privacy while enabling training.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The automated annotation system acts as an intermediary between the confidential unannotated data and the training process. This intermediary processes the data algorithmically without human intervention, eliminating the privacy risk of human annotators accessing sensitive information while still producing the necessary annotated training data.

Inventive Principle:
Principle #24Intermediary (Mediator)

4Reliability

If large amounts of annotated data are collected through manual processes, then the effectiveness of machine learning models is improved, but the productivity of the annotation process is too low

Engineering Contradiction:
Improvemodel effectivenessVSAvoiddata generation rate
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent uses automated copying and transformation of existing data to generate large volumes of annotated training data rapidly. This computational approach can produce thousands of annotated examples in minutes, vastly outpacing manual annotation productivity while providing sufficient data quantity for effective model training.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system autonomously generates training data at high speed through automated algorithms, achieving productivity levels impossible for human annotators. This self-service data generation capability enables rapid production of large annotated datasets that effectively train machine learning models without being constrained by human annotation throughput.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS12321693B2Unsupervised method to generate annotations for natural language understanding tasks
Publication Date: 2025.06.03 MICROSOFT TECHNOLOGY LICENSING LLC
  • US12321693B2 patent drawing
  • US12321693B2 patent drawing
  • US12321693B2 patent drawing

AI summary

A method for training a machine learning model with parallel annotations of source instances and while facilitating security of the source instances can be performed by a system that generates a coupled machine learning model from (1) a first machine learning model trained on a first set of training data comprising unannotated natural language and (2) a second machine learning model trained on populated target templates which are populated with a plurality of vocabulary words. Once formed, the coupled machine learning model is configured to transform unannotated natural language into annotated machine-readable text.