Synthetic Utterance Generation for Enterprise NLP Customization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing natural language search engines are expensive and time-consuming to customize for specific enterprises, limiting their accessibility and efficiency in meeting the needs of various data-driven applications.

Innovation Solution

A server computing device equipped with a processor that generates a glossary file from a knowledge graph, allowing for the creation of utterance templates with replaceable fields, which are then filled with ontology entities to generate training utterances, enabling efficient natural language query processing and database search functionality.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If existing natural language search engines are customized to meet specific enterprise needs, then the search functionality becomes more accurate and useful, but the customization process becomes expensive and time-consuming

Engineering Contradiction:
Improvecustomization to enterprise needsVSAvoidcustomization time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent applies preliminary action by pre-generating large volumes of synthetic training utterances using template-based synthesis before the natural language processing model needs to be trained. Instead of collecting training data in real-time during deployment, the system prepares diverse training examples in advance by combining templates with entity data from knowledge graphs, thereby eliminating the time-consuming data collection phase during customization.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent uses copying by creating synthetic training utterances through template instantiation. Rather than requiring actual user queries or manual data entry, the system copies and adapts template structures, filling them with entities from knowledge graphs to generate realistic training examples that mimic actual search interactions.

Inventive Principle:
Principle #26Copying

2Adaptability or versatility

If existing natural language search engines are customized to meet specific enterprise needs, then the search functionality becomes more accurate and useful, but the customization cost increases

Engineering Contradiction:
Improvecustomization to enterprise needsVSAvoidcustomization cost
Core Design Contradiction:
Adaptability or versatilityVSEase of manufacture

Solution Approach 1:

The patent uses copying by creating synthetic training utterances through template instantiation. Rather than requiring actual user queries or manual data entry, the system copies and adapts template structures, filling them with entities from knowledge graphs to generate realistic training examples that mimic actual search interactions.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system applies self-service by automatically generating its own training data using the available knowledge graph and template resources. The synthetic data generation process is self-contained, requiring minimal external input or manual intervention, thereby reducing the need for expensive external data collection services or extensive human annotation efforts.

Inventive Principle:
Principle #25Self-service

3Measurement precision

If natural language search requires specialized knowledge of query languages, then search precision improves, but user accessibility decreases

Engineering Contradiction:
Improvesearch precisionVSAvoiduser accessibility
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The patent applies mechanics substitution by replacing the need for users to manually construct complex query language syntax with a natural language interface. Instead of requiring users to learn and write formal query languages, the system uses trained NLP models that automatically interpret natural language inputs and translate them into appropriate search queries, thereby maintaining search precision while dramatically improving accessibility.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS10916237B2Training utterance generation
Publication Date: 2021.02.09 MICROSOFT TECHNOLOGY LICENSING LLC
  • US10916237B2 patent drawing
  • US10916237B2 patent drawing
  • US10916237B2 patent drawing

AI summary

A server computing device, including memory storing a knowledge graph including a plurality of ontology entities connected by a plurality of edges. The server computing device may further include a processor configured to generate a glossary file based on the knowledge graph. The glossary file may include a plurality of ontology entities included in the knowledge graph. The processor may receive a plurality of utterance templates. Each utterance template may include an utterance and a predefined intention. For each utterance template, the processor may generate one or more utterance template copies in which one or more ontology entities included in the utterance are replaced with one or more utterance template fields. The processor may generate a plurality of training utterances at least in part by filling the one or more utterance template fields of the one or more utterance template copies with respective ontology entities included in the glossary file.