Synthetic Utterance Generation for Enterprise NLP Customization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing natural language search engines are expensive and time-consuming to customize for specific enterprises, limiting their accessibility and efficiency in meeting the needs of various data-driven applications.
Innovation Solution
A server computing device equipped with a processor that generates a glossary file from a knowledge graph, allowing for the creation of utterance templates with replaceable fields, which are then filled with ontology entities to generate training utterances, enabling efficient natural language query processing and database search functionality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If existing natural language search engines are customized to meet specific enterprise needs, then the search functionality becomes more accurate and useful, but the customization process becomes expensive and time-consuming
Solution Approach 1:
The patent applies preliminary action by pre-generating large volumes of synthetic training utterances using template-based synthesis before the natural language processing model needs to be trained. Instead of collecting training data in real-time during deployment, the system prepares diverse training examples in advance by combining templates with entity data from knowledge graphs, thereby eliminating the time-consuming data collection phase during customization.
Solution Approach 2:
The patent uses copying by creating synthetic training utterances through template instantiation. Rather than requiring actual user queries or manual data entry, the system copies and adapts template structures, filling them with entities from knowledge graphs to generate realistic training examples that mimic actual search interactions.
2Adaptability or versatility
If existing natural language search engines are customized to meet specific enterprise needs, then the search functionality becomes more accurate and useful, but the customization cost increases
Solution Approach 1:
The patent uses copying by creating synthetic training utterances through template instantiation. Rather than requiring actual user queries or manual data entry, the system copies and adapts template structures, filling them with entities from knowledge graphs to generate realistic training examples that mimic actual search interactions.
Solution Approach 2:
The system applies self-service by automatically generating its own training data using the available knowledge graph and template resources. The synthetic data generation process is self-contained, requiring minimal external input or manual intervention, thereby reducing the need for expensive external data collection services or extensive human annotation efforts.
3Measurement precision
If natural language search requires specialized knowledge of query languages, then search precision improves, but user accessibility decreases
Solution Approach 1:
The patent applies mechanics substitution by replacing the need for users to manually construct complex query language syntax with a natural language interface. Instead of requiring users to learn and write formal query languages, the system uses trained NLP models that automatically interpret natural language inputs and translate them into appropriate search queries, thereby maintaining search precision while dramatically improving accessibility.
Data Source
AI summary
A server computing device, including memory storing a knowledge graph including a plurality of ontology entities connected by a plurality of edges. The server computing device may further include a processor configured to generate a glossary file based on the knowledge graph. The glossary file may include a plurality of ontology entities included in the knowledge graph. The processor may receive a plurality of utterance templates. Each utterance template may include an utterance and a predefined intention. For each utterance template, the processor may generate one or more utterance template copies in which one or more ontology entities included in the utterance are replaced with one or more utterance template fields. The processor may generate a plurality of training utterances at least in part by filling the one or more utterance template fields of the one or more utterance template copies with respective ontology entities included in the glossary file.


