Multi-Intent Dataset Generation via Dynamic Connective Selection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing multi-intent utterance datasets, such as MixATIS and MixSNIPS, lack diversity in connectives and rely on simple concatenation rules, making it easy for models to detect intents based on superficial patterns rather than understanding the complexity of real-world conversations.
Innovation Solution
A method and apparatus for generating multi-intent datasets by concatenating utterances using various conjunctions and complex patterns, preserving the meanings and structures of single-intent utterances, and selecting utterances to be merged based on cosine similarity or randomized selection.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If single-intent utterances are merged using simple concatenation rules with fixed connectives (AND variants), then the dataset generation process is simple and efficient, but the diversity and complexity of the resulting multi-intent datasets are insufficient
Solution Approach 1:
The patent applies dynamics by making the connective selection dynamic rather than static. Instead of always using the same AND variants, the system randomly selects from multiple types of connectives (conjunctions, punctuation marks, and implicit connection patterns) during dataset generation. This dynamic selection process increases connective diversity while maintaining generation efficiency, directly resolving the technical contradiction between simple generation and diverse output.
2Ease of manufacture
If multi-intent datasets are generated using only AND variants and basic merging patterns, then the generation process is easy to implement, but models can easily detect intents by counting conjunctions or recognizing commas rather than understanding real conversation complexity
Solution Approach 1:
The patent applies parameter changes by varying multiple parameters of the merging process: different types of connectives (conjunctions like 'and', 'or'; punctuation marks like commas, semicolons; and implicit connections), different merging patterns (random selection, semantic similarity-based selection, hierarchical structure), and different utterance combinations. These parameter variations create more realistic and complex multi-intent utterances that prevent models from relying on superficial pattern matching, thereby improving intent detection accuracy while maintaining implementation feasibility.
3Productivity
If existing multi-intent datasets use naïve merging patterns with limited connectives, then the datasets are easy to generate, but they fail to reflect the complexity and diversity of real-world conversations
Solution Approach 1:
The patent applies preliminary action by pre-defining multiple types of connectives and merging patterns before the actual dataset generation process. The system prepares a library of conjunctions, punctuation marks, and connection strategies in advance, then randomly or strategically selects from these pre-prepared options during generation. This preliminary preparation enables the system to generate complex, diverse multi-intent utterances efficiently without sacrificing generation speed, as the complexity is built into the pre-configured merging patterns rather than computed in real-time.
Data Source
AI summary
A computable-implementable method for generating multi-intent datasets includes collecting single-intent datasets; preprocessing the collected single-intent datasets while preserving meanings and structures of utterances in the collected single-intent datasets; selecting a plurality of single-intent utterances to be merged from the preprocessed single-intent datasets; and merging the plurality of selected single-intent utterances into one multi-intent utterance.


