Dialogue Agent Knowledge Dictionary Automatic Expansion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Maintaining an up-to-date knowledge information dictionary for speech semantic analysis in dialogue agents is costly and challenging due to the difficulty in ensuring accuracy, especially when relying on manual updates or external databases.
Innovation Solution
An information processing device with a tagging unit, semantic analysis unit, and dictionary expansion unit that assigns category tags to input speech terms, estimates the intended domain, extracts relevant phrases, and automatically registers them in the knowledge information dictionary, even if a category tag is not initially assigned, using a hierarchical structure for response generation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If manual updates or external databases are used to maintain the knowledge information dictionary, then the dictionary can be kept up-to-date, but the cost and difficulty of maintenance increase significantly
Solution Approach 1:
The system performs self-service by automatically extracting noun phrases from search documents and registering them in the thesaurus dictionary without human intervention. The dictionary expansion unit autonomously acquires new terms and expands the knowledge base, eliminating the need for manual updates while maintaining accuracy through systematic extraction processes.
Solution Approach 2:
The system uses feedback mechanisms where search conditions and results are continuously analyzed to identify new noun phrases that should be added to the dictionary. The expansion process is driven by actual search usage patterns, ensuring that the dictionary evolves based on real-world application needs while maintaining relevance and accuracy.
2Ease of manufacture
If web pages are crawled to automatically update the knowledge information dictionary, then maintenance cost is reduced, but the accuracy of the information cannot be ensured
Solution Approach 1:
Instead of directly crawling web pages, the system uses search documents as an intermediary source. These documents are obtained through controlled search operations using the existing thesaurus dictionary, providing a filtered and validated source of new terms. This intermediary approach ensures accuracy while enabling automatic expansion.
Solution Approach 2:
The system performs preliminary actions by first conducting searches using existing dictionary terms to gather relevant documents, then extracting new noun phrases from these pre-filtered sources. This preliminary filtering through search operations ensures that only relevant and accurate terms are candidate for addition to the dictionary.
3Ease of manufacture
If external databases are imported to update the knowledge information dictionary, then automatic updates are achieved, but the system becomes dependent on other parties and required information may not always be available
Solution Approach 1:
The system achieves self-sufficiency by generating its own training data through search operations. Instead of relying on external databases controlled by other parties, the system independently extracts noun phrases from search documents it generates itself, ensuring continuous operation without external dependencies.
Solution Approach 2:
The system uses the thesaurus dictionary for multiple purposes: as a search tool to find relevant documents, as a source of search conditions for extraction, and as the target for expansion. This multi-functional use of the same resource eliminates the need for separate external databases while maintaining the ability to automatically expand knowledge.
Data Source
AI summary
Automatic expansion of a knowledge information dictionary for speech semantic analysis and generation of responses for a dialogue agent are performed in a favorable manner. Category tags are assigned to each term in input speech for all categories when terms are registered in the knowledge information dictionary. A domain of speech content intended by the input speech is estimated, and terms pertaining to the estimated domain are extracted from the input speech as a phrase of a predetermined entity. A response is generated on the basis of the domain of the speech content intended by the input speech and the phrase of the predetermined entity. When a category tag is not assigned to the phrase of a predetermined entity, the phrase of the predetermined entity is registered for the category corresponding to the predetermined entity in the knowledge information dictionary. The knowledge information dictionary has a hierarchical structure, and the application unit generates the response using the hierarchical structure.


