Technical Question Generation via Domain Graphs
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current automated question generation techniques are inadequate for technical text, as they rely on general-purpose datasets, focus on named entities, and struggle to generate questions with 'No' as an answer or span multiple words, failing to capture the essence of technical domains.
Innovation Solution
A processor-implemented method and system that extracts structure and linguistic information from technical documents to create domain-specific graphs, identifying technical terms and relationships, and generates questions using semantic templates, independent of named entities.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Extent of automation
If deep learning-based automated question generation techniques are used, then question generation capability is improved, but the system requires large training datasets and is not suitable for technical domains
Solution Approach 1:
The patent introduces an intermediary component - a technical term extraction module that uses keyword extraction algorithms and domain-specific term clustering algorithms to identify and extract technical terms from documents. This intermediary process creates structured technical term data that can be used to generate questions without requiring large annotated training datasets, thus resolving the contradiction between automation capability and data requirements
Solution Approach 2:
The patent performs preliminary actions by extracting and organizing technical terms, relationships, and attributes from technical documents before question generation. This includes creating technical term dictionaries, relationship graphs, and attribute databases in advance, which enables subsequent automated question generation to proceed without needing large training datasets
2Extent of automation
If existing deep learning techniques are used, then general purpose question generation is improved, but the system fails to capture technical domain essence and focuses on named entities
Solution Approach 1:
The patent applies local quality by making the question generation system domain-specific through technical term extraction and relationship identification. Instead of using generic named entity recognition, the system extracts domain-specific technical terms, their relationships, and attributes, thereby ensuring the generated questions are relevant to technical domains rather than general-purpose text
Solution Approach 2:
The patent changes the parameters of question generation by shifting from named entity-based approaches to technical term-based approaches. It introduces new parameters such as technical term relationships, domain-specific attributes, and technical concept hierarchies, which fundamentally alter how questions are generated to better suit technical domains
3Ease of manufacture
If traditional question generation approaches are used, then simple question generation is achieved, but the system cannot generate questions with multi-word answers or 'No' as an answer
Solution Approach 1:
The patent introduces dynamics by making the question generation process adaptable to different question types and answer formats. The system dynamically selects question templates and generation strategies based on the extracted technical terms, relationships, and attributes, enabling it to generate diverse question types including those with multi-word answers and yes/no questions, rather than being limited to simple answer formats
Data Source
AI summary
Questions play a central role in assessment of a candidate's expertise during an interview or examination. However, generating such questions from input text documents manually needs specialized expertise and experience. Further, techniques that are available for automated question generation require input sentence as well as an answer phrase in that sentence to generate question. This in-turn requires large training datasets consisting tuples of input sentence answer-phrase and the corresponding question. Additionally, training datasets are available are for general purpose text, but not for technical text. Present application provides systems and methods for generating technical questions from technical documents. The system extracts meta information and linguistic information of text data present in technical documents. The system then identifies relationships that exist in provided text data. The system further creates one or more graphs based on the identified relationships. The created graphs are the used by the system to generate technical questions.


