ML Domain Classification for Electronic Survey Deduplication
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional systems for administering electronic surveys are inefficient due to the storage and transmission of duplicative data, inflexibility in generating and administering survey content, and excessive use of computing resources, leading to inaccurate and irrelevant response data.
Innovation Solution
The use of machine-learning models to deduplicate electronic survey questions in real-time by mapping questions to domain classifications and utilizing natural language processing to identify semantically similar questions, thereby removing duplicates and optimizing survey content.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional systems store and transmit survey data, then complete information is captured, but data volume increases and processing efficiency decreases
Solution Approach 1:
The system performs deduplication preprocessing on survey questions before administration. By identifying and removing duplicate questions in advance using machine learning models, the system reduces the total number of questions respondents need to answer, thereby decreasing data volume and processing requirements while maintaining information completeness.
Solution Approach 2:
The system extracts and removes duplicate survey questions from the survey pool. By identifying questions with similar semantic meanings using natural language processing and machine learning, the system separates redundant content from necessary content, keeping only unique questions to be administered.
2Ease of operation
If conventional systems use fixed survey templates, then survey administration is simple, but flexibility in adapting to different domains is reduced
Solution Approach 1:
The system dynamically adapts survey content based on the target domain. Machine learning models classify the survey domain and automatically select or generate appropriate questions, allowing the survey system to flexibly adapt to different domains (e.g., healthcare, education, business) while maintaining ease of administration through automated processes.
Solution Approach 2:
The system changes survey parameters such as question selection, domain classification, and deduplication thresholds based on the specific survey context. By adjusting these parameters dynamically according to the domain and survey objectives, the system achieves both simplicity and adaptability.
3Device complexity
If conventional systems process all survey questions uniformly, then processing is straightforward, but computing resource consumption increases
Solution Approach 1:
The system segments survey questions into different domains using machine learning classification. By dividing the survey content into domain-specific segments (e.g., technical, administrative, behavioral), the system can apply optimized processing strategies to each segment, reducing overall computing resource consumption while maintaining straightforward processing within each segment.
Solution Approach 2:
The system performs preliminary domain classification and deduplication of survey questions before the actual survey administration. This preprocessing step groups similar questions together and identifies duplicates, so that during survey delivery, the system only needs to handle already-processed, de-duplicated content, significantly reducing real-time computing requirements.
4Reliability
If conventional systems administer surveys to all potential respondents, then comprehensive data collection is achieved, but administration time increases
Solution Approach 1:
The system administers a reduced set of deduplicated questions that capture the essential information needed. By removing redundant questions while retaining core survey objectives, the system achieves comprehensive data collection with fewer questions, thereby reducing administration time without sacrificing data completeness.
Data Source
AI summary
Methods, systems, and non-transitory computer readable storage media are disclosed for utilizing machine-learning models to deduplicate electronic survey questions of electronic surveys or questionnaires in real-time. Specifically, the disclosed system maps electronic survey questions to specific domain classifications by utilizing a machine-learning model to classify portions of electronic surveys based on context within the portions of the electronic surveys. Additionally, the disclosed system utilizes the mappings of electronic survey questions to domain classifications to determine whether to deduplicate specific questions that are semantically similar and within the same domain classifications. For instance, the disclosed system utilizes natural language processing to find semantically similar questions across a plurality of electronic surveys and deduplicate the similar questions if their domain classifications are the same.


