ML Domain Classification for Electronic Survey Deduplication

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional systems for administering electronic surveys are inefficient due to the storage and transmission of duplicative data, inflexibility in generating and administering survey content, and excessive use of computing resources, leading to inaccurate and irrelevant response data.

Innovation Solution

The use of machine-learning models to deduplicate electronic survey questions in real-time by mapping questions to domain classifications and utilizing natural language processing to identify semantically similar questions, thereby removing duplicates and optimizing survey content.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional systems store and transmit survey data, then complete information is captured, but data volume increases and processing efficiency decreases

Engineering Contradiction:
Improvecompleteness of information captureVSAvoiddata processing efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system performs deduplication preprocessing on survey questions before administration. By identifying and removing duplicate questions in advance using machine learning models, the system reduces the total number of questions respondents need to answer, thereby decreasing data volume and processing requirements while maintaining information completeness.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system extracts and removes duplicate survey questions from the survey pool. By identifying questions with similar semantic meanings using natural language processing and machine learning, the system separates redundant content from necessary content, keeping only unique questions to be administered.

Inventive Principle:
Principle #2Taking out (Extraction)

2Ease of operation

If conventional systems use fixed survey templates, then survey administration is simple, but flexibility in adapting to different domains is reduced

Engineering Contradiction:
Improvesimplicity of survey administrationVSAvoidflexibility in domain adaptation
Core Design Contradiction:
Ease of operationVSAdaptability or versatility

Solution Approach 1:

The system dynamically adapts survey content based on the target domain. Machine learning models classify the survey domain and automatically select or generate appropriate questions, allowing the survey system to flexibly adapt to different domains (e.g., healthcare, education, business) while maintaining ease of administration through automated processes.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes survey parameters such as question selection, domain classification, and deduplication thresholds based on the specific survey context. By adjusting these parameters dynamically according to the domain and survey objectives, the system achieves both simplicity and adaptability.

Inventive Principle:
Principle #35Parameter changes

3Device complexity

If conventional systems process all survey questions uniformly, then processing is straightforward, but computing resource consumption increases

Engineering Contradiction:
Improvesimplicity of processingVSAvoidcomputing resource consumption
Core Design Contradiction:
Device complexityVSUse of energy by moving object

Solution Approach 1:

The system segments survey questions into different domains using machine learning classification. By dividing the survey content into domain-specific segments (e.g., technical, administrative, behavioral), the system can apply optimized processing strategies to each segment, reducing overall computing resource consumption while maintaining straightforward processing within each segment.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary domain classification and deduplication of survey questions before the actual survey administration. This preprocessing step groups similar questions together and identifies duplicates, so that during survey delivery, the system only needs to handle already-processed, de-duplicated content, significantly reducing real-time computing requirements.

Inventive Principle:
Principle #10Preliminary action

4Reliability

If conventional systems administer surveys to all potential respondents, then comprehensive data collection is achieved, but administration time increases

Engineering Contradiction:
Improvecomprehensiveness of data collectionVSAvoidsurvey administration time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system administers a reduced set of deduplicated questions that capture the essential information needed. By removing redundant questions while retaining core survey objectives, the system achieves comprehensive data collection with fewer questions, thereby reducing administration time without sacrificing data completeness.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS20250029127A1Deduplicating electronic survey questions utilizing machine-learning model domain classification
Publication Date: 2025.01.23 ONETRUST LLC
  • US20250029127A1 patent drawing
  • US20250029127A1 patent drawing
  • US20250029127A1 patent drawing

AI summary

Methods, systems, and non-transitory computer readable storage media are disclosed for utilizing machine-learning models to deduplicate electronic survey questions of electronic surveys or questionnaires in real-time. Specifically, the disclosed system maps electronic survey questions to specific domain classifications by utilizing a machine-learning model to classify portions of electronic surveys based on context within the portions of the electronic surveys. Additionally, the disclosed system utilizes the mappings of electronic survey questions to domain classifications to determine whether to deduplicate specific questions that are semantically similar and within the same domain classifications. For instance, the disclosed system utilizes natural language processing to find semantically similar questions across a plurality of electronic surveys and deduplicate the similar questions if their domain classifications are the same.