Mobile App Noise Filtering for NER Data Collection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Natural Language Processing (NLP) tasks require large amounts of annotated or transcribed language data, which are expensive to create and often suffer from noise and reduced quality when using crowdsourcing, making it challenging to maintain high data quality.

Innovation Solution

A system and method that utilizes unmanaged crowds to collect utterances through mobile devices, employing a mobile app to filter out noise by auditing user responses and configuring campaigns with parameters for ambient noise levels, ensuring high-quality data collection for Named Entity Recognition (NER) models.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of energy

If crowdsourcing is used to collect language data, then costs are reduced, but data quality deteriorates due to noise from unreliable workers

Engineering Contradiction:
ImprovecostVSAvoiddata quality
Core Design Contradiction:
Loss of energyVSReliability

Solution Approach 1:

The system performs preliminary actions by configuring campaigns with specific parameters (ambient noise levels, calibration requirements, audit checks) before data collection begins. This pre-configuration establishes quality thresholds and filtering mechanisms that prevent low-quality data from being collected in the first place, rather than filtering afterward

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system implements feedback mechanisms through audit checks that evaluate user responses and provide information about data quality. The mobile app configures itself based on campaign configuration parameters, creating a feedback loop that maintains quality standards while using crowdsourcing

Inventive Principle:
Principle #23Feedback

2Reliability

If manual annotation and transcription are used, then data quality is maintained, but costs increase significantly

Engineering Contradiction:
Improvedata qualityVSAvoidcost
Core Design Contradiction:
ReliabilityVSLoss of energy

Solution Approach 1:

The mobile application performs self-configuration based on campaign parameters, automatically setting up audit checks, calibration procedures, and quality thresholds. This automation reduces the need for manual intervention and expert annotation while maintaining quality standards through systematic self-regulation

Inventive Principle:
Principle #25Self-service

3Productivity

If unmanaged crowds are used for data collection, then costs are reduced and scalability increases, but noise and data quality deteriorate

Engineering Contradiction:
Improvedata collection efficiencyVSAvoiddata quality
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The system changes parameters by configuring campaigns with specific thresholds for ambient noise levels, calibration requirements, and audit check criteria. These parameter settings allow the system to accommodate large numbers of crowd workers while maintaining consistent quality standards through systematic filtering and evaluation

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS9772993B2System and method of recording utterances using unmanaged crowds for natural language processing
Publication Date: 2017.09.26 VOICEBOX TECH CORP
  • US9772993B2 patent drawing
  • US9772993B2 patent drawing
  • US9772993B2 patent drawing

AI summary

A system and method of recording utterances for building Named Entity Recognition (“NER”) models, which are used to build dialog systems in which a computer listens and responds to human voice dialog. Utterances to be uttered may be provided to users through their mobile devices, which may record the user uttering (e.g., verbalizing, speaking, etc.) the utterances and upload the recording to a computer for processing. The use of the user's mobile device, which is programmed with an utterance collection application (e.g., configured as a mobile app), facilitates the use of crowd-sourcing human intelligence tasking for widespread collection of utterances from a population of users. As such, obtaining large datasets for building NER models may be facilitated by the system and method disclosed herein.