Speech-to-Text Input for Crowdsourced Task Processing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Remote workers on crowdsourcing platforms often face productivity and motivation issues due to inadequate typing skills, necessitating alternative data entry methods.
Innovation Solution
A method and system that convert audio inputs from crowdworkers into phrases, allowing them to respond to tasks via speech-to-text conversion, with mode selection based on worker parameters, enabling hands-free operation and improved accessibility.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If remote workers use traditional text-based data entry methods, then task completion is possible, but productivity is hampered due to inadequate typing skills
Solution Approach 1:
The patent replaces the mechanical typing system with an acoustic/speech-based input system. Crowdworkers speak their responses instead of typing, and the system converts speech to text using audio processing and speech recognition technologies. This substitution eliminates the need for typing skills while maintaining task completion capability.
Solution Approach 2:
The patent changes the input modality parameter from text-based to speech-based. By accepting audio inputs and converting them to text, the system fundamentally alters how workers interact with the platform, making it accessible to those without typing skills while improving productivity.
2Adaptability or versatility
If speech-to-text conversion is implemented, then accessibility is improved for workers with inadequate typing skills, but system complexity increases
Solution Approach 1:
The patent introduces speech recognition technology as an intermediary between the worker's speech and the task system. This intermediary automatically converts spoken responses into text, handling the complexity of audio processing, speech-to-text conversion, and text normalization, thereby making the system accessible without requiring workers to manage the underlying complexity.
3Ease of operation
If audio input mode is selected based on worker parameters, then user experience is optimized, but processing complexity increases
Solution Approach 1:
The patent performs preliminary actions by pre-defining multiple audio input modes (e.g., character-wise speech, word-wise speech) and automatically selecting the appropriate mode based on worker parameters before task execution. This preliminary configuration optimizes the user experience for each worker without requiring real-time complex decision-making during task performance.
Data Source
AI summary
The disclosed embodiments illustrate methods and systems for processing one or more crowdsourced tasks. The method comprises converting an audio input received from a crowdworker to one or more phrases by one or more processors in at least one computing device. The audio input is at least a response to a crowdsourced task. A mode of the audio input is selected based on one or more parameters associated with the crowdworker. Thereafter, the one or more phrases are presented on a display of the at least one computing device by the one or more processors. Finally, one of the one or more phrases is selected by the crowdworker as a correct response to the crowdsourced task.


