Crowd-Sourced Language Model Refinement via User Tagging
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional approaches to refining language models are labor-intensive and costly, especially for minority languages, and struggle to effectively exclude undesirable words like profanity, misspellings, and technical terms from crowd-sourced models.
Innovation Solution
Users can categorize and delete undesirable words from their individual language models by tagging them as profanity, misspelled, or out-of-language, with the system aggregating these tags to refine the language model based on crowd-sourced data, allowing for more accurate reflection of user language use.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If manual refinement of language models is performed, then language model quality is improved, but labor cost and time consumption increase significantly
Solution Approach 1:
The system enables users to automatically refine language models by tagging words as profanity, misspelled, or out-of-language during normal usage. The crowd-sourced tagging system allows the language model to self-improve without manual intervention, resolving the contradiction between quality improvement and time consumption.
Solution Approach 2:
The system implements feedback loops where user taggings are aggregated and used to automatically update language models. This continuous feedback mechanism enables rapid refinement of language models based on real-world usage data, eliminating the need for time-consuming manual refinement processes.
2Adaptability or versatility
If crowd-sourced language models are built, then language coverage is improved, but undesirable words like profanity and misspellings are included
Solution Approach 1:
The system extracts and removes undesirable words from crowd-sourced language models by allowing users to tag specific words as profanity, misspelled, or out-of-language. This extraction process separates harmful content from the beneficial crowd-sourced language data, resolving the contradiction between language coverage and undesirable word inclusion.
Solution Approach 2:
The system changes the state of words in the language model by applying tags that modify their properties. Words tagged as profanity, misspelled, or out-of-language have their status parameters changed, allowing the system to maintain broad language coverage while filtering out undesirable content through parameter modification.
3Measurement precision
If language models are continuously refined manually, then predictive accuracy is improved, but resource consumption increases
Solution Approach 1:
The system enables language models to automatically refine themselves through crowd-sourced user taggings. This self-service mechanism eliminates the need for continuous manual refinement resources while maintaining and improving predictive accuracy through automated processing of user feedback.
Solution Approach 2:
The system implements continuous refinement through automated processing of user taggings in real-time. This continuous useful action maintains predictive accuracy without the need for periodic manual intervention, reducing resource consumption while sustaining model quality improvements.
Data Source
AI summary
Technology is described for refining a language model for a language recognition system based on aggregating and analyzing word tag metadata from multiple users of the language. The technology allows a user to mark a word or phrase in a selected language (e.g., as offensive or misspelled, or as a part of speech or other category), combines information collected from multiple users of the selected language, and updates the user's language model based on the combined information from multiple users of the selected language.


