Crowd-Sourced Language Model Refinement via User Tagging

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional approaches to refining language models are labor-intensive and costly, especially for minority languages, and struggle to effectively exclude undesirable words like profanity, misspellings, and technical terms from crowd-sourced models.

Innovation Solution

Users can categorize and delete undesirable words from their individual language models by tagging them as profanity, misspelled, or out-of-language, with the system aggregating these tags to refine the language model based on crowd-sourced data, allowing for more accurate reflection of user language use.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If manual refinement of language models is performed, then language model quality is improved, but labor cost and time consumption increase significantly

Engineering Contradiction:
Improvelanguage model qualityVSAvoidrefinement time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The system enables users to automatically refine language models by tagging words as profanity, misspelled, or out-of-language during normal usage. The crowd-sourced tagging system allows the language model to self-improve without manual intervention, resolving the contradiction between quality improvement and time consumption.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system implements feedback loops where user taggings are aggregated and used to automatically update language models. This continuous feedback mechanism enables rapid refinement of language models based on real-world usage data, eliminating the need for time-consuming manual refinement processes.

Inventive Principle:
Principle #23Feedback

2Adaptability or versatility

If crowd-sourced language models are built, then language coverage is improved, but undesirable words like profanity and misspellings are included

Engineering Contradiction:
Improvelanguage coverageVSAvoidundesirable words
Core Design Contradiction:
Adaptability or versatilityVSObject-affected harmful factors

Solution Approach 1:

The system extracts and removes undesirable words from crowd-sourced language models by allowing users to tag specific words as profanity, misspelled, or out-of-language. This extraction process separates harmful content from the beneficial crowd-sourced language data, resolving the contradiction between language coverage and undesirable word inclusion.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system changes the state of words in the language model by applying tags that modify their properties. Words tagged as profanity, misspelled, or out-of-language have their status parameters changed, allowing the system to maintain broad language coverage while filtering out undesirable content through parameter modification.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If language models are continuously refined manually, then predictive accuracy is improved, but resource consumption increases

Engineering Contradiction:
Improvepredictive accuracyVSAvoidrefinement resources
Core Design Contradiction:
Measurement precisionVSLoss of energy

Solution Approach 1:

The system enables language models to automatically refine themselves through crowd-sourced user taggings. This self-service mechanism eliminates the need for continuous manual refinement resources while maintaining and improving predictive accuracy through automated processing of user feedback.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system implements continuous refinement through automated processing of user taggings in real-time. This continuous useful action maintains predictive accuracy without the need for periodic manual intervention, reducing resource consumption while sustaining model quality improvements.

Inventive Principle:
Principle #20Continuity of useful action

Data Source

PatentUS10073828B2Updating language databases using crowd-sourced input
Publication Date: 2018.09.11 MICROSOFT TECHNOLOGY LICENSING LLC
  • US10073828B2 patent drawing
  • US10073828B2 patent drawing
  • US10073828B2 patent drawing

AI summary

Technology is described for refining a language model for a language recognition system based on aggregating and analyzing word tag metadata from multiple users of the language. The technology allows a user to mark a word or phrase in a selected language (e.g., as offensive or misspelled, or as a part of speech or other category), combines information collected from multiple users of the selected language, and updates the user's language model based on the combined information from multiple users of the selected language.