Automated Text Evaluation Using String-Structure Similarity
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for filtering inappropriate language in user-generated content, such as blacklisting and manual human evaluation, are inefficient and ineffective, especially in languages like Arabic with varying dialects and no unified dictionary for informal language.
Innovation Solution
An automated text-evaluation service using a computer system with an artificial neural network (ANN) that processes input text through a string-structure similarity measure to generate evaluation scores, capable of identifying inappropriate language across different dialects and reducing dictionary size for improved efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual human evaluation is used to filter inappropriate language, then accuracy can be maintained, but the process becomes very slow and cannot keep up with live commenting
Solution Approach 1:
The system uses automated text evaluation services that process user-generated content without human intervention. The service receives input text, processes it through string-structure similarity measures, and automatically returns evaluation scores indicating the degree of inappropriate language, enabling high-speed processing while maintaining filtering capability
Solution Approach 2:
The patent replaces manual human evaluation (mechanical process) with an automated computer-based text evaluation service. The system uses computational methods including string-structure similarity measures and predefined dictionaries to automatically evaluate text, substituting human labor with automated processing that can handle live commenting speeds
2Ease of manufacture
If blacklisting methods are used for filtering, then the process is simple to implement, but it is not effective due to different dialects and ways words are pronounced or written throughout different countries
Solution Approach 1:
The system changes the parameter of word comparison from exact matching (blacklisting) to string-structure similarity measurement. By evaluating the degree of similarity between input words and dictionary words rather than requiring exact matches, the system can identify inappropriate language across different dialects, pronunciations, and spellings while maintaining implementation simplicity
Solution Approach 2:
The text evaluation service is designed to handle multiple languages and dialects universally. The system uses string-structure similarity measures that work across different linguistic variations, making the filtering mechanism effective for various countries and dialects without requiring separate blacklists for each variant
3Measurement precision
If a comprehensive dictionary is used to cover all dialects and informal language, then filtering accuracy improves, but the dictionary size and processing complexity increase significantly
Solution Approach 1:
The system changes the approach from storing complete comprehensive dictionaries to using string-structure similarity measures. By calculating similarity degrees between input words and dictionary words rather than requiring exact matches, the system achieves broad language coverage with a more manageable dictionary size, reducing processing complexity while maintaining filtering accuracy
Data Source
AI summary
A method for an automated text-evaluation service, and more particularly a method and apparatus for automatically evaluating text and returning a score which represents a degree of inappropriate language. The method is implemented in a computer infrastructure having computer executable code tangibly embodied in a computer readable storage medium having programming instructions. The programming instructions are configured to: receive an input text which comprises an unstructured message at a first computing device; process the input text according to a string-structure similarity measure which compares each word of the input text to a predefined dictionary to indicate whether there is similarity in meaning, and generate an evaluation score for each word of the input text and send the evaluation score to another computing device. The evaluation score for each input message is based on the string-structure similarity measure between each word of the input text and the predefined dictionary.


