Short Message Age Classification via Linguistic Feature Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data mining techniques are inadequate for analyzing short messages due to their brevity, making it difficult to extract useful user information such as age from large volumes of text data generated via social networking services.

Innovation Solution

A classification system that determines keyword feature information from short messages using features like vowel/consonant ratio, capitalization, emoticon usage, punctuation, texting slang, and keyword frequency, employing classifiers like Naïve Bayesian and maximum entropy classifiers to estimate user age.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If current data mining techniques are used to analyze short messages, then the analysis can be performed using existing tools, but the brevity of short messages makes it difficult to extract useful user information

Engineering Contradiction:
Improveadaptability of data mining techniques to short messagesVSAvoidprecision of user information extraction
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent changes the parameters used for analysis from traditional text mining approaches to linguistic features specific to short messages. It extracts features such as vowel/consonant ratios, capitalization patterns, emoticon usage, punctuation styles, texting slang, and keyword frequency - these are parameter changes tailored to the brevity constraint of short messages, enabling effective user information extraction despite the limited text length

Inventive Principle:
Principle #35Parameter changes

2Quantity of substance

If the volume of short messages is increased to improve data quantity, then more data is available for analysis, but the brevity of each message makes information extraction more difficult

Engineering Contradiction:
Improvevolume of text dataVSAvoidinformation extraction effectiveness
Core Design Contradiction:
Quantity of substanceVSLoss of information

Solution Approach 1:

The patent extracts specific linguistic features from short messages that are most indicative of user demographics. By taking out and analyzing specific features such as vowel/consonant ratios, capitalization patterns, emoticon usage, punctuation styles, texting slang, and keyword frequency, the system can effectively extract user information even from the limited text content of short messages, preventing information loss despite the brevity constraint

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS9063927B2Short message age classification
Publication Date: 2015.06.23 ADVANCE MAGAZINE PUBLISHERS
  • US9063927B2 patent drawing
  • US9063927B2 patent drawing
  • US9063927B2 patent drawing

AI summary

Systems and methods for short message age classification in accordance with embodiments of the invention are disclosed. In one embodiment of the invention, classifying messages using a classifier includes determining keyword feature information for a message using the classifier, classifying the determined feature information using the classifier, and estimating user age using the classifier.