Text Mining Attribute Analysis for Internet Media Users
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for attribute analysis of Internet media users are limited in their ability to accurately and comprehensively identify and analyze user attributes, as they primarily rely on rough analysis and do not effectively handle noise and outliers in data.
Innovation Solution
A text mining-based attribute analysis method that involves establishing and updating a label main corpus and feature corpus through steps like data cleaning, clustering, semantic analysis, noise reduction, and model classification, using techniques such as dynamic and fuzzy clustering, and noise level threshold comparisons to generate accurate classified texts.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If article sample mining-based analysis is used, then the analysis process is simple, but the accuracy and comprehensiveness of user attribute identification is insufficient
Solution Approach 1:
The patent transforms the analysis approach by changing parameters from simple keyword matching to multi-dimensional text mining parameters including semantic analysis, clustering algorithms, and noise level thresholds. This enables comprehensive extraction of user attributes while maintaining systematic processing.
Solution Approach 2:
The patent introduces an intermediary classification system with labeled corpora and feature corpora that mediate between raw article samples and final user attribute identification. This intermediary layer enables accurate attribute extraction while keeping the overall process structured and manageable.
2Speed
If traditional analysis methods are used, then the processing speed is fast, but the ability to handle noise and outliers is poor
Solution Approach 1:
The patent performs preliminary noise reduction and data cleaning before main analysis. By pre-processing the text data to remove noise and outliers, the system maintains fast processing speed while ensuring high reliability in attribute identification through subsequent clustering and classification.
Solution Approach 2:
The patent implements feedback mechanisms through iterative clustering and classification processes. The system continuously refines its analysis by comparing results against established corpora and adjusting parameters to optimize both speed and reliability in handling noisy data.
3Use of energy by moving object
If rough analysis is used, then the computational resources required are low, but the comprehensiveness of attribute analysis is limited
Solution Approach 1:
The patent segments the analysis process into distinct modules: text preprocessing, clustering, classification, and attribute extraction. Each module processes specific aspects of the data independently, enabling comprehensive attribute analysis while optimizing computational resource usage through targeted processing.
Solution Approach 2:
The patent applies partial action by focusing computational resources on the most significant attributes and features identified through clustering. Rather than analyzing all possible attributes equally, the system concentrates resources on high-value insights while maintaining comprehensive coverage of key user characteristics.
Data Source
AI summary
A text mining-based attribute analysis method for Internet media users comprises a first steps of establishing a label main corpus and a feature corpus sequentially, and updating and maintaining the label main corpus and the feature corpus respectively, and a second step of extracting all history article samples of Internet users, and cleaning out videos, audios and pictures in the samples. The text mining-based attribute analysis method can form attributes of browsed sample articles for each Internet media user, and analyze accurately weights of interesting categories, to identify deeply, analyze, and mine the user attributes, and the basic attributes of the Internet users can also be analyzed.


