Microblog Garbage Template Identification via Multi-Feature Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The rapid proliferation of microblog 'garbage template articles' overwhelms platforms, with manual identification being inefficient and unable to handle the large data volume, wasting search engine resources and affecting user experience.

Innovation Solution

A method and apparatus that extract features such as punctuation, topic, bracket, link, and account name features from microblog articles to generate article features, which are then compared to a garbage template list to identify and filter out repetitive and redundant content automatically.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual identification method is used to determine garbage template articles, then identification accuracy can be maintained, but identification speed and efficiency are too low to handle large data volume

Engineering Contradiction:
Improveidentification accuracyVSAvoididentification speed
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent segments the article identification process into multiple feature dimensions: punctuation features, bracket features, link features, account name features, and content features. Each feature is extracted and analyzed independently, then combined to make a comprehensive judgment. This segmentation enables automated processing while maintaining accuracy by examining multiple aspects of each article.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent replaces the manual mechanical identification process with an automated computer-based system that extracts and analyzes features programmatically. The system uses algorithms to identify punctuation patterns, bracket structures, link characteristics, and content similarities, substituting human manual review with automated computational analysis that can process large volumes of articles efficiently.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If manual identification of every microblog article is performed, then all garbage template articles can be identified, but the process is impossible to complete due to huge data amount

Engineering Contradiction:
Improveidentification completenessVSAvoidtime consumption
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent implements a multi-feature analysis approach that examines multiple aspects of each article (punctuation, brackets, links, account names, content) to achieve reliable identification. By analyzing several features simultaneously rather than relying on a single criterion, the system maintains high identification completeness while enabling automated processing that can handle the huge data volume of microblog platforms.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The patent transforms the identification problem from a single-dimensional manual review into a multi-dimensional automated analysis by changing the parameters of examination. The system extracts and analyzes multiple features (punctuation patterns, bracket structures, link characteristics, account name formats, content similarity) simultaneously, enabling comprehensive identification of garbage template articles through automated parameter comparison rather than time-consuming manual inspection.

Inventive Principle:
Principle #35Parameter changes

3Adaptability or versatility

If garbage template articles are not identified and filtered, then all articles remain available for search, but search engine resources are wasted and user experience is seriously affected

Engineering Contradiction:
Improvearticle availabilityVSAvoidsearch engine resource waste
Core Design Contradiction:
Adaptability or versatilityVSLoss of energy

Solution Approach 1:

The patent implements preliminary identification and filtering of garbage template articles before they enter the search index. By extracting and analyzing multiple features of articles in advance, the system identifies and filters out repetitive template content proactively, preventing these articles from consuming search engine resources. This preliminary action maintains adaptability by keeping legitimate articles available while eliminating waste from garbage content.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS9330075B2Method and apparatus for identifying garbage template article
Publication Date: 2016.05.03 TENCENT TECHNOLOGY (SHENZHEN) CO LTD
  • US9330075B2 patent drawing
  • US9330075B2 patent drawing
  • US9330075B2 patent drawing

AI summary

Method and apparatus for identifying garbage template articles in network communication field are disclosed. The method includes: extracting a feature from an eligible microblog article to generate an article feature including a punctuation feature, a topic feature, a bracket feature, a link feature and an account name feature; acquiring a garbage template list including garbage template feature, i.e. an article feature whose frequency reaches a preset threshold, wherein they are extracted in a same way; identifying the microblog article as a garbage template article when the article feature is the same as the garbage template feature. The apparatus includes: a feature extracting module, an acquiring module, and an identifying module. Features of a microblog article are extracted to determine whether the microblog article is a garbage template article, so that garbage template articles in the present microblog platform can be identified effectively and search engine resources are saved.