Tag Mining via Seed Rule Generalization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for mining tags are limited, as they rely on vertical websites, which are not suitable for unpopular fields and often result in common noun tags that do not meet specific user demands, and are dependent on abundant text attributes that may not include subjective user tags.

Innovation Solution

A method and apparatus that use a tag seed rule with placeholders to match and generalize tags based on historical search information, constructing new search sequences and iteratively updating the tag seed rule until convergence, allowing for the mining of comprehensive and specific tags across all fields without relying on vertical websites.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If tags are extracted based on structuralization of vertical websites, then common noun tags can be obtained, but the method cannot be widely used for unpopular fields without vertical websites and cannot meet specific question-answer requirements

Engineering Contradiction:
Improveapplicability to different fieldsVSAvoidtag specificity
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent applies universality by creating a tag mining method that works across all fields and webpage types without requiring field-specific vertical websites. The system uses general-purpose text attributes (title, abstract, content) that exist on all webpages to extract tags, making the method universally applicable to popular and unpopular fields alike, thereby resolving the contradiction between broad adaptability and specific tag quality

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent introduces text attributes (title, abstract, content) as intermediaries between the webpage and the tag extraction process. These intermediary text attributes serve as the basis for mining both common noun tags and subjective user tags, enabling the system to overcome the limitation of relying on vertical website structures while still producing meaningful tags across diverse fields

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If tags are extracted based on text attributes of the entity, then some subjective tags can be mined, but text attributes are not abundant enough to mine all user tags

Engineering Contradiction:
Improvetag specificityVSAvoidmissing user tags
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The patent applies dimensionality change by utilizing multiple text attribute dimensions (title, abstract, content) simultaneously rather than relying on a single attribute. This multi-dimensional approach enriches the information available for tag extraction, enabling the system to mine both subjective tags from the abstract and additional tags from the title and content, thereby reducing information loss about user intent

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Productivity

If a unified flow is used to mine tags for various types of webpages, then development time is reduced, but the method must handle diverse webpage structures and content types

Engineering Contradiction:
Improvedevelopment efficiencyVSAvoidprocessing complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent achieves productivity improvement through universality by designing a single tag mining flow that handles all webpage types. The system uses a unified approach that extracts text attributes (title, abstract, content) from any webpage and processes them through the same tag extraction algorithm, eliminating the need for separate processing pipelines for different webpage structures and thereby reducing development time while managing complexity through standardization

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11409813B2Method and apparatus for mining general tag, server, and medium
Publication Date: 2022.08.09 BEIJING BAIDU NETCOM SCI & TECH CO LTD
  • US11409813B2 patent drawing
  • US11409813B2 patent drawing
  • US11409813B2 patent drawing

AI summary

A method and apparatus for mining a general tag, a server and a medium are disclosed. The method can comprise: matching a tag seed rule containing a tag placeholder and an attribute of the tag placeholder with historical search information to determine a matching tag; combining the existing tag seed rule and the matching tag to construct a new search sequence set; and performing a generalization process on search sequences included in the new search sequence set to obtain a new tag seed rule, and returning to perform the operation of matching the new tag seed rule with the historical search information to determine a new tag until the tag and the tag seed rule satisfy a convergence condition. A more comprehensive and profound tag can be mined, and the entire flow of mining the tag can not be dependent on a vertical website.