Markup Schema Generation for Content Classification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Businesses face the challenge of monitoring and responding to unfavorable comments on social networking sites and other online platforms due to the vast amount of user-generated content, where existing systems rely solely on keyword classification, which is inefficient in identifying relevant content.

Innovation Solution

A method is developed to automatically generate a mark-up language schema by receiving training samples, comparing candidate schemas, and selecting those that match a predetermined threshold, allowing for the extraction of relevant data elements from online resources, which can then be analyzed and categorized using a support vector machine algorithm.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If keyword-based classification is used to monitor social networks, then the system can process large volumes of user-generated content, but the precision in identifying relevant content deteriorates

Engineering Contradiction:
Improvevolume of content processedVSAvoidprecision in identifying relevant content
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent introduces mark-up language schemas as intermediary structures between raw text and classification. These schemas define hierarchical relationships and semantic meaning, allowing the system to maintain high processing volume while improving precision through structured semantic analysis rather than simple keyword matching

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system changes the parameter of content analysis from simple keyword presence to structured mark-up language elements with hierarchical relationships. By transforming unstructured text into structured mark-up representations, the system achieves both high productivity through automated processing and high precision through semantic structure analysis

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If manual monitoring of all comments and messages is performed, then the precision in identifying unfavorable content is improved, but the time and resources required deteriorate

Engineering Contradiction:
Improveprecision in identifying unfavorable contentVSAvoidtime for monitoring content
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system enables automated self-service monitoring where the mark-up language schemas automatically structure and classify content without human intervention. The schemas define rules for identifying unfavorable content, allowing the system to serve itself in monitoring large volumes of content with precision previously achievable only through manual review

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent applies preliminary action by pre-defining mark-up language schemas that encode knowledge about unfavorable content patterns. These schemas are prepared in advance and automatically applied to incoming content, eliminating the need for time-consuming manual analysis while maintaining high precision through pre-encoded semantic rules

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If automated schema generation is implemented, then the adaptability to different content structures is improved, but the device complexity deteriorates

Engineering Contradiction:
Improveadaptability to content structuresVSAvoidcomplexity of schema generation system
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system implements feedback mechanisms where the automated schema generation process learns from classified content and refines schemas iteratively. The classification results feed back into schema improvement, allowing the system to adapt to different content structures automatically while the feedback loop manages complexity by focusing refinements on actual performance needs rather than exhaustive schema coverage

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS9460231B2System of generating new schema based on selective HTML elements
Publication Date: 2016.10.04 BRITISH TELECOM PLC
  • US9460231B2 patent drawing
  • US9460231B2 patent drawing
  • US9460231B2 patent drawing

AI summary

The present invention provides a system which is able to detect similar web page elements which are described in mark-up language, such that the content of those elements can be captured. Text content may then be sent to a text classifier for further analysis.