Browser Extension for Automatic Web Data Tagging and Classification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional internet data collection methods are inconvenient as they require users to manually copy and paste data, store different file types, and cannot retrieve web page content simultaneously, making data collection and classification cumbersome.

Innovation Solution

An internet data collection method that allows users to mark target data on a web page, retrieve web addresses and location information, and store them as tags, enabling automatic recording and classification, with features like screen capture and text recognition for inaccessible data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If users manually copy and paste data through browser application, then data collection can be performed, but the operation process becomes cumbersome and time-consuming

Engineering Contradiction:
Improvedata collection efficiencyVSAvoidoperation convenience
Core Design Contradiction:
ProductivityVSEase of operation

Solution Approach 1:

The system performs automatic data collection and tagging without requiring manual copy-paste operations. The browser extension automatically captures web page content, extracts target data, and stores it with metadata tags, allowing the system to serve itself rather than requiring continuous user intervention for each data collection step.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

A browser extension acts as an intermediary between the web page and the user's data collection needs. The extension intercepts web page content, processes it to extract target data, and automatically stores it with associated metadata, eliminating the need for manual copying while maintaining ease of use through a simple user interface.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If users store data by different file types for different formats, then data can be collected in various formats, but the classification and management process becomes complex

Engineering Contradiction:
Improvedata format compatibilityVSAvoidclassification complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system uses a universal storage format (JSON) that can represent multiple data types (text, images, videos, audio) within a single structured framework. All data regardless of original format is converted to this universal representation with appropriate metadata tags, allowing versatile data collection without requiring separate file management systems for each data type.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system changes the parameter representation of data by converting various file types into a standardized JSON structure with metadata parameters. Instead of managing different file types separately, all data is transformed into a common parameter-based representation that includes type information, source URL, and other metadata, simplifying classification and management.

Inventive Principle:
Principle #35Parameter changes

3Loss of information

If users retrieve web page data, then data can be collected, but the content of the accessed website or page cannot be retrieved simultaneously

Engineering Contradiction:
Improvedata completenessVSAvoiddata retrieval efficiency
Core Design Contradiction:
Loss of informationVSProductivity

Solution Approach 1:

The system performs preliminary action by capturing the complete web page content including HTML, text, images, videos, and audio simultaneously when the page is loaded. This preliminary capture ensures all content is available before any processing or selection occurs, preventing information loss while maintaining efficiency through batch processing.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system segments the web page content into different data types (text, images, videos, audio) while maintaining their associations. By dividing the content into manageable segments with clear metadata tags, the system can retrieve and process each type efficiently while preserving the complete information set and their relationships.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11630872B2Internet data collection method
Publication Date: 2023.04.18 ASUSTEK COMPUTER INC
  • US11630872B2 patent drawing
  • US11630872B2 patent drawing
  • US11630872B2 patent drawing

AI summary

An internet data collection method includes steps of receiving a collecting instruction, the collecting instruction corresponds to target data that marked on a web page; retrieving a web address corresponding to the web page and the location information of the target data on the web page; and storing the web address and the location information as a tag to an operating end.