Web Content Collection Tool for Structured Data Export

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Users face difficulties in staying organized and retrieving web content across multiple browsing sessions due to the inefficiencies of traditional bookmarking systems, which result in long lists of text-based URLs that are hard to navigate and manage.

Innovation Solution

A web browser-integrated content collection tool that allows users to identify webpage types, extract relevant content, and save it in a collection pane, which can be interacted with, shared, and exported to other productivity applications, while automatically updating when source content changes.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If traditional bookmarking systems are used to save web content, then users can store multiple webpages, but the content becomes difficult to organize and retrieve

Engineering Contradiction:
Improvenumber of stored webpagesVSAvoidease of organization and retrieval
Core Design Contradiction:
Quantity of substanceVSEase of operation

Solution Approach 1:

The patent segments web content into structured data fields (title, URL, description, tags, metadata) rather than storing complete URLs. This segmentation allows systematic organization through tags and categories, making large collections manageable and retrievable through multiple access points.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary processing layer that transforms raw web content into structured data representations. This intermediary system extracts and organizes content attributes, enabling efficient search and retrieval without users directly managing raw URLs.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If users manually manage bookmarks across multiple browsing sessions, then they can access web content, but productivity decreases due to time spent organizing and locating content

Engineering Contradiction:
Improveefficiency of web researchVSAvoidtime spent organizing and retrieving content
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent performs preliminary organization of web content during the saving process itself. Content is automatically structured into fields and tagged during initial collection, eliminating the need for later manual organization and enabling immediate efficient retrieval.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system provides feedback mechanisms including search functionality and organized display of collected content, allowing users to quickly locate previously saved webpages through multiple access points rather than manually browsing through lists.

Inventive Principle:
Principle #23Feedback

3Loss of information

If complete web content is saved for future reference, then all information is preserved, but storage requirements and data management complexity increase

Engineering Contradiction:
Improvecompleteness of saved contentVSAvoidcomplexity of data management
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent extracts only the essential and useful portions of web content into structured fields (title, URL, description, key metadata) rather than storing complete webpage copies. This extraction maintains informational value while dramatically reducing storage requirements and management complexity.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent transforms unstructured web content into structured data with defined parameters and fields. This parameterization organizes content into manageable, queryable units that are easier to store, retrieve, and manage while preserving the essential information.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11586698B2Transforming collections of curated web data
Publication Date: 2023.02.21 MICROSOFT TECHNOLOGY LICENSING LLC
  • US11586698B2 patent drawing
  • US11586698B2 patent drawing
  • US11586698B2 patent drawing

AI summary

In non-limiting examples of the present disclosure, systems, methods and devices for surfacing collected web content are presented. A collection of web content may be maintained, wherein the collection of web content is divided into a plurality of sections, each of the plurality of sections comprising a subset of web content from a different webpage. An indication to export the collection of web content to a productivity application may be received. A plurality of attributes that each of the plurality of sections have a value for may be identified. A productivity application document may be populated with the plurality of attributes and the corresponding values from each of the sections.