Automatic Web Application Crawling and Curation System

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current systems lack an efficient method to automatically integrate and curate web applications and related content from the Internet into online marketplaces, limiting the availability of digital goods to users.

Innovation Solution

A computer-implemented method and system that generates criteria for web applications, translates these criteria into rules, crawls metadata from websites, and determines if they include executable features, allowing for the generation and display of icons and content as selectable listings in an online application store, using metrics like usage, installs, and user ratings to filter and update listings.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If manual curation methods are used to integrate web applications into online marketplaces, then quality control can be maintained, but the productivity and speed of application integration are limited

Engineering Contradiction:
Improveapplication integration speedVSAvoidautomation system complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system enables automatic self-curation of web applications by crawling, analyzing, and integrating applications without human intervention. The automated system evaluates applications based on predefined criteria including quality metrics, user ratings, and relevance, allowing the marketplace to maintain quality control while significantly improving integration productivity

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system transforms the application integration process by changing parameters from manual evaluation to automated metric-based assessment. It uses quantitative parameters such as user ratings, quality scores, and engagement metrics to objectively evaluate and select applications for integration, replacing subjective manual curation with data-driven automation

Inventive Principle:
Principle #35Parameter changes

2Quantity of substance

If automated crawling is implemented to discover web applications, then the quantity of available applications increases, but the difficulty of detecting and measuring application quality increases

Engineering Contradiction:
Improvenumber of web applicationsVSAvoidapplication quality assessment
Core Design Contradiction:
Quantity of substanceVSDifficulty of detecting and measuring

Solution Approach 1:

The system introduces intermediary metrics and evaluation mechanisms that bridge the gap between automated crawling and quality assessment. It uses intermediate parameters such as user ratings, quality scores, and engagement metrics as mediators to objectively evaluate application quality, making it feasible to assess large numbers of applications automatically without direct human intervention

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system implements feedback mechanisms by continuously monitoring user interactions, ratings, and engagement metrics with crawled applications. This feedback loop enables the system to learn from user behavior and refine its quality assessment criteria, improving the accuracy of quality detection as the quantity of applications grows

Inventive Principle:
Principle #23Feedback

3Reliability

If comprehensive metadata is collected from websites, then the reliability of application selection improves, but the loss of time required for data processing increases

Engineering Contradiction:
Improveapplication selection accuracyVSAvoidmetadata processing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system extracts only the most relevant metadata parameters needed for reliable application selection, such as user ratings, quality metrics, and engagement statistics. By selectively extracting only the essential data elements rather than processing all available metadata, the system maintains high selection accuracy while minimizing data processing time

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system applies partial action by focusing on collecting and processing only the critical subset of metadata that most significantly impacts application selection reliability. Rather than comprehensively analyzing all website data, it prioritizes key indicators that provide the highest value for quality assessment, reducing processing time while maintaining reliability

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS11055369B2Automatic crawling of applications
Publication Date: 2021.07.06 GOOGLE LLC
  • US11055369B2 patent drawing
  • US11055369B2 patent drawing
  • US11055369B2 patent drawing

AI summary

Systems and methods are described for generating criteria for a plurality of web applications in an online application store, translating the criteria into at least one rule, the at least one rule based on predefined categories defined by the online application store, obtaining, metadata associated with a plurality of websites, determining, using the metadata and the at least one rule, whether any of the websites in the plurality of websites, includes code that executes a feature associated with the at least one rule, and displaying the icon as a selectable listing in the online application store.