Web Browsing Robot Using Visual Recognition for Task Automation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Humans spend considerable time browsing the web for tasks like website testing, data gathering, and online shopping, as existing systems lack efficiency in automating these processes.

Innovation Solution

A system of robots that navigate webpages by analyzing content rather than code, using a knowledge base to learn and execute tasks similar to human interactions, including accessing databases, recognizing text, and using coupons, to efficiently accomplish goals such as online purchases.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If traditional web browsing automation methods are used, then web tasks can be automated, but the time required to complete tasks remains considerable and efficiency is low

Engineering Contradiction:
Improveweb task completion efficiencyVSAvoidtime spent on web browsing tasks
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent replaces traditional code-based web automation mechanisms with a vision-based system. The robot uses camera imaging to capture webpage content and optical character recognition (OCR) to identify and interact with web elements, substituting mechanical code parsing with visual perception similar to human browsing. This enables more natural and efficient web interaction.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables the robot to independently perform web browsing tasks without requiring human intervention or pre-programmed code for each task. The robot autonomously navigates webpages, identifies elements through visual recognition, and completes tasks such as online shopping and data gathering on its own, significantly improving productivity while reducing time loss.

Inventive Principle:
Principle #25Self-service

2Adaptability or versatility

If code-based automation systems are used, then web tasks can be performed, but the systems lack adaptability to changing webpage layouts and content

Engineering Contradiction:
Improveadaptability to webpage changesVSAvoidcomplexity of automation system
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent replaces complex code-based automation mechanisms with a simpler vision-based system. By using camera imaging and OCR, the system naturally adapts to different webpage layouts and content without requiring complex code parsing and interpretation, reducing system complexity while improving adaptability.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The vision-based system provides universal adaptability across different websites and webpage layouts. The robot can recognize and interact with various web elements (buttons, links, text fields) through visual patterns rather than site-specific code, making the system versatile and adaptable to changing webpage designs without increasing complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Ease of operation

If human users perform web browsing tasks manually, then tasks can be completed with understanding and judgment, but considerable time and effort are required

Engineering Contradiction:
Improveease of web task executionVSAvoidtime spent on web tasks
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The robot performs web browsing tasks autonomously without requiring human operation or intervention. It independently navigates webpages, identifies elements through visual recognition, makes decisions based on task goals, and completes actions such as purchasing items or gathering data, eliminating the time and effort humans would otherwise spend on these repetitive tasks.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system copies human browsing behavior by using visual perception and recognition mechanisms similar to how humans read and interact with webpages. The robot captures images of webpages, recognizes text and elements through OCR, and interacts with the interface in ways that mirror human actions, providing ease of operation while dramatically reducing time investment.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS11429686B2Web browsing robot system and method
Publication Date: 2022.08.30 VM ROBOT
  • US11429686B2 patent drawing
  • US11429686B2 patent drawing
  • US11429686B2 patent drawing

AI summary

A method for using a robot on the web is disclosed. The method may include assigning a goal to a robot. The robot may then direct a web browser to code corresponding to a URL. Using the code, the web browser may render a webpage comprising a plurality of rendered elements. The robot may identify each rendered element by using OCR or an OCR equivalent or by positioning a virtual mouse in a plurality of locations on the webpage and obtaining, from the code, element-identification information corresponding to each location. The robot may map each rendered elements with an element type selected from a closed set of element types stored within a knowledge base accessible by the robot. The robot may further select, from a set of possible actions, an action corresponding to each rendered element that is most likely to lead toward the goal and implement each such action.