Web Browsing Robot Using Visual Recognition for Task Automation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Humans spend considerable time browsing the web for tasks like website testing, data gathering, and online shopping, as existing systems lack efficiency in automating these processes.
Innovation Solution
A system of robots that navigate webpages by analyzing content rather than code, using a knowledge base to learn and execute tasks similar to human interactions, including accessing databases, recognizing text, and using coupons, to efficiently accomplish goals such as online purchases.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If traditional web browsing automation methods are used, then web tasks can be automated, but the time required to complete tasks remains considerable and efficiency is low
Solution Approach 1:
The patent replaces traditional code-based web automation mechanisms with a vision-based system. The robot uses camera imaging to capture webpage content and optical character recognition (OCR) to identify and interact with web elements, substituting mechanical code parsing with visual perception similar to human browsing. This enables more natural and efficient web interaction.
Solution Approach 2:
The system enables the robot to independently perform web browsing tasks without requiring human intervention or pre-programmed code for each task. The robot autonomously navigates webpages, identifies elements through visual recognition, and completes tasks such as online shopping and data gathering on its own, significantly improving productivity while reducing time loss.
2Adaptability or versatility
If code-based automation systems are used, then web tasks can be performed, but the systems lack adaptability to changing webpage layouts and content
Solution Approach 1:
The patent replaces complex code-based automation mechanisms with a simpler vision-based system. By using camera imaging and OCR, the system naturally adapts to different webpage layouts and content without requiring complex code parsing and interpretation, reducing system complexity while improving adaptability.
Solution Approach 2:
The vision-based system provides universal adaptability across different websites and webpage layouts. The robot can recognize and interact with various web elements (buttons, links, text fields) through visual patterns rather than site-specific code, making the system versatile and adaptable to changing webpage designs without increasing complexity.
3Ease of operation
If human users perform web browsing tasks manually, then tasks can be completed with understanding and judgment, but considerable time and effort are required
Solution Approach 1:
The robot performs web browsing tasks autonomously without requiring human operation or intervention. It independently navigates webpages, identifies elements through visual recognition, makes decisions based on task goals, and completes actions such as purchasing items or gathering data, eliminating the time and effort humans would otherwise spend on these repetitive tasks.
Solution Approach 2:
The system copies human browsing behavior by using visual perception and recognition mechanisms similar to how humans read and interact with webpages. The robot captures images of webpages, recognizes text and elements through OCR, and interacts with the interface in ways that mirror human actions, providing ease of operation while dramatically reducing time investment.
Data Source
AI summary
A method for using a robot on the web is disclosed. The method may include assigning a goal to a robot. The robot may then direct a web browser to code corresponding to a URL. Using the code, the web browser may render a webpage comprising a plurality of rendered elements. The robot may identify each rendered element by using OCR or an OCR equivalent or by positioning a virtual mouse in a plurality of locations on the webpage and obtaining, from the code, element-identification information corresponding to each location. The robot may map each rendered elements with an element type selected from a closed set of element types stored within a knowledge base accessible by the robot. The robot may further select, from a set of possible actions, an action corresponding to each rendered element that is most likely to lead toward the goal and implement each such action.


