Automated UI Testing via Language Model Visual Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing user interface testing methods lack efficiency and accuracy in automating the testing of user interfaces across software applications, particularly in verifying target outcomes and executing sequences of actions within webpages.
Innovation Solution
A method that involves accessing a test statement, capturing screenshots of webpage regions, transforming webpage code into contextual tags, generating prompts for a language model, and executing sequences of actions based on the model's responses to verify target outcomes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional automated testing methods are used, then testing can be automated to some extent, but accuracy in verifying target outcomes and executing sequences of actions is insufficient
Solution Approach 1:
The patent introduces a language model as an intermediary between the automated testing system and the webpage content. The language model processes visual information from screenshots and code information, then generates accurate responses about target outcomes. This intermediary enables high-accuracy verification by translating complex visual and code data into reliable test results, while the overall system remains fully automated.
Solution Approach 2:
The patent replaces traditional mechanical automated testing approaches with an AI-based language model that processes information more accurately. Instead of relying on rigid automated scripts that struggle with accuracy, the system uses the language model's natural language processing capabilities to interpret webpage content and verify outcomes with high precision while maintaining automation.
2Productivity
If manual testing methods are used, then accuracy in verifying target outcomes can be maintained, but efficiency and productivity are reduced
Solution Approach 1:
The patent implements a self-service automated testing system where the language model independently processes screenshots and code information to verify target outcomes without human intervention. The system automatically captures webpage content, transforms code into contextual tags, queries the language model, and executes sequences of actions based on model responses, achieving both high efficiency and accuracy through full automation.
Solution Approach 2:
The patent substitutes manual testing operations with an automated system powered by a language model. The language model processes visual and code information to achieve accuracy previously only attainable through manual testing, while the automated execution of sequences of actions provides the efficiency that manual testing cannot deliver.
3Reliability
If comprehensive testing of sequences of actions is implemented, then testing coverage is improved, but complexity of the testing system increases
Solution Approach 1:
The patent employs a universal language model that can handle multiple testing scenarios and sequences of actions through its natural language processing capabilities. Rather than requiring separate specialized systems for different test types, the language model provides a multi-functional platform that improves testing coverage across various webpage interactions while maintaining relatively simple system architecture.
Solution Approach 2:
The language model serves as an intermediary that simplifies the complexity of comprehensive testing by processing diverse test scenarios through a unified interface. It transforms complex sequences of actions into manageable processing steps, enabling improved testing coverage without proportionally increasing system complexity.
Data Source
AI summary
One variation of a method includes: accessing a test statement defining a target outcome for a target webpage; capturing a screenshot of the target webpage depicting a set of target content; accessing a set of webpage code defined for the target webpage; transforming the set of webpage code into a sequence of contextual tags, corresponding to the set of target content depicted in the screenshot, to generate a textual representation of the target webpage; generating a prompt including the test statement, the screenshot, and the textual representation; accessing a language model configured to generate responses to test statements based on visual and textual content extracted from corresponding prompts; based on the language model and the prompt, generating a textual response to the test statement representing occurrence of the target outcome at the target webpage; and serving the textual response to a user associated with the organization.


