Screen Interaction Mapping for Accurate UI Element Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing computer vision techniques fail to accurately map timestamp and coordinate point information for user interactions on a computing system screen, leading to inaccurate analysis of changes and screen elements.
Innovation Solution
A method and system that utilize text detection, contouring, and custom filtering techniques to determine regions of interest on a screen, identify user interactions as keyboard or mouse type, and extract content from screen elements using pattern recognition and OCR, while mapping timestamps to events for precise change identification.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If existing computer vision techniques are used to identify user interactions on screen, then basic interaction detection is achieved, but timestamp and coordinate point information mapping accuracy deteriorates
Solution Approach 1:
The patent segments the screen into multiple regions of interest based on coordinate information from user interactions. By dividing the screen into specific regions where interactions occur, the system can apply targeted analysis to each region, improving the accuracy of timestamp and coordinate mapping while maintaining reliable change analysis.
Solution Approach 2:
The patent introduces an intermediary processing layer that receives raw video frames and interaction data, then processes them through multiple stages including region identification, change detection, and coordinate mapping. This intermediary layer acts as a mediator between raw data and final analysis, enhancing both mapping accuracy and analysis reliability.
2Loss of information
If comprehensive screen analysis is performed on all areas, then complete interaction information is captured, but processing time and computational resources increase
Solution Approach 1:
The patent extracts and focuses analysis only on regions of interest where user interactions occur, rather than analyzing the entire screen. By taking out and isolating the relevant regions based on coordinate information, the system captures complete interaction information while significantly reducing processing time and computational resources.
Solution Approach 2:
The patent applies partial action by performing comprehensive analysis only on specific regions of interest rather than the entire screen. This selective approach ensures that all interaction information in relevant areas is captured while avoiding unnecessary processing of empty or irrelevant screen areas, thus reducing overall processing time.
3Measurement precision
If video frame comparison is used to detect changes, then basic change detection is achieved, but coordinate point information mapping accuracy deteriorates
Solution Approach 1:
The patent replaces the traditional mechanical frame-by-frame comparison method with a coordinate-based region identification system. By substituting the mechanical comparison approach with coordinate-driven region selection and pattern recognition, the system achieves both accurate coordinate mapping and precise change detection.
Solution Approach 2:
The patent changes the analysis parameters by shifting from global frame comparison to localized region analysis based on coordinate information. By changing the parameter of analysis scope from entire screen to specific regions of interest, and by using pattern recognition algorithms, the system achieves improved coordinate mapping accuracy while maintaining high change detection precision.
Data Source
AI summary
The present disclosure relates to method and system for extracting information based on user interactions performed on screen of a computing system. Method includes receiving processed input data comprising video capturing screen and user interactions events occurring from user interactions, and co-ordinate information. Thereafter, computing system determines Regions of Interest (RoI) on screen based on captured user interactions, events and co-ordinate information, and identifies type of user interaction performed with at least one screen element in the RoI, as either keyboard or mouse type interaction. Subsequently, computing system performs one of: determining type of screen element as a text box or a table based on pattern recognition when type of user interaction is identified as keyboard type interaction or determining type of screen element to be selectable User Interface elements and extracting content and label of selectable UI elements, when interaction is identified as mouse type interaction.


