OCR-Based Content Selection for Touch Interfaces
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional touch screen devices are inefficient and unintuitive for performing tasks like copy and paste, with varying techniques across applications and inaccuracies in text selection, making it difficult for users to select and manipulate content effectively.
Innovation Solution
A computer-implemented method that receives user selection signals, identifies content attributes, determines a content entity, and provides relevant actions for display, using techniques such as OCR and machine learning to enhance content selection and interaction within user interfaces.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional touch screen text selection techniques are used, then text can be selected, but the selection is inaccurate and does not fully capture all desired text
Solution Approach 1:
The patent replaces manual mechanical text selection (dragging fingers across the screen) with an automated optical recognition system. The system captures an image of the display, uses OCR to identify text boundaries and content, and automatically determines the complete text entity based on the user's initial touch point, eliminating the need for manual dragging while improving selection accuracy.
Solution Approach 2:
The patent introduces an intermediary system between the user's touch input and the text selection outcome. This intermediary uses image capture, OCR processing, and entity recognition to bridge the gap between a simple touch gesture and accurate text selection, allowing users to select text by touching anywhere near it rather than precisely at its boundaries.
2Adaptability or versatility
If multiple different techniques are used for different applications, then text selection can be performed across various contexts, but the user must learn multiple techniques making the system complex
Solution Approach 1:
The patent implements a universal text selection system that works across all applications by capturing the display image and processing it through a unified OCR and entity recognition pipeline. Instead of requiring application-specific techniques, the system uses a single multi-functional approach that adapts to any content displayed on the screen, whether in messaging apps, browsers, documents, or other contexts.
3Productivity
If conventional copy and paste techniques are used, then text manipulation is possible, but the process is inefficient and slow
Solution Approach 1:
The patent performs preliminary text selection automatically based on the user's initial touch gesture. By using OCR to identify text boundaries and entity recognition to determine the complete text unit, the system prepares the text for copying before the user even initiates the copy command, eliminating the time-consuming manual selection process and enabling immediate copy-paste operations.
4Ease of operation
If manual text selection is used, then user control is maintained, but the process is unintuitive and difficult to perform
Solution Approach 1:
The patent replaces the complex mechanical interaction of dragging fingers across text with a simple tap gesture. The system uses image capture and OCR to automatically detect text boundaries and perform entity recognition, substituting the difficult manual measuring and selecting process with an intuitive touch interface that requires minimal user effort while maintaining precision.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Systems and methods of providing content selection are provided. For instance, one or more signals indicative of a user selection of an object displayed within a user interface can be received. Responsive to receiving the one or more signals, a content attribute associated with one or more objects displayed within the user interface can be identified. A content entity can be determined based at least in part on the content attribute and the user selection. One or more relevant actions can then be determined based at least in part on the determined content entity. Data indicative of the relevant actions can then be provided for display.