Layered OCR Text Extraction for Discontinuous Content
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for extracting discontinuous content from a user interface require repeated entry and exit for each segment, leading to low text extraction efficiency.
Innovation Solution
A method involving screen capture, character recognition, and drag operations to extract text from a user interface, with features like screenshot stitching and character highlighting to facilitate simultaneous extraction of multiple pieces of content.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If the user extracts discontinuous content by repeatedly entering and exiting the extraction for each segment, then each segment can be extracted accurately, but the text extraction efficiency is low
Solution Approach 1:
The patent merges multiple extraction operations into a single continuous operation. By allowing users to select multiple discontinuous text segments through a unified selection interface and perform extraction in one action, the system eliminates the need to repeatedly enter and exit extraction mode for each segment, thereby improving productivity and reducing time loss.
Solution Approach 2:
The patent segments the user interface into multiple layers (original interface layer and screenshot layer) and allows independent selection of text segments within the screenshot. This segmentation enables users to select discontinuous content segments efficiently without affecting the underlying interface, resolving the contradiction between accurate segment extraction and extraction efficiency.
2Ease of operation
If the screenshot is displayed at the original layer, then the interface layout is simple, but the user cannot select text without triggering interface interactions
Solution Approach 1:
The patent introduces a new dimension by creating a screenshot layer above the original interface layer. This additional layer provides a dedicated text selection environment where users can select text without triggering interface interactions, while the layered structure manages the complexity of having both the original interface and screenshot visible simultaneously.
Solution Approach 2:
The screenshot layer acts as an intermediary between the user and the original interface. It captures the interface content and presents it in a form that allows text selection without side effects, mediating between the need for simple interface layout and the need for reliable text selection capability.
3Manufacturing precision
If the screenshot area includes the complete preset control, then the control is not truncated, but the screenshot area becomes larger
Solution Approach 1:
The patent dynamically adjusts screenshot parameters (area, position, size) based on the detected preset control. By calculating the optimal screenshot area that includes the complete control while minimizing unnecessary coverage, the system achieves high screenshot quality without excessive area, adapting parameters to the specific content being captured.
Data Source
AI summary
A text extraction method and a related device are provided. The method includes performing screen capture on a user interface of a first application, to obtain a screenshot of the first application, displaying the screenshot at a target layer, where the target layer is above a layer at which the user interface is located, and performing character recognition on the screenshot. When a character selection operation for the screenshot displayed at the target layer is detected, the method includes highlighting a selected character at the target layer, and when a drag operation for the selected character is detected, dragging the selected character to a second application.


