Document Content Point-and-Select Method for Untagged PDFs
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The current point-and-select method for untagged documents, such as PDF files, is cumbersome and time-consuming, as users need to accurately pull and select document content, making it difficult to select entire sentences or tables, and only character-based copying is possible.
Innovation Solution
A document content point-and-select method that determines document location information and uses document structure recognition to select and highlight word or sentence text upon point-and-click operations, enabling direct selection and editing of document content without the need for pulling and selecting, and allowing for the automatic copying of entire tables.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of time
If users use the traditional click and pull back method to select document content, then the selection operation can be performed, but the operation becomes cumbersome and time-consuming when the document content is long
Solution Approach 1:
The patent segments the selection operation into distinct click actions with different functions: first click to start selection, second click to end selection. This segmentation allows the system to automatically determine selection boundaries without requiring continuous user dragging, thus reducing selection time while maintaining ease of operation.
Solution Approach 2:
The patent implements preliminary action by having the system prepare and recognize document structure elements (words, sentences, paragraphs, tables) before the user performs the selection operation. When the user clicks, the system can immediately identify and select the appropriate pre-recognized element, eliminating the need for manual dragging and significantly reducing selection time.
2Measurement precision
If users manually pull and select document content to ensure accuracy, then the selection precision can be maintained, but the operation becomes cumbersome and time-consuming
Solution Approach 1:
The patent replaces the mechanical dragging operation with an intelligent recognition system. Instead of relying on the user's manual dragging action to define selection boundaries, the system uses document structure recognition to automatically identify words, sentences, paragraphs, and tables based on the user's click position, thereby maintaining high selection accuracy while dramatically reducing the time and effort required.
Solution Approach 2:
The patent introduces an intermediary document structure recognition system that acts as a mediator between the user's click action and the final selection result. This intermediary layer automatically identifies the intended selection target (word, sentence, paragraph, or table) based on the click position and document structure, ensuring accurate selection without requiring precise manual dragging by the user.
3Extent of automation
If users use traditional selection method in untagged documents, then character-based copying is possible, but entire sentences and tables cannot be automatically selected and copied
Solution Approach 1:
The patent implements universality by creating a multi-functional selection system that can automatically identify and select multiple types of document elements (words, sentences, paragraphs, and tables) based on the user's click position. This universal selection capability allows the system to adapt to different selection needs without requiring separate operations for each element type, thereby enhancing both automation extent and selection scope.
Solution Approach 2:
The patent applies preliminary action by pre-processing the untagged document to recognize and mark the structural boundaries of words, sentences, paragraphs, and tables before the user performs selection. This preliminary structure recognition enables the system to automatically identify and select complete sentences and tables when the user clicks, providing automated selection capability across multiple element types and expanding the versatility of the selection function.
Data Source
AI summary
Embodiments of the present disclosure disclose a document content point-and-select method, device, electronic apparatus, medium and program product. One implementation of the method includes: in response to detecting a point-and-click operation acting on an untagged document, determining document location information of the point-and-click operation; determining a document structure recognition result of the document content at a document location characterized by the document location information in the untagged document; in response to determining that the point-and-click operation is a first point-and-click operation, selecting a word text corresponding to the document location information from the document structure recognition result as a target word, and highlighting in an area corresponding to the target word; in response to determining that the point-and-click operation is a second point-and-click operation, selecting a sentence text corresponding to the document location information from the document structure recognition result as a target sentence.


