Document Reading Support via Segmented Summarization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional document reading support methods, such as displaying patent specifications and drawings side by side, often require extensive time and may fail to provide an accurate understanding, especially when using language models with transformer architectures that are limited by character or memory constraints.
Innovation Solution
A document reading support method that segments documents into manageable sections, allows users to select specific parts, and inputs these selections along with a summarization prompt into a language model, ensuring the number of tokens does not exceed a predetermined value to generate accurate summaries.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If the entire document is input to a language model to generate a summary, then the summary comprehensiveness is improved, but the character or memory limit is exceeded making it impossible to process
Solution Approach 1:
The patent divides the entire document into multiple sections or paragraphs, processes each section separately through the language model to generate individual summaries, and then combines these summaries to form a comprehensive summary of the whole document. This segmentation approach allows the system to work within character limits while maintaining summary comprehensiveness.
Solution Approach 2:
The patent transitions from a single-dimension approach (processing the entire document as one block) to a multi-dimensional approach by organizing the document into hierarchical sections, allowing simultaneous processing of multiple segments in parallel, thus overcoming the character limit constraint.
2Measurement precision
If manual sentence extraction is performed to prepare input for the language model, then the input quality is improved, but the processing time and effort increase significantly
Solution Approach 1:
The system automatically performs sentence extraction and selection without requiring manual intervention. The language model itself or automated preprocessing scripts identify and extract relevant sentences from the document, eliminating the need for manual extraction while maintaining input quality.
Solution Approach 2:
The patent implements automated preprocessing steps that prepare the document input before it reaches the language model. This includes automatic sentence segmentation, relevance filtering, and formatting, which are performed in advance to ensure high-quality input without manual effort.
3Loss of information
If the document is displayed in a conventional format without segmentation, then the complete document is visible, but the reading time increases and understanding becomes difficult
Solution Approach 1:
The patent segments the document into clearly defined sections, paragraphs, or chunks that are displayed separately on the interface. This segmentation allows readers to navigate and understand the document in manageable portions while the system generates comprehensive summaries that cover all segments, thus reducing reading time without losing information completeness.
4Ease of operation
If drawings and descriptions are displayed separately as in conventional patent documents, then the layout is clear, but the reader fails to find related information efficiently
Solution Approach 1:
The patent merges the display of drawings and their corresponding descriptions by using the language model to generate integrated summaries that reference both visual and textual elements. The system creates cross-references and contextual links between drawings and their descriptions, allowing readers to efficiently locate related information without sacrificing layout clarity.
Data Source
AI summary
A document reading support method using a language model is provided. The document reading support method includes the steps of: displaying a segmented document; receiving selection of a part of the document; inputting the part and an instruction sentence for summarizing the part to a language model; determining whether the number of tokens of the part is less than or equal to a predetermined value; and obtaining a summary of the part determined to have the tokens of less than or equal to the predetermined value.


