Page Segmentation of Vector Graphics Using Neural Networks
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing solutions for page segmentation in unstructured vector graphics documents are inefficient due to reliance on complex heuristic rules that require manual correction and fail to analyze embedded images, lacking self-correction and considering information beyond the document itself.
Innovation Solution
The use of machine learning algorithms, such as deep convolutional neural networks and convolutional neural networks, to generate and classify element proposals in unstructured vector graphics documents, enabling automatic page segmentation by identifying and categorizing page elements like text blocks, tables, and figures, and resolving overlapping proposals.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If heuristic rules are used for page segmentation, then implementation is straightforward, but accuracy and adaptability deteriorate due to inability to self-correct and handle variations
Solution Approach 1:
The patent replaces manual heuristic rule-based systems with machine learning models (convolutional neural networks and recurrent neural networks) that automatically learn segmentation patterns from training data. This substitution enables the system to adapt to document variations without manual intervention while maintaining high accuracy, resolving the contradiction between ease of implementation and segmentation accuracy.
2Adaptability or versatility
If complex heuristic rules are used to handle all document variations, then coverage improves, but system complexity and manual maintenance requirements worsen
Solution Approach 1:
The patent implements self-service through machine learning models that automatically learn and adapt to document variations during training and deployment. The models self-correct by learning from training data and self-adjust by adapting to new document types without requiring manual intervention to add corner cases, thereby reducing system complexity while maintaining versatility.
Solution Approach 2:
The patent uses parameter changes by training machine learning models with adjustable parameters (weights and biases) that are optimized during training. This allows the system to adapt to different document variations by changing internal parameters rather than modifying the algorithm structure, reducing complexity while improving coverage.
3Productivity
If heuristic algorithms are deployed to end users, then immediate availability is achieved, but inability to update for new cases worsens service quality
Solution Approach 1:
The patent applies dynamics by making the segmentation system adaptable and updatable through machine learning model retraining. The models can be retrained with new data to handle emerging document variations, allowing the system to evolve over time while maintaining deployment efficiency. This dynamic capability ensures continuous improvement of service quality without sacrificing deployment speed.
4Speed
If analysis is performed only on vector graphics documents, then processing speed is maintained, but information completeness deteriorates due to inability to analyze embedded images
Solution Approach 1:
The patent applies segmentation by separating the analysis of vector graphics and embedded images within a unified framework. The system processes vector graphics elements directly while using machine learning models to analyze embedded images, combining both results for complete document understanding. This approach maintains processing speed for vector elements while recovering information from images that would otherwise be lost.
Data Source
AI summary
Disclosed systems and methods generate page segmented documents from unstructured vector graphics documents. The page segmentation application executing on a computing device receives as input an unstructured vector graphics document. The application generates an element proposal for each of many areas on a page of the input document tentatively identified as being page elements. The page segmentation application classifies each of the element proposals into one of a plurality of defined type of categories of page elements. The page segmentation application may further refine at least one of the element proposals and select a final element proposal for each element within the unstructured vector document. One or more of the page segmentation steps may be performed using a neural network.


