Machine Learning Document Compression via Content Relevance
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The increasing volume of graphic-based documents in cloud storage poses a challenge for organizations, as traditional compression methods fail to effectively reduce storage footprints, leading to overwhelming storage demands.
Innovation Solution
A machine learning-based system that splits documents into text and image parts, applies optical character recognition, and uses a content decision engine with various machine learning components to assess relevance, allowing for tailored compression or elimination of non-relevant parts, thereby optimizing storage usage.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If traditional compression methods are used on graphic-based documents, then the storage footprint is reduced to some extent, but the compression effectiveness is insufficient and storage demands remain overwhelming
Solution Approach 1:
The patent segments documents into distinct components (text regions, image regions, tables, charts) and applies different processing strategies to each segment. The content decision engine evaluates each segment independently to determine whether to compress, discard, or retain it, enabling more effective storage reduction compared to treating the entire document as a single unit.
Solution Approach 2:
The patent applies different compression qualities and strategies to different parts of the document based on their relevance. The content decision engine assigns different relevance scores to various document segments and applies tailored compression approaches - highly relevant segments are retained with minimal compression, while less relevant segments undergo aggressive compression or discarding.
2Productivity
If machine learning-based content analysis is applied to assess document relevance, then compression effectiveness is improved, but system complexity increases
Solution Approach 1:
The patent introduces a content decision engine as an intermediary component that bridges the gap between raw document content and compression decisions. This engine uses machine learning models to analyze document segments and generate relevance scores, which then guide the compression process. This intermediary layer manages the complexity by centralizing the decision-making logic.
Solution Approach 2:
The machine learning models within the content decision engine automatically analyze and evaluate document segments without requiring manual intervention. The system self-adjusts compression strategies based on the learned relevance patterns, reducing the need for complex manual configuration and management.
3Reliability
If documents are stored in graphic-based file formats to preserve visual content, then content fidelity is maintained, but storage space consumption increases
Solution Approach 1:
The patent extracts and separates different content types from graphic-based documents, identifying text regions, image regions, tables, and charts as distinct extractable elements. By extracting text content from graphical representations, the system can store text in more space-efficient formats while preserving only essential graphical elements, thereby reducing overall storage space while maintaining content fidelity.
Data Source
AI summary
In an example embodiment, machine learning is used to intelligently compress documents to reduce the overall footprint of storing large amounts of files for an organization. Specifically, a document is split into parts, with each part representing a grouping of text or an image. Optical character recognition is performed to identify the text in images. Machine learning techniques are then applied to a part of a document in order to determine how relevant the document is for the organization. The parts that are deemed to be not relevant may then be reduced in size, either by omitting them completely or by summarizing them. This allows for the compression to be tailored specifically to the organization, resulting in the ability to compress or eliminate parts of documents that other organizations might have found relevant (and thus would not have been compressed or eliminated through traditional means).


