Intelligent File Conversion via Layout Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing technologies face challenges in converting file data between different formats, particularly when importing files from fixed formats like PDF into applications that require live, free-flowing content, as they fail to accurately translate layout and formatting attributes, leading to loss of content characteristics.

Innovation Solution

The development of a conversion data model that analyzes digital file attributes such as content, layout, and formatting to generate intelligent authoring inferences, enabling the creation of a tailored representation suitable for the target application/service, which includes identifying file types, applying inference determination rules, and surfacing the converted file data in a user-friendly format.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If file data is converted between different formats using traditional methods, then the conversion process is simple and fast, but the layout and formatting attributes are lost or not accurately translated

Engineering Contradiction:
Improvelayout and formatting translation accuracyVSAvoidconversion process complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The system performs preliminary analysis of the source file's layout and formatting attributes before conversion. It evaluates properties such as content structure, styling, and visual characteristics in advance, then uses this pre-analyzed information to guide the conversion process, ensuring accurate reproduction of formatting in the target format

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary conversion process that acts as a mediator between the source and target formats. This intermediary layer analyzes the source file attributes, translates them into equivalent representations in the target format, and handles complex formatting mappings that cannot be achieved through direct conversion

Inventive Principle:
Principle #24Intermediary (Mediator)

2Loss of information

If applications/services evaluate file attributes during import, then content characteristics are preserved, but the processing time and computational resources increase

Engineering Contradiction:
Improvecontent characteristics preservationVSAvoidfile import processing time
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The system performs preliminary evaluation of file attributes during the import process, analyzing content characteristics, layout properties, and formatting information before the actual conversion. This pre-evaluation enables the system to make informed decisions about how to best preserve and translate content characteristics, reducing the need for repeated analysis during conversion

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces traditional mechanical file conversion methods with an intelligent system that uses attribute evaluation and inference. Instead of simple format transformation, the system analyzes file properties, generates inferences about content characteristics, and uses these inferences to guide the conversion process, thereby preserving more information

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Productivity

If traditional file conversion is used, then processing is fast and simple, but identification of data types such as titles, headers, and captions is not readily available

Engineering Contradiction:
Improvedata type identification capabilityVSAvoidfile analysis complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system uses feedback mechanisms where the analysis of file attributes informs the conversion process. By evaluating properties such as text styling, position, and structure, the system generates feedback about the likely data types (titles, headers, captions, body text), which then guides how the content should be structured in the target format

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The conversion system performs self-service by automatically analyzing file attributes and identifying data types without requiring manual intervention. The system evaluates formatting properties, infers content hierarchy, and automatically tags or structures the converted content according to the identified data types, making the process autonomous

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS11354489B2Intelligent inferences of authoring from document layout and formatting
Publication Date: 2022.06.07 MICROSOFT TECHNOLOGY LICENSING LLC
  • US11354489B2 patent drawing
  • US11354489B2 patent drawing
  • US11354489B2 patent drawing

AI summary

Non-limiting examples of the present disclosure describe processing that generates intelligent inferences of authoring from analysis of attributes associated with a digital file being imported in an application/service. Examples described herein are configured to work with any type of application/service including an authoring application/service. For instance, a request to import a digital file is received in an application/service. The application/service may be configured to analyze the digital file and generate authoring inferences based on an analysis of attributes of the digital file. For example, a conversion data model may be utilized to identify a file type of the digital file, analyze attributes of the identified digital file (e.g. content portions, layout, formatting, metadata, etc.) and output file data in a format that is tailored for the application/service based on authoring inferences. A converted representation of the digital file is surfaced in the application/service based on output of the file data.