Content editable document format conversion method and system based on visual identification
By introducing visual big model and OCR technology into document format conversion, combined with the three-column visual interface and real-time proofreading function, the problems of insufficient accuracy of complex layout recognition and low proofing efficiency are solved, and a high accuracy and flexibility document conversion process is achieved.
Patent Information
- Application Number
- CN202510344040.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-23
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art lacks the accuracy of identifying complex layouts during document format conversion, and lacks real-time proofreading and visual editing capabilities, resulting in content breakage, style loss and proofreading efficiency after conversion.
The document format conversion method based on visual big model and OCR technology is adopted. By generating a three-column visual interface, real-time proofreading and editing functions are provided, which supports users to modify and adjust the recognition results during the conversion process, and improve proofreading efficiency through intelligent auxiliary functions.
It significantly improves the recognition accuracy of complex layouts, realizes visual comparison between content and originals, reduces the difficulty of proofreading, and improves the flexibility and proofreading efficiency of document conversion.
Smart Images

Figure CN120218015A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of document processing, and particularly to a document format conversion method and system based on a large vision model and OCR technology, which is particularly suitable for realizing content-editable and visual interactive proofreading during document format conversion. Background Art
[0002] In the field of document processing, format conversion is a common requirement. For example, converting PDF or scanned documents into editable formats such as Word and Excel. Traditional methods usually use OCR technology to directly recognize the content of the document and output the target format, but there are the following problems: Insufficient recognition accuracy: Existing OCR technology has insufficient recognition accuracy for complex layouts (such as multi-page tables, vector graphics, tables, formulas, multi-column layout), resulting in broken content or lost styles after conversion; Low proofreading efficiency and one-way processing: Mainstream tools adopt a linear process of "input-conversion-output", lacking the ability to visually edit intermediate structured data. Users can only check and correct errors item by item in the target format, lacking an intuitive comparison with the original manuscript and being prone to missing problems; Poor editing flexibility: Traditional methods cannot adjust the recognition results in real time during format conversion, cannot compare the original document and the conversion result in real time, and the modification operation needs to be completed in an independent editor, resulting in low efficiency and a large amount of subsequent editing work.
[0003] In view of the above problems, the present invention proposes a brand-new document format conversion method, which significantly improves the accuracy of document conversion and the proofreading efficiency by introducing an innovative solution of a three-column visual interface and a real-time proofreading mechanism.
[0004] The present invention aims to solve the following technical problems: ● How to improve the recognition accuracy of complex layouts (such as tables, formulas, multi-column layout) during document format conversion; ● How to provide an intuitive interactive interface that enables users to proofread and adjust the recognition results in real time during conversion; ● How to realize the visual comparison between the recognized content and the original manuscript to reduce the proofreading difficulty; ● How to improve the flexibility of document conversion and support users to perform secondary editing on the content before formal output. Summary of the Invention
[0005] To achieve the above objectives, the present invention provides a content-editable document format conversion method based on visual recognition, including the following steps.
[0006] The document format conversion system receives a document whose format needs to be changed, including but not limited to PDF files, image files or scanned documents.
[0007] Document Content Extraction and Structuring ● Use a visual large model or OCR technology to parse the original document (such as PDF, scanned document), extract elements such as text, tables, images, etc., and generate intermediate structured data; ● Enhance the parsing of complex layouts (such as tables, formulas, multi-column layouts) to ensure the integrity and accuracy of the recognition results.
[0008] Three-column Visual Interface Generation Generate an interactive interface with three columns, which display the following content respectively: ● Left column: Display the layout rendering or the original file image of the original document for intuitive comparison with the recognition results; ● Middle column: Present the structured text and layout elements that are the same as the original file image after being parsed by a visual large model or OCR. All text blocks are displayed in an editable form, supporting users to manually modify and edit; ● Right column: Generate a real-time preview of the final output style based on the template engine to ensure that users can see the actual effect after modification.
[0009] Interactive Proofreading and Content Adjustment ● Users perform editing operations (such as text correction, table structure adjustment, formula modification, vector graph adjustment, etc.) on the recognition results in the middle column, and the system updates the output preview in the right column in real time; ● Provide a difference comparison function, visually mark the modified areas in the left column (such as highlighting or positioning anchors) to help users quickly locate problems; ● Support operations such as text block dragging, merging, and splitting, and the system automatically processes the layout adjustment of related content.
[0010] Intelligent Assistant Functions ● Intelligent Revision Suggestions: When it is detected that the edited content in the middle column deviates too much from the semantics of the original document, the system automatically pushes it to the relevant original text paragraphs for reference; ● Context-aware Toolbar: Dynamically load a dedicated editing instruction set according to the type of the currently selected object (such as text, table, formula), and provide targeted operation options; ● Error Detection and Reminder: Detect logical conflicts (such as inconsistent table data, incorrect formula format) through the semantic analysis module, and generate difference reminder marks in the interface; ● Operation Backtracking Function: Record the historical versions of all modification operations, support users to roll back to any modification node at any time, and ensure the controllability of the editing process.
[0011] Output and Adaptation ● Responsive template library based on CSS3, automatically adapting to the layout constraints of different output formats (such as Word, Excel, HTML); ● Support high-quality export of multiple document formats, ensuring that the converted documents retain the original layout and editability; ● Provide a verification function for the export results, supporting users to batch compare the output files with the original manuscript to further improve accuracy. Description of the Drawings
[0012] Figure 1 Flowchart of Document Format Conversion Figure 2 Layout and Simple Interaction Schematic Diagram of a Three-column Visual Interface Interaction Platform Figure 3 Schematic Diagram of Version Rollback Step 1 of the S01 document format conversion process, receiving the document to be converted Preview diagram of the final output style in the S0303 three-column visual interface interaction platform Step 2 of the S02 document format conversion process, recognition and parsing S0304 version traceability button Step 3 of the S03 document format conversion process, interactive platform of the generated three-column visual interface S030101 The original document after recognition and parsing Step 4 of the S04 document format conversion process, interactive editing and revision of the document content Diagram of the final output format style compared with the original document Step 5 of the S05 document format conversion process, output of the determined content and format Highlight display of the positioning and difference reminder of the original document after editing Layout rendering diagram of the original document or original file image in the S0301 three-column visual interface interaction platform S030201 Synchronized to the final output style preview after editing Structured text and layout elements after recognition and parsing in the S0302 three-column visual interface interaction platform S030401 - S030408 Traceable versions during the editing process Detailed Implementation Manner
[0013] In order to enable those skilled in the art of this technology to better understand the solution of this application, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of this application.
[0014] The document format conversion system receives a document that needs to change the format, including but not limited to PDF files, image files or scanned documents.
[0015] Perform parsing and structuring determination of the received document.
[0016] First, use the sliding window mechanism to perform high-resolution regional OCR recognition on the original document image to ensure the integrity and continuity of cross-page tables and document structures.
[0017] Furthermore, when the original document contains vector graphics, after recognition, the vector graphics are converted to SVG format and the Bezier curve control points are retained.
[0018] Furthermore, perform enhanced parsing on complex layouts (such as tables, formulas, multi-column layouts, distorted scanned documents, multi-language structures, etc.) to ensure the integrity and accuracy of the recognition results.
[0019] Furthermore, parse the recognized document content, extract elements such as text, tables, and images, and generate intermediate structured data.
[0020] Furthermore, the generated intermediate structured data contains ternary data of <text block, coordinates, style label>.
[0021] Three-column visual interface generation.
[0022] First, generate an interactive interface with three columns, which respectively display the left column, the middle column, and the right column.
[0023] Left column: Displays the layout rendering diagram or the original file image of the original document for intuitive comparison with the recognition results.
[0024] Middle column: Presents the same structured text and layout elements as the original file image after being parsed by the vision large model or OCR. All text blocks are displayed in an editable form, supporting users to manually modify and edit.
[0025] Right column: Generates a real-time preview of the final output style based on the template engine to ensure that users can see the actual effect after modification.
[0026] Furthermore, while generating the content of the middle column, different colors are used to standardize the generated content through the built-in confidence heat map.
[0027] Interactive proofreading and content adjustment.
[0028] First, users perform editing operations (such as text correction, table structure adjustment, formula modification, vector graph adjustment, etc.) on the recognition results in the middle column. The document type discriminator automatically selects the processing pipeline, the semantic deviation detection triggers intelligent revision suggestions, the context-aware toolbar dynamically loads the editing instruction set, and the system real-time updates the output preview in the right column.
[0029] Furthermore, a difference comparison function is provided. The dynamic programming algorithm realizes paragraph-level difference comparison, and visual marks (such as highlighting or positioning anchors) are made on the modified areas in the left column to help users quickly locate problems.
[0030] Furthermore, the generative adversarial network separates printed text from handwritten annotations.
[0031] Furthermore, operations such as text block dragging, merging, and splitting are supported, and the system automatically processes the layout adjustment of related content.
[0032] Use of intelligent assistance functions.
[0033] First, intelligent revision suggestions: When it is detected that the editing content in the middle column deviates too much from the semantics of the original document, the system automatically pushes the relevant original text paragraphs for reference.
[0034] Furthermore, the context-aware toolbar: Dynamically loads a dedicated editing instruction set according to the type of the currently selected object (such as text, table, formula), providing targeted operation options.
[0035] Further, error detection and reminder: Detect logical conflicts (such as inconsistent table data, incorrect formula format) through the semantic analysis module, and generate difference reminder marks in the interface. For example, when editing a table, synchronously highlight the original document area, and a hue value ΔH > 30 triggers a visual reminder.
[0036] Further, the operation backtracking function: Record the historical versions of all modification operations, support users to roll back to any modification node at any time, and ensure the controllability of the editing process.
[0037] Output and adaptation.
[0038] First, based on the CSS3 responsive template library, automatically adapt to the layout constraints of different output formats (such as Word, Excel, HTML).
[0039] Support high-quality export of multiple document formats to ensure that the converted document retains the original layout and editability.
[0040] Provide a verification function for the export result, support users to batch compare the output file with the original manuscript, and further improve the accuracy.
[0041] The above has introduced in detail a method and system for converting a content-editable document format based on visual recognition provided by the embodiments of the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method of the present application. At the same time, for those of ordinary skill in the art, according to the method of the present application, there will be changes in the specific implementation manner and application scope.
[0042] The above has introduced in detail a method and system for converting a content-editable document format based on visual recognition provided by the embodiments of the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method of the present application and is not used as the basis for the loss of rights of the invention in other fields.
[0043] In summary, the content of this specification should not be construed as a limitation of the present application. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present application should be covered by the protection scope of the present application. On the basis of the implementation manners provided in the above aspects of the present application, further combinations can be made to provide more implementation manners.
Claims
1. A method and system for converting content-editable document formats based on visual recognition, characterized in that: The method comprises: S1, the document format conversion system receives a document that needs to be formatted; S2, using image recognition and text analysis technology to extract content elements from the original document to form intermediate structured data; S3, generates an interactive platform with three columns of visual interface, the three columns are: a. The left column displays the layout rendering of the original document or the original file image; b. The middle column presents structured text and layout elements parsed by the visual model, wherein the structured text exists in the form of editable text blocks; c. The right column shows a preview of the final output style generated based on the template engine; S4, in response to the user's manual modification and editing of the content in the middle column, synchronously triggers the following mechanisms: a. The text semantic analysis module performs logical conflict detection on the modified content and compares it with the visual features of the original document to generate difference reminder marks; b. The template adaptation engine dynamically updates the style rendering results of the right column based on the corrected structured data; c. The visual positioning anchor point is displayed in the corresponding modified area in the original document in the left column; S5, outputting the finalized output style as the converted document.
2. The method according to claim 1, characterized in that The operations of the middle column include: a. Free reorganization and editing of text blocks: change the position relationship and content of text blocks through drag-and-drop and editing operations, triggering the layout rearrangement of upstream and downstream related texts; b. Composite Editing Mode: When you initiate the "Merge Cells" or "Insert Separator" command on the table structure, the visual features of the corresponding area in the original document are highlighted and updated synchronously.
3. The method according to claim 1, characterized in that: The training of the large visual model includes: a. Adopt the domain adaptation strategy to fine-tune the general document recognition model into a dedicated model for the target business scenario; b. The explainability module embedded with the attention mechanism generates a confidence heat map in the middle column for the recognition results of complex elements such as formulas and tables.
4. The method according to claim 1, characterized in that: Step S2 specifically includes: a. Perform OCR recognition of high-resolution document images in different regions through a sliding window mechanism to eliminate the problem of broken tables across pages; b. Perform vectorization tracing on vector graphics inserted into the document, retaining scalable vector graphics (SVG) editing capabilities in the middle column; c. End-to-end LSTM-CRF text correction network to solve the problems of uneven lighting and wrinkle interference when scanning documents.
5. An interactive verification system for document format conversion, characterized in that: include: a. Visual computing engine: integrates multimodal document understanding models to perform image-text separation, layout analysis, and semantic annotation; b. Difference comparison module: Use dynamic programming algorithm to align the original manuscript and intermediate structured data paragraph by paragraph; c. Style adaptive components: deploy a responsive template library based on CSS3 to automatically adapt to the layout constraints of different output formats; d. Operation backtracking unit: records the version tree of all correction operations in the middle column, and supports cross-validation of output results and original modification records.
6. The system according to claim 5, characterized in that The visual computing engine comprises: a. Document type discriminator: automatically selects the processing pipeline such as PDF / image / scanned document based on the file binary header information and visual features; b. Out-of-domain data processing module: Generative Adversarial Network (GAN) is used to perform printed-handwritten separation on documents containing handwritten annotations.
7. An intelligent document processing terminal device, characterized in that: include: a processor configured to execute the method according to any one of claims 1-4; b. Display unit, supporting touch-screen support for multi-touch zooming and cross-column dragging of the text block in the middle column; c. Hardware acceleration unit, which can integrate GPU resources or professional acceleration chips for real-time reasoning calculations of large visual models.
8. The method according to claim 1, characterized in that The interactive platform further comprises: a. Intelligent revision suggestion component: When it is detected that the semantic deviation between the edited content in the middle column and the original document exceeds the threshold, the original reference paragraph based on the attention weight is pushed; b. Context-aware toolbar: Dynamically loads a dedicated editing command set based on the type of the currently selected object (text / table / formula).
Citation Information
Cited By
File conversion method, device, equipment and program product
CN120471022A
Method, system and equipment for correcting and combining bidding documents of road and bridge engineering and medium
CN121029705A
Interactive big and small model collaborative public service complex document analysis method and device
CN121415423A
Document revision method, document revision device and storage medium
CN121435933A
Document processing method and device based on concurrent editing and electronic equipment
CN121659906A