Document layout restructuring method and apparatus therefor

CN122547446APending Publication Date: 2026-08-11NANJING VIVO SOFTWARE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-13
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

然而由于移动设备的屏幕尺寸通常小于标准文档页面,导致PDF、扫描图像等格式的文档在显示时往往无法呈现一整页的内容,用户需要频繁进行缩放、平移或滚动等操作才能阅读完整内容,导致阅读操作繁琐、效率低下

Benefits of technology

[0009] In a sixth aspect, embodiments of this application provide a computer program product stored in a storage medium, which is executed by at least one processor to implement the method as described in the first aspect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122547446A_ABST
    Figure CN122547446A_ABST
Patent Text Reader

Abstract

This application discloses a document layout reconstruction method and apparatus, belonging to the field of electronic device technology. The method includes: in response to a user's first input, extracting N page images from an original document; based on the N page images, performing layout analysis on each page image to determine the reading order of each document content area in the N page images; the document content area includes different types of media elements, including at least one of text, titles, images, tables, charts, and headers / footers, where N is a positive integer; determining the document's logical structure based on the content and format information of the document content area; reconstructing the original document's page layout into a flowing layout based on the reading order and the document's logical structure; using typesetting parameters, performing formatting and rendering on the flowing layout to generate a reconstructed document; the typesetting parameters are determined based on the screen parameters of a first electronic device; and displaying the reconstructed document.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of electronic equipment technology, specifically relating to a document layout reconstruction method and apparatus. Background Technology

[0002] With the increasing popularity of digital office work and mobile reading, users often need to view various documents on mobile devices of different sizes and resolutions, such as mobile phones, tablets, and e-readers. However, because the screen size of mobile devices is usually smaller than that of a standard document page, documents in formats such as PDFs and scanned images often cannot display a full page of content. Users need to frequently zoom, pan, or scroll to read the complete content, resulting in cumbersome and inefficient reading operations. Summary of the Invention

[0003] The purpose of this application is to provide a document layout reconstruction method and apparatus that can reconstruct the original document into a flow layout document suitable for reading on mobile terminals, thereby simplifying reading operations and improving reading efficiency.

[0004] In a first aspect, embodiments of this application provide a document layout reconstruction method, executed by a first electronic device, the method comprising: In response to the user's first input, extract N page images from the original document; Based on N page images, a layout analysis is performed on each page image to determine the reading order of each document content area in the N page images; the document content area includes different types of media elements, including at least one of text, title, image, table, chart, header and footer, where N is a positive integer; Determine the document's logical structure based on the content and formatting information of the document's content areas; Based on the reading order and the document's logical structure, the original document's page layout is reconstructed into a flow layout; Using layout parameters, the fluid layout is formatted and rendered to generate a restructured document; the layout parameters are determined based on the screen parameters of the first electronic device. Display the refactored document.

[0005] Secondly, embodiments of this application provide a document layout reconstruction apparatus, the apparatus comprising: The extraction module is used to extract N page images from the original document in response to the user's first input; The first determining module is used to perform layout analysis on each of the N page images to determine the reading order of each document content area in the N page images; the document content area includes different types of media elements, including at least one of text, title, image, table, chart, header and footer, where N is a positive integer; The second determining module is used to determine the logical structure of the document based on the content and format information of the document content area; The layout reconstruction module is used to reconstruct the page layout of the original document into a flow layout based on the reading order and the document's logical structure. The rendering module is used to format and render the fluid layout using layout parameters to generate a reconstructed document; the layout parameters are determined based on the screen parameters of the first electronic device. The display module is used to display the refactored document.

[0006] Thirdly, embodiments of this application provide an electronic device including a processor and a memory, the memory storing a program or instructions that can run on the processor, the program or instructions implementing the steps of the method as described in the first aspect when executed by the processor.

[0007] Fourthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, and when the program or instructions are executed by a processor, they implement the steps of the method as described in the first aspect.

[0008] Fifthly, embodiments of this application provide a chip, which includes a processor and a communication interface. The communication interface and the processor are coupled, and the processor is used to run programs or instructions to implement the method as described in the first aspect.

[0009] In a sixth aspect, embodiments of this application provide a computer program product stored in a storage medium, which is executed by at least one processor to implement the method as described in the first aspect.

[0010] In this embodiment, in response to the user's first input, N page images are extracted from the original document; based on the N page images, layout analysis is performed on each page image to determine the reading order of each document content area in the N page images; the document content area includes different types of media elements, including at least one of text, title, image, table, chart, header and footer, where N is a positive integer; the document logical structure is determined according to the content information and format information of the document content area; the page layout of the original document is reconstructed into a flow layout according to the reading order and the document logical structure; the flow layout is formatted and rendered using layout parameters to generate the reconstructed document; the layout parameters are determined based on the screen parameters of the first electronic device; and the reconstructed document is displayed.

[0011] This allows for the reconstruction of a fixed-layout original document into a fluid, screen-adaptive layout. By combining layout analysis of page images to determine the reading order with content and format analysis to determine the document's logical structure, the reconstructed document is not only visually smooth and conforms to reading habits, but also maintains the original text structure, ensuring logical clarity. Furthermore, rendering with typesetting parameters makes the reconstructed document's display effect more closely match the screen display requirements of the primary electronic device, thereby simplifying reading operations and improving reading efficiency. Attached Figure Description

[0012] Figure 1 This is a flowchart illustrating a document layout reconstruction method provided in some embodiments of this application; Figure 2 This is one of the interface diagrams of the document layout reconstruction method provided in some embodiments of this application; Figure 3 This is the second schematic diagram of the interface of the document layout reconstruction method provided in some embodiments of this application; Figure 4 This is a schematic flowchart of a scenario embodiment of the document layout reconstruction method provided in some embodiments of this application; Figure 5 This is a schematic diagram of the structure of a document layout reconstruction apparatus provided in some embodiments of this application; Figure 6 These are schematic diagrams of the structure of electronic devices provided in some embodiments of this application; Figure 7 These are schematic diagrams of the hardware structure of electronic devices provided in some embodiments of this application. Detailed Implementation

[0013] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0014] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0015] The document layout reconstruction method provided in this application will be described in detail below with reference to the accompanying drawings, through specific embodiments and application scenarios.

[0016] Figure 1 This is a flowchart illustrating the document layout reconstruction method provided in an embodiment of this application. The document layout reconstruction method is executed by a first electronic device, and the method may include: Step 101: In response to the user's first input, extract N page images from the original document.

[0017] In step 101, an Artificial Intelligence (AI) multimodal typesetting engine can be deployed on the first electronic device to achieve intelligent document transformation. For example, such as... Figure 2 As shown, after receiving the original document uploaded or downloaded by the user, the first electronic device displays the original document on the document management interface 200. At the same time, the document management interface 200 may also include functional controls for initiating intelligent document layout conversion, such as the "AI Reading" control 201. The user can activate the intelligent document layout conversion function by clicking the "AI Reading" control 201. That is, at this time, the first electronic device can reconstruct the original document into a flow layout to facilitate user reading.

[0018] The original document can be in various formats such as Word, PPT, Excel, and PDF; no specific limitation is made here. The following explanation will use scanned or photocopied PDF documents as an example. Understandably, since scanned or photocopied PDF documents are essentially image formats, the reading experience is rather poor when displayed on small-screen electronic devices such as mobile phones.

[0019] Based on this, the first electronic device can parse the original document and extract each page of the original document as a high-resolution image in the order of page number, thus obtaining N page images of the original document.

[0020] In some examples, metadata information of the original document can also be obtained simultaneously, including page size, number of pages, creation time, etc., to establish a basic data structure for subsequent content recognition and reordering.

[0021] In some embodiments, the method may further include: Extract N initial images from the original document; Detect the rotation angle of the original document; Based on the rotation angle, angle correction processing is performed on N initial images to obtain N page images.

[0022] In this embodiment, as mentioned above, each page of the original document can be extracted sequentially according to the page number order to obtain N initial images of the original document. At the same time, the rotation angle of the original document can be detected, and the angle can be automatically corrected based on the rotation angle to ensure that the orientation of the document content is correct, thereby obtaining N corrected page images.

[0023] In this way, by automatically detecting and correcting the rotation angle, the subsequent layout analysis is based on the correct visual coordinate system, avoiding the failure of the entire analysis process or the confusion of results caused by incorrect document orientation, thus enhancing the practicality and robustness of the method.

[0024] In some embodiments, the method may further include: If the original document is an encrypted document that is restricted from being opened, a password input interface is displayed in response to the user's second input; In response to a third input from the user on the password input interface, verify the password associated with the third input; If the password verification is successful, extract N page images from the original document.

[0025] In this embodiment, it is understood that the original document may still be an encrypted document.

[0026] Therefore, when parsing the original document, we can first check whether the original document contains encryption flags and permission restriction flags. Based on the encryption type, it can be divided into two categories: one is user password encryption, where the original document is an encrypted document restricted from being opened; the other is permission password encryption, where the original document is a document restricted from being edited or copied.

[0027] For encrypted documents that are restricted from being opened, a password input interface can be displayed in response to a second input that triggers a document layout reconstruction by the user. For example, a password input box can pop up on the current screen to prompt the user to enter a password.

[0028] It can respond to a third input from the user at the password input interface, obtain the password associated with the third input, and then verify the password. If the verification is successful, the parsing process can continue to extract N page images of the original document.

[0029] In some examples, if a user enters the wrong password three times in a row, the parsing of the original document will be terminated, and the message "Incorrect password, unable to open this document" may be displayed.

[0030] In other examples, if the original document uses an encryption algorithm that is not supported by the first electronic device, a message can be displayed saying "The current document encryption format is not supported. Please contact the document provider." The original document will then be marked as skipped, without affecting the normal processing flow of reconstructing the layout of other documents in the batch.

[0031] For documents that restrict editing or copying, pages can be extracted directly in read-only mode without affecting content reading, resulting in N page images of the original document.

[0032] This allows for password verification of encrypted documents, enabling the extraction of page images from the encrypted documents to perform subsequent layout reconstruction processes. While ensuring the security of the original document, it also ensures that layout reconstruction can proceed smoothly, thus broadening the applicability of document layout reconstruction methods.

[0033] Step 102: Based on N page images, perform layout analysis on each page image to determine the reading order of each document content area in the N page images; the document content area includes different types of media elements, including at least one of text, title, image, table, chart, header and footer, where N is a positive integer.

[0034] In step 102, computer vision technology can be used to analyze the layout of each page image, identify the regions of different types of media elements in the page image, obtain the document content regions, and then analyze the reading order of these document content regions. The media elements can include at least one of the following: text, titles, images, tables, charts, headers, and footers.

[0035] It is understandable that the reading order of document content areas can include the reading order of content within a single document content area, as well as the reading order between multiple document content areas.

[0036] In some embodiments, based on N page images, performing layout analysis on each page image to determine the reading order of each document content area in the N page images may include: Based on N page images, layout analysis is performed on each page image to identify the regions corresponding to different types of media elements in the N page images, resulting in M ​​document content regions; M is a positive integer. Determine the spatial relationships and layout structure of M document content areas; Determine the reading order of each document content area based on spatial relationships and layout structure.

[0037] In this embodiment, computer vision technology can be used to analyze the layout of each page image and identify the regions corresponding to different types of media elements in N page images. For example, object detection algorithms can be used to locate the positions and boundaries of various media elements such as text, titles, images, tables, charts, headers, and footers to obtain the document content areas.

[0038] Analyze the spatial relationships between these document content areas to identify complex layout structures such as multi-column layouts and text wrapping around images. Based on the spatial relationships and layout structures, determine the reading order of each document content area.

[0039] For example, areas with spatial relationships preceding each other should be read in a higher order than areas with spatial relationships following each other.

[0040] For example, the two-column format commonly used in academic papers can accurately distinguish between the left and right columns and determine the correct reading flow by converting the left and right column areas into a top-bottom arrangement according to the left-right order.

[0041] In this way, by first identifying different types of media element areas, then analyzing their spatial relationships and layout structure, and finally deriving the reading order, the layout analysis process becomes more refined, improving the accuracy and robustness of determining the reading order, and laying a solid foundation for generating a fluid layout that conforms to human reading habits.

[0042] In some embodiments, when the layout structure of the first document content area is text wrapping around an image, determining the reading order between the document content areas based on spatial relationships and layout structure may include: Determine the text wrapping direction based on the relative positions of images and text within the first document's content area; Convert the text into linearly arranged text based on the text wrapping direction; The reading order of the first document's content area is determined as follows: linearly arranged text is located below or above images; The first document content area is any one of the M document content areas.

[0043] In this embodiment, after identifying a document content area with a text wrapping structure of an image, the coordinates of the vertices around the image and the area coordinates of the surrounding text can be extracted to determine the relative positional relationship between the image and the text in the document content area, and thus determine whether the text wrapping direction is left wrapping, right wrapping, or wrapping around all four sides.

[0044] Text can be converted into linearly arranged text based on its wrapping direction. For example, if the text wraps to the left or right, a coordinate system can be constructed using a single text region. The text within that region is then arranged linearly according to its vertical coordinates, taking into account the screen width of the electronic device. If the text wraps around all four sides, multiple text regions can be constructed within the same coordinate system. The text within these regions is then arranged linearly according to its vertical coordinates, taking into account the screen width of the electronic device.

[0045] At this point, the reading order of the first document content area can be determined as follows: linearly arranged text is positioned below or above the images. In other words, when reconstructing the flow layout, images within this document content area can be converted into block-level independent elements. The image width can adapt to the screen width while maintaining its original aspect ratio and centering. Linearly arranged text can be placed below or above the images to ensure a consistent logical order of text and images in the reconstructed document displayed on the first electronic device.

[0046] In this way, by analyzing the relative positions of images and text to determine the wrapping direction, and converting the wrapping text into a linear arrangement, the vertical order of the text and images can be clearly defined. This can effectively restore the user's reading intention in the area where text wraps around images, improve the separation and misalignment of images and corresponding explanatory text after reconstruction, and enhance the coherence and accuracy of content expression.

[0047] In some embodiments, when N is an integer greater than 1, based on N page images, layout analysis is performed on each page image to identify regions corresponding to different types of media elements in the N page images, resulting in M ​​document content regions, which may include: Based on N page images, layout analysis is performed on each page image to identify the regions corresponding to different types of media elements in each page image, resulting in K document content regions corresponding to each page image; K is a positive integer. Perform continuity detection on the last document content region in the i-th page image and the first document content region in the (i+1)-th page image to obtain the detection result; i is a positive integer, and i+1 is less than or equal to N; If the detection result indicates that the last document content region in the i-th page image is continuous with the first document content region in the (i+1)-th page image, the last document content region in the i-th page image and the first document content region in the (i+1)-th page image are merged to obtain M document content regions corresponding to N page images.

[0048] In this embodiment, K document content regions can be the number of document content regions included in a single page image, and M document content regions are the total number of document content regions included in all page images.

[0049] During the page layout analysis phase, continuity detection can be performed on the boundary regions of adjacent pages. If the last document content region in the i-th page image is continuous with the first document content region in the (i+1)-th page image, the last document content region in the i-th page image and the first document content region in the (i+1)-th page image are merged. Then, the M merged document content regions are further processed to determine the reading order of each document content region.

[0050] For example, cross-page tables and cross-page paragraphs can be identified by comparing the table border features at the bottom of the current page and the top of the next page, as well as the integrity of the text paragraphs. Cross-page tables merge adjacent page segments into a complete logical table based on column width alignment. Cross-page paragraphs determine continuity by detecting whether the last character is a terminating punctuation mark, concatenating them, and then rearranging them.

[0051] In this way, by detecting the continuity of the beginning and end areas between adjacent pages and merging continuous content, the problem of table and long paragraph fragmentation caused by physical pagination can be effectively improved, the integrity of the content of the reconstructed fluid layout document can be improved, and the smoothness of reading long documents can be enhanced.

[0052] Step 103: Determine the document's logical structure based on the content and formatting information of the document's content area.

[0053] In step 103, the logical structure of the document can be determined based on the content and format information of the document content area. For example, it can be determined which content is at different levels such as first-level headings, second-level headings, body text, citations, and footnotes, which content belongs to the same topic, and which are parallel chapters.

[0054] In some embodiments, determining the document logical structure based on the content information and format information of the document content area may include: Based on the recognition strategies corresponding to different types of media elements, identify the content and format information of each document content area; Perform semantic analysis on the content information to identify the topic categories and hierarchical relationships of the text; Determine the document's logical structure based on topic categories, hierarchical relationships, and formatting information.

[0055] In this embodiment, corresponding recognition strategies are adopted for the regions corresponding to different types of media elements in order to identify the content information and format information of each document content region.

[0056] For example, for document content areas where the media element is text, a high-precision optical character recognition (OCR) engine is used to recognize the text content, while extracting formatting information such as font, font size, color, bold, and italics. For document content areas where the media element is an image or chart, the image itself is extracted, and related text such as image titles and captions is recognized. For document content areas where the media element is a table, the table's row and column structure is recognized, and cell content and its logical relationships are extracted. Furthermore, it can also recognize mathematical formulas, special symbols, and other specialized content, ensuring the completeness and accuracy of the recognition.

[0057] It can perform semantic analysis on content information, identify the topic categories and hierarchical relationships of the text, and determine the logical structure of the document based on topic categories, hierarchical relationships and format information.

[0058] For example, natural language processing (NLP) techniques are used to perform semantic analysis on the identified content information to understand semantic logic and obtain the topic categories and hierarchical relationships of the identified text. Then, by analyzing the characteristics of formatting information such as font size, color, and position, combined with topic categories and hierarchical relationships, it is determined which content belongs to different levels such as first-level headings, second-level headings, body text, citations, and footnotes. Furthermore, the logical relationships between paragraphs are identified, determining which content belongs to the same topic and which are parallel chapters. For structured content such as lists and numbering, their hierarchical relationships are accurately extracted. In this way, the logical structure of the original document can be determined.

[0059] In this way, by acquiring content and format information through multimodal recognition strategies and combining them with semantic analysis, the logical structure of the document can be inferred from both visual format and text content levels. This allows the reconstructed document to retain the hierarchy and information organization of the original document, thereby improving the reliability of the fluid layout reconstruction.

[0060] Step 104: Based on the reading order and the document's logical structure, reconstruct the original document's page layout into a flow layout.

[0061] In step 104, the fixed page layout of the original document can be converted into a flowing layout suitable for mobile devices. The identified content is reorganized according to the reading order, with elements such as titles, body text, images, and tables arranged vertically according to the document's logical structure. Users can read continuously by simply swiping up and down without needing to move left or right.

[0062] For example, for multi-column content in the original document, the layout is converted from horizontal multi-column to vertical single-column scrolling layout to maintain content continuity. For paragraphs or tables that span multiple pages, they are automatically joined to eliminate the sense of disjointedness at page boundaries, achieving a smooth reading experience similar to web articles.

[0063] In some examples, the scaling ratio for converting a fixed page layout to a fluid layout can be calculated based on the pixel ratio of the first electronic device, the effective screen width, and the page width and left and right margins of the original document. The calculation formula can be shown in formula (1): Where Sfinal is the streaming layout rendering scaling ratio, Wdevice is the effective screen width of the first electronic device, Wpdf is the page width of the original document, M is the left and right margins of the page, and DPR is the pixel ratio of the first electronic device.

[0064] After the width is scaled proportionally, the height of the page content changes accordingly. The height calculation formula is as shown in Formula (2): Among them, Sfinal is the rendering scaling ratio of the flow layout, Hnew is the content height of the reconstructed flow layout, and Hpdf is the page height of the original document.

[0065] It can be understood that the constraint conditions of the above Formula (1) and Formula (2) are: 0 < Sfinal ≤ 1 and (Wdevice 2 × M) > 0. That is, the scaling ratio does not exceed the original size, and the effective content width must be greater than zero.

[0066] Step 105: Use the layout parameters to perform format layout rendering on the flow layout to generate a reconstructed document; the layout parameters are determined based on the screen parameters of the first electronic device.

[0067] In Step 105, the layout parameters matching the first electronic device can be determined first based on the screen parameters of the first electronic device. The layout parameters may include text attribute parameters such as the size of the main text font, the ratio of the title font size, the line spacing, the paragraph spacing, the page margin, and the scaling ratio of non-text media elements, and may also include personalized attribute parameters such as brightness, color, theme background, and page turning method.

[0068] Among them, the brightness can default to display the brightest value, and the brightest value follows the brightness of the current screen state. The adjustment process does not affect the screen brightness display. The color can include white, yellow, green, black, etc., and can default to display the display mode. If it is the light mode, it can default to display white; if it is the dark mode, it can default to display black. The theme background can adjust the theme background setting according to the global memory. After adjusting the theme background setting, it is default to be memorized, and the subsequent entries will all be based on the last adjustment parameter. The translation method can include up and down sliding page turning, left and right page turning, automatic page turning, etc.

[0069] In some examples, the corresponding relationship between the screen parameters and the text attribute parameters in the layout parameters can be determined in advance, and the corresponding text attribute parameters can be directly determined by looking up the table based on the screen parameters of the first electronic device. In other examples, the text attribute parameters such as the size of the main text font, the ratio of the title font size, the line spacing, the paragraph spacing, the page margin, and the scaling ratio of non-text media elements can also be calculated based on the screen parameters of the first electronic device. The personalized attribute parameters such as brightness, color, theme background, and page turning method can be default parameters preset by the first electronic device and can be changed in response to the user's adjustment input.

[0070] In some embodiments, before generating the restructured document by formatting and rendering the fluid layout using layout parameters, the method may further include: Obtain the screen parameters of the first electronic device; In response to the user's fourth input, obtain the configuration parameters corresponding to the fourth input; Determine the layout parameters based on the screen parameters and configuration parameters.

[0071] In this embodiment, screen parameters such as screen width, pixel ratio, and preset margin of the first electronic device can be obtained.

[0072] It can respond to a fourth user input and retrieve the configuration parameters corresponding to that fourth input. For example... Figure 3 As shown, a layout parameter configuration interface 300 can be displayed, where users can input corresponding configuration parameters. For example, users can customize font size, font format, line spacing, paragraph spacing, brightness, color, theme background, page turning method, etc. Setting the font size can be done by directly setting the body text font size, setting the number of characters per line, or setting the font size level. For example, if the first electronic device pre-sets the number of characters per line to be between 25 and 35, the corresponding number of characters per line can be matched according to the size level.

[0073] The appropriate body text font size can be calculated based on the screen width, preset margins, and pixel ratio of the first electronic device. The font size of each level of heading can be calculated based on the body text font size and the preset heading level coefficient. The line spacing can be calculated based on the body text font size and the preset line spacing coefficient.

[0074] For example, the formula for calculating the body text font size is shown in formula (3): in, It refers to the body text font size. M is the screen width, and M is the page margin. This refers to the target number of characters per line. It's the pixel ratio.

[0075] The formula for calculating the title font size ratio is shown in formula (4): in, This is the font size for the nth level heading; This is the scaling factor for the nth level heading. The scaling factor can be set according to actual needs; no specific limitation is made here. For example... =2.0, =1.5, =1.25.

[0076] The formula for calculating line spacing is shown in formula (5): in, It refers to line spacing; This refers to the line spacing factor, which can be set according to actual needs. No specific limit is specified here. For example... The value can be 1.5 1.8).

[0077] The scaling formula for non-text media elements such as images and tables is shown in formula (6): in, This refers to the scaling ratio of non-text media elements; This is the original width of the non-text media element; It is the original height of the non-text media element; It is the scaled height of a non-text media element.

[0078] The constraints for calculating the above typesetting parameters can be set as follows: 25≤ ≤35, 1.5≤ ≤1.8 and 0< ≤1. This ensures that the number of characters per line in the reconstructed document is within a comfortable range, the font size is not too large or too small, the line spacing is within a comfortable range, it is not too dense to affect reading, and it is not too sparse to waste the effective display area of ​​the screen, and it also ensures that the media content of non-text media elements does not exceed the original size.

[0079] In this way, by combining the device's objective screen parameters with the user's subjective configuration preferences to determine the final layout parameters, the reconstructed document can automatically adapt to the hardware while respecting the user's personalized reading habits, achieving a balance between intelligent adaptation and user customization, and improving the user's reading experience.

[0080] Layout parameters can be used to format and render the fluid layout, generating a restructured document. This transforms the original document into a restructured document with a fluid layout adapted to the layout parameters. The restructured document can apply fonts suitable for mobile reading, ensuring clarity and readability on small screens. Headings are bolded and appropriately enlarged to create a clear visual hierarchy with the body text. Important content retains its original color markings or highlighting effects. Line spacing and paragraph spacing are automatically adjusted to avoid overcrowding, and appropriate white space is added.

[0081] In some examples, for long documents, an interactive table of contents can be generated based on the heading information in the document's logical structure, allowing users to quickly jump to chapters of interest.

[0082] The refactored document can be output in various formats, including HTML5 web page format, Epub ebook format, or the application's native reading format.

[0083] In other examples, the output refactored document supports interactive features such as text selection, copying, searching, and annotation. Users can highlight, annotate, and bookmark the content of the refactored document, and these markings are saved synchronously. It also provides a reading progress memory function, automatically resetting the user's reading position the next time they open the document.

[0084] Step 106: Display the refactored document.

[0085] In step 106, the reconstruction document can be displayed on the screen of the first electronic device for the user's convenience.

[0086] For example, after a user clicks the "AI Reading" control, the AI ​​multimodal typesetting engine of the first electronic device can execute the document layout reconstruction method described above. Before the reconstructed document is generated, a full-screen loading page can be displayed, showing the "Smart Loading" status text and the loading progress. If the user closes the application or loading is interrupted for other reasons, the loading progress should be maintained upon the next entry. After the reconstructed document is generated, loading is complete, and a cover and title can be generated and displayed. Responding to the user's reading operation, the reconstructed document can be read on the first electronic device.

[0087] In this embodiment, the document layout reconstruction method responds to a user's first input by extracting N page images from the original document; based on the N page images, it performs layout analysis on each page image to determine the reading order of each document content area in the N page images; the document content area includes different types of media elements, including at least one of text, title, image, table, chart, header and footer, where N is a positive integer; the document logical structure is determined according to the content and format information of the document content area; the page layout of the original document is reconstructed into a fluid layout according to the reading order and the document logical structure; the fluid layout is formatted and rendered using layout parameters to generate the reconstructed document; the layout parameters are determined based on the screen parameters of the first electronic device; and the reconstructed document is displayed.

[0088] This allows for the reconstruction of a fixed-layout original document into a fluid, screen-adaptive layout. By combining layout analysis of page images to determine the reading order with content and format analysis to determine the document's logical structure, the reconstructed document is not only visually smooth and conforms to reading habits, but also maintains the original text structure, ensuring logical clarity. Furthermore, rendering with typesetting parameters makes the reconstructed document's display effect more closely match the screen display requirements of the primary electronic device, thereby simplifying reading operations and improving reading efficiency.

[0089] In some embodiments, when the fluid layout includes non-text media elements, before formatting and rendering the fluid layout using typography parameters to generate the reconstructed document, the method may further include: Perform processing operations on non-text media elements in a flow layout; the processing operations include at least one of the following: Scaling the image; Convert the chart to vector format; Determine how the table will be rendered.

[0090] In this embodiment, after the fluid layout is reconstructed, processing operations can also be performed on non-text media elements in the fluid layout to optimize the display of media elements such as images, charts, and tables in the document.

[0091] For example, images can be scaled proportionally to the screen width to ensure full display and clear visibility, with users able to zoom in for details. Image titles and captions can be automatically placed below or above the image to maintain context. Spacing between text and images can also be optimized to ensure clear visual hierarchy.

[0092] For complex charts, you can choose to maintain the original resolution and support local zoom, or you can convert the chart to an interactive vector format.

[0093] The table can intelligently select its rendering method based on the number of columns and the complexity of the content, determining whether the rendered table is displayed with an adaptive width to show the full table, displayed horizontally with scrolling, or converted to a card layout to facilitate user reading.

[0094] In this way, by performing targeted processing operations on images, charts, tables, etc. in advance, it is ensured that these elements can be presented clearly, beautifully, and fully functionally in the flow layout, improving the distortion, information loss, or inconvenience of interaction caused by simple stretching or compression, and enhancing the overall presentation quality of the document.

[0095] In some embodiments, when the non-text media element is a table, performing processing operations on the non-text media element in the flow layout may include: Get the number of columns and content complexity of the table; The rendering method of the table is determined based on the number of columns and the complexity of the content; the rendering method includes any of the following: adaptive width rendering, horizontal scrolling rendering, and card layout rendering.

[0096] In this embodiment, the number of columns and content complexity of the table can be obtained. The content complexity can be determined based on the cell content length and / or the header level.

[0097] The table rendering method can be determined based on the number of columns and the complexity of the content. For example, if the number of columns does not exceed the first threshold (e.g., 3-5 columns) and the average number of characters per cell does not exceed the second threshold (e.g., 8-12 characters), it is considered a simple table and can be rendered directly using adaptive width rendering.

[0098] If the number of columns is greater than the first threshold and less than or equal to the third threshold (e.g., 6 to 8 columns), a horizontal scrolling rendering method can be used for rendering, and the resulting table can be displayed horizontally.

[0099] If the number of columns exceeds the third threshold, or if there are multiple levels of headers, a card layout will be used for rendering. The rendered table will be automatically converted to a card layout, and each row of data will be presented in the form of vertical key-value pairs of "field name - field value" to ensure readability on narrow screens.

[0100] In some examples, when rendering tables using horizontal scrolling, the first column of the table can be fixed in place, and the header row can remain suspended on top as the user scrolls. In other words, for tables with horizontal scrolling enabled, the first column, which serves as the row header, can be fixed and snapped to the left side of the screen, remaining visible as the user scrolls horizontally. The top header row automatically floats on top during vertical scrolling, ensuring that users can refer to the row and column headers while browsing any cell, avoiding loss of context due to scrolling and improving the reading efficiency of multidimensional table data.

[0101] In this way, by evaluating the number of columns and the complexity of the content, the most suitable rendering method is dynamically selected to render the table, which improves the problem of wide tables being incomplete or deformed and compressed on small screens, making them difficult to read, and achieves optimal visualization and operability of table content on different screens.

[0102] In some embodiments, after displaying the reconstructed document, the method may further include: Receive user input regarding the shared refactoring document; In response to shared input, the refactored document is sent to a second electronic device.

[0103] In this embodiment, it is possible to quickly share the refactored document after the flow layout reconstruction. In response to the user's sharing input, the refactored document can be sent to the second electronic device so that the user of the second electronic device can directly and quickly read the refactored document on the second electronic device.

[0104] In this way, reconstructed documents can be easily transferred to other devices, which not only improves the reading experience for individual users but also promotes the dissemination of high-quality reading formats.

[0105] To facilitate understanding of the document layout reconstruction method provided in the above embodiments, the following describes the document layout reconstruction method using a specific scenario embodiment. Figure 4 This is a schematic flowchart of a scenario embodiment of the document layout reconstruction method provided in this application.

[0106] like Figure 4 As shown, taking a PDF document as an example, this scenario implementation can specifically include the following steps: Step 401: Parse the PDF document and extract page images; Step 402: Perform layout analysis on the page image, segment out the document content area, and determine the reading order based on spatial relationships and layout format; Step 403: Perform multimodal content recognition on the document content area to obtain content information and format information; Step 404: Perform semantic analysis on the content information and determine the document's logical structure by combining it with the format information; Step 405: Reconstruct the flow layout based on the reading order and document logical structure; Step 406: Calculate the layout parameters adapted to the device; Step 407: Perform processing operations on non-text media elements to optimize the display effect; Step 408: Render the reconstructed flow layout based on the layout parameters to generate the reconstructed document; Step 409: Output and sharing of the reconstructed document.

[0107] In this scenario, the reading experience and efficiency of PDF documents on mobile devices are improved. Intelligent reflow simplifies the previously frequent zooming and horizontal swiping operations to a single vertical swipe. This lowers the barrier to entry for using PDF documents on mobile devices, converting them into a readable format while preserving the original document's structure, diagrams, formatting, and other important information, ensuring the integrity and logic of the content. Users no longer need to switch between computers and mobile phones; they can read any PDF document anytime, anywhere, breaking down device limitations.

[0108] The document layout reconstruction method provided in this application can be executed by a document layout reconstruction device. This application uses an example of a document layout reconstruction device performing document layout reconstruction to illustrate the document layout reconstruction device provided in this application.

[0109] like Figure 5 As shown in the figure, this application embodiment provides a document layout reconstruction device 500, which may include: Extraction module 501 is used to extract N page images from the original document in response to the user's first input; The first determining module 502 is used to perform layout analysis on each of the N page images based on the N page images, and determine the reading order of each document content area in the N page images; the document content area includes different types of media elements, and the media elements include at least one of text, title, image, table, chart, header and footer, where N is a positive integer; The second determining module 503 is used to determine the logical structure of the document based on the content information and format information of the document content area; The layout reconstruction module 504 is used to reconstruct the page layout of the original document into a flow layout based on the reading order and the document's logical structure. Rendering module 505 is used to perform formatting and rendering of the flow layout using layout parameters to generate a reconstructed document; the layout parameters are determined based on the screen parameters of the first electronic device. Display module 506 is used to display the refactored document.

[0110] This allows for the reconstruction of a fixed-layout original document into a fluid, screen-adaptive layout. By combining layout analysis of page images to determine the reading order with content and format analysis to determine the document's logical structure, the reconstructed document is not only visually smooth and conforms to reading habits, but also maintains the original text structure, ensuring logical clarity. Furthermore, rendering with typesetting parameters makes the reconstructed document's display effect more closely match the screen display requirements of the primary electronic device, thereby simplifying reading operations and improving reading efficiency.

[0111] In some embodiments, the first determining module 501 may specifically be used for: Based on N page images, layout analysis is performed on each page image to identify the regions corresponding to different types of media elements in the N page images, resulting in M ​​document content regions; M is a positive integer. Determine the spatial relationships and layout structure of M document content areas; Determine the reading order of each document content area based on spatial relationships and layout structure.

[0112] In this way, by first identifying different types of media element areas, then analyzing their spatial relationships and layout structure, and finally deriving the reading order, the layout analysis process becomes more refined, improving the accuracy and robustness of determining the reading order, and laying a solid foundation for generating a fluid layout that conforms to human reading habits.

[0113] In some embodiments, when the layout structure of the first document content area is text wrapping around an image, the first determining module 501 may specifically be used for: Determine the text wrapping direction based on the relative positions of images and text within the first document's content area; Convert the text into linearly arranged text based on the text wrapping direction; The reading order of the first document's content area is determined as follows: linearly arranged text is located below or above images; The first document content area is any one of the M document content areas.

[0114] In this way, by analyzing the relative positions of images and text to determine the wrapping direction, and converting the wrapping text into a linear arrangement, the vertical order of the text and images can be clearly defined. This can effectively restore the user's reading intention in the area where text wraps around images, improve the separation and misalignment of images and corresponding explanatory text after reconstruction, and enhance the coherence and accuracy of content expression.

[0115] In some embodiments, when N is an integer greater than 1, the first determining module 501 may specifically be used to: Based on N page images, layout analysis is performed on each page image to identify the regions corresponding to different types of media elements in each page image, resulting in K document content regions corresponding to each page image; K is a positive integer. Perform continuity detection on the last document content region in the i-th page image and the first document content region in the (i+1)-th page image to obtain the detection result; i is a positive integer, and i+1 is less than or equal to N; If the detection result indicates that the last document content region in the i-th page image is continuous with the first document content region in the (i+1)-th page image, the last document content region in the i-th page image and the first document content region in the (i+1)-th page image are merged to obtain M document content regions corresponding to N page images.

[0116] In this way, by detecting the continuity of the beginning and end areas between adjacent pages and merging continuous content, the problem of table and long paragraph fragmentation caused by physical pagination can be effectively improved, the integrity of the content of the reconstructed fluid layout document can be improved, and the smoothness of reading long documents can be enhanced.

[0117] In some embodiments, the second determining module 502 may specifically be used for: Based on the recognition strategies corresponding to different types of media elements, identify the content and format information of each document content area; Semantic analysis of the content information is performed to identify the topic categories and hierarchical relationships of the text. Determine the document's logical structure based on topic categories, hierarchical relationships, and formatting information.

[0118] In this way, by acquiring content and format information through multimodal recognition strategies and combining them with semantic analysis, the logical structure of the document can be inferred from both visual format and text content levels. This allows the reconstructed document to retain the hierarchy and information organization of the original document, thereby improving the reliability of the fluid layout reconstruction.

[0119] The document layout reconstruction device in this application embodiment can be an electronic device or a component within an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the device.

[0120] The document layout reconstruction device in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit it.

[0121] The document layout reconstruction apparatus provided in this application embodiment can implement the various processes implemented in the method embodiment, and will not be described again here to avoid repetition.

[0122] Optionally, such as Figure 6As shown, this application embodiment also provides an electronic device 600, including a processor 601 and a memory 602. The memory 602 stores a program or instructions that can run on the processor 601. When the program or instructions are executed by the processor 601, they implement the various steps of the above-described document layout reconstruction method embodiment and can achieve the same technical effect. To avoid repetition, they will not be described again here.

[0123] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.

[0124] Figure 7 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application.

[0125] The electronic device 700 includes, but is not limited to, components such as: radio frequency unit 701, network module 702, audio output unit 703, input unit 704, sensor 705, display unit 706, user input unit 707, interface unit 708, memory 709, and processor 710.

[0126] Those skilled in the art will understand that the electronic device 700 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 710 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 7 The electronic device structure shown does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.

[0127] The processor 710 can be used for: In response to the user's first input, extract N page images from the original document; Based on N page images, a layout analysis is performed on each page image to determine the reading order of each document content area in the N page images; the document content area includes different types of media elements, including at least one of text, title, image, table, chart, header and footer, where N is a positive integer; Determine the document's logical structure based on the content and formatting information of the document's content areas; Based on the reading order and the document's logical structure, the original document's page layout is reconstructed into a flow layout; Using layout parameters, the fluid layout is formatted and rendered to generate a restructured document; the layout parameters are determined based on the screen parameters of the first electronic device. Display unit 706 can be used to display reconstructed documents.

[0128] This allows for the reconstruction of a fixed-layout original document into a fluid, screen-adaptive layout. By combining layout analysis of page images to determine the reading order with content and format analysis to determine the document's logical structure, the reconstructed document is not only visually smooth and conforms to reading habits, but also maintains the original text structure, ensuring logical clarity. Furthermore, rendering with typesetting parameters makes the reconstructed document's display effect more closely match the screen display requirements of the primary electronic device, thereby simplifying reading operations and improving reading efficiency.

[0129] In some embodiments, the processor 710 can also be used for: Based on N page images, layout analysis is performed on each page image to identify the regions corresponding to different types of media elements in the N page images, resulting in M ​​document content regions; M is a positive integer. Determine the spatial relationships and layout structure of M document content areas; Determine the reading order of each document content area based on spatial relationships and layout structure.

[0130] In this way, by first identifying different types of media element areas, then analyzing their spatial relationships and layout structure, and finally deriving the reading order, the layout analysis process becomes more refined, improving the accuracy and robustness of determining the reading order, and laying a solid foundation for generating a fluid layout that conforms to human reading habits.

[0131] In some embodiments, when the layout structure of the first document content area is text wrapping around an image, the processor 710 can also be used for: Determine the text wrapping direction based on the relative positions of images and text within the first document's content area; Convert the text into linearly arranged text based on the text wrapping direction; The reading order of the first document's content area is determined as follows: linearly arranged text is located below or above images; The first document content area is any one of the M document content areas.

[0132] In this way, by analyzing the relative positions of images and text to determine the wrapping direction, and converting the wrapping text into a linear arrangement, the vertical order of the text and images can be clearly defined. This can effectively restore the user's reading intention in the area where text wraps around images, improve the separation and misalignment of images and corresponding explanatory text after reconstruction, and enhance the coherence and accuracy of content expression.

[0133] In some embodiments, when N is an integer greater than 1, the processor 710 can also be used for: Based on N page images, layout analysis is performed on each page image to identify the regions corresponding to different types of media elements in each page image, resulting in K document content regions corresponding to each page image; K is a positive integer. Perform continuity detection on the last document content region in the i-th page image and the first document content region in the (i+1)-th page image to obtain the detection result; i is a positive integer, and i+1 is less than or equal to N; If the detection result indicates that the last document content region in the i-th page image is continuous with the first document content region in the (i+1)-th page image, the last document content region in the i-th page image and the first document content region in the (i+1)-th page image are merged to obtain M document content regions corresponding to N page images.

[0134] In this way, by detecting the continuity of the beginning and end areas between adjacent pages and merging continuous content, the problem of table and long paragraph fragmentation caused by physical pagination can be effectively improved, the integrity of the content of the reconstructed fluid layout document can be improved, and the smoothness of reading long documents can be enhanced.

[0135] In some embodiments, the processor 710 can also be used for: Based on the recognition strategies corresponding to different types of media elements, identify the content and format information of each document content area; Perform semantic analysis on the content information to identify the topic categories and hierarchical relationships of the text; Determine the document's logical structure based on topic categories, hierarchical relationships, and formatting information.

[0136] In this way, by acquiring content and format information through multimodal recognition strategies and combining them with semantic analysis, the logical structure of the document can be inferred from both visual format and text content levels. This allows the reconstructed document to retain the hierarchy and information organization of the original document, thereby improving the reliability of the fluid layout reconstruction.

[0137] It should be understood that, in this embodiment, the input unit 704 may include a graphics processing unit (GPU) 7041 and a microphone 7042. The GPU 7041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 706 may include a display panel 7061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 707 includes at least one of a touch panel 7071 and other input devices 7072. The touch panel 7071 is also called a touch screen. The touch panel 7071 may include a touch detection device and a touch controller. Other input devices 7072 may include, but are not limited to, a physical keyboard, function keys (such as volume control buttons, power buttons, etc.), a trackball, a mouse, and a joystick, which will not be described in detail here.

[0138] The memory 709 can be used to store software programs and various data. The memory 709 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 709 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 709 in the embodiments of this application includes, but is not limited to, these and any other suitable types of memory.

[0139] Processor 710 may include one or more processing units; optionally, processor 710 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 710.

[0140] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described document layout reconstruction method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.

[0141] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0142] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described document layout reconstruction method embodiment and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0143] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0144] This application provides a computer program product that is stored in a storage medium and executed by at least one processor to implement the various processes of the document layout reconstruction method embodiment described above, and achieves the same technical effect. To avoid repetition, it will not be described again here.

[0145] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0146] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0147] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. A method of document layout restructuring, characterized by, Performed by a first electronic device, the method includes: In response to the user's first input, extract N page images from the original document; Based on the N page images, a layout analysis is performed on each page image to determine the reading order of each document content area in the N page images; the document content area includes different types of media elements, and the media elements include at least one of text, title, image, table, chart, header and footer, where N is a positive integer; The document's logical structure is determined based on the content and formatting information of the document's content area; Based on the reading order and the document's logical structure, the page layout of the original document is reconstructed into a flow layout; Using layout parameters, the flow layout is formatted and rendered to generate a reconstructed document; the layout parameters are determined based on the screen parameters of the first electronic device. Display the restructured document.

2. The method of claim 1, wherein, The step of performing layout analysis on each of the N page images to determine the reading order of each document content area in the N page images includes: Based on the N page images, layout analysis is performed on each page image to identify the regions corresponding to different types of media elements in the N page images, resulting in M ​​document content regions; M is a positive integer. Determine the spatial relationships and layout structure of the M document content areas; Based on the spatial relationships and the layout structure, determine the reading order of each document content area.

3. The method of claim 2, wherein, When the layout structure of the first document content area is text wrapping around an image, determining the reading order between each document content area based on the spatial relationship and the layout structure includes: The text wrapping direction is determined based on the relative positional relationship between the image and text in the first document content area; Based on the text wrapping direction, the text is converted into linearly arranged text; The reading order of the first document content area is determined as follows: the linearly arranged text is located below or above the image; The first document content area is any one of the M document content areas.

4. The method of claim 2, wherein, When N is an integer greater than 1, based on the N page images, layout analysis is performed on each page image to identify the regions corresponding to different types of media elements in the N page images, resulting in M ​​document content regions, including: Based on the N page images, layout analysis is performed on each page image to identify the regions corresponding to different types of media elements in each page image, thereby obtaining K document content regions corresponding to each page image; K is a positive integer; Perform continuity detection on the last document content region in the i-th page image and the first document content region in the (i+1)-th page image to obtain the detection result; i is a positive integer, and i+1 is less than or equal to N; If the detection result indicates that the last document content region in the i-th page image is continuous with the first document content region in the (i+1)-th page image, the last document content region in the i-th page image and the first document content region in the (i+1)-th page image are merged to obtain M document content regions corresponding to the N page images.

5. The method of claim 1, wherein, The step of determining the document logical structure based on the content and format information of the document content area includes: Based on the recognition strategies corresponding to different types of media elements, identify the content and format information of each document content area; Perform semantic analysis on the content information to identify the topic categories and hierarchical relationships of the text; The document's logical structure is determined based on the topic category, the hierarchical relationship, and the format information.

6. A document layout restructuring apparatus characterized by comprising: The device includes: The extraction module is used to extract N page images from the original document in response to the user's first input; The first determining module is used to perform layout analysis on each of the N page images based on the N page images, and determine the reading order of each document content area in the N page images; the document content area includes different types of media elements, and the media elements include at least one of text, title, image, table, chart, header and footer, where N is a positive integer; The second determining module is used to determine the logical structure of the document based on the content information and format information of the document content area; The layout reconstruction module is used to reconstruct the page layout of the original document into a flow layout based on the reading order and the document's logical structure. The rendering module is used to perform formatting and rendering of the fluid layout using layout parameters to generate a reconstructed document; the layout parameters are determined based on the screen parameters of the first electronic device. The display module is used to display the reconstructed document.

7. The apparatus of claim 6, wherein, The first determining module is specifically used for: Based on the N page images, layout analysis is performed on each page image to identify the regions corresponding to different types of media elements in the N page images, resulting in M ​​document content regions; M is a positive integer. Determine the spatial relationships and layout structure of the M document content areas; Based on the spatial relationships and the layout structure, determine the reading order of each document content area.

8. The apparatus of claim 7, wherein, When the layout structure of the first document content area is text wrapping around an image, the first determining module is specifically used for: The text wrapping direction is determined based on the relative positional relationship between the image and text in the first document content area; Based on the text wrapping direction, the text is converted into linearly arranged text; The reading order of the first document content area is determined as follows: the linearly arranged text is located below or above the image; The first document content area is any one of the M document content areas.

9. The apparatus of claim 7, wherein, When N is an integer greater than 1, the first determining module is specifically used for: Based on the N page images, layout analysis is performed on each page image to identify the regions corresponding to different types of media elements in each page image, thereby obtaining K document content regions corresponding to each page image; K is a positive integer; Perform continuity detection on the last document content region in the i-th page image and the first document content region in the (i+1)-th page image to obtain the detection result; i is a positive integer, and i+1 is less than or equal to N; If the detection result indicates that the last document content region in the i-th page image is continuous with the first document content region in the (i+1)-th page image, the last document content region in the i-th page image and the first document content region in the (i+1)-th page image are merged to obtain M document content regions corresponding to the N page images.

10. The apparatus according to claim 6, characterized in that, The second determining module is specifically used for: Based on the recognition strategies corresponding to different types of media elements, identify the content and format information of each document content area; Perform semantic analysis on the content information to identify the topic categories and hierarchical relationships of the text; The document's logical structure is determined based on the topic category, the hierarchical relationship, and the format information.