Layout analysis method and device
By decoupling the layout analysis task into a logical sequence and a parallel element localization stage, and utilizing multimodal large models and object detection technology, a structured text sequence is generated and coordinate positions are processed in parallel. This solves the problems of time-consuming reasoning and low efficiency in handling complex scenarios in existing technologies, and achieves efficient and accurate layout analysis.
Patent Information
- Application Number
- CN202511453411.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-13
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2045-10-13
AI Technical Summary
In existing document layout analysis methods, the autoregressive mechanism leads to excessively long output sequences and significant time consumption in inference, making it difficult to meet the requirements of efficient and real-time layout analysis. Furthermore, traditional methods struggle to balance the accuracy of logical order and processing efficiency when dealing with complex scenarios.
The layout analysis task is decoupled into two stages: "logical order and content parsing" and "parallel element localization". A structured text sequence without coordinates is generated, and the logical order and coordinate position are processed separately through multimodal large model and object detection technology to achieve parallel processing.
It improves the reasoning efficiency of layout analysis, and can improve the efficiency of processing complex documents while preserving the logical order and understanding accurately, thus solving the technical problem of balancing efficiency and accuracy in traditional methods.
Smart Images

Figure CN120932258A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer technology, and in particular to a layout analysis method and apparatus. Background Technology
[0002] Layout analysis refers to the technology of processing document images to automatically identify and locate different content elements such as text blocks, headings, tables, and images, and to analyze the spatial layout and logical hierarchy among them. Because it can transform unstructured document images into structured, computer-processable data, it is widely used in many scenarios such as office automation, financial document recognition, intelligent document retrieval, and knowledge base construction.
[0003] Currently, document layout analysis primarily employs end-to-end methods based on multimodal large language models. These methods typically input the document image into a large multimodal model and, through autoregression, generate a structured sequence that integrates the text content and coordinate positions of all layout elements (such as text blocks and tables) within the document. However, the autoregressive mechanism requires generating all content tags and coordinate position tags individually, resulting in excessively long output sequences and significant inference time consumption, making it difficult to meet the demands of efficient and real-time layout analysis applications. Summary of the Invention
[0004] This invention provides a layout analysis method and apparatus to address the deficiencies in the prior art.
[0005] This invention provides a layout analysis method, comprising the following steps: Extract image features from the document to be analyzed and text features from the layout analysis prompts; Using the image features and the text features, a structured sequence containing the text content of each element in the document and the logical order between the elements is generated; Extract the feature representations corresponding to each element during the generation of the structured sequence; Based on the feature representations of each element, target detection is performed on each element to obtain the coordinate positions of each element.
[0006] According to a layout analysis method provided by the present invention, the step of generating a structured sequence containing text content of each element within a document and the logical order between elements by utilizing the image features and the text features includes: The image features and text features are fused to generate a multimodal input representation; Autoregressive decoding is performed based on the multimodal input representation to map the visual layout order of each element provided by the image features to the logical order between elements, and to identify and generate the text content of each element, thereby obtaining the structured sequence.
[0007] According to a layout analysis method provided by the present invention, the step of mapping the visual layout order of each element provided by the image features to a logical order between elements, and identifying and generating the text content of each element to obtain the structured sequence includes: During the process of generating the text content of each element in the logical order, predefined structured markers are embedded at the corresponding positions in the content stream to wrap the generated text content, thereby defining the hierarchy and scope of each element and obtaining the structured sequence.
[0008] According to a layout analysis method provided by the present invention, the step of extracting the feature representations corresponding to each element during the generation of the structured sequence includes: Extract the embedding vectors corresponding to each word segmentation during the structured sequence generation process; The embedding vectors corresponding to each word segment are input into a mapping layer consisting of a multilayer perceptron and an activation function for processing, generating a mapped feature vector. Based on the element inclusion relationship defined by the start and end markers in the structured sequence, all mapped feature vectors belonging to the same element are merged to obtain the feature representation of the corresponding element.
[0009] According to a layout analysis method provided by the present invention, the step of performing target detection on each of the elements based on the feature representation of each element to obtain the coordinate position of each element includes: Combine all the features of the same element into query features; Based on the query features of each element, target detection is performed on each element to obtain the coordinate position of each element.
[0010] According to a layout analysis method provided by the present invention, the step of combining all feature representations of the same element into query features includes: For any element defined by the start marker and the end marker in the structured sequence, extract the feature representations corresponding to the start marker, the end marker, and all word segments located between the start marker and the end marker; All extracted feature representations are merged to generate the query feature for any of the elements.
[0011] According to a layout analysis method provided by the present invention, the step of extracting image features of the document to be analyzed includes: The image of the document to be analyzed is divided into multiple image blocks of predetermined size; Each image block is flattened into a vector, and the vectors corresponding to all image blocks are combined with the position encoding vector to obtain an initial embedding representation. The position encoding vector is used to characterize the position of each image block in the image of the document to be analyzed. The initial embedding representation is encoded to obtain the image features of the document to be analyzed.
[0012] The present invention also provides a layout analysis device, comprising the following modules: The first extraction unit is used to extract the image features of the document to be analyzed and the text features of the layout analysis prompt text; The element parsing unit is used to generate a structured sequence containing the text content of each element in the document and the logical order between the elements by utilizing the image features and the text features. The second extraction unit is used to extract the feature representations corresponding to each element during the generation of the structured sequence; The feature localization unit is used to perform target detection on each of the features based on the feature representation of each feature, and obtain the coordinate position of each feature.
[0013] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the layout analysis method as described above.
[0014] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the layout analysis method as described above.
[0015] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the layout analysis method as described above.
[0016] The layout analysis method and apparatus provided by this invention decouples the layout analysis task into two stages: "logical order and content parsing" and "parallel element localization." First, it generates a structured text sequence without coordinates. Then, it extracts feature representations from this generation process and uses these feature representations to output the coordinate positions of all elements in parallel, achieving a unity of logical order recovery and efficient reasoning. Because this invention separates the most time-consuming coordinate generation task in traditional methods from the serial autoregressive decoding process and innovatively transforms it into a parallelizable target detection task, it significantly improves the reasoning efficiency of the entire layout analysis process while fully preserving the accurate understanding of the logical order of complex document layouts (such as multi-column, mixed text and image layouts, and nested tables). This effectively solves the technical problem of balancing efficiency and accuracy in traditional methods. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0018] Figure 1 This is a flowchart illustrating the layout analysis method provided by the present invention.
[0019] Figure 2 This is a flowchart illustrating another layout analysis method provided by the present invention.
[0020] Figure 3 This is a schematic diagram of the layout analysis device provided by the present invention.
[0021] Figure 4 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0022] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0023] Currently, document layout analysis primarily employs autoregressive large language models, object detection-based methods, and traditional rule- or template-based methods. While autoregressive large language model methods can maintain the logical reading order of document content through sequential generation, their character-by-character or token-by-tag decoding nature results in slow inference speeds, making them unsuitable for real-time parsing. Object detection-based methods, although achieving efficient inference through parallel prediction, produce unordered layout elements (such as text blocks and tables), losing the original hierarchical structure and reading logic of the document. Traditional rule- or template-based methods heavily rely on prior knowledge and struggle to adapt to complex scenarios such as multi-column layouts, irregular tables, or mixed text and image layouts.
[0024] Therefore, traditional text layout analysis methods cannot simultaneously ensure both the accuracy of logical order and high processing efficiency in layout analysis tasks. To address this, this invention provides a layout analysis method that can be applied to the automated structural analysis and information extraction of various document images. For example, it can process scanned academic papers, financial statements, legal contracts, manuscripts, or various invoices to identify and locate different layout elements such as titles, paragraphs, tables, and images. Figure 1 This is a flowchart illustrating the layout analysis method provided by the present invention, as shown below. Figure 1 As shown, the method includes steps 110, 120, 130 and 140.
[0025] Step 110: Extract the image features of the document to be analyzed and the text features of the layout analysis prompt text.
[0026] Here, the document to be analyzed can be understood as a document that requires layout analysis. It can exist in image form or be an electronic document in other formats, such as a digital native document containing structured information such as text, vector graphics, or embedded images (e.g., PDF, Word documents). During processing, such electronic documents can be rendered or converted into image formats for subsequent image feature extraction. This embodiment does not specifically limit this. The document to be analyzed can be a scanned copy of a document obtained through a scanner or document scanner, a photograph of a document taken using a mobile phone or camera, or obtained from electronic files such as PDF or PPT. For example, the document to be analyzed can be a single page of an academic paper image containing multi-column layout, mixed text and images, and nested tables.
[0027] Image features refer to vector representations extracted from a document to be analyzed that characterize its visual content (such as texture, shape, color, layout, etc.). These features can be extracted using backbone networks such as Convolutional Neural Networks (CNNs) or Vision Transformers (ViT). For example, a pre-trained ResNet model can be used, taking the document as input to obtain feature maps at different levels. These feature maps can then be further processed (e.g., flattening, pooling) to generate the final image features. Considering the advantages of ViT models in capturing global dependencies in images, this embodiment preferably uses ViT as the visual encoder. Specifically, ViT is used to segment the image into a series of image patches, and each patch is flattened and combined with a positional encoding vector before being fed into the Transformer encoder, thereby obtaining image features that can finely characterize global layout information.
[0028] Page layout analysis prompt text refers to a piece of natural language text used to guide the model in performing a specific page layout analysis task. This prompt text can be a prompt message. The prompt text clarifies the analysis objective and the desired output format. For example, for the page layout analysis of a paper, the prompt text could be "Please analyze the page structure of this document and identify the header, footer, headings, and body paragraphs"; for an invoice, the prompt text could be "Please extract all table items from the invoice."
[0029] Text features refer to the vector representations that characterize the semantic meaning of the layout analysis prompt text obtained by processing it using a text encoder. These features can be extracted using a text encoder, such as BERT (Bidirectional Encoder Representations from Transformers), or a dedicated prompt encoder. For example, inputting the prompt text "Please analyze the layout structure of this document" into the BERT model will yield the corresponding text features.
[0030] The text features can be obtained based on the following formula:
[0031] This represents the text features obtained after encoder processing. This indicates a layout analysis prompt. This indicates that the text used for layout analysis is encoded.
[0032] Step 120: Using image features and text features, generate a structured sequence containing the text content of each element in the document and the logical order between the elements.
[0033] Specifically, each element within a document can be understood as a component unit of the document being analyzed that has an independent function in terms of layout and content. For example, headers, footers, titles, paragraphs, tables, figures, and formulas can all be considered document elements.
[0034] The logical order between elements refers to the order that conforms to human reading habits or the inherent hierarchical relationship of a document, rather than a simple physical layout order from top to bottom or left to right. For example, for a two-column layout page, the logical order could be to read all the content in the left column first, and then read all the content in the right column; for a title and its paragraphs, the logical order is the title first, followed by the paragraphs.
[0035] A structured sequence is a special type of text sequence that not only contains the textual content of each element, but also explicitly represents the category, scope, and logical and hierarchical relationships between the elements through predefined structured markers. For example, a structured sequence can be...<doc_s><header_s> Paper title< / header_e><page_s> This is the first paragraph.< / page_e><footer_s> Page 1_e>< / doc_e> .in,<doc_s> and<doc_e> These indicate the beginning and end of the document, respectively.<header_s> and<header_e> The header elements have been defined.
[0036] As an optional implementation, step 120 can be implemented using a multimodal large model. This model receives the image features and text features extracted in step 110 as input and effectively combines the two information through an internal fusion mechanism (such as cross-attention). Subsequently, the model generates tokens one by one in an autoregressive manner, ultimately forming a complete structured sequence. Because the model's internal mechanism can simultaneously understand the visual layout of the image and the instructions of the prompt text when generating the structured sequence, it can correctly map the visual layout relationship into a logical order that conforms to reading habits, and simultaneously complete the recognition of text within each element. The key to this step is that the generated structured sequence contains the text content of each element within the document and a structured sequence of the logical order between elements, but does not contain the coordinate information of each element. This greatly shortens the length of the output sequence, thereby significantly improving the inference speed.
[0037] The structured sequence can be obtained based on the following formula:
[0038] in, Represents a structured sequence. Representing text features, Representing image features, This represents a multimodal large model.
[0039] Step 130: Extract the feature representations corresponding to each element during the generation of the structured sequence.
[0040] Here, feature representation can be understood as the hidden state or embedding vector within the multimodal large model of step 120, corresponding one-to-one with each token (whether content segmentation or structured token) in the generated structured sequence. This feature representation is a highly condensed information carrier that integrates the image visual information, text prompt information, and preceding content information relied upon when generating the token.
[0041] As an optional embodiment, during the autoregressive generation process in step 120, the hidden state vector of the top layer or a specific layer corresponding to each token generated by the decoder can be cached and recorded. For example, when the model generates a sequence...<header_s> Paper title< / header_e> At that time, it is possible to extract the corresponding data.<header_s> "Paper", "Title"< / header_e> The internal feature vectors corresponding to these four tokens are used as the feature representations that constitute the "header" element.
[0042] Since these original feature vectors may have high dimensionality or not be optimized for subsequent object detection tasks, this embodiment preferably processes these vectors through a projection layer after extraction. For example, this projection layer can consist of one or more multilayer perceptrons (MLPs) and activation functions (such as ReLU or GeLU), which map the original hidden state vectors to a feature space more suitable as an object detection query, resulting in mapped feature vectors.
[0043] Step 140: Based on the feature representation of each element, perform target detection on each element to obtain the coordinate position of each element.
[0044] Specifically, coordinate position usually refers to the bounding box used to accurately locate the feature region on the original image. It can be represented by the coordinates of the top left and bottom right corners (x1, y1, x2, y2) or by the coordinates of the top left corner and the width and height (x, y, w, h).
[0045] The feature representations of each element are used as "query features" to drive object detection. Each feature representation contains rich semantic information such as the element's category and content, which can accurately guide the detection model "where to find this specific element in the image".
[0046] As an optional implementation, a query-based object detection architecture, such as DETR (DEtection TRansformer) or the DINO model, can be used to detect objects in each element. Specifically, for any element (such as a header) defined by start and end markers in a structured sequence, the feature representations corresponding to all tokens within it can be merged (e.g., summed, averaged, or concatenated) to form a single feature representation representing the complete element. Then, the feature representations of all elements are input in parallel into the decoder of the object detection model, and the model can simultaneously output the coordinate positions corresponding to all elements.
[0047] Since the feature representations of different elements are independent of each other, the object detection process can be fully parallelized, which contrasts sharply with traditional methods that require sequential autoregressive generation of each coordinate value. Considering the superior performance of the DINO model in handling the detection of small and dense objects, this embodiment preferably adopts DINO as the basic architecture of the feature localization module in order to accurately locate small lines of text or closely arranged table cells that may exist in the document.
[0048] The layout analysis method provided in this embodiment decouples the layout analysis task into two stages: "logical order and content parsing" and "parallel element localization." First, it generates a structured text sequence without coordinates. Then, it extracts feature representations from this generation process and uses these feature representations to output the coordinate positions of all elements in parallel, achieving a unity of logical order recovery and efficient reasoning. Because this embodiment innovatively separates the most time-consuming coordinate generation task in traditional methods from the serial autoregressive decoding process and transforms it into a parallelizable object detection task, it significantly improves the reasoning efficiency of the entire layout analysis process while fully preserving the accurate understanding of the logical order of complex document layouts (such as multi-column, mixed text and image layouts, and nested tables). This effectively solves the technical challenge of balancing efficiency and accuracy in traditional methods.
[0049] Based on the above embodiments, using image features and text features, a structured sequence containing the text content of each element within a document and the logical order between elements is generated, including: Image features are fused with text features to generate a multimodal input representation; Autoregressive decoding based on multimodal input representation maps the visual layout order of each element provided by image features to the logical order between elements, and identifies and generates the text content of each element to obtain a structured sequence.
[0050] Specifically, a multimodal input representation can be understood as a sequence of hybrid feature vectors that includes both visual information from images and semantic instructions from text. By fusing image features with text features, the resulting multimodal input representation contains both image features representing the visual content of the document to be analyzed and text features representing the instructions for the layout analysis task.
[0051] As an optional embodiment, the fusion process can be to concatenate the image feature vector sequence obtained in step 110 with the text feature vector sequence. For example, if the dimension of the image features is [N, D] and the dimension of the text features is [M, D], then the dimension of the multimodal input representation obtained after concatenation is [N+M, D].
[0052] Furthermore, to achieve deeper interaction and alignment between the two modalities, this embodiment preferably employs a more refined fusion strategy. For example, image features and text features can be jointly input into a multimodal fusion module (such as a cross-attention layer of one or more Transformers). In this module, text features can act as a query to "focus" on relevant regions in image features; conversely, image features can also act as a query to "focus" on key instructions in text features. Ultimately, the module outputs a deeply fused and aligned multimodal input representation.
[0053] Next, after obtaining the multimodal input representation, autoregressive decoding is performed based on the multimodal input representation to map the visual layout order of each element provided by the image features into the logical order between elements, and to identify and generate the text content of each element, thus obtaining a structured sequence.
[0054] Here, autoregressive decoding is a sequential generative process, meaning that when generating the output (a token) at the current moment, it depends on all previously generated outputs and the initial multimodal input representation. This process is like building the final sequence "word by word," with each step's decision based on "images seen" and "previous text written."
[0055] In this embodiment, a multimodal large model decoder can be used for autoregressive decoding. That is, after receiving the fused multimodal input representation, the multimodal large model decoder starts to perform autoregressive decoding. Its core tasks include logical order mapping and text content generation.
[0056] Regarding logical order mapping, the model has learned rich knowledge of page layout through pre-training on massive amounts of document data. Therefore, when it "sees" the visual typography information contained in the multimodal input representation (e.g., an image of a two-column document), it does not mechanically generate text from top to bottom and left to right. Instead, it understands the inherent reading flow and intelligently maps the visual typography order to a logical order that conforms to human reading habits. For example, the model will first generate all the content of the left column and then begin generating the content of the right column, even if the top of the right column may be physically higher than some parts of the left column.
[0057] For text content generation, as the model "moves" along a predetermined logical order, it focuses on the corresponding region of the image and performs a function similar to optical character recognition (OCR), recognizing the pixel information in the image and converting it into the corresponding text tokens, which are then generated one by one.
[0058] In this way, the model simultaneously completes the structural understanding, sequence reconstruction, and content recognition of the document in a single autoregressive decoding process, ultimately outputting a structured sequence that contains both accurate textual content and reflects the correct logical order.
[0059] This embodiment deeply fuses image features and text features, and utilizes the autoregressive decoding capability of a multimodal large model to achieve a deep understanding of the document's visual layout and accurate recognition of text content in one step. This end-to-end approach not only simplifies the processing flow and avoids the need for separate modules such as layout analysis, region sorting, and OCR in traditional methods, but more importantly, it can handle complex scenarios that were previously difficult to solve (such as irregular text flows and cross-column tables), ensuring that the final generated structured sequence has high accuracy in both logical order and text content dimensions.
[0060] Based on any of the above embodiments, the visual layout order of each element provided by image features is mapped to the logical order between elements, and the text content of each element is identified and generated to obtain a structured sequence, including: During the process of generating the text content of each element in a logical order, predefined structured markers are embedded at the corresponding positions in the content flow to wrap the generated text content, thereby defining the hierarchy and scope of each element and obtaining a structured sequence.
[0061] Here, predefined structured tokens can be understood as a set of special tokens added to the multimodal large model vocabulary. These tokens themselves do not carry the meaning of natural language words, but rather serve as meta-instructions or tags specifically used to represent the structural information of a document. For example, we can define...<doc_s> and<doc_e> These indicate the beginning and end of the document, respectively.<header_s> and<header_e> These represent the beginning and end of the header element, respectively; similarly, there can also be...<page_s> / <page_e> (Main text page)<footer_s> / <footer_e> (footer), / (Table) and other paired markers.
[0062] By using paired start and end markers, the entire text content contained in any element can be precisely defined. For example, all elements located within...<header_s> and< / header_e> The text tokens between them are all explicitly attributed to the same header element.
[0063] By nesting tags, the inclusion relationship between elements can be clearly expressed. For example, the header element (composed of...)<header_s> ...< / header_e> Package) and footer elements (by<footer_s> ..._e>wrap) can be nested within a single page element (by<page_s> ...< / page_e> Within the package, the hierarchical tree structure of the document is accurately reflected.
[0064] As an optional implementation, when a multimodal large model begins to generate a logically new feature (e.g., a document header) based on its understanding of the image layout, it first generates a start marker for that feature (e.g., ...).<header_s> The model then focuses on the corresponding header area in the image and generates the text content within that area one by one (e.g., "Chapter 1 Introduction"). Once all the text content for that element has been generated, the model immediately generates its corresponding end marker (e.g., ...).< / header_e> After that, the model moves to the next element (e.g., the main text paragraph) according to the logical order, and repeats the process of "generating start marker -> generating content -> generating end marker".
[0065] For example, for a simple document page containing a header, body, and footer, the final generated structured sequence might look like this: out =<doc_s><header_s> …< / header_e><page_s> …< / page_e><footer_s> …_e>< / doc_e> .
[0066] This embodiment intelligently embeds pairs of structured tags into the text content stream, transforming the document's two-dimensional physical layout and implicit logical hierarchy into a one-dimensional, self-contained, machine-readable text sequence. This not only makes the boundaries and affiliations of elements clear and unambiguous, but also provides the necessary foundation for the subsequent accurate extraction and aggregation of feature representations of each element.
[0067] Based on any of the above embodiments, the feature representations corresponding to each element in the structured sequence generation process are extracted, including: Extract the embedding vectors corresponding to each word segmentation during the structured sequence generation process; The embedding vectors corresponding to each word segment are input into a mapping layer consisting of a multilayer perceptron and an activation function for processing, generating a mapped feature vector. Based on the element inclusion relationship defined by the start and end markers in the structured sequence, all mapped feature vectors belonging to the same element are merged to obtain the feature representation of the corresponding element.
[0068] Specifically, the embedding vector corresponding to each word segment can be understood as the internal hidden state vector generated by the multimodal large model during the autoregressive decoding process to generate each token (which can be understood as "word segment") in the structured sequence.
[0069] As an optional implementation, the hidden state vector corresponding to each token can be extracted as an embedding vector from the output of the last layer or several layers of the multimodal large model decoder. For example, for a sequence...<header_s> Paper title< / header_e> This step will extract the relevant data.<header_s> "Paper", "Title"< / header_e> These four tokens correspond one-to-one to four embedding vectors, denoted as [embed_h_s, embed_t1, embed_t2, embed_h_e].
[0070] After obtaining the embedding vectors corresponding to each word segment, the embedding vectors corresponding to each word segment are input into a mapping layer consisting of a multilayer perceptron and an activation function for processing to generate a mapped feature vector.
[0071] Here, the mapping layer can be understood as a lightweight neural network module whose main function is to transform and refine the feature space of the original embedding vectors. The original embedding vectors primarily serve the text generation task, and their feature distribution may not be entirely suitable as query input for subsequent object detection tasks. The mapping layer, through nonlinear transformations, can adjust these vectors into a more discriminative feature space that is more suitable for the localization task.
[0072] The mapping layer can be structured as a multilayer perceptron (MLP) consisting of one or more fully connected layers, with activation functions (such as ReLU, GeLU, Sigmoid, etc.) used between layers to introduce non-linearity. The feature mapping can be expressed based on the following formula:
[0073] in, This represents the mapped feature vector. When a large-scale language model (LLM) decodes and generates structured sequences, it is related to the first... The embedding vector corresponding to each token. When a large-scale language model (LLM) decodes and generates structured sequences, it is related to the first... The embedding vector corresponding to each token. When a large-scale language model (LLM) decodes and generates structured sequences, it is related to the first... The embedding vector corresponding to each token. This represents a multilayer perceptron. This represents the activation function.
[0074] For example, feeding the four embedding vectors [embed_h_s, embed_t1, embed_t2, embed_h_e] into the mapping layer one by one will result in a new set of optimized feature vectors [Z_h_s, Z_t1, Z_t2, Z_h_e].
[0075] Finally, based on the feature inclusion relationship defined by the start and end markers in the structured sequence, all mapped feature vectors belonging to the same feature are merged to obtain the feature representation of the corresponding feature.
[0076] Here, the element inclusion relationship is defined by pairs of structured tags (such as...).<header_s> and< / header_e> As clearly defined, these markers clearly distinguish which consecutive tokens together constitute a complete element.
[0077] Considering that subsequent target detection requires a single, fixed-dimensional query feature to characterize each element to be located, and that an element itself is described by multiple, variable-number feature vectors corresponding to word segmentation, this embodiment merges all mapped feature vectors belonging to the same element, so that the resulting feature representation can comprehensively summarize and characterize the complete information of the element as a whole.
[0078] The above-mentioned merging operation can be to add all mapped feature vectors belonging to the same element element by element, or to calculate the average value of all mapped feature vectors belonging to the same element, or to take the maximum value from all mapped feature vectors belonging to the same element in each dimension, or to connect all vectors in order into a longer vector.
[0079] This embodiment transforms the fine-grained token-level features generated during the decoding process into a global element-level feature representation that can macroscopically characterize each independent page element. This element-level feature representation not only integrates the semantic information of all content within the element, but also implicitly includes the element's type (through the features of the token) and scope information.
[0080] Based on any of the above embodiments, target detection is performed on each element based on its feature representation to obtain the coordinate position of each element, including: Combine all the features of the same element into query features; Based on the query characteristics of each element, target detection is performed on each element to obtain the coordinate position of each element.
[0081] Here, query features refer to feature vectors that uniquely represent elements in the document being analyzed, guiding the localization process in object detection. Each query feature acts as a "search instruction," asking the model, "Please find the object (i.e., the layout element) that matches this feature in the image." A high-quality query feature should contain sufficiently rich semantic information so that the model can accurately match it with a specific region in the image.
[0082] In this process, all mapped feature vectors belonging to the same element (including feature vectors corresponding to content segmentation and structured tokens) can be aggregated into a single vector through methods such as summation and averaging. This vector represents the query feature of the corresponding element. For example, the query feature of a header element is formed by merging the mapped feature vectors of all its internal tokens.
[0083] Next, based on the query features of each element, target detection is performed on each element to obtain the coordinate position of each element. This step can be implemented using an object detection model. Considering the large differences in element size and the possibility of dense arrangement or even nesting in the document layout, this embodiment preferably uses a query-based detection model, such as DINO (DETR with Improved Denoising Anchor Boxes).
[0084] Specifically, query features (e.g., [Query_header, Query_page, Query_footer,...]) are input into the decoder of the DINO model in parallel, while the extracted image features are also input into the encoder and decoder of the DINO model.
[0085] Within DINO's decoder, each query feature interacts multiple times with image features through a cross-attention mechanism. During this process, the query feature gradually "focuses" on the region in the image that best matches it.
[0086] After multiple layers of decoding, each query feature will eventually output two prediction results: one is the prediction of the category of the feature (for example, confirming that it is a "header"), and the other is the prediction of the coordinate position of the feature in the image, which is usually represented by the four coordinate values of a bounding box.
[0087] Because the DINO model's decoder can process all query features simultaneously, the coordinates of all elements in the document are calculated in parallel during a single forward propagation, unlike autoregressive models which generate coordinate values one by one, thus greatly improving inference efficiency.
[0088] This embodiment constructs corresponding query features for each element and uses these features for target detection, thereby enabling the precise location of all elements to be found simultaneously in a single calculation, avoiding the time bottleneck caused by serial decoding in traditional coordinate generation methods.
[0089] Based on any of the above embodiments, combining all feature representations of the same element into a query feature includes: For any element defined by the start and end markers in the structured sequence, extract the feature representations corresponding to the start marker, the end marker, and all word segments located between the start and end markers; All extracted feature representations are merged to generate query features for any element.
[0090] Specifically, the start marker and the end marker (such as...)<header_s> and< / header_e> The corresponding feature representation mainly carries the "category" information of the element. For example,<header_s> The feature vectors semantically point to "this is a header", which is crucial for distinguishing different types of elements.
[0091] The feature representations corresponding to all the word segments located between the two (such as "paper title") mainly carry the content information of the elements. This part of the information plays a decisive role in distinguishing elements with different content but the same type (e.g., two different paragraphs).
[0092] Based on this, the feature representation of the start marker provides a clear category identity for the element, the feature representation of the end marker defines the endpoint of the element content, and the feature representation of all the word segments in between fills in the specific semantic details of the element, thus enabling the construction of an overall profile of the element that is complete in information dimensions and has both category and content specificity.
[0093] After obtaining all feature representations, considering that subsequent target detection requires a single, fixed-dimensional vector as query input to locate a complete element, this embodiment merges all extracted feature representations, thereby aggregating multiple discrete feature information of an element into a unified overall feature representation. That is, the obtained query feature is a high-dimensional aggregated vector that simultaneously contains the category identity and specific content semantics of the element.
[0094] For example, when it is necessary to locate the entire document page, you can target the page by...<doc_s> and<doc_e> The outermost element defined is extracted.<doc_s> ,<doc_e> This process involves taking the feature representations of all tokens corresponding to all other elements contained within them (such as headers, body text, footers, etc.), and then generating a query feature representing the entire page through a merging operation (e.g., summation). This process can be schematically represented as: query_doc = Z_<doc_s> + Z_<header_s> + ... + Z_<footer_e> + Z_<doc_e> , where Z represents the mapped feature vector.
[0095] Based on any of the above embodiments, image features of the document to be analyzed are extracted, including: The image of the document to be analyzed is divided into multiple image blocks of predetermined size; Each image patch is flattened into a vector, and the vectors corresponding to all image patches are combined with the position encoding vectors to obtain the initial embedding representation. The position encoding vectors are used to characterize the position of each image patch in the image of the document to be analyzed. The initial embedding representation is encoded to obtain the image features of the document to be analyzed.
[0096] Here, a pre-defined image block refers to a series of small image regions obtained by applying a fixed-size sliding window (without overlap) to the corresponding image in the document. For example, for a 224×224 pixel image, if the pre-defined size is set to 16×16 pixels, the image will be divided into (224 / 16)×(224 / 16)=196 image blocks.
[0097] Considering that after obtaining multiple image patches, the subsequent flattening operation will lose the original two-dimensional spatial arrangement information of each image patch, and the model that encodes the initial embedding representation itself does not have the ability to process the sequence order, it is necessary to add positional information to the sequence composed of each image patch. The positional encoding vector is a learnable or fixed vector with the same dimension as the flattened image patch vector, and its value uniquely represents the absolute or relative position of each image patch in the original image grid. Combining all the flattened image patch vectors with the corresponding positional encoding vectors (usually element-wise addition) injects the spatial positional information of each image patch into its content representation, so that the resulting initial embedding representation can simultaneously represent the local visual content of each image patch and its positional information in the global image. Here, flattening into a vector means directly stretching each two-dimensional image patch into a one-dimensional long vector.
[0098] Since the initial embedding representation contains both the visual content of each local region (i.e., image patch) of the image and explicitly injects the spatial location information of each region in the global image through positional encoding, the image features obtained after encoding the initial embedding representation can fully capture the long-distance dependency between any two regions in the image through the multi-head self-attention mechanism. This forms a comprehensive perception of the overall layout structure of the document (such as columns, margins, and text-image position relationships), thereby providing richer and more contextual visual input for subsequent multimodal fusion and logical order parsing, significantly improving the accuracy of analyzing complex layout documents.
[0099] Wherein, the initial embedding representation It can be obtained based on the following formula:
[0100]
[0101] in, Indicates the segmented first... Image blocks, Indicates the segmented first... Image blocks, Indicates the segmented first... Image blocks, Indicates the total number of image patches. This represents the embedding matrix, which is used to map flattened image patches from pixel space to a high-dimensional feature space. Indicates to give one × An image patch of pixels and C channels (e.g., C=3 in an RGB image) is flattened into one After obtaining a C-dimensional vector, it is then projected into a D-dimensional feature vector. This represents the positional encoding vector.
[0102] The initial embedding representation can be encoded using the following formula:
[0103]
[0104]
[0105] in, Indicates the current level. This indicates the total number of layers in the Transformer encoder. Indicates the first The output of the layer, that is, the first layer The input of the layer. The layer normalization operation is used to stabilize the training process and accelerate model convergence. MSA stands for Multi-Head Self-Attention. Indicates after the first The intermediate output obtained after computation by the attention layer. MLP stands for Multilayer Perceptron, which typically consists of two linear layers and a nonlinear activation function (such as GELU). Indicates the first The final output of the Transformer encoder layer. Indicates the final output sequence The 0th vector in. This represents the final output sequence obtained after processing by all L layers of encoders. This represents the final image features output by the entire visual encoder module, which represent the entire document image to be analyzed.
[0106] Based on any of the above embodiments Figure 2 This is a flowchart illustrating another layout analysis method provided by the present invention, as shown below. Figure 2 As shown, the method includes: First, the image of the document to be analyzed is input into a vision encoder. This encoder segments the image into multiple image patches and converts them into a series of vision tokens containing visual information about the image, i.e., image features. Simultaneously, the layout analysis prompt text used to guide the layout analysis task is encoded into text tokens, i.e., text features.
[0107] Next, the vision token and text token are input into a large-scale language model (LLM), which begins generating text sequences in an autoregressive manner. When generating each token, previous generation results and fused multimodal features are referenced. During the generation process, the model embeds predefined structured tokens into the identified text content stream, such as...<doc_s> (Document begins)<header_s> (Header begins)< / header_e> (End of header) etc.
[0108] Ultimately, the output of this stage is a complete sequence of structured text, such as out =<doc_s><header_s> ...< / header_e> ...<doc_e> The structured sequence contains the textual content of all elements, and preserves the logical relationships between elements through the nesting and order of tags.
[0109] Furthermore, embedding vectors generated by the LLM during the generation of structured sequences are extracted, and these embedding vectors correspond one-to-one with each word in the output. These embedding vectors are then processed by a projection layer (usually an MLP), which transforms them into a feature space more suitable for the target localization task, generating mapped feature vectors.
[0110] Based on the inclusion relationships of the structured sequences, the mapped feature vectors are grouped. For example, all from...<header_s> arrive< / header_e> The corresponding feature vectors are grouped together to represent the "header" element. All feature vectors within each group are merged into a single, highly condensed feature vector through a fusion operation (such as summation or averaging). This final vector is the feature representation of the element, i.e., the query feature.
[0111] The query features of all elements are input in parallel into a query-based object detection model (such as DINO). DINO uses these query features as "object-finding instructions" to search for the region in the image features that best matches each query. Since all queries are processed simultaneously, DINO can output the coordinate positions of all elements at once, such as doc_pos, header_pos, and page_pos on the right side of the figure.
[0112] The layout analysis device provided by the present invention is described below. The layout analysis device described below and the layout analysis method described above can be referred to in correspondence.
[0113] Based on any of the above embodiments Figure 3 This is a schematic diagram of the layout analysis device provided by the present invention, as shown below. Figure 3 As shown, the device includes: The first extraction unit 310 is used to extract the image features of the document to be analyzed and the text features of the layout analysis prompt text; The feature parsing unit 320 is used to generate a structured sequence containing the text content of each feature in the document and the logical order between the features by utilizing image features and text features. The second extraction unit 330 is used to extract the feature representations corresponding to each element during the structured sequence generation process; The feature localization unit 340 is used to perform target detection on each feature based on the feature representation of each feature, and obtain the coordinate position of each feature.
[0114] Based on any of the above embodiments, using image features and text features, a structured sequence containing the text content of each element within a document and the logical order between elements is generated, including: Image features are fused with text features to generate a multimodal input representation; Autoregressive decoding based on multimodal input representation maps the visual layout order of each element provided by image features to the logical order between elements, and identifies and generates the text content of each element to obtain a structured sequence.
[0115] Based on any of the above embodiments, the visual layout order of each element provided by image features is mapped to the logical order between elements, and the text content of each element is identified and generated to obtain a structured sequence, including: During the process of generating the text content of each element in a logical order, predefined structured markers are embedded at the corresponding positions in the content flow to wrap the generated text content, thereby defining the hierarchy and scope of each element and obtaining a structured sequence.
[0116] Based on any of the above embodiments, the feature representations corresponding to each element in the structured sequence generation process are extracted, including: Extract the embedding vectors corresponding to each word segmentation during the structured sequence generation process; The embedding vectors corresponding to each word segment are input into a mapping layer consisting of a multilayer perceptron and an activation function for processing, generating a mapped feature vector. Based on the element inclusion relationship defined by the start and end markers in the structured sequence, all mapped feature vectors belonging to the same element are merged to obtain the feature representation of the corresponding element.
[0117] Based on any of the above embodiments, target detection is performed on each element based on its feature representation to obtain the coordinate position of each element, including: Combine all the features of the same element into query features; Based on the query characteristics of each element, target detection is performed on each element to obtain the coordinate position of each element.
[0118] Based on any of the above embodiments, combining all feature representations of the same element into a query feature includes: For any element defined by the start and end markers in the structured sequence, extract the feature representations corresponding to the start marker, the end marker, and all word segments located between the start and end markers; All extracted feature representations are merged to generate query features for any element.
[0119] Based on any of the above embodiments, image features of the document to be analyzed are extracted, including: The image of the document to be analyzed is divided into multiple image blocks of predetermined size; Each image block is flattened into a vector, and the vectors corresponding to all image blocks are combined with the position encoding vector to obtain the initial embedding representation. The position encoding vector is used to characterize the position of each image block in the image of the document to be analyzed. The initial embedding representation is encoded to obtain the image features of the document to be analyzed.
[0120] Figure 4 This is a schematic diagram of the structure of the electronic device provided by the present invention, such as... Figure 4 As shown, the electronic device may include a processor 410, a communications interface 420, a memory 430, and a communication bus 440, wherein the processor 410, communications interface 420, and memory 430 communicate with each other via the communication bus 440. The processor 410 can call logical instructions in the memory 430 to execute a layout analysis method, which includes: extracting image features of the document to be analyzed and text features of layout analysis prompt text; using the image features and text features, generating a structured sequence containing the text content of each element within the document and the logical order between the elements; extracting feature representations corresponding to each element during the generation of the structured sequence; and performing target detection on each element based on the feature representations of each element to obtain the coordinate positions of each element.
[0121] Furthermore, the logical instructions in the aforementioned memory 430 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0122] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the layout analysis method provided by the above methods. The method includes: extracting image features of the document to be analyzed and text features of layout analysis prompt text; using the image features and the text features, generating a structured sequence containing text content of each element in the document and the logical order between the elements; extracting feature representations corresponding to each element during the generation of the structured sequence; and performing target detection on each element based on the feature representations of each element to obtain the coordinate position of each element.
[0123] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the layout analysis method provided by the above methods. The method includes: extracting image features of the document to be analyzed and text features of layout analysis prompt text; using the image features and the text features, generating a structured sequence containing text content of each element in the document and the logical order between the elements; extracting feature representations corresponding to each element during the generation of the structured sequence; and performing target detection on each element based on the feature representations of each element to obtain the coordinate position of each element.
[0124] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0125] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0126] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A layout analysis method, characterized in that, include: Extract image features from the document to be analyzed and text features from the layout analysis prompts; Using the image features and the text features, a structured sequence containing the text content of each element in the document and the logical order between the elements is generated; Extract the feature representations corresponding to each element during the generation of the structured sequence; Based on the feature representations of each element, target detection is performed on each element to obtain the coordinate positions of each element.
2. The layout analysis method according to claim 1, characterized in that, The step of generating a structured sequence containing text content and logical order between elements within a document, using the image features and text features, includes: The image features and text features are fused to generate a multimodal input representation; Autoregressive decoding is performed based on the multimodal input representation to map the visual layout order of each element provided by the image features to the logical order between elements, and to identify and generate the text content of each element, thereby obtaining the structured sequence.
3. The layout analysis method according to claim 2, characterized in that, The process of mapping the visual layout order of the elements provided by the image features to a logical order between elements, and identifying and generating the text content of each element to obtain the structured sequence, includes: During the process of generating the text content of each element in the logical order, predefined structured markers are embedded at the corresponding positions in the content stream to wrap the generated text content, thereby defining the hierarchy and scope of each element and obtaining the structured sequence.
4. The layout analysis method according to any one of claims 1 to 3, characterized in that, The extraction of feature representations corresponding to each element during the generation of the structured sequence includes: Extract the embedding vectors corresponding to each word segmentation during the structured sequence generation process; The embedding vectors corresponding to each word segment are input into a mapping layer consisting of a multilayer perceptron and an activation function for processing, generating a mapped feature vector. Based on the element inclusion relationship defined by the start and end markers in the structured sequence, all mapped feature vectors belonging to the same element are merged to obtain the feature representation of the corresponding element.
5. The layout analysis method according to any one of claims 1 to 3, characterized in that, The step of performing target detection on each of the elements based on their feature representations to obtain the coordinate positions of each element includes: Combine all the features of the same element into query features; Based on the query features of each element, target detection is performed on each element to obtain the coordinate position of each element.
6. The layout analysis method according to claim 5, characterized in that, The method of combining all feature representations of the same element into query features includes: For any element defined by the start marker and the end marker in the structured sequence, extract the feature representations corresponding to the start marker, the end marker, and all word segments located between the start marker and the end marker; All extracted feature representations are merged to generate the query feature for any of the elements.
7. The layout analysis method according to any one of claims 1 to 3, characterized in that, The extraction of image features from the document to be analyzed includes: The image of the document to be analyzed is divided into multiple image blocks of predetermined size; Each image block is flattened into a vector, and the vectors corresponding to all image blocks are combined with the position encoding vector to obtain an initial embedding representation. The position encoding vector is used to characterize the position of each image block in the image of the document to be analyzed. The initial embedding representation is encoded to obtain the image features of the document to be analyzed.
8. A layout analysis device, characterized in that, include: The first extraction unit is used to extract the image features of the document to be analyzed and the text features of the layout analysis prompt text; The element parsing unit is used to generate a structured sequence containing the text content of each element in the document and the logical order between the elements by utilizing the image features and the text features. The second extraction unit is used to extract the feature representations corresponding to each element during the generation of the structured sequence; The feature localization unit is used to perform target detection on each of the features based on the feature representation of each feature, and obtain the coordinate position of each feature.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the layout analysis method as described in any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the layout analysis method as described in any one of claims 1 to 7.
11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the layout analysis method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Layout analysis method and device, computer equipment and storage medium
CN113807218A
Text information generation method and device, model training method and device and electronic equipment
CN118587729A
Systems and methods for automatically extracting data from electronic documents including tables
US20110249905A1