Method for determining document reading sequence, electronic equipment and storage medium
By combining visual features of document images with text region location information and utilizing an encoder and decoder architecture, the reading order of text regions is predicted, solving the problems of insufficient accuracy and generalization ability in complex document layouts in existing technologies, and achieving higher accuracy in reading order detection.
Patent Information
- Application Number
- CN202510820148.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-11-14
AI Technical Summary
In existing technologies, methods for determining the reading order of documents have low accuracy, especially in complex documents where they lack generalization ability and cannot effectively handle complex layouts such as multi-column layouts and mixed text and images.
By extracting visual features from document images and positional information of text regions, and utilizing an encoder-decoder architecture that combines self-attention and cross-attention mechanisms, the reading order of text regions is predicted.
It improves the accuracy and generalization of document reading order, adapts to various complex document layouts, and enhances the user experience.
Smart Images

Figure CN120954034A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of document processing. More specifically, it relates to a method for determining the reading order of documents, an electronic device, a storage medium, and a computer program product. Background Technology
[0002] With the development of technology, automated document understanding has been widely applied in various fields, such as invoice recognition, document digitization, contract review, and financial auditing. Determining the document reading order is crucial for automated document understanding. The document reading order refers to the logical order in which the various text blocks in the document are read, and it forms the basis for downstream tasks such as document structure understanding, information extraction, and natural language processing.
[0003] However, related technologies typically determine the reading order of documents based on heuristic rules, textual information, and their positional information, which has low accuracy and cannot handle complex documents. Summary of the Invention
[0004] This application is made in consideration of the above-mentioned issues.
[0005] According to one aspect of this application, a method for determining the document reading order is provided, comprising:
[0006] Obtain the target document;
[0007] Extract the first visual features of the document image corresponding to the target document;
[0008] Obtain the first position information of each text region in the target document, wherein each text region is obtained through layout analysis of the target document; and
[0009] The reading order of each text region is determined based on the first visual features and the first positional information.
[0010] For example, extracting the first visual features of the document image corresponding to the target document includes: inputting the document image into an encoder, and using the encoder to extract the first visual features of the document image; determining the reading order of each text region based on the first visual features and first position information includes: for any text region, inputting the first position information of the text region into a feature encoding module, and using the feature encoding module to encode the first position information to obtain the corresponding first position feature; concatenating the first visual features with the first position features corresponding to each text region to obtain a concatenation result; inputting the concatenation result into a decoder, and using the decoder to determine the reading order of each text region based on the concatenation result.
[0011] For example, using a decoder, the reading order of each text region is determined based on the concatenation result, including: using a self-attention mechanism, obtaining a first attention feature based on the first position features of each first text region in the target text set, wherein the first text region is the text region whose reading order has been determined among the text regions; using a cross-attention mechanism, obtaining a second attention feature based on the first attention feature and the concatenation result; and based on the second attention feature, determining the next text region to be read after the target first text region, and adding the next text region to the target text set in reading order, wherein the reading order of each text region is obtained when the target text set includes all text regions in each text region, and the target first text region refers to the first text region in the target text set that is read last.
[0012] For example, the first location information is represented by the text bounding box corresponding to the text region. The first location information includes the coordinate information of the key points of the text bounding box. For any text region, the first location information of the text region is input into the feature encoding module, and the first location information is encoded by the feature encoding module to obtain the corresponding first location feature. This includes: for any text region, performing feature encoding on the coordinate information of the key points corresponding to the text region to obtain the coordinate feature of each coordinate information; and fusing the coordinate features of multiple coordinate information of the text region to obtain the first location feature of the text region.
[0013] For example, key points include the vertices of the text bounding box and / or the center point of the text bounding box.
[0014] For example, fusing the coordinate features of multiple coordinate information of the text region to obtain the first position feature of the text region includes: linearly combining the coordinate features of all coordinate information of the text region to obtain a combined vector; and fusing the combined vector with the identifier encoding of the identifier information of the text region to obtain the first position feature of the text region.
[0015] For example, the decoder is obtained through the following training operations: inputting the second visual features of the document image corresponding to the training document and the second positional features of each training text region in the training document into the initial decoder; determining the reading order of each training text region and the category of the text content in the training text region through the initial decoder, the category including at least one of title, body text, header, and footer; calculating the loss function value based on the determined reading order of the training text regions and the marked reading order of the training text regions, as well as the determined category of the text content in the training text regions and the marked category of the text content in the training text regions; adjusting the parameters of the initial decoder based on the loss function value until the preset conditions are met and training stops.
[0016] According to another aspect of this application, an electronic device is provided, comprising: a processor and a memory, wherein the memory stores computer program instructions, which, when executed by the processor, are used to perform the method described above for determining the document reading order.
[0017] According to another aspect of this application, a storage medium is provided on which program instructions are stored, which, when executed by a processor, are used to perform the method described above for determining the document reading order.
[0018] According to another aspect of this application, a computer program product is provided, including computer program instructions that, when executed by a processor, are used to perform the method described above for determining the document reading order.
[0019] The above technical solution determines the reading order of each text region by combining the first visual features of the target document with the first position information of each text region. The visual features of the target document contain important information for determining the reading order, such as font style, color, background, and lines. Combining the first visual features with the first position information of each text region can more accurately analyze the reading order of each text region, and has stronger generalization ability, enabling it to adapt to various complex documents.
[0020] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description
[0021] The above and other objects, features, and advantages of this application will become more apparent from the more detailed description of the embodiments of this application in conjunction with the accompanying drawings. The accompanying drawings are used to provide a further understanding of the embodiments of this application and form part of the specification. They are used together with the embodiments of this application to explain this application and do not constitute a limitation thereof.
[0022] Figure 1 A schematic flowchart illustrating a method for determining document reading order according to an embodiment of this application is shown;
[0023] Figure 2 A schematic flowchart illustrating the training operation of a decoder according to one embodiment of this application is shown;
[0024] Figure 3 A schematic diagram of the training operation of a decoder according to an embodiment of this application is shown;
[0025] Figure 4A schematic diagram illustrating the determination of the target document reading order according to an embodiment of this application is shown;
[0026] Figure 5 A schematic block diagram of an electronic device according to one embodiment of this application is shown. Detailed Implementation
[0027] It should be noted that the document data obtained in this application is accessed, collected, stored and used for subsequent analysis and processing after the user or relevant data owner has been clearly informed of the content of the data collection, the purpose of the data, the processing method and other information, and with the consent and authorization of the user or relevant data owner. Furthermore, the application can provide the user or relevant data owner with the means to access, correct or delete the data, as well as the method to revoke consent or authorization.
[0028] To make the objectives, technical solutions, and advantages of this application more apparent, exemplary embodiments according to this application will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are merely a part of the embodiments of this application.
[0029] Among related technologies, methods for determining document reading order include heuristic rule-based methods. These methods typically rely on manually defined rules, such as sorting according to top-down or left-to-right layout. While simple to implement and perform reasonably well in structurally sound documents, their accuracy drops significantly and their generalization ability is insufficient when faced with complex layouts, multi-column layouts, documents with interspersed illustrations, or documents with mixed structures, specifically documents with many tables, magazine-style multi-column layouts, or mixed text and images. Another type of related technology is model-based methods for determining document reading order. These methods use text information and its positional information as input data to the model, which then outputs the document reading order. This type of method avoids the limitation of relying on pre-defined rules, but its accuracy still needs improvement.
[0030] To at least address the aforementioned technical problems, this application provides a method for determining the reading order of a document. This detection method considers not only the positional information of text regions within the document but also the document's visual features, such as font style, color, background, and images. These visual features carry important clues to the document's reading order. By comprehensively considering both the positional information of text regions and visual features, the accuracy of the detected reading order can be significantly improved.
[0031] Figure 1 A schematic flowchart illustrating a method for determining document reading order according to an embodiment of this application is shown. Figure 1 As shown, the method for determining the document reading order includes the following steps S1100 to S1400.
[0032] In step S1100, the target document is obtained.
[0033] In step S1200, the first visual features of the document image corresponding to the target document are extracted.
[0034] In step S1300, the first position information of each text region in the target document is obtained, wherein each text region is obtained by performing layout analysis on the target document.
[0035] In step S1400, the reading order of each text region is determined based on the first visual feature and the first position information.
[0036] For example, in step S1100, the target document is obtained. The target document can be a document for which reading order detection is required. The format of the target document can include existing or future document formats such as DOC, DOCX, PDF, and TXT. The system can receive target documents sent by the user. The target document can be pre-stored, and pre-stored target documents can be retrieved.
[0037] In step S1200, the first visual features of the document image corresponding to the target document are extracted.
[0038] Visual features refer to the visual characteristics of a document image corresponding to a target document, used to describe, distinguish, and identify different objects. Visual features can include color features, texture features, edge features, keypoint features, etc., of the document image. Visual features can fully consider visually relevant features such as font style, color, background, lines, and images in the document image. The target document can be converted from a document format to an image format to obtain the corresponding document image.
[0039] The first visual features of a document image can be extracted using models based on convolutional neural networks, such as Residual Neural Networks (ResNet) and Convolutional Next (ConvNeXt). Alternatively, models with transformer structures, such as VisionTransformers (ViT) and Swing Transformers (a hierarchical visual self-attention model based on moving windows), can also be used. Any existing or future feature extraction method can be employed. Optionally, the document image can be preprocessed according to the format of the target document and the requirements of the method for extracting the first visual features before extraction. For example, preprocessing may include converting the document image to grayscale and normalizing the resulting grayscale values.
[0040] In step S1300, the first position information of each text region in the target document is obtained. Each text region is obtained through layout analysis of the target document. A text region can represent the area containing text content such as paragraphs, tables, headings, and headers in the target document. For example, a paragraph in the target document can be a text region. It can be understood that a target document can include multiple paragraphs, tables, headings, headers, and other text content. Each text content can have its corresponding text region. Therefore, a target document can include multiple text regions.
[0041] For example, layout analysis models can be used to determine text regions in a target document. These models can be deep learning models, such as convolutional recurrent neural networks (CRNNs) or transformer-based models. Alternatively, traditional image processing techniques can be used to perform layout analysis to determine text regions in a target document. For instance, the document image can first be thresholded, and then the thresholded segmentation results can be processed using connected component analysis and morphological operations. Finally, the text regions are determined based on the distances between connected components in the image. Any existing or future-developed technology can be used to determine text regions in a target document.
[0042] After identifying multiple text regions in the target document, the initial positional information of each text region can be determined based on their distribution within the document. For example, a coordinate system can be established based on the target document to determine the initial positional information of each text region. This initial positional information might include the coordinates of each vertex of the text region. If the text region is rectangular, the initial positional information could include the coordinates of one vertex and the length and width of the text region. Alternatively, the initial positional information could be the coordinates of any two opposite vertices.
[0043] In step S1400, the reading order of each text region is determined based on the first visual feature and the first position information.
[0044] The reading order described above can be represented by the sequential arrangement of multiple text regions. For example, if a target document has three text regions, the reading order is text region 3, text region 1, and text region 2. That is, first read the text content in text region 3, then read the text content in text region 1, and then read the text content in text region 2. The reading order can be determined by combining the first visual features of the target document and the first position information of each of the multiple text regions. Any model used to detect the order of text regions can be used to fuse the first visual features and the first position information to determine the reading order. The above model can be based on a multimodal feature fusion framework. For example, a model used to implement the Next Token Prediction task can be used, taking the first visual features and the first position information of each text region as input data, to predict the order of the text regions and thus determine the reading order of each text region.
[0045] The above technical solution determines the reading order of each text region by combining the first visual features of the target document with the first position information of each text region. The visual features of the target document contain important information for determining the reading order, such as font style, color, background, and lines. Combining the first visual features with the first position information of each text region can more accurately analyze the reading order of each text region, and has stronger generalization ability, enabling it to adapt to various complex documents.
[0046] For example, step S1200 extracts the first visual features of the document image corresponding to the target document, including step S1210. In step S1210, the document image is input into an encoder, and the encoder is used to extract the first visual features of the document image. The encoder can be used to extract global and / or local visual features as the first visual features from the document image corresponding to the target document. The encoder can be implemented using a residual neural network (ResNet) based on a convolutional neural network or a next-generation convolutional model (ConvNeXt). Preferably, the encoder can also be implemented using a model based on a transformer structure, such as a visual transformer (ViT) or a hierarchical visual self-attention model based on a moving window (Swin Transformer), which can extract richer multi-scale features.
[0047] Step S1400 determines the reading order of each text region based on the first visual features and the first position information, including steps S1410, S1420, and S1430. In step S1410, for any text region, the first position information of the text region is input into the feature encoding module, which encodes the first position information to obtain the corresponding first position feature. The feature encoding module encodes the first position information of the text region into a numerical form of the first position feature. Specifically, the feature encoding module can perform an embedding operation on the first position information to vectorize it and obtain the first position feature. The feature encoding module can encode the first position information into a form that the decoder can decode. The first position feature facilitates the decoder in capturing the relationship between the position of the text region and the reading order, enhancing the decoder's ability to interpret the influence of the text region's position on the reading order.
[0048] In step S1420, the first visual feature is concatenated with the first positional feature corresponding to each text region to obtain the concatenation result.
[0049] The dimensions of the first positional features of all text regions can be the same, while the dimensions of the first visual features can be different from those of the first positional features. The dimensions of the first visual features can be processed to be the same as those of the first positional features before concatenating the processed first visual features with the first positional features. For example, if the dimension of the first visual feature is higher than that of the first positional feature, a dimensionality reduction operation can be performed on the first visual feature; conversely, if the dimension of the first visual feature is lower than that of the first positional feature, a dimensionality increase operation can be performed on the first visual feature. For instance, if the first visual feature is n*256*1024 and the first positional feature is 256*1024, the first visual feature can be reduced from n*256*1024 to 256*1024. In some embodiments, after processing the first visual features and the first positional features to the same dimension, the first positional features of all text regions can be concatenated in any order to form a sequence. In other embodiments, after processing the first visual feature and the first positional feature to the same dimension, the first positional features of all text regions can be concatenated in a specific order to form a sequence, such as from top to bottom, from left to right, etc. The first visual feature can be concatenated before or after the sequence composed of multiple first positional features. The concatenation result is input into the decoder to determine the reading order of the text content in the multiple text regions.
[0050] In step S1430, the splicing result is input into the decoder, and the reading order of each text region is determined based on the splicing result using the decoder.
[0051] In the above scheme, the method for determining the document reading order can be implemented based on an encoder-decoder architecture. It can be understood that in step S1210, the encoder extracts feature vectors, i.e., first visual features, from the text image. These first visual features can be in a form that the decoder can decode. Furthermore, in step S1410, the feature encoding module converts the first position information of each text region into feature vectors, i.e., first position features. The decoder can be implemented using models such as Transformer-Decoder or Mamba. The decoder can be used to predict the next text region based on the first position features and first visual features of multiple text regions in the target document, i.e., to sort multiple text regions and determine the reading order of each text region. For example, in the concatenation result, the sequence composed of first position features can be used as a query, and the first visual features can be used as keys and values. Inputting the concatenation result into the decoder allows it to predict the order of the corresponding multiple text regions in the concatenation result, which serves as the reading order of each text region.
[0052] For example, a model for determining the document reading order may include an encoder, a feature encoding module, and a decoder. The target document may contain first positional information of each text region. In other words, the first positional information can be acquired simultaneously when acquiring the target document. The image corresponding to the target document can be input into the encoder to obtain first visual features, and the first positional information of each text region of the target document can be input into the feature encoding module to obtain corresponding first positional features. The first visual features and the first positional features of each text region can be concatenated into a sequence and input into the decoder to obtain the reading order of each text region.
[0053] The above technical solution utilizes an encoder to determine the first visual features of the document image corresponding to the target document, a feature encoding module to determine the first positional features of the text regions, and a decoder to determine the reading order of each text region based on the first positional features and the first visual features. Therefore, by adopting a decoder-encoder architecture, visual features and positional features are fused, and reading order detection is constructed as a text region order prediction task, improving the accuracy of reading order detection. This approach can be effectively applied to reading order detection in documents with various complex layouts, thereby enhancing the user experience.
[0054] For example, step S1430 inputs the splicing result into the decoder, and uses the decoder to determine the reading order of each text region based on the splicing result, including steps S1431, S1432 and S1433.
[0055] In step S1431, a first attention feature is obtained based on the first position features of each first text region in the target text set using a self-attention mechanism. Here, the first text region is the text region whose reading order has been determined among the various text regions.
[0056] In one specific embodiment, the decoder can employ a Transformer architecture. Through the synergistic effect of self-attention and cross-attention mechanisms, it dynamically predicts the reading order of multiple text regions. The decoder generates an ordered sequence iteratively, predicting the next most likely text region to be read (i.e., the next reading position) in each iteration step.
[0057] The self-attention mechanism is used to enable the decoder to dynamically focus on each first text region in the target text set when predicting the next text region among multiple text regions. This first text region is the text region whose reading order has been determined among the various text regions. The target text set can be a set of text regions whose reading order has been determined. Text regions with a determined reading order can be stored in the target text set according to their reading order. In other words, the arrangement order of text regions in the target text set can correspond to the reading order. For example, when determining the i-th text region, the self-attention mechanism causes the decoder to consider the first positional features of the previous (i-1) text regions whose reading order has been determined, thus obtaining a first attention feature. It can be understood that the reading order of these previous (i-1) text regions in the text region has been determined, and they are earlier in the reading order. The first attention feature considers the first positional features of the text regions whose reading order has been determined; therefore, when determining the i-th text region, the first positional features of the text regions whose reading order has been determined will also be used as a reference. For the case where the first text region needs to be determined, i.e., when none of the text regions have yet had their reading order determined, the first attention feature can be obtained based on the initial features. The initial feature can be any custom character, for example <bos> 、 <eos>A custom character can be pre-stored in the target text set.
[0058] Specifically, when the i-th text region is determined, a self-attention mechanism can be used to calculate the relevance weights of the first positional features of the preceding (i-1) text regions to obtain the first attention feature. For multi-column documents, a self-attention mechanism can be used to associate text regions that are far apart but whose text content is logically continuous, such as the text region at the bottom of one column and the text region at the top of the next column. This avoids the errors caused by the fixed scanning direction in traditional heuristic-based prediction methods.
[0059] In step S1432, a second attention feature is obtained based on the first attention feature and the splicing result using a cross-attention mechanism.
[0060] Cross-attention mechanisms enable the decoder to dynamically associate the first visual feature and the first positional features of each text region when predicting the next text region from multiple text regions, based on the first attention feature. For example, the first attention feature can be used as a query, the first visual feature as a key and value, and the first positional feature as a key and value to obtain the second attention feature. Thus, the second attention feature considers both the positional features of text regions whose reading order has been determined and integrates visual and positional features. For instance, when determining the j-th text region, if two text regions are close to the (j-1)-th text region, but one of the text regions has a larger font, the decoder tends to identify the text region with the larger font as the j-th text region. This allows for effective interaction between multiple first positional features and first visual features of text regions to determine the reading order.
[0061] In step S1433, based on the second attention feature, the next text region to be read after the target first text region is determined, and the next text region is added to the target text set in reading order. When the target text set includes all text regions from each text region, the reading order of each text region is obtained. The target first text region refers to the first text region in the target text set that is read last. For example, based on the second attention feature, the probability of each text region being the next text region can be calculated, and the text region with the higher probability can be determined as the next text region to be read after the target first text region. It can be understood that text regions in the target text set can be stored in the order of reading. The text region read first in the target text set can be placed at the beginning, and the text region read later can be placed at the end. The target first text region can be the last text region in the target text set, that is, the last text region read among the text regions whose reading order has been determined. For example, if a target document contains n text regions, based on the second focus feature, the next text region after the (k-1)th text region is determined, i.e., the kth text region is determined. This kth text region can then be added to the target text set according to the reading order, thus determining the reading order of the k text regions. If k equals n, meaning the target text set includes all text regions from each of the previous text regions, then the reading order of multiple text regions is obtained. If k is less than n, meaning there are still text regions whose reading order has not been determined, it is necessary to determine the (k+1)th text region based on the existing k text regions, and so on, until the reading order of all text regions is determined.
[0062] The above technical solution utilizes self-attention and cross-attention mechanisms in the decoder to determine the reading order based on both first positional features and first visual features. Therefore, the self-attention mechanism ensures a more continuous reading order for multiple text regions determined by the decoder, aiding in determining the initial reading order. The cross-attention mechanism guides the decoder to predict the next text region based on the first visual and first positional features extracted by the encoder. Thus, this technical solution can more effectively predict the reading order of irregularly laid-out documents, such as documents with mixed text and images, documents with multiple columns, and documents with nested tables, based on positional and visual features. The dual-attention mechanism effectively overcomes the limitations of existing technologies in cross-regional transitions, improving the accuracy of the reading order.
[0063] For example, the first positional information is represented by a text bounding box corresponding to the text region, and the first positional information includes the coordinate information of the key points of the text bounding box. The text region can be obtained by performing layout analysis on the target document. When performing layout analysis on the text region, a text bounding box can be generated at the location of the text region in the target document, and the text content in the text region can be enclosed by the text bounding box. The shape of the text bounding box can be determined according to the text region. For example, the text bounding box can be a polygon, and the length and number of the sides of the polygon can be determined according to the text region. Optionally, the text bounding box can be a quadrilateral, i.e., a rectangle. For each text bounding box, its positional information includes the coordinate information of its key points. A coordinate system can be established on the target document to determine the coordinate information. For example, if the text bounding box can be a rectangle, the first positional information can be the coordinate information of at least any point in the text bounding box. Optionally, the first positional information can include the coordinate information of the four vertices of the rectangle. Preferably, in order to reduce the complexity of the first positional information, the first positional information can include the maximum and minimum values of the horizontal and vertical coordinates of the horizontal and vertical coordinates of the multiple vertices of the text bounding box. This can both determine the location of the text region in the target document and reduce the complexity of the first location information.
[0064] For example, key points include the vertices of the text bounding box and / or the center point of the text bounding box. The first positional information includes the coordinates of the vertices of the text bounding box; for example, if the text bounding box is a rectangle, the first positional information includes the coordinates of the four vertices of the rectangle. The first positional information includes the coordinates of the center point of the text bounding box; for example, the coordinates of the center point can be determined by the average of the horizontal and vertical coordinates within the text bounding box. The first positional information may include the coordinates of the vertices and the center point of the text bounding box.
[0065] In the above technical solution, key points include the vertices of the text bounding box. Therefore, the first positional information describes both the position of the text bounding box in the target document and its outline features. Key points also include the center point of the text bounding box. Thus, the first positional information can also concisely describe the distribution of the text bounding box in the target document. Since key points include both the vertices and the center point of the text bounding box, it can describe both the outline features and the distribution of the text bounding box in the target document.
[0066] For the coordinate information of key points in each text region, steps S1411 and S1412 included in step S1410 can be executed.
[0067] In step S1411, the coordinate information of the key points corresponding to the text region is feature encoded to obtain the coordinate features of each coordinate information.
[0068] In the example where the coordinate information of key points in the text region includes the maximum and minimum values of the horizontal and vertical axes, the maximum, minimum, maximum, and minimum values of the horizontal and vertical axes constitute four coordinate information points. These four coordinate information points can be feature-encoded separately to obtain the coordinate features of each. Feature encoding can be implemented using methods such as direct normalization, fully connected layers, and multilayer perceptrons. For example, the maximum, minimum, maximum, and minimum values of the horizontal and vertical axes can be input into their respective independent multilayer perceptrons to perform feature encoding for each coordinate information point and obtain the corresponding coordinate features. Independent multilayer perceptrons do not need to share parameters. Optionally, for coordinate information from different text regions, coordinate information of the same type can be input into the same multilayer perceptron to obtain the corresponding coordinate features. For example, the maximum values of multiple horizontal axis coordinates of a text region can be input to multilayer perceptron 1, the minimum values of the horizontal axis coordinates can be input to multilayer perceptron 2, the maximum values of the vertical axis coordinates can be input to multilayer perceptron 3, and the minimum values of the vertical axis coordinates can be input to multilayer perceptron 4, thereby obtaining their respective coordinate features.
[0069] In the example where the coordinate information of the key points in the text area includes the horizontal and vertical coordinates of the key points, feature encoding can be performed on the horizontal and vertical coordinates of each key point to obtain the coordinate features of the coordinate information of each key point.
[0070] In step S1412, the coordinate features of multiple coordinate information of the text region are fused to obtain the first position feature of the text region.
[0071] For the text region, after obtaining the coordinate features corresponding to each coordinate information of the text region, all coordinate features of the text region can be fused to obtain the first position feature of the text region. In some embodiments, multiple coordinate features can be concatenated in a preset order to serve as the first position feature. In other embodiments, multiple coordinate features can be summed, or the multiple coordinate features can be averaged, weighted, or otherwise fused to obtain the first position feature of the text region.
[0072] The above technical solution encodes the coordinate information of key points in a text region to obtain the coordinate features of each coordinate information, and then fuses all the coordinate features of the text region to determine the first position feature of the text region. Therefore, the first position feature, which incorporates the features of each coordinate information of the text region, can fully represent the position of the text region in the target document, thereby improving the accuracy of reading order detection.
[0073] For example, step S1412 fuses the coordinate features of multiple coordinate information of the text region to obtain the first position feature of the text region, including steps S1412A and S1412B.
[0074] In step S1412A, the coordinate features of all coordinate information in the text region are linearly combined to obtain a combined vector. Coordinate features can be represented by vectors, and the dimensions of the vectors representing coordinate features can be the same. In some embodiments, vectors representing different coordinate information can be directly added together to obtain a combined vector. In other embodiments, vectors representing different coordinate information can be weighted and summed to obtain a combined vector.
[0075] In step S1412B, the combined vector is fused with the identifier code of the text region's identifier information to obtain the first positional feature of the text region. The identifier code of the identifier information can be used to identify the position of the first positional feature of the text region in the concatenation result. The decoder's operation on the first positional feature can be undirected, so each summed vector can be fused with its corresponding identifier code to obtain the first positional feature. Thus, the decoder can determine the relationship between different first positional features.
[0076] The identifier encoding can be determined based on the number of text regions in the target document. In other words, each text region has a different identifier encoding. For example, if the target document contains three text regions, then the identifier encodings for these three text regions can be 1, 2, and 3, respectively. Optionally, the identifier encoding can be one-dimensional, for example, implemented using integer form, values within the range of 0 to 1, etc. Alternatively, the identifier encoding can also be multi-dimensional, for example, randomly generating multiple different identifier encodings, and the dimension of the identifier encoding can be the same as the dimension of the combined vector mentioned above.
[0077] The fusion operation between the combined vector and the identifier encoding can be an addition operation, a concatenation operation, a gated feature fusion operation, etc.
[0078] The above technical solution linearly combines the coordinate features of all coordinate information in the text region to obtain a combined vector, and then fuses the combined vector with the corresponding identifier encoding to obtain the first positional feature of the text region. Thus, the first positional feature includes both the coordinate features of each coordinate information (absolute positional information) and the identifier encoding, which introduces more complex spatial relationships (relative or global positional information). Therefore, the first positional feature is more representative of the text region, enhancing the decoder's understanding of the layout of text regions in the target document and helping it output a more accurate reading order of the text regions.
[0079] For example, before step S1410, which inputs the first position information of any text region into the feature encoding module and encodes the first position information using the feature encoding module to obtain the corresponding first position feature, the method for determining the document reading order further includes step S1500. Step S1500 can be executed before step S1410.
[0080] In step S1500, the coordinate information of each text region is integerized. The coordinate information of the text region can be in floating-point form, which may lead to inaccurate calculations when determining the first positional feature. Therefore, the coordinate information can be integerized to improve the accuracy of subsequent data processing used to obtain the first positional feature. Methods such as direct integerization, rounding up, rounding down, and rounding can be used to integerize each coordinate information. In some embodiments, the first positional feature can be determined for a text region after all its coordinate information has been integerized. In other embodiments, the first positional feature can be determined for each of all text regions separately after all text regions have had their coordinate information integerized.
[0081] The above technical solution, before encoding the first position information using the feature encoding module to obtain the corresponding first position feature, performs integer processing on multiple coordinate information of the text region. Therefore, the integerized coordinate information can be accurately calculated, thereby improving the accuracy of the first position feature and consequently improving the accuracy of reading order detection.
[0082] For example, the decoder described above can be obtained through a training operation. Figure 2 A schematic flowchart illustrating the training operation of a decoder according to one embodiment of this application is shown. Figure 2 As shown, the training operation may include steps S2100 and S2200.
[0083] In step S2100, the second visual features of the document image corresponding to the training document and the second positional features of each training text region in the training document are input into the initial decoder. The initial decoder determines the reading order of each training text region and the category of the text content in the training text region. The category includes at least one of title, body text, header, and footer.
[0084] The training document can be processed using the same processing method as the target document to obtain the second visual features of the document image corresponding to the training document and the second positional features of each training text region. According to embodiments of this application, the initial decoder may include a first module for determining the reading order of text content in the training text regions based on the second visual features and the second positional features of each training text region. The first module is the model used in step S1430 above to determine the reading order of multiple text regions, which can be implemented using models such as Transformer-Decoder or Mamba. In addition to the first module, the decoder may also include a second module for determining the category of text content in the training text regions based on the second positional features of each training text region. Exemplarily, the category of the text region includes at least one of title, body text, header, and footer. The second module can be implemented using any classifier. In some embodiments, after determining the reading order of the training text regions using the first module of the initial decoder, the second module can be used to classify each training text region based on the second positional features of each training text region to determine the category of text content in the training text. In other embodiments, determining the reading order of the training text regions and determining the category of text content in the training text regions can be performed simultaneously.
[0085] In step S2200, based on the determined reading order of the training text region and the marked reading order of the training text region, as well as the determined categories of the text content in the training text region and the marked categories of the text content in the training text region, the loss function value is calculated, and the parameters of the initial decoder are adjusted based on the loss function value until the preset conditions are met and training stops.
[0086] In some embodiments, manual annotation can be used to determine the reading order and category of the training text regions. In other embodiments, a layout analysis model can be used to automatically annotate the categories of multiple training text regions in the training document. In still other embodiments, a heuristic method can be used to initially determine the reading order of the training text regions, and then the reading order can be manually reviewed and modified to determine the final reading order of the training text regions. The second visual features and the second positional features of the training text regions can be input into the initial decoder. Based on the difference between the reading order of the training text regions output by the initial decoder and the reading order of the training text regions marked by the training text regions, as well as the difference between the category of the text content in the training text regions output by the decoder and the category of the marking, a loss function value is calculated. The parameters of the initial decoder are adjusted based on the loss function value to train the decoder. For example, the loss function can be a cross-entropy loss function, etc. The preset condition can include the loss function value being less than the loss threshold. When the calculated loss function is less than the loss threshold, the decoder used in step S1430 above can be obtained. The preset condition can also include the number of training iterations of the initial decoder being equal to the number of iterations threshold. When the number of training iterations is equal to the number of iterations threshold, the training can be terminated. Optionally, if the loss function value is greater than the loss function threshold when the number of training iterations equals the threshold, the decoder can be retrained.
[0087] Figure 3 A schematic diagram illustrating the training operation of a decoder according to one embodiment of this application is shown. Figure 3 As shown, the document image corresponding to the training document is input into the encoder to obtain the second visual features. The positional information of each training text region of the training document is input into the feature encoding module to obtain the second positional features. The second positional features and the second visual features are concatenated and then input into the decoder, which outputs the reading order and category in the training text region. Finally, the decoder parameters can be adjusted based on the decoder's output and the labeled reading order and category of the training text region to make its output as close as possible to the labeled data of the training text region, thus completing the training operation.
[0088] The above technical solution trains a decoder based on the determined reading order and labeled reading order of the training text regions, as well as the determined categories and labeled categories of the training text regions. Thus, the labeled categories of the training text regions serve as auxiliary supervision signals for decoder training to determine document reading order, improving the decoder's understanding of structural semantics and enhancing the accuracy of reading order detection.
[0089] Preferably, when performing reading order detection on the target document, the second module used to determine the category of text regions can be removed from the decoder. That is, the decoder does not need to determine the category of the text content in the text region, but only the reading order of the text region. This reduces processing time and improves efficiency. Figure 4 A schematic diagram illustrating the determination of the target document reading order according to one embodiment of this application is shown. Figure 4 As shown, the document image corresponding to the target document is input into the encoder to obtain the first visual features. The position information of each text region of the target document is input into the feature encoding module to obtain the first position features. The first position features and the first visual features are concatenated and then input into the decoder to obtain the reading order of the text regions. Figure 3 and Figure 4 As shown, when the decoder is applied to the target document, the category of the text region no longer needs to be output.
[0090] By way of example, according to another aspect of this application, an apparatus for determining a document reading order is also provided. The apparatus includes a processor. The processor is configured to perform the method for determining a document reading order according to any of the above embodiments.
[0091] By way of example, according to another aspect of this application, an electronic device is also provided. Figure 5 A schematic block diagram of an electronic device 500 according to an embodiment of this application is shown. The electronic device 500 includes a processor 510 and a memory 520. The memory 520 stores computer program instructions that, when executed by the processor 510, are used to perform the method described above for determining the document reading order.
[0092] By way of example, according to another aspect of this application, a storage medium is also provided, on which program instructions are stored, which, when executed, are used to perform the method described above for determining the document reading order. The storage medium may, for example, include an erasable programmable read-only memory (EPROM), a portable read-only memory (CD-ROM), a USB memory, or any combination of the above storage media. The storage medium may be any combination of one or more computer-readable storage media.
[0093] By way of example, according to another aspect of this application, a computer program product is also provided, including computer program instructions that, when run, are used to perform the method described above for determining the document reading order.
[0094] Those skilled in the art can understand the specific implementation schemes and beneficial effects of the above-described methods for determining document reading order by reading the relevant descriptions. For the sake of brevity, they will not be repeated here.
[0095] Although exemplary embodiments have been described herein with reference to the accompanying drawings, it should be understood that the above exemplary embodiments are merely illustrative and are not intended to limit the scope of this application. Various changes and modifications can be made therein by those skilled in the art without departing from the scope and spirit of this application. All such changes and modifications are intended to be included within the scope of this application as claimed in the appended claims.
[0096] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0097] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed.
[0098] Numerous specific details are set forth in the specification provided herein. However, it will be understood that embodiments of this application may be practiced without these specific details. In some instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of this specification.
[0099] Similarly, it should be understood that, in order to simplify this application and aid in understanding one or more aspects of the application, various features of this application may sometimes be grouped together in a single embodiment, figure, or description thereof in the description of exemplary embodiments of this application. However, this approach should not be construed as reflecting an intention that the claimed application requires more features than are expressly recited in each claim. Rather, as reflected in the corresponding claims, the point of application is that the corresponding technical problem can be solved with fewer features than all of a single disclosed embodiment. Therefore, the claims following the detailed description are hereby expressly incorporated into that detailed description, wherein each claim itself is a separate embodiment of this application.
[0100] Those skilled in the art will understand that, apart from the mutual exclusion of features, all features disclosed in this specification (including the accompanying claims, abstract, and drawings) and all processes or units of any method or apparatus so disclosed can be combined in any combination. Unless otherwise expressly stated, each feature disclosed in this specification (including the accompanying claims, abstract, and drawings) may be replaced by an alternative feature that serves the same, equivalent, or similar purpose.
[0101] Furthermore, those skilled in the art will understand that although some embodiments described herein include certain features but not others included in other embodiments, combinations of features from different embodiments are intended to be within the scope of this application and form different embodiments. For example, in the claims, any one of the claimed embodiments can be used in any combination.
[0102] The various component embodiments of this application can be implemented in hardware, or as software modules running on one or more processors, or a combination thereof. Those skilled in the art will understand that microprocessors or digital signal processors (DSPs) can be used in practice to implement some or all of the functions of some modules in the apparatus for determining the document reading order according to embodiments of this application. This application can also be implemented as an apparatus program (e.g., a computer program and computer program product) for performing part or all of the methods described herein. Such an implementation of this application can be stored on a computer-readable medium, or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, provided on a carrier signal, or provided in any other form.
[0103] It should be noted that the above embodiments are illustrative of this application and not limiting of it, and that those skilled in the art can devise alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses should not be construed as limiting the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. This application can be implemented by means of hardware comprising several different elements and by means of a suitably programmed computer. In the unit claims enumerating several means, several of these means may be embodied by the same item of hardware. The use of the words first, second, and third, etc., does not indicate any order. These words can be interpreted as names.
[0104] The above description is merely a specific embodiment or illustration of the embodiments of this application. The scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. The scope of protection of this application shall be determined by the scope of the claims.< / eos> < / bos>
Claims
1. A method for determining the document reading order, characterized in that, include: Obtain the target document; Extract the first visual features of the document image corresponding to the target document; Obtain the first position information of each text region in the target document, wherein each text region is obtained by performing layout analysis on the target document; as well as The reading order of each text region is determined based on the first visual feature and the first location information.
2. The method for determining document reading order according to claim 1, characterized in that, The step of extracting the first visual features of the document image corresponding to the target document includes: The document image is input into the encoder, and the encoder is used to extract the first visual features of the document image; Determining the reading order of each text region based on the first visual feature and the first location information includes: For any text region, the first position information of the text region is input into the feature encoding module, and the first position information is encoded by the feature encoding module to obtain the corresponding first position feature; The first visual feature is concatenated with the first positional feature corresponding to each text region to obtain the concatenation result; The splicing result is input into the decoder, and the decoder is used to determine the reading order of each text region based on the splicing result.
3. The method for determining document reading order according to claim 2, characterized in that, The step of using the decoder to determine the reading order of each text region based on the splicing result includes: Using a self-attention mechanism, a first attention feature is obtained based on the first position features of each first text region in the target text set, wherein the first text region is the text region whose reading order has been determined among the text regions. Using a cross-attention mechanism, a second attention feature is obtained based on the first attention feature and the concatenation result; Based on the second attention feature, the next text region to be read after the target first text region is determined, and the next text region is added to the target text set in reading order. When the target text set includes all text regions of each text region, the reading order of each text region is obtained. The target first text region refers to the first text region in the target text set that is read last.
4. The method for determining document reading order according to claim 2 or 3, characterized in that, The first location information is represented by a text bounding box corresponding to the text region. The first location information includes the coordinate information of key points of the text bounding box. For any text region, the first location information of the text region is input into a feature encoding module, and the first location information is encoded using the feature encoding module to obtain the corresponding first location feature, including: For any text region, The coordinate information of the key points corresponding to the text region is feature-encoded to obtain the coordinate features of each coordinate information; The coordinate features of multiple coordinate information of the text region are fused to obtain the first position feature of the text region.
5. The method for determining document reading order according to claim 4, characterized in that, The key points include the vertices of the text bounding box and / or the center point of the text bounding box.
6. The method for determining document reading order according to claim 4, characterized in that, The step of fusing the coordinate features of multiple coordinate information of the text region to obtain the first position feature of the text region includes: The coordinate features of all coordinate information in the text region are linearly combined to obtain a combined vector. The combined vector is fused with the identifier encoding of the identifier information of the text region to obtain the first positional feature of the text region.
7. The method for determining document reading order according to claim 2 or 3, characterized in that, The decoder is obtained through the following training operation: The second visual features of the document image corresponding to the training document and the second positional features of each training text region in the training document are input into the initial decoder. The initial decoder determines the reading order of each training text region and the category of the text content in the training text region. The category includes at least one of title, body text, header, and footer. Based on the determined reading order of the training text region and the marked reading order of the training text region, as well as the determined categories of the text content in the training text region and the marked categories of the text content in the training text region, a loss function value is calculated. The parameters of the initial decoder are adjusted based on the loss function value until a preset condition is met and training stops.
8. An electronic device, comprising: Processor and memory, characterized in that, The memory stores computer program instructions, which, when executed by the processor, are used to perform the method for determining the document reading order as described in any one of claims 1 to 7.
9. A storage medium on which program instructions are stored, characterized in that, The program instructions, when executed by the processor, are used to perform the method for determining the document reading order as described in any one of claims 1 to 7.
10. A computer program product comprising computer program instructions, characterized in that, The computer program instructions, when executed by a processor, are used to perform the method for determining the document reading order as described in any one of claims 1 to 7.
Citation Information
Cited By
Recognition method for reading sequence of document layout
CN121686495A