Reading Order Prediction Method, Training Method and Device for Reading Order Prediction Model

By extracting detection box and positional encoding features from text images, the method improves the accuracy of predicting reading order in diverse document layouts using deep learning, addressing the limitations of existing technologies.

CN116092090BActive Publication Date: 2025-07-15BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310135899.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-10
Publication Date
2025-07-15
Estimated Expiration
2043-02-10

AI Technical Summary

Technical Problem

In the prior art, in the prediction of document reading order, methods based on heuristic rules are poor in generalization, methods based on deep learning lack global information modeling, and it is difficult to effectively distinguish the connection relationship between text lines, resulting in insufficient prediction accuracy.

Method used

By combining the visual features, detection frame features and position encoding features of text images, multi-dimensional feature extraction and encoding methods are used to determine the target encoding features of text elements, and the reading order between network predicted text elements is determined through cascading features and reading order.

Benefits of technology

It improves the prediction accuracy of document reading order, can better handle the document logic structure in e-commerce network pictures, and enhances the modeling ability of global information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116092090B_ABST
    Figure CN116092090B_ABST
Patent Text Reader

Abstract

The present disclosure provides a reading order prediction method, a training method and device for a reading order prediction model, which relates to the field of artificial intelligence technology, specifically to the fields of deep learning, image processing, and computer vision technology, and can be applied to scenarios such as OCR. The specific implementation solution is as follows: according to the visual features of a text image, determine the detection box features of at least two text elements in the text image; according to the visual features of the text image, the detection box features of the text elements, and the position encoding features, determine the target encoding features of the text elements; according to the target encoding features of the text elements, determine the reading order between different text elements. Through the above technical solution, the reading order between text elements in a document can be accurately predicted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technologies, specifically to the fields of deep learning, image processing, and computer vision technologies, and can be applied to scenarios such as OCR. Background Art

[0002] With the development of artificial intelligence technologies, for the document reading prediction task, it is required to output the detected and recognized document content in the reading order to obtain text content that can be directly read. Therefore, how to achieve end-to-end prediction of the document reading order is crucial. Summary of the Invention

[0003] The present disclosure provides a method for predicting the reading order, a method for training a reading order prediction model, and an apparatus.

[0004] According to one aspect of the present disclosure, there is provided a method for predicting the reading order, the method comprising:

[0005] Determining detection box features of at least two text elements in the text image according to visual features of the text image;

[0006] Determining target encoding features of the text elements according to the visual features of the text image, the detection box features of the text elements, and position encoding features;

[0007] Determining the reading order between different text elements according to the target encoding features of the text elements.

[0008] According to another aspect of the present disclosure, there is provided a method for training a reading order prediction model, the method comprising:

[0009] Determining detection box features of at least two text elements in the text image according to visual features of the text image;

[0010] Determining target encoding features of the text elements according to the visual features of the text image, the detection box features of the text elements, and position encoding features;

[0011] Determining the reading order between different text elements according to the target encoding features of the text elements;

[0012] Training a reading order prediction model according to the reading order of the text elements and label data.

[0013] According to another aspect of the present disclosure, there is provided a reading order prediction apparatus, the apparatus comprising:

[0014] A detection box feature determination module, configured to determine detection box features of at least two text elements in the text image according to visual features of the text image;

[0015] A target encoding feature determination module, configured to determine the target encoding feature of the text element according to the visual feature of the text image, the detection box feature of the text element, and the position encoding feature;

[0016] A reading order determination module, configured to determine the reading order between different text elements according to the target encoding feature of the text element.

[0017] According to another aspect of the present disclosure, there is provided a training device for a reading order prediction model, the device includes:

[0018] A detection box feature determination module, configured to determine the detection box features of at least two text elements in the text image according to the visual feature of the text image;

[0019] A target encoding feature determination module, configured to determine the target encoding feature of the text element according to the visual feature of the text image, the detection box feature of the text element, and the position encoding feature;

[0020] A reading order determination module, configured to determine the reading order between different text elements according to the target encoding feature of the text element;

[0021] A model training module, configured to train the reading order prediction model according to the reading order of the text element and the label data.

[0022] According to another aspect of the present disclosure, there is provided an electronic device, the electronic device includes:

[0023] At least one processor; and

[0024] A memory communicatively connected to the at least one processor; wherein,

[0025] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the reading order prediction method provided in any embodiment of the present disclosure, or the training method of the reading order prediction model.

[0026] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the reading order prediction method provided in any embodiment of the present disclosure, or the training method of the reading order prediction model.

[0027] According to another aspect of the present disclosure, there is provided a computer program product, including a computer program, where the computer program, when executed by a processor, implements the reading order prediction method provided in any embodiment of the present disclosure, or the training method of the reading order prediction model.

[0028] According to the technology of the present disclosure, the prediction accuracy of the document reading order can be improved.

[0029] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. Description of the Drawings

[0030] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0031] Figure 1 is a flowchart of a method for predicting the reading order provided according to an embodiment of the present disclosure;

[0032] Figure 2 is a flowchart of another method for predicting the reading order provided according to an embodiment of the present disclosure;

[0033] Figure 3 is a flowchart of yet another method for predicting the reading order provided according to an embodiment of the present disclosure;

[0034] Figure 4A is a flowchart of still another method for predicting the reading order provided according to an embodiment of the present disclosure;

[0035] Figure 4B is a schematic diagram of a process for predicting the reading order provided according to an embodiment of the present disclosure;

[0036] Figure 5 is a flowchart of a method for training a reading order prediction model provided according to an embodiment of the present disclosure;

[0037] Figure 6 is a schematic structural diagram of a device for predicting the reading order provided according to an embodiment of the present disclosure;

[0038] Figure 7 is a schematic structural diagram of a device for training a reading order prediction model provided according to an embodiment of the present disclosure;

[0039] Figure 8 is a block diagram of an electronic device for implementing the method for predicting the reading order or the method for training the reading order prediction model according to an embodiment of the present disclosure. Detailed Embodiments

[0040] The exemplary embodiments of the present disclosure will be described below with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, the description of well-known functions and structures is omitted below.

[0041] It should be noted that in the description and claims of the present invention and the above-mentioned accompanying drawings, terms such as "target" and "candidate" are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0042] In addition, it should also be noted that in the technical solution of the present invention, the collection, storage, use, processing, transmission, provision, and disclosure of text images, text elements, etc. comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0043] Existing reading order prediction schemes are divided into methods based on heuristic rules and methods based on deep learning. The methods based on heuristic rules can only target specific data types and design manual rules for specific data scenarios, and they do not perform well when dealing with other document data; the methods based on machine learning only use a small number of samples of a certain type for training, and it is difficult for the model to learn effective feature representations. For documents with multiple layout formats, such as the case of two columns or three columns; at the same time, there are also special document layout structures for e-commerce network pictures. Analyzing the logical structure of this type of document poses a greater challenge, and it is difficult for the methods based on heuristic rules to have good generalization. The methods based on deep learning use graph networks for feature encoding. For cases with a large number of text lines, this encoding method usually constructs local subgraphs and selects nodes that are relatively close through the spatial L2 distance for feature encoding. This method lacks the modeling of global information and cannot effectively model the global information of the layout; at the same time, for the document scenario, the visual features of different handwritten text lines are similar, and it is difficult to effectively distinguish the connection relationship of text lines relying on visual features. Therefore, this method also has certain limitations.

[0044] Figure 1It is a flowchart of a reading order prediction method provided according to an embodiment of the present disclosure. This embodiment is applicable to the situation of predicting the reading order of a document. This method can be executed by a reading order prediction device, which can be implemented in software and / or hardware and integrated into an electronic device with the function of predicting the reading order, such as a server. As Figure 1 shown, the reading order prediction method of this embodiment may include:

[0045] S101, according to the visual features of the text image, determine the detection box features of at least two text elements in the text image.

[0046] In this embodiment, the text image may be an image containing text, any text image for which the reading order of text elements needs to be predicted. The text image may include at least one type of text element, where the granularity of each type of text element may be at the text line level, text paragraph level, text column level, etc. It should be noted that the granularities of at least two text elements in the text image may be the same or different; for example, at least two text elements in the text image are both text lines; another example is that at least two text elements in the text image include text lines and text paragraphs, etc.

[0047] The visual feature refers to the image feature used to characterize the text image, which can be represented in the form of a matrix or a vector. The detection box feature refers to the feature used to characterize the detection box of the text element, which can be represented in the form of a matrix or a vector.

[0048] Specifically, a detection box feature extraction network may be used to extract the features of the visual features of the text image to obtain the detection box features of at least two text elements in the text image. Among them, the detection box feature extraction network may be a convolutional neural network with any structure.

[0049] S102, according to the visual features of the text image, the detection box features of the text elements, and the position encoding features, determine the target encoding features of the text elements.

[0050] In this embodiment, the position encoding feature refers to the feature used to characterize the position of the text element, which can be represented in the form of a matrix or a vector. The target encoding feature refers to the final feature used to multi-dimensionally characterize the text element, which can be represented in the form of a matrix or a vector.

[0051] An optional method may be to use an encoding feature determination network to process the visual features of the text image, the detection box features of the text elements, and the position encoding features to obtain the target encoding features of each text element.

[0052] In another optional approach, for each text element, the visual features of the text image, the detection box features of the text element, and the position encoding features can also be fused based on a preset fusion method to obtain the target encoding features of the text element. For example, the visual features of the text image, the detection box features of the text element, and the position encoding features can be concatenated, and the concatenated result can be used as the target encoding features of the text element. Another example is that the visual features of the text image, the detection box features of the text element, and the position encoding features can be added together, and the result after addition can be used as the target encoding features of the text element.

[0053] S103. Determine the reading order between different text elements according to the target encoding features of the text elements.

[0054] An optional approach is to construct the graph features between the text elements in the text image according to the target encoding features of the text elements, and then the reading order between different elements in the text image can be determined based on the graph features.

[0055] In yet another optional approach, the target features of the text elements in the text image can also be predicted based on a reading order prediction model to obtain the reading order between different text elements. Among them, the reading order prediction model can be obtained based on a deep learning algorithm.

[0056] The technical solution provided by the embodiments of the present disclosure determines the detection box features of at least two text elements in the text image according to the visual features of the text image, and then determines the target encoding features of the text elements according to the visual features of the text image, the detection box features of the text elements, and the position encoding features, and further determines the reading order between different text elements according to the target encoding features of the text elements. The above technical solution combines the visual features of the text image, the detection box features of the text elements, and the position encoding features to determine the features of the text elements in multiple dimensions, making the prediction of the reading order between different text elements more accurate.

[0057] Figure 2 It is a flowchart of another reading order prediction method provided by the embodiments of the present disclosure. Based on the above embodiments, this embodiment further optimizes "determining the reading order between different text elements according to the target encoding features of the text elements" and provides an optional implementation scheme. As Figure 2 shown, the reading order prediction method of this embodiment may include:

[0058] S201. Determine the detection box features of at least two text elements in the text image according to the visual features of the text image.

[0059] S202. Determine the target encoding features of the text elements according to the visual features of the text image, the detection box features of the text elements, and the position encoding features.

[0060] S203. Cascade the target encoding features of different text elements two by two to obtain cascaded features.

[0061] S204. Determine the reading order between different text elements according to the cascaded features.

[0062] In this embodiment, the cascaded feature refers to the feature after cascading the target encoding features of different text elements, and can be represented in the form of a matrix or a vector. The reading order refers to the front-back connection relationship between different text elements.

[0063] Specifically, cascade the target encoding features of any two different text elements in the text image to obtain the cascaded features corresponding to the two text elements. Then, a reading order determination network can be used to determine the front-back connection relationship between the two text elements according to the cascaded features corresponding to the two text elements. Furthermore, according to the front-back connection relationships between each pair of text elements, determine the reading order between different text elements. Among them, the reading order determination network can be a neural network composed of fully connected layers.

[0064] The technical solution provided by the embodiments of the present disclosure determines the detection frame features of at least two text elements in the text image according to the visual features of the text image, then determines the target encoding features of the text elements according to the visual features of the text image, the detection frame features of the text elements, and the position encoding features, and further cascades the target encoding features of different text elements two by two to obtain cascaded features, and determines the reading order between different text elements according to the cascaded features. The above technical solution determines the reading order between different text elements based on the cascaded features of pairwise text elements, and can quickly and accurately determine the reading order of different text elements in the entire text image.

[0065] Figure 3 It is a flowchart of another reading order prediction method provided according to the embodiments of the present disclosure. On the basis of the above embodiments, this embodiment further optimizes "determine the detection frame features of at least two text elements in the text image according to the visual features of the text image" and provides an optional implementation solution. As Figure 3 shown, the reading order prediction method of this embodiment may include:

[0066] S301. Extract visual features from the text image to obtain the visual features of the text image.

[0067] Specifically, a visual feature extraction network can be used to extract visual features from the text image to obtain the visual features of the text image. Among them, the visual feature extraction network can be a convolutional neural network with any structure, for example, it can be a ResNet50 network.

[0068] S302. Determine the detection box features of at least two text elements in the text image according to the visual features of the text image.

[0069] An optional method is to perform screening detection on the visual features of the text image, screen out the detection boxes corresponding to the text elements, and then according to the detection boxes, screen out the visual features corresponding to the text elements from the visual features of the text image as the detection box features of the text elements.

[0070] S303. Determine the target encoding features of the text elements according to the visual features of the text image, the detection box features of the text elements, and the position encoding features.

[0071] S304. Determine the reading order between different text elements according to the target encoding features of the text elements.

[0072] The technical solution provided by the embodiments of the present disclosure extracts the visual features of the text image to obtain the visual features of the text image, and determines the detection box features of at least two text elements in the text image according to the visual features of the text image. Then, according to the visual features of the text image, the detection box features of the text elements, and the position encoding features, the target encoding features of the text elements are determined. Furthermore, according to the target encoding features of the text elements, the reading order between different text elements is determined. Through the above technical solution, by screening out the detection box features corresponding to each text element from the visual features of the text image, the influence of non-text elements in the text image on the prediction of the reading order of text elements can be avoided, thus laying a foundation for improving the accuracy of the reading order of text elements.

[0073] On the basis of the above embodiments, as an optional method of the present disclosure, determining the detection box features of at least two text elements in the text image according to the visual features of the text image may also be to detect the visual features of the text image to obtain at least two candidate detection boxes of the text elements; normalize the candidate detection boxes to obtain the target detection boxes of the text elements; encode the target detection boxes of the text elements to obtain the detection box features of the text elements.

[0074] Among them, the candidate detection box refers to the detection boxes of different scale sizes and different orientation angles of the text elements. The target detection box refers to the detection box obtained by performing scale normalization processing on the candidate detection box of the text element, and may include the coordinate information of the four vertices of the rectangle.

[0075] Specifically, a rotated region candidate network can be used to detect the visual features of a text image to obtain candidate detection boxes for at least two text elements. Subsequently, a rotated region feature pooling network can be used to map the candidate detection boxes with orientation angles into rectangular boxes, and then scale normalization is performed on the scales of the rectangular boxes to obtain the target detection boxes of the text elements. Furthermore, the target detection boxes of the text elements can be encoded. For example, a detection box encoding network can be used to encode the coordinate information of the target detection boxes to obtain the detection box features of the text elements. Alternatively, the coordinate information of the target detection boxes can also be directly concatenated to obtain the detection box features of the text elements. Among them, the rotated region candidate network can be an object detection network based on deep learning; the rotated region feature pooling network can be a neural network including a pooling layer.

[0076] It can be understood that by normalizing the candidate detection boxes, the detection box coordinates of the more accurate target detection boxes can be obtained, so that the determined detection box features of the text elements are more accurate and reasonable.

[0077] Figure 4A It is a flowchart of yet another reading order prediction method provided according to an embodiment of the present disclosure. On the basis of the above embodiment, this embodiment further optimizes "determining the target encoding features of text elements according to the visual features, detection box features, and position encoding features of the text image" and provides an optional implementation solution. As Figure 4A shown, the reading order prediction method of this embodiment may include:

[0078] S401, according to the visual features of the text image, determine the detection box features of at least two text elements in the text image.

[0079] S402, according to the detection box features and position encoding features of the text elements, determine the coordinate encoding features of the text elements.

[0080] In this embodiment, the position encoding feature refers to the position feature of the text element in the text image and can be a one-dimensional vector; optionally, a word embedding method can be used to determine the position encoding feature of the text element. The coordinate encoding feature refers to the feature used to represent the position information of the text element and can be represented in the form of a matrix or a vector.

[0081] Specifically, the detection box features and position encoding features of the text elements can be superimposed to obtain the coordinate encoding features of the text elements.

[0082] S403, according to the candidate detection boxes of the text elements and the visual features of the text image, determine the visual encoding features of the text elements.

[0083] In this embodiment, the visual coding feature refers to the image feature corresponding to the text element, which can be represented in the form of a matrix or a vector.

[0084] An optional method is to screen out the visual feature corresponding to the text element from the visual features of the text image according to the candidate detection box of the text element, and use it as the visual coding feature of the text element.

[0085] Another optional method is to determine the candidate coding feature of the text element from the visual features of the text image according to the candidate detection box of the text element; normalize the candidate coding feature of the text element to obtain the visual coding feature of the text element. Specifically, map the candidate detection box of the text element to a certain area on the visual feature of the text image, and use the visual feature corresponding to this area as the candidate coding feature of the text element. Then, a rotated region feature pooling network can be used to normalize the candidate coding feature of the text element to obtain the visual coding feature of the text element.

[0086] It can be understood that normalizing the visual features of the text elements makes the scales of the visual coding features of each text element the same, which is convenient for predicting the reading order of different text elements later.

[0087] S404. Determine the target coding feature of the text element according to the visual coding feature and the coordinate coding feature of the text element.

[0088] Specifically, the visual coding feature and the coordinate coding feature of the text element can be superimposed to obtain a superimposed feature, and then the superimposed feature can be input into the coding network to obtain the target coding feature of the text element. Among them, the coding network can be a Transformer coding network; the Transformer coding layer consists of six layers of attention modules, and each attention module includes two sub-layers. The input feature (superimposed feature) first passes through the multi-head self-attention mechanism. In this process, the query vector (Q), key vector (K), and value vector (V) of the input are specified. For the self-attention mechanism, the input Q, K, and V are the same feature. In the calculation process of the self-attention mechanism, first calculate the inner product of Q and K, then divide it by the scaling coefficient, and obtain the weight through the Softmax function to weight V. The output feature dimension of the self-attention module is [N, 256], and then it is respectively input into the regularization layer and the residual connection layer. The output feature of the first layer is input into the feed-forward fully connected layer, regularization layer, and residual layer of the second sub-layer, and the output obtains the feature dimension [N, 256]. Input the output of the previous layer of attention module into the next layer, and the final output coding feature (target coding feature) dimension is [N, 256].

[0089] S405. Determine the reading order between different text elements according to the target coding feature of the text element.

[0090] In the technical solution provided by the embodiment of the present disclosure, according to the visual features of the text image, the detection box features of at least two text elements in the text image are determined, and according to the detection box features and position encoding features of the text elements, the coordinate encoding features of the text elements are determined. Then, according to the candidate detection boxes of the text elements and the visual features of the text image, the visual encoding features of the text elements are determined. Furthermore, according to the visual encoding features and coordinate encoding features of the text elements, the target encoding features of the text elements are determined. Finally, according to the target encoding features of the text elements, the reading order between different text elements is determined. In the above technical solution, by combining the features of different dimensions such as the position and vision of the text elements to determine the target encoding features of the text elements, the target encoding features of the text elements are made more abundant, thus laying a foundation for determining the reading order of different text elements.

[0091] Based on the above embodiment, combined with Figure 4B the schematic diagram of the reading order prediction process shown, the prediction process of the reading order of different text elements is specifically described. Optionally, the reading order prediction model includes a visual feature extraction network, a rotated region candidate network, a rotated region feature pooling network, an encoding network, and a reading order determination network. The text image is input into the visual feature extraction network to obtain the visual features of the text image; then, the rotated region candidate network is used to detect the visual features of the text image to obtain the candidate detection boxes of the text elements; furthermore, the rotated region feature pooling network is used to normalize the candidate detection boxes of the text elements to obtain the target detection boxes of the text elements, and the detection box features of the text elements are determined according to the target detection boxes. According to the candidate detection boxes of the text elements and the visual features of the text image, the visual encoding features of the text elements are determined; then, according to the detection box features and position encoding features of the text elements, the coordinate encoding features of the text elements are determined; the encoding network is used to determine the target encoding features of the text elements according to the visual encoding features and coordinate encoding features of the text elements. Finally, the reading order determination network is used to determine the reading order between different text elements according to the target encoding features of the text elements.

[0092] Figure 5 is a flowchart of a method for training a reading order prediction model provided by an embodiment of the present disclosure. This embodiment is applicable to the situation of predicting the reading order of a document. This method can be executed by a training device of the reading order prediction model, and the device can be implemented in a software and / or hardware manner and can be integrated into an electronic device with the training function of the reading order prediction model, such as a server. As Figure 5 shown, the method for training the reading order prediction model in this embodiment may include:

[0093] S501. Determine the detection box features of at least two text elements in the text image according to the visual features of the text image.

[0094] In this embodiment, the text image can be an image containing text, any text image for which the reading order prediction of text elements is required. The text image may include at least one type of text element. Among them, the granularity of each text element can be at the text line level, text paragraph level, text column level, etc. It should be noted that the granularity of at least two text elements in the text image can be the same or different; for example, at least two text elements in the text image are both text lines; another example is that at least two text elements in the text image include text lines and text paragraphs, etc.

[0095] The visual feature refers to the image feature used to characterize the text image, which can be represented in the form of a matrix or a vector. The detection box feature refers to the feature of the detection box used to characterize the text element, which can be represented in the form of a matrix or a vector.

[0096] An optional method is to use a detection box feature extraction network to extract the visual features of the text image to obtain the detection box features of at least two text elements in the text image. Among them, the detection box feature extraction network can be a convolutional neural network with any structure.

[0097] Another optional method is to use the visual feature extraction network in the reading order prediction model to extract the visual features of the text image to obtain the visual features of the text image; determine the detection box features of at least two text elements in the text image according to the visual features of the text image. Further, determining the detection box features of at least two text elements in the text image according to the visual features of the text image can also be to detect the visual features of the text image to obtain at least two candidate detection boxes of the text elements; normalize the candidate detection boxes to obtain the target detection boxes of the text elements; encode the target detection boxes of the text elements to obtain the detection box features of the text elements. Among them, the visual feature extraction network can be a convolutional neural network with any structure, such as the ResNet50 network.

[0098] S502. Determine the target encoding features of the text elements according to the visual features of the text image, the detection box features of the text elements, and the position encoding features.

[0099] In this embodiment, the position encoding feature refers to the feature used to characterize the position of the text element, which can be represented in the form of a matrix or a vector. The target encoding feature refers to the final feature used to multi-dimensionally characterize the text element, which can be represented in the form of a matrix or a vector.

[0100] An optional method is to use an encoding feature determination network to process the visual features of the text image, the detection box features of the text elements, and the position encoding features to obtain the target encoding features of each text element.

[0101] Another optional method is that for each text element, based on a preset fusion method, the visual features of the text image, the detection box features of this text element, and the position encoding features can be fused to obtain the target encoding features of this text element. For example, the visual features of the text image, the detection box features of this text element, and the position encoding features can be concatenated, and the concatenated result can be used as the target encoding features of this text element. Another example is that the visual features of the text image, the detection box features of this text element, and the position encoding features can be added together, and the added result can be used as the target encoding features of this text element.

[0102] Another optional method is to determine the coordinate encoding features of the text element according to the detection box features and position encoding features of the text element; determine the visual encoding features of the text element according to the candidate detection box of the text element and the visual features of the text image; then use an encoding network to determine the target encoding features of the text element according to the visual encoding features and coordinate encoding features of the text element; where the encoding network can be a Transformer encoding network.

[0103] Furthermore, determining the visual encoding features of the text element according to the candidate detection box of the text element and the visual features of the text image can also be that the rotation region candidate network in the reading order prediction model is used to determine the candidate encoding features of the text element from the visual features of the text image according to the candidate detection box of the text element; the rotation region feature pooling network in the reading order prediction model is used to normalize the candidate encoding features of the text element to obtain the visual encoding features of the text element. Among them, the rotation region candidate network can be an object detection network based on deep learning; the rotation region feature pooling network can be a neural network including a pooling layer.

[0104] S503. Determine the reading order between different text elements according to the target encoding features of the text elements.

[0105] An optional method is to construct the graph features between the text elements in the text image according to the target encoding features of each text element, and then the reading order between different elements in the text image can be determined based on the graph features.

[0106] Another optional method is to also predict the target features of each text element in the text image based on the reading order prediction model to obtain the reading order between different text elements. Among them, the reading order prediction model can be obtained based on a deep learning algorithm.

[0107] Another optional approach is to perform pairwise concatenation on the target encoding features of different text elements to obtain concatenated features. Subsequently, the reading order determination network of the reading order prediction model can be used to determine the reading order between different text elements based on the concatenated features. Among them, the reading order determination network can be a neural network composed of fully connected layers.

[0108] S504. Train the reading order prediction model according to the reading order of the text elements and the label data.

[0109] Among them, the label data refers to the data of the true reading order of the text elements, that is, the data used to annotate the front-back connection relationship of the text elements. For example, for any two text elements, if they are adjacent, the label data is 1; if not, the label data is 0.

[0110] An optional approach is to preset a loss function, calculate the training loss according to the reading order of the text elements and the label data, and then use the training loss to train the reading order prediction model until the training stop condition is met, and the training of the model is stopped. Among them, the training stop condition can be that the training loss is stable within a set range, or the number of iterations meets the set number; the set range and the set number can be set by those skilled in the art according to the actual situation. It should be noted that the preset loss function can be a cross-entropy loss function.

[0111] Furthermore, in order to more effectively optimize the reading loss prediction model, during the training process, through OHEM, the top K items with the largest loss in the negative samples are selected, and the positive-negative sample ratio is maintained at a set value such as 1:3 to mine difficult-to-recognize samples.

[0112] Another optional approach is to, during the training process of the reading order prediction model, respectively determine the losses corresponding to the visual feature extraction network, the rotated region candidate network, the rotated region feature pooling network, the encoding network, and the reading order determination network in the reading order prediction model, and then sum and average the losses to obtain the training loss, and use the training loss to train the reading order prediction model.

[0113] The technical solution provided by the embodiments of the present disclosure determines the detection box features of at least two text elements in a text image according to the visual features of the text image, and then determines the target encoding features of the text elements according to the visual features of the text image, the detection box features of the text elements, and the position encoding features. Furthermore, according to the target encoding features of the text elements, the reading order between different text elements is determined. Finally, according to the reading order of the text elements and the label data, the reading order prediction model is trained. The above technical solution combines the visual features of the text image, the detection box features of the text elements, and the position encoding features to determine the features of the text elements in multiple dimensions, making the prediction of the reading order between different text elements more accurate.

[0114] Figure 6 FIG. is a schematic structural diagram of a reading order prediction device provided by an embodiment of the present disclosure. This embodiment is applicable to the situation of predicting the reading order of a document. The device can be implemented in software and / or hardware and can be integrated into an electronic device with the function of predicting the reading order, such as a server. As Figure 6 shown, the reading order prediction device 600 of this embodiment may include:

[0115] A detection box feature determination module 601, configured to determine the detection box features of at least two text elements in the text image according to the visual features of the text image;

[0116] A target encoding feature determination module 602, configured to determine the target encoding features of the text elements according to the visual features of the text image, the detection box features of the text elements, and the position encoding features;

[0117] A reading order determination module 603, configured to determine the reading order between different text elements according to the target encoding features of the text elements.

[0118] The technical solution provided by the embodiments of the present disclosure determines the detection box features of at least two text elements in a text image according to the visual features of the text image, and then determines the target encoding features of the text elements according to the visual features of the text image, the detection box features of the text elements, and the position encoding features. Furthermore, according to the target encoding features of the text elements, the reading order between different text elements is determined. The above technical solution combines the visual features of the text image, the detection box features of the text elements, and the position encoding features to determine the features of the text elements in multiple dimensions, making the prediction of the reading order between different text elements more accurate.

[0119] Further, the reading order determination module 603 is specifically configured to:

[0120] Perform pairwise concatenation on the target encoding features of different text elements to obtain concatenated features;

[0121] Determine the reading order between different text elements according to the cascaded features.

[0122] Furthermore, the detection box feature determination module 601 includes:

[0123] A visual feature determination unit for extracting visual features of the text image to obtain the visual features of the text image;

[0124] A detection box feature determination unit for determining the detection box features of at least two text elements in the text image according to the visual features of the text image.

[0125] Furthermore, the detection box feature determination unit is specifically used for:

[0126] Detect the visual features of the text image to obtain candidate detection boxes for at least two text elements;

[0127] Normalize the candidate detection boxes to obtain the target detection boxes of the text elements;

[0128] Encode the target detection boxes of the text elements to obtain the detection box features of the text elements.

[0129] Furthermore, the target encoding feature determination module 602 includes:

[0130] A coordinate encoding feature determination unit for determining the coordinate encoding features of the text elements according to the detection box features and position encoding features of the text elements;

[0131] A visual encoding feature determination unit for determining the visual encoding features of the text elements according to the candidate detection boxes of the text elements and the visual features of the text image;

[0132] A target encoding feature determination unit for determining the target encoding features of the text elements according to the visual encoding features and coordinate encoding features of the text elements.

[0133] Furthermore, the visual encoding feature determination unit is specifically used for:

[0134] Determine the candidate encoding features of the text elements from the visual features of the text image according to the candidate detection boxes of the text elements;

[0135] Normalize the candidate encoding features of the text elements to obtain the visual encoding features of the text elements.

[0136] Figure 7It is a schematic structural diagram of a training device for a reading order prediction model provided according to an embodiment of the present disclosure. This embodiment is applicable to the situation of predicting the reading order of documents. The device can be implemented in software and / or hardware and can be integrated into an electronic device with the training function of the reading order prediction model, such as a server. As Figure 7 shown, the training device 700 of the reading order prediction model in this embodiment may include:

[0137] A detection box feature determination module 701, configured to determine the detection box features of at least two text elements in the text image according to the visual features of the text image;

[0138] A target encoding feature determination module 702, configured to determine the target encoding features of the text elements according to the visual features of the text image, the detection box features of the text elements, and the position encoding features;

[0139] A reading order determination module 703, configured to determine the reading order between different text elements according to the target encoding features of the text elements;

[0140] A model training module 704, configured to train the reading order prediction model according to the reading order of the text elements and the label data.

[0141] The technical solution provided by the embodiment of the present disclosure determines the detection box features of at least two text elements in the text image according to the visual features of the text image, then determines the target encoding features of the text elements according to the visual features of the text image, the detection box features of the text elements, and the position encoding features, and further determines the reading order between different text elements according to the target encoding features of the text elements. Finally, the reading order prediction model is trained according to the reading order of the text elements and the label data. The above technical solution combines the visual features of the text image, the detection box features of the text elements, and the position encoding features to determine the features of the text elements in multiple dimensions, making the prediction of the reading order between different text elements more accurate.

[0142] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0143] Figure 8 It is a block diagram of an electronic device for implementing the reading order prediction method or the training method of the reading order prediction model according to the embodiment of the present disclosure. Figure 8FIG. shows a schematic block diagram of an exemplary electronic device 800 that may be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as, for example, personal digital processors, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0144] As Figure 8 shown, the electronic device 800 includes a computing unit 801 that can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the electronic device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0145] Multiple components in the electronic device 800 are connected to the I / O interface 805, including: an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, an optical disc, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the electronic device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0146] The computing unit 801 may be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 executes the various methods and processes described above, such as the reading order prediction method or the training method of the reading order prediction model. For example, in some embodiments, the reading order prediction method or the training method of the reading order prediction model may be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program may be loaded and / or installed onto the electronic device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the reading order prediction method or the training method of the reading order prediction model described above may be executed. Alternatively, in other embodiments, the computing unit 801 may be configured to execute the reading order prediction method or the training method of the reading order prediction model by any other suitable means (e.g., by means of firmware).

[0147] Various embodiments of the systems and techniques described above in this document may be implemented in digital electronic circuitry, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), systems-on-a-chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include: being implemented in one or more computer programs that may be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a special-purpose or general-purpose programmable processor, and which may receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0148] The program code for implementing the methods of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general purpose computer, a special purpose computer, or other programmable data processing apparatus, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on the remote machine or server.

[0149] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0150] In order to provide interaction with a user, the systems and techniques described herein may be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and the input from the user may be received in any form (including acoustic input, speech input, or tactile input).

[0151] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end, middleware, or front-end components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.

[0152] A computer system can include clients and servers. The clients and servers are generally far from each other and typically interact through a communication network. The client-server relationship is created by computer programs that run on the respective computers and have a client-server relationship with each other. The server can be a cloud server, a server of a distributed system, or a server integrated with a blockchain.

[0153] Artificial intelligence is a discipline that studies the use of computers to simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, planning, etc.), and it has both hardware-level and software-level technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing; artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech recognition technology, natural language processing technology, machine learning / deep learning technology, big data processing technology, and knowledge graph technology.

[0154] Cloud computing refers to a technical system in which elastic and scalable shared physical or virtual resource pools are accessed through a network. The resources can include servers, operating systems, networks, software, applications, and storage devices, etc., and the resources can be deployed and managed in a on-demand and self-service manner. Through cloud computing technology, it is possible to provide efficient and powerful data processing capabilities for the application and model training of technologies such as artificial intelligence and blockchain.

[0155] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.

[0156] The above specific embodiments do not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub - combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present disclosure shall be included within the protection scope of the present disclosure.

Claims

1. A reading order prediction method, comprising: Determining detection box features of at least two text elements in the text image according to visual features of the text image; Determining coordinate encoding features of the text elements according to the detection box features and position encoding features of the text elements; Determining visual encoding features of the text elements according to candidate detection boxes of the text elements and visual features of the text image; Determining target encoding features of the text elements according to the visual encoding features and coordinate encoding features of the text elements; wherein the target encoding features refer to the final features for multi-dimensionally characterizing text elements; Determining the reading order between different text elements according to the target encoding features of the text elements.

2. The method according to claim 1, wherein, The determining the reading order between different text elements according to the target encoding features of the text elements includes: Performing two-by-two concatenation on the target encoding features of different text elements to obtain concatenated features; Determining the reading order between different text elements according to the concatenated features.

3. The method according to claim 1, wherein, The determining detection box features of at least two text elements in the text image according to visual features of the text image includes: Performing visual feature extraction on the text image to obtain visual features of the text image; Determining detection box features of at least two text elements in the text image according to the visual features of the text image.

4. The method according to claim 3, wherein The determining detection box features of at least two text elements in the text image according to visual features of the text image includes: Detecting visual features of the text image to obtain candidate detection boxes of at least two text elements; Normalizing the candidate detection boxes to obtain target detection boxes of the text elements; Encoding the target detection boxes of the text elements to obtain detection box features of the text elements.

5. The method according to claim 1, wherein, The determining visual encoding features of the text elements according to candidate detection boxes of the text elements and visual features of the text image includes: Determining candidate encoding features of the text elements from the visual features of the text image according to the candidate detection boxes of the text elements; Normalizing the candidate encoding features of the text elements to obtain visual encoding features of the text elements.

6. A training method for a reading order prediction model, comprising: Determining detection box features of at least two text elements in the text image according to visual features of the text image; Determining coordinate encoding features of the text elements according to the detection box features and position encoding features of the text elements; Determining visual encoding features of the text elements according to candidate detection boxes of the text elements and visual features of the text image; Determining target encoding features of the text elements according to the visual encoding features and coordinate encoding features of the text elements; wherein the target encoding features refer to the final features for multi-dimensionally characterizing text elements; Determining the reading order between different text elements according to the target encoding features of the text elements; Training a reading order prediction model according to the reading order of the text elements and label data.

7. A reading order prediction device, comprising: A detection box feature determination module, configured to determine the detection box features of at least two text elements in the text image according to the visual features of the text image; A target encoding feature determination module, including: A coordinate encoding feature determination unit, configured to determine the coordinate encoding features of the text element according to the detection box features and position encoding features of the text element; A visual encoding feature determination unit, configured to determine the visual encoding features of the text element according to the candidate detection box of the text element and the visual features of the text image; A target encoding feature determination unit, configured to determine the target encoding features of the text element according to the visual encoding features and coordinate encoding features of the text element; wherein, the target encoding feature refers to the final feature for multi-dimensionally representing the text element; A reading order determination module, configured to determine the reading order between different text elements according to the target encoding features of the text elements.

8. The device according to claim 7, wherein Specifically, the reading order determination module is configured to: Perform pairwise cascading on the target encoding features of different text elements to obtain cascaded features; Determine the reading order between different text elements according to the cascaded features.

9. The device according to claim 7, wherein, The detection box feature determination module includes: A visual feature determination unit, configured to extract visual features of the text image to obtain the visual features of the text image; A detection box feature determination unit, configured to determine the detection box features of at least two text elements in the text image according to the visual features of the text image.

10. The device according to claim 9, wherein, Specifically, the detection box feature determination unit is configured to: Detect the visual features of the text image to obtain candidate detection boxes of at least two text elements; Normalize the candidate detection boxes to obtain the target detection boxes of the text elements; Encode the target detection boxes of the text elements to obtain the detection box features of the text elements.

11. The apparatus according to claim 7, wherein, Specifically, the visual encoding feature determination unit is configured to: Determine the candidate encoding features of the text element from the visual features of the text image according to the candidate detection box of the text element; Normalize the candidate encoding features of the text element to obtain the visual encoding features of the text element.

12. A training device for a reading order prediction model, including: A detection box feature determination module, configured to determine the detection box features of at least two text elements in the text image according to the visual features of the text image; A target encoding feature determination module, configured to determine the coordinate encoding features of the text element according to the detection box features and position encoding features of the text element; Determine the visual encoding features of the text element according to the candidate detection box of the text element and the visual features of the text image; Determine the target encoding features of the text element according to the visual encoding features and coordinate encoding features of the text element; wherein, the target encoding feature refers to the final feature for multi-dimensionally representing the text element; A reading order determination module, configured to determine the reading order between different text elements according to the target encoding features of the text elements; A model training module, configured to train the reading order prediction model according to the reading order of the text elements and the label data.

13. An electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the reading order prediction method according to any one of claims 1-5, or the training method of the reading order prediction model according to claim 6.

14. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause a computer to execute the reading order prediction method according to any one of claims 1-5, or the training method of the reading order prediction model according to claim 6.

15. A computer program product, comprising a computer program which, when executed by a processor, implements the reading order prediction method according to any one of claims 1-5, or the training method of the reading order prediction model according to claim 6.

Citation Information

Patent Citations

  • Document layout analysis method and device, model training method and device, and equipment

    CN113378580A

  • Text content processing method and device, computer equipment and storage medium

    CN113822283A