Layout information determination method and apparatus, electronic device, and storage medium
Patent Information
- Application Number
- CN202211730625.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-30
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2042-12-30
AI Technical Summary
[0010]应当理解,本部分所描述的内容并非旨在标识本公开的实施例的关键或重要特征,也不用于限制本公开的范围。本公开的其它特征将通过以下的说明书而变得容易理解。
Smart Images

Figure CN116416638B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence technology, specifically deep learning, image processing, large model, and computer vision technology, and can be applied to scenarios such as optical character recognition (OCR). In particular, it relates to a method, device, electronic device, and storage medium for determining layout information. Background Technology
[0002] Artificial intelligence (AI) is the study of enabling computers to simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). It involves both hardware and software technologies. AI hardware technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, and big data processing. AI software technologies mainly include computer vision, speech recognition, natural language processing, as well as machine learning, deep learning, big data processing, and knowledge graph technologies.
[0003] In related technologies, the task of extracting layout information involves detecting areas in a document that fall into different layout categories, such as text paragraphs, headings, tables, illustrations, and stamps. The detected layout categories are then provided to other downstream tasks (e.g., layout restoration tasks). Summary of the Invention
[0004] This disclosure provides a method, apparatus, electronic device, storage medium, and computer program product for determining layout information.
[0005] According to a first aspect of this disclosure, a method for determining layout information is provided, comprising: acquiring a document image, wherein the document image includes a text region image; acquiring a sampled feature map corresponding to the text region image, wherein the sampled feature map includes at least one feature point; and determining layout information corresponding to the text region image based on the sampled feature map and the at least one feature point.
[0006] According to a second aspect of this disclosure, a layout information determination apparatus is provided, comprising: a first acquisition module for acquiring a document image, wherein the document image includes a text region image; a second acquisition module for acquiring a sampling feature map corresponding to the text region image, wherein the sampling feature map includes at least one feature point; and a determination module for determining layout information corresponding to the text region image based on the sampling feature map and the at least one feature point.
[0007] According to a third aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method of the first aspect of this disclosure.
[0008] According to a fourth aspect of this disclosure, a non-transitory computer-readable storage medium storing computer instructions for causing a computer to perform a method according to an embodiment of the first aspect of this disclosure is provided.
[0009] According to a fifth aspect of this disclosure, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the method of the first aspect of this disclosure.
[0010] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0011] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:
[0012] Figure 1 This is a schematic diagram based on the first embodiment of the present disclosure;
[0013] Figure 2 This is a schematic diagram according to the second embodiment of the present disclosure;
[0014] Figure 3 This is a schematic diagram according to the third embodiment of the present disclosure;
[0015] Figure 4 This is a schematic diagram according to the fourth embodiment of the present disclosure;
[0016] Figure 5 This is a schematic diagram of the model structure for the document layout detection and fine-tuning stage in this embodiment of the disclosure;
[0017] Figure 6 This is a schematic diagram according to the fifth embodiment of the present disclosure;
[0018] Figure 7 This is a schematic diagram of the model structure during the general model pre-training stage in this embodiment of the disclosure;
[0019] Figure 8 This is a schematic diagram according to the sixth embodiment of the present disclosure;
[0020] Figure 9 This is a schematic diagram according to the seventh embodiment of the present disclosure;
[0021] Figure 10 A schematic block diagram of an example electronic device is shown that can be used to implement the layout information determination method of embodiments of the present disclosure. Detailed Implementation
[0022] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0023] Figure 1 This is a schematic diagram according to the first embodiment of the present disclosure.
[0024] It should be noted that the execution subject of the layout information determination method in this embodiment is a layout information determination device. This device can be implemented by software and / or hardware. This device can be configured in an electronic device, which may include, but is not limited to, a terminal, a server, etc.
[0025] This disclosure relates to the field of artificial intelligence technology, specifically deep learning, image processing, large models, and computer vision technology, and can be applied to scenarios such as optical character recognition (OCR).
[0026] Artificial Intelligence (AI) is a new technological science that studies and develops theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence.
[0027] Deep learning learns the inherent patterns and hierarchical representations of sample data. The information gained during this learning process greatly aids in interpreting data such as text, images, and sound. Its ultimate goal is to enable machines to possess analytical and learning capabilities like humans, allowing them to recognize data such as text, images, and sound.
[0028] Image processing is the technique of using computers to analyze images to achieve desired results. It is also known as image processing. Image processing generally refers to digital image processing. A digital image is a large two-dimensional array obtained by capturing images using equipment such as industrial cameras, video cameras, and scanners. The elements of this array are called pixels, and their values are called grayscale values. Image processing techniques generally include three parts: image compression, enhancement and restoration, and matching, description, and recognition.
[0029] Large-scale AI models are a milestone technology for artificial intelligence to move towards general intelligence. They possess both the attributes of "large scale" and "pre-training". Before modeling for real-world tasks, they need to be pre-trained on massive amounts of general data, which can significantly improve the generalization, versatility and practicality of artificial intelligence (AI).
[0030] Computer vision is a technology that uses computers to simulate the human visual process, enabling them to perceive their environment and perform human visual functions. It integrates image processing, artificial intelligence, and pattern recognition technologies. It primarily uses computers to simulate human visual functions, extracting information from images of objective things, processing and understanding it for practical detection, measurement, and control.
[0031] The layout information extraction task is to detect areas in a document that have different layout categories, including text paragraphs, headings, tables, illustrations, and stamps, and to provide the detected layout categories to other downstream tasks (e.g., layout restoration tasks).
[0032] In related technologies, deep learning-based methods mostly employ general object detection techniques, treating different layout categories in document images as different content categories. However, for some layout categories, the generated visual features are similar, such as titles, footers, footnotes, and page numbers, making it difficult to accurately distinguish them based solely on visual features.
[0033] Therefore, in order to solve the above-mentioned technical problems, this embodiment of the present disclosure obtains a document image, which includes a text region image, and obtains a sampling feature map corresponding to the text region image, wherein the sampling feature map includes at least one feature point, and determines the layout information corresponding to the text region image based on the sampling feature map and at least one feature point. This can effectively improve the accuracy of the layout information determination and significantly improve the detection effect for ambiguous layout information.
[0034] like Figure 1 As shown, the method for determining this layout information includes:
[0035] S101: Obtain a document image, wherein the document image includes: a text region image.
[0036] The image from which the layout information is to be extracted can be called a document image. This document image may include text, pictures, seals, signatures, and other content with specific layout information, and there are no restrictions on this.
[0037] Among them, text region image refers to a local region image of text that may contain specific layout information, obtained by pre-identifying the document image.
[0038] For example, if a document image includes a table containing multiple rows and columns of text, then the local area of the table in the document image can be called a text region image; if a document image includes a line of text containing multiple text characters, then the local area of the line of text in the document image can also be called a text region image, without any limitation.
[0039] In this embodiment of the disclosure, the document image may be pre-processed with initial recognition, such as OCR recognition, to identify one or more text region images from the document image, and then the layout information may be recognized based on the text region images.
[0040] In this embodiment of the disclosure, the document image may include one or more text region images. Alternatively, the entire document image may be detected and recognized to identify the features corresponding to each text region image.
[0041] S102: Obtain a sampled feature map corresponding to the text region image, wherein the sampled feature map includes at least one feature point.
[0042] The text region image has corresponding features, such as color features and depth features. The image used to represent the corresponding features of the text region image can be called a feature map. A sampled feature map refers to a feature map obtained by sampling the feature map. Each sampled feature map can include several pixels, and each pixel can be regarded as a feature point. Accordingly, in this embodiment of the disclosure, all pixels contained in the sampled feature map can be used as feature points, or some pixels can be sampled from all pixels as feature points. The obtained feature points can be used for subsequent detection and determination of the layout information of the text region.
[0043] In some embodiments, feature identification can be performed on each text region image to obtain a feature map, and then the feature map can be downsampled to obtain a sampled feature map. Alternatively, a reference region image corresponding to the text region image can be obtained (assuming that the similarity between the text region image and the reference region image meets the condition). Then, the sampled feature map pre-annotated to the reference region image can be used as the sampled feature map corresponding to the text region image. Of course, any other possible methods can be used to obtain the sampled feature map corresponding to the text region image, and there are no restrictions on this.
[0044] For example, the document image has dimensions (H, W, 3), where H represents the height of the document image, W represents the width of the document image, and 3 represents the three color channels in the RGB color mode. The RGB color mode is an industry color standard that uses the red (R), green (G), and blue (B) color channels to downsample the document image (the downsampling rate can be 1 / 4) to obtain a sampled feature map corresponding to each text region image. The sampled feature map can be represented as (H / 4, W / 4, 128), where the height of the sampled feature map is H / 4, the width is W / 4, and 128 represents 128 dimensions of features. In addition to the color features of the RGB color mode, it can also include edge information features, depth features, etc., without any restrictions.
[0045] S103: Determine the layout information corresponding to the text region image based on the sampled feature map and at least one feature point.
[0046] The information used to describe the layout of the text region image can be called layout information. Layout information can include, for example, layout category, and the position information of the text region image in the overall document image, in order to effectively extend downstream tasks.
[0047] The above-mentioned document image acquisition method includes a text region image, and a sampling feature map corresponding to the text region image is acquired. The sampling feature map includes at least one feature point. Then, the layout information corresponding to the text region image can be determined by combining the sampling feature map and at least one feature point.
[0048] For example, the information of the text detection box to which each feature point belongs can be parsed from the sampled feature map, and the layout information can be predicted based on the information of the text detection box. Alternatively, the sampled feature map and at least one feature point can be processed in any other possible way to determine the layout information corresponding to the text region image, such as by using a neural network model or an engineering approach.
[0049] In this embodiment, by acquiring a document image, which includes a text region image, and acquiring a sampling feature map corresponding to the text region image, wherein the sampling feature map includes at least one feature point, and determining the layout information corresponding to the text region image based on the sampling feature map and at least one feature point, the accuracy of the layout information determination can be effectively improved, and the detection effect can be significantly improved for ambiguous layout information.
[0050] In some embodiments of this disclosure, the number of feature points in the sampled feature map corresponding to the text region image can be multiple. When determining the layout information corresponding to the text region image based on the sampled feature map and at least one feature point, the layout category and candidate position information of the candidate text box to which each feature point belongs can be determined based on the sampled feature map. Furthermore, the positional offset information between the feature point and the candidate text box can be determined based on the sampled feature map and the candidate position information of the candidate text box. Based on the positional offset information, the target text box can be determined from multiple candidate text boxes. The layout category and candidate position information of the target text box can be used together as layout information. This enables accurate and rapid identification of the layout information corresponding to each text region image in the document image. Moreover, it can not only accurately identify the layout category but also accurately identify the position information of the text region image in the document image, effectively extending the downstream tasks of layout information determination and improving the practicality of the layout information determination method.
[0051] Among them, the feature point can be, for example, a pixel in the sampled feature map; the candidate text box to which the feature point belongs can be predicted based on the position information of the feature point and the text box it may belong to in the document image; the layout category expresses the layout classification to which the candidate text box may belong, such as title, footer, footnote, page number, etc.; and the candidate position information expresses the position information of the candidate text box relative to the document image, such as, for example, the horizontal and vertical coordinate values of the four vertices of the candidate text box.
[0052] Feature points have corresponding positional information. The relative positional offset between feature points and candidate text boxes can be called positional offset information. Positional offset information can be specifically, for example, the distance information between the feature point and each vertex of the candidate text box, without any restrictions.
[0053] The above-mentioned positional offset information between each feature point and the candidate text box is determined. Since there are multiple feature points, multiple candidate text boxes will be predicted accordingly. The target text box can be determined from the multiple candidate text boxes based on the positional offset information. Since the layout category and candidate position information of the target text box have been predicted, the layout information of the text region image to which the target text box belongs can be directly determined.
[0054] Specifically, determining the target text box from multiple candidate text boxes based on the position offset information can be achieved by referencing the position offset information to deduplicate multiple candidate text boxes, or by selecting the best candidate text box to determine the target text box from multiple candidate text boxes.
[0055] For example, by combining the regression value (position offset information) of each feature point (i.e., pixel) in the sampled feature map with the layout category, the target text box is selected from multiple candidate text boxes by using the Non-Maximum Suppression (NMS) algorithm. The layout category and candidate position information of the target text box are used as the layout information of the detected text region image. If there are multiple text region images, the output value of the detection result can be represented as (N, 4, 2), where N represents the number of target detection boxes, 4 represents that the target detection box has four vertices, and 2 represents the position information of each vertex, which can be represented by two coordinate values, such as the X coordinate value and the Y coordinate value, without restriction.
[0056] Figure 2 This is a schematic diagram according to the second embodiment of the present disclosure.
[0057] like Figure 2 As shown, the method for determining this layout information includes:
[0058] S201: Obtain a document image, wherein the document image includes: a text region image.
[0059] For a detailed description of S201, please refer to the above embodiments, which will not be repeated here.
[0060] S202: Input the document image into the target sampling network model and obtain the sampling feature map corresponding to the text region image output by the target sampling network model. The target sampling network model has learned the mapping relationship between the document image and the sampling feature map corresponding to the text region image in the document image during the document layout detection fine-tuning stage. The sampling feature map includes at least one feature point.
[0061] The target sampling network model can be used to downsample each text region image in a document image to output a sampled feature map corresponding to the text region image.
[0062] The target sampling network model in this embodiment can be configured and trained on a lightweight initial backbone network model based on a pre-trained backbone network model, i.e., a visual feature extraction model, during the document layout detection fine-tuning stage. The pre-trained backbone network model can be trained during the general model pre-training stage, as detailed in the following embodiments.
[0063] The document layout detection fine-tuning stage refers to the model training stage where the target sampling network model is fine-tuned.
[0064] For example, the target model can be initialized based on the target model parameters determined in the general model pre-training stage, and then fine-tuned and iteratively trained in the document layout detection fine-tuning stage based on a low learning rate to obtain the target sampling network model.
[0065] Therefore, since the target sampling network model has learned the mapping relationship between the document image and the sampling feature map corresponding to the text region image in the document image during the document layout detection fine-tuning stage, when processing one or more text region images in the document image based on the target sampling network model, it can realize the extraction of sampling feature map based on the fine-tuned target sampling network model, so as to effectively improve the accuracy of subsequent layout information extraction and effectively enhance the expression accuracy of sampling feature map.
[0066] S203: Determine the layout information corresponding to the text region image based on the sampled feature map and at least one feature point.
[0067] For a detailed description of S203, please refer to the above embodiments, which will not be repeated here.
[0068] In this embodiment, since the target sampling network model has learned the mapping relationship between the document image and the sampling feature maps corresponding to the text region images in the document image during the document layout detection fine-tuning stage, when processing one or more text region images in the document image based on the target sampling network model, it can extract sampling feature maps based on the fine-tuned target sampling network model. This effectively improves the accuracy of subsequent layout information extraction and enhances the expressive accuracy of the sampling feature maps. It can effectively improve the accuracy of layout information determination and significantly improve the detection effect for ambiguous layout information.
[0069] Figure 3 This is a schematic diagram according to the third embodiment of the present disclosure.
[0070] like Figure 3 As shown, the method for determining this layout information includes:
[0071] S301: During the document layout detection and fine-tuning stage, sample images are acquired, including: sample region images, which have corresponding labeled feature maps.
[0072] The document layout detection fine-tuning stage refers to the model training stage where the target sampling network model is fine-tuned.
[0073] For example, the initial sampling network model can be initialized based on the target model parameters determined in the general model pre-training stage to obtain the sampling network model to be trained. Then, in the document layout detection fine-tuning stage, the sampling network model to be trained is fine-tuned and iteratively trained based on a low learning rate to obtain the target sampling network model.
[0074] Among them, the sample region image refers to a local region image of text that may contain specific layout information, which is obtained by pre-identifying the sample image. The sample region image can also be interpreted as the document region image used to train the target sampling network model. The sample region image can be a document image identified from the sample image and used to train the target sampling network model, and can be called a sample image.
[0075] The sample image may include text, images, seals, signatures, etc. with specific layout information, and there are no restrictions on this.
[0076] Among them, the feature map used to iteratively train the sampled network model to be trained and to determine whether the sampled network model has converged during the iterative training process, and which serves as a reference, can be called the labeled feature map.
[0077] S302: Obtain the sampled network model to be trained.
[0078] The sampling network model to be trained can be referred to as the sampling network model to be trained. The sampling network model has the function of extracting visual features from document images. The sampling network model to be trained can be a network model in artificial intelligence, such as a machine learning model or a neural network model, without any restrictions.
[0079] In the embodiments of this disclosure, the sampling network model to be trained can be obtained by initializing the initial sampling network model based on the target model parameters determined in the general model pre-training stage.
[0080] The target model parameters refer to the pre-determined optimal network model parameters, such as the weights, number of connection layers, number of pooling layers, number of convolutional layers, etc., in the network model, and there are no restrictions on these parameters.
[0081] In some embodiments of this disclosure, when performing the step of obtaining the sampled network model to be trained, an initial sampled network model can be obtained, and target model parameters can be obtained. The target model parameters are determined based on sample images during the general model pre-training stage. The initial sampled network model is configured according to the target model parameters to obtain the sampled network model to be trained. This enables fine-tuning of the sampled network model to be trained based on a lower learning rate during the document layout detection fine-tuning stage. This allows the AI model required for layout information determination to be endowed with both "large-scale" and "pre-trained" attributes, which can greatly improve the generalization, versatility, and practicality of the target sampled network model. This enables accurate extraction of sampled feature maps based on the fine-tuned target sampled network model.
[0082] S303: Input the sample image into the sampling network model to be trained, and obtain the predicted feature map output by the sampling network model to be trained, which corresponds to the sample region image in the sample image.
[0083] In the document layout detection fine-tuning stage, sample region images are obtained from the sample images. These sample region images have corresponding labeled feature maps. After initializing the sample network model to be trained, the sample images can be directly input into the sample network model to be trained, and the predicted feature maps output by the sample network model to be trained can be obtained.
[0084] For example, the sample image has dimensions (H, W, 3), where H represents the height of the sample image, W represents the width of the sample image, and 3 represents the three color channels in the RGB color mode. The RGB color mode is an industry color standard that uses the red (R), green (G), and blue (B) color channels to downsample the sample image (the downsampling rate can be 1 / 4) to obtain a sampled feature map corresponding to each text region image. The sampled feature map can be represented as (H / 4, W / 4, 128), where the height of the sampled feature map is H / 4, the width is W / 4, and 128 represents 128 dimensions of features. In addition to the color features of the RGB color mode, it can also include edge information features, depth features, etc., without any restrictions.
[0085] S304: Iteratively train the sampling network model to be trained based on the labeled feature map and the predicted feature map until it is determined that the sampling network model obtained by iterative training satisfies the convergence condition, and use the sampling network model obtained by iterative training as the target sampling network model.
[0086] In the document layout detection fine-tuning stage, sample region images are obtained from the sample images. The sample region images have corresponding labeled feature maps. After initializing the sample network model to be trained, the sample images can be directly input into the sample network model to be trained, and the predicted feature maps output by the sample network model to be trained can be obtained. Then, the sample network model to be trained can be iteratively trained based on the labeled feature maps and the predicted feature maps.
[0087] For example, the loss value between the labeled feature map and the predicted feature map can be determined. If the loss value is less than the loss threshold, it is determined that the sampled network model obtained by iterative training meets the convergence condition, and the sampled network model obtained by iterative training is used as the target sampled network model. If the loss value is greater than or equal to the loss threshold, it is determined that the sampled network model obtained by iterative training does not meet the convergence condition. Then, the target model parameters of the sampled network model obtained by iterative training are continuously fine-tuned until the sampled network model obtained by iterative training meets the convergence condition, and the sampled network model obtained by iterative training is used as the target sampled network model.
[0088] Therefore, in this embodiment of the present disclosure, during the document layout detection fine-tuning stage, sample images are acquired, including sample region images with corresponding labeled feature maps. A sampling network model to be trained is also acquired. The sample images are input into the sampling network model to be trained, and the predicted feature maps output by the sampling network model corresponding to the sample region images in the sample images are obtained. The sampling network model to be trained is iteratively trained based on the labeled feature maps and predicted feature maps until the iteratively trained sampling network model satisfies the convergence condition. The iteratively trained sampling network model is then used as the target sampling network model. This enables rapid downsampling of sample images during online layout information determination to determine the sampling feature maps corresponding to each text region image in the sample images. Furthermore, during the document layout detection fine-tuning stage, the sampling network model to be trained can be fine-tuned based on a lower learning rate, resulting in higher training efficiency for the sampling network model.
[0089] S305: Acquire a document image, wherein the document image includes: a text region image.
[0090] S306: Input the text region image into the target sampling network model and obtain the sampling feature map output by the target sampling network model, wherein the sampling feature map includes at least one feature point.
[0091] S307: Determine the layout information corresponding to the text region image based on the sampled feature map and at least one feature point.
[0092] For a detailed description of S305-S307, please refer to the above embodiments, which will not be repeated here.
[0093] In this embodiment, since the target sampling network model has learned the mapping relationship between text region images and sampling feature maps during the document layout detection fine-tuning stage, when processing one or more text region images in a document image based on the target sampling network model, it can extract sampling feature maps based on the fine-tuned target sampling network model. This effectively improves the accuracy of subsequent layout information extraction and enhances the expressive accuracy of the sampling feature maps. It can effectively improve the accuracy of layout information determination, and significantly improve detection performance for ambiguous layout information. By assigning the AI model required for layout information determination with both "large-scale" and "pre-trained" attributes, it can significantly improve the generalization, versatility, and practicality of the target sampling network model, enabling accurate extraction of sampling feature maps based on the fine-tuned target sampling network model. It can support rapid downsampling of sample images during online layout information determination to determine the sampling feature map corresponding to each text region image in the sampled image. It also supports fine-tuning the sample network model to be trained based on a low learning rate during the document layout detection fine-tuning stage, thus achieving high training efficiency for the sampling network model.
[0094] Figure 4 This is a schematic diagram according to the fourth embodiment of the present disclosure.
[0095] like Figure 4 As shown, the method for determining this layout information includes:
[0096] S401: Obtain a document image, wherein the document image includes: a text region image.
[0097] For a detailed description of S401, please refer to the above embodiments, which will not be repeated here.
[0098] S402: Input the document image into the target sampling network model and obtain the sampling feature map corresponding to the text region image output by the target sampling network model. The sampling feature map includes multiple feature points.
[0099] Among them, the target sampling network model has learned the mapping relationship between the document image and the sampling feature map corresponding to the text region image in the document image during the document layout detection fine-tuning stage.
[0100] The target sampling network model can be used to downsample each text region image in a document image to output a sampled feature map corresponding to the text region image.
[0101] The target sampling network model in this embodiment can be configured and trained on a lightweight initial backbone network model based on a pre-trained backbone network model, i.e., a visual feature extraction model, during the document layout detection fine-tuning stage. The pre-trained backbone network model can be trained during the general model pre-training stage, as detailed in the following embodiments.
[0102] The document layout detection fine-tuning stage refers to the model training stage where the target sampling network model is fine-tuned.
[0103] For example, the target model can be initialized based on the target model parameters determined in the pre-training stage of the general model. Then, in the document layout detection fine-tuning stage, the target model can be fine-tuned and iteratively trained based on a low learning rate to obtain the target sampling network model.
[0104] In this embodiment of the disclosure, two model branches can be added to the target sampling network model, including a text region prediction branch and a text position regression branch. Each branch represents a branch model. The branch model used for text region prediction can be called the target text region prediction model, and the branch model used for text position regression can be called the target text position regression model.
[0105] In other words, in this embodiment of the present disclosure, it is also supported to jointly train the target sampling network model and its two model branches during the document layout detection fine-tuning stage. After training, the sampling feature map output by the target sampling network model can be input into different model branches respectively. The mask region (candidate text box) of the text region image is predicted based on the branch model used for text region prediction, and the distance from the feature point in the text box to the four vertices of the text box is regressed based on the branch model used for text position regression. The regressed distance can be called position offset information.
[0106] In this embodiment of the disclosure, based on the model function of the branch model for text region prediction, it can predict the candidate text box to which each feature point belongs based on the sampled feature map. In the document layout detection fine-tuning stage, the sample feature map can be input into the branch model for text region prediction. A block of labeled detection boxes with marked layout categories in the sample feature map can be masked. Then, the processed sample feature map is input into the branch model for text region prediction. Based on the branch model for text region prediction, the masked region of the processed sample feature map is predicted to obtain the layout category and candidate position information of the candidate text box. Then, the branch model for text region prediction is iteratively trained based on the layout category and candidate position information of the candidate text box, the labeled layout category, and the labeled position information of the labeled detection box.
[0107] S403: Input the sampled feature map into the target text region prediction model, and obtain the layout category and candidate position information of the candidate text box to which each feature point output by the target text region prediction model belongs.
[0108] Among them, the target text region prediction model has learned the mapping relationship between each feature point in the sampled feature map, the layout category of the candidate text box to which it belongs, and the candidate position information.
[0109] In other words, it can support the joint training of the target sampling network model and its two model branches during the document layout detection fine-tuning stage. After training, the sampling feature maps output by the target sampling network model can be input into the target text region prediction model, and the layout category and candidate position information of the candidate text box to which each feature point belongs can be obtained. This can effectively improve the prediction efficiency and accuracy of the layout category and candidate position information of the candidate text box to which each feature point belongs, and can significantly improve the detection effect for ambiguous layout categories.
[0110] For example, a target text region prediction model can contain three repeating unit structures. Each unit structure includes: a convolutional layer -> a regularization layer -> an activation function (ReLU, representing the name of an activation function) layer. The number of channels in the convolutional layers in different unit structures are 64, 64, and 128, respectively. The last layer is a convolutional layer that connects to each unit structure. The number of channels output by the last convolutional layer includes: the number of layout categories + 1 (1 represents the background class). The output feature map dimension is represented as (H / 4, W / 4, number of categories + 1), carrying the layout category and candidate position information of the candidate text boxes.
[0111] S404: Input the sampled feature map and the candidate position information of the candidate text box into the target text position regression model, and obtain the position offset information between each feature point output by the target text position regression model and the candidate text box.
[0112] Among them, the target text position regression model has learned the mapping relationship between each feature point in the sampled feature map, the candidate position information of the candidate text box, and the position offset information between the feature point and the candidate text box.
[0113] In other words, it can support the joint training of the target sampling network model and its two model branches during the document layout detection fine-tuning stage. After training, the sampled feature maps output by the target sampling network model can be input into the target text position regression model to obtain the position offset information between each feature point output by the target text position regression model and the candidate text box. This can effectively improve the prediction efficiency and accuracy of the position offset information between each feature point and the candidate text box, making the predicted position offset information have high reference value.
[0114] For example, a target text location regression model can contain three repeating unit structures. Each unit structure includes: a convolutional layer -> a regularization layer -> an activation function (ReLU, representing the name of an activation function) layer. The number of channels in the convolutional layers in different unit structures are 64, 64, and 128, respectively. The last layer is a convolutional layer that connects to each unit structure. The number of channels output by the last convolutional layer is 8 (assuming the candidate text box contains four vertices, the horizontal and vertical offset distances from the feature point to one of the vertices form two offset distances, and correspondingly, four vertices correspond to eight offset distances). The output feature map has dimensions of (H / 4, W / 4, 8) and contains the positional offset information between the feature point and the candidate text box.
[0115] Therefore, in this embodiment, the target model parameters are determined during the general model pre-training stage, and the initial sampling network model is initialized to obtain the sampling network model to be trained. Then, in the document layout detection fine-tuning stage, the sampling network model to be trained is fine-tuned and iteratively trained based on a low learning rate to obtain the target sampling network model. In the document layout detection fine-tuning stage, the target sampling network model and its two model branches are jointly trained. After training, the sampling feature map output by the target sampling network model can be input into different model branches. The branch model for text region prediction predicts the mask region (candidate text box) of the text region image, and the branch model for text position regression regresses the distance from the feature points within the text box to the four vertices of the text box. The regressed distance can be called the position offset information. The training label of the branch model for text region prediction can be generated based on the text contour in the sample image. The mask value of the layout category is generated within the text contour. This branch can be iteratively optimized using the cross-entropy loss function. The branch model used for text position regression can supervise only the position regression values (position offset information) of candidate text boxes (positive samples, i.e., candidate text boxes in which feature points are located) of positive samples. Each feature point has 8 output channels, which correspond to the position offset information of the feature point from the four vertices of the candidate text box.
[0116] S405: Based on the position offset information, determine the target text box from multiple candidate text boxes, and use the layout category of the target text box and the candidate position information together as layout information.
[0117] For example, by combining the regression value (position offset information) of each feature point (i.e., pixel) in the sampled feature map with the layout category, the target text box is selected from multiple candidate text boxes by using the Non-Maximum Suppression (NMS) algorithm. The layout category and candidate position information of the target text box are used as the layout information of the detected text region image. If there are multiple text region images, the output value of the detection result can be represented as (N, 4, 2), where N represents the number of target detection boxes, 4 represents that the target detection box has four vertices, and 2 represents the position information of each vertex, which can be represented by two coordinate values, such as the X coordinate value and the Y coordinate value, without restriction.
[0118] like Figure 5 As shown, Figure 5 This is a schematic diagram of the model structure for the document layout detection and fine-tuning stage in this embodiment of the disclosure. Wherein, Figure 5 The document image contains one or more text region images. The document image can be input into a target sampling network model and two branch models: a target text region prediction model and a target text position regression model. The target sampling network model processes the document image to obtain a sampling feature map. Then, the sampling feature map is provided to the target text region prediction model and the target text position regression model respectively. The target text region prediction model and the target text position regression model process the sampling feature map respectively, and fuse the output of the two branch models to obtain the layout information corresponding to each text region image in the document image.
[0119] In this embodiment, since the target sampling network model has learned the mapping relationship between text region images and sampling feature maps during the document layout detection fine-tuning stage, when processing one or more text region images in a document image based on the target sampling network model, it can extract sampling feature maps based on the fine-tuned target sampling network model. This effectively improves the accuracy of subsequent layout information extraction and enhances the expressive accuracy of the sampling feature maps. During the document layout detection fine-tuning stage, the target sampling network model and its two model branches can be jointly trained. After training, the sampling feature maps output by the target sampling network model can be input into the target text region prediction model, obtaining the layout category and candidate position information of the candidate text boxes to which each feature point belongs. This effectively improves the prediction efficiency and accuracy of the layout category and candidate position information of the candidate text boxes to which each feature point belongs, significantly improving the detection effect for ambiguous layout categories. During the document layout detection fine-tuning stage, it supports joint training of the target sampling network model and its two model branches. After training, the sampled feature maps output by the target sampling network model can be input into the target text position regression model to obtain the position offset information between each feature point output by the target text position regression model and the candidate text box. This can effectively improve the prediction efficiency and accuracy of the position offset information between each feature point and the candidate text box, making the predicted position offset information have high reference value.
[0120] This disclosure also provides a technical solution for determining target model parameters based on sample images during the general model pre-training stage. The determined target model parameters are used to initialize the initial sampling network model in the document layout detection fine-tuning stage. This can be achieved by obtaining a reference model to be trained during the general model pre-training stage, and determining sample text line information from sample images, where each sample image has corresponding labeled text line information. The sample images are then segmented to obtain multiple sample image blocks, where each labeled text line information has corresponding labeled image block information. The reference model to be trained is iteratively trained based on the sample text line information, multiple sample image blocks, labeled text line information, and labeled image block information until the iteratively trained reference model meets the convergence condition. The model parameters of the iteratively trained reference model are then used as the target model parameters. This allows for accurate modeling of the target model parameters during the general model pre-training stage and supports subsequent initialization of the initial sampling network model based on the target model parameters during the document layout detection fine-tuning stage. This effectively improves the modeling accuracy of the target model parameters and enhances the overall layout information extraction accuracy.
[0121] The model parameters of the reference model to be trained obtained from iterative training can be, for example, parameters such as weights, number of connection layers, number of pooling layers, and number of convolutional layers in the network model, and there are no restrictions on them.
[0122] The model trained in the general model pre-training stage can be called the reference model to be trained, and the model parameters of the reference model to be trained obtained by iterative training can be used as the target model parameters.
[0123] The sample text line information can refer to information related to the text lines contained in the sample image, such as the text line number, size, and position of the text line in the sample image.
[0124] Among them, a sample image block refers to an image block obtained by segmenting a sample image.
[0125] Among them, the labeled text line information and labeled image patch information can be reference annotation information used to determine the convergence timing of the reference model to be trained. The labeled text line information includes, for example, visual features and text embedding features obtained by annotating sample text lines. The labeled image patch information includes, for example, information such as the size, sequence number, encoding, and position of image patches obtained by annotating sample image patches.
[0126] Figure 6 This is a schematic diagram according to the fifth embodiment of the present disclosure.
[0127] like Figure 6 As shown, the method for determining this layout information includes:
[0128] S601: In the general model pre-training stage, obtain the reference model to be trained, which includes: language recognition sub-model, visual feature extraction sub-model, and image-text alignment sub-model.
[0129] The model to be trained, used to determine the parameters of the target model, can be called the reference model to be trained. This reference model to be trained can be a network model in artificial intelligence, such as a machine learning model or a neural network model, without any restrictions.
[0130] The reference model to be trained can include three parts: a language recognition sub-model, a visual feature extraction sub-model, and an image-text alignment sub-model. The language recognition sub-model is used to identify text features related to text lines in the sample image. The visual feature extraction sub-model is used to extract visual features related to text lines in the sample image. The image-text alignment sub-model is used to align the text features and visual features related to text lines. The aligned features can be called target image-text alignment features.
[0131] S602: Determine sample text line information from the sample image, wherein the sample image has corresponding labeled text line information.
[0132] Among them, one or more text lines identified from the sample image can be called sample text lines. The information used to describe the sample text lines can be called sample text line information. Sample text line information can refer to information related to the text lines contained in the sample image, such as the text line number, size, and position of the text line in the sample image.
[0133] S603: Segment the sample image to obtain multiple sample image blocks, wherein the labeled text line information has corresponding labeled image block information.
[0134] Here, a sample image patch refers to an image patch obtained by segmenting a sample image. Annotated text lines and labeled image patch information can be reference annotation information used to determine the convergence timing of the reference model to be trained. Annotated text lines include visual features and text embedding features obtained by annotating sample text lines. Annotated image patch information includes information such as the size, sequence number, encoding, and position of the image patch obtained by annotating sample image patches.
[0135] The above process involves determining sample text line information from sample images and segmenting the sample images to obtain multiple sample image blocks, which can be used for alignment processing. The visual features of the aligned target images, the labeled text line information, and the labeled image block information can be used for iterative training of the reference model to be trained.
[0136] S604: Input the sample text line information into the language recognition sub-model and obtain the text embedding features output by the language recognition sub-model.
[0137] After determining the sample text line information from the sample image, the sample text line information can be input into the language recognition sub-model to obtain the text embedding features output by the language recognition sub-model. The text embedding features refer to the text features represented by real-valued vectors based on strings.
[0138] The iterative training process of the language recognition sub-model can be called Masked Language Modeling (MLM). This means that the sample image may contain some masked text lines, and the recognized sample text line information is the unmasked part. The unmasked part is input into the language recognition sub-model so that the language recognition sub-model models the text embedding features of the entire sample text line. The labeled text line information includes the information of the complete sample text line in the document image. Thus, the ability of the language recognition sub-model to predict the text embedding features of any sample text line in the sample image can be iteratively trained.
[0139] S605: Input the sample image block into the visual feature extraction sub-model and obtain the initial image visual features output by the visual feature extraction sub-model.
[0140] After segmenting the sample image to obtain multiple sample image blocks, the sample image blocks can be input into the visual feature extraction sub-model to obtain the initial image visual features output by the visual feature extraction sub-model. The initial image visual features refer to the visual dimension features of the image obtained by extracting visual features from each sample image block.
[0141] The iterative training process of the visual feature extraction sub-model can be called Mask Image Modeling (MIM). This means that the sample image may contain some masked sample image patches and some unmasked sample image patches. All sample image patches are input into the visual feature extraction sub-model so that the visual feature extraction sub-model models the overall initial image visual features of all sample image patches. The labeled image patch information includes the information of all sample image patches in the document image. Thus, the visual feature extraction sub-model can be iteratively trained to predict the initial image visual features of any sample image patch in the sample image.
[0142] S606: Process text embedding features to obtain target text features, and process initial image visual features to obtain target image visual features.
[0143] In this embodiment, after extracting the text embedding features and the initial image visual features, the text embedding features and the initial image visual features can be optimized respectively to effectively expand the feature expression dimension of the text embedding features and the initial image visual features. The text embedding features can also be processed to obtain the target text features, and the initial image visual features can be processed to obtain the target image visual features.
[0144] In some embodiments of this disclosure, in order to effectively expand the feature representation dimension of the text embedding feature and represent the position of the sample text line, when processing the text embedding feature to obtain the target text feature, the position encoding information of the sample text line can be determined based on the sample text line information, and the text embedding feature and the position encoding information can be concatenated to obtain the target text feature.
[0145] For example, document images can first be recognized using OCR technology to determine sample text line information from the document images. Then, the sample text line information can be processed using a Bidirectional Encoder Representation from Transformers (BERT) sub-model to obtain high-dimensional text embedding features. One-dimensional (1D) and two-dimensional (2D) positional codes are added to the text embedding features. The 1D and 2D positional codes are used together as the positional coding information of the sample text lines. The 1D positional code is, for example, the sequence number of the sample text lines, and the 2D positional code is, for example, the geometric information of the sample text lines.
[0146] In some embodiments of this disclosure, in order to effectively expand the feature representation dimension of the initial image visual features of the sample text line and represent the position of the sample image block, when processing the initial image visual features to obtain the target image visual features, the position encoding information corresponding to the sample image block can also be determined, and the initial image visual features and the position encoding information can be concatenated to obtain the target image visual features.
[0147] For example, a document image (H, W, 3) is divided into (PxP) patches. Multiple sample image patches can form a sequence. A fully connected layer is used to obtain high-dimensional initial image visual features of the sample image patches. In order to represent the position of the sample image patches, 1D position coordinates are added to the initial image visual features (the 1D position coordinates are the position encoding information corresponding to the sample image patches).
[0148] S607: Align the target text features and the target image visual features based on the image-text alignment sub-model to obtain the target image-text alignment features.
[0149] After obtaining the target text features and target image visual features, the target text features and target image visual features can be aligned based on the image-text alignment sub-model, and the aligned features can be used as the target image-text alignment features.
[0150] Alignment, for example, can be achieved by matching the predicted target text features with the visual features of the target image in the visual dimension, so that the matched target text features and the visual features of the target image correspond to the same text content.
[0151] The process of iteratively training the image-image alignment sub-model can be called WordPatch Alignment (WPA).
[0152] S608: Iteratively train the reference model to be trained based on the loss values between the target image-text alignment features, the labeled text line information, and the labeled image block information, until the reference model to be trained obtained from the iterative training meets the convergence condition.
[0153] In other words, in this embodiment, the three sub-models are iteratively trained with different pre-training objectives. During the training process, a self-supervised approach can be adopted, including: Masked Language Modeling (MLM), Masked Image Modeling (MIM), and Text-Image Patch Alignment (WPA).
[0154] For example, such as Figure 7 As shown, Figure 7 This is a schematic diagram of the model structure during the pre-training stage of the general model in this embodiment. It includes: Masked Language Modeling (MLM), Masked Image Modeling (MIM), and Text-Image Patch Alignment (WPA). During the pre-training stage of the general model, effective representations can be learned through three self-supervised methods. Examples of these three self-supervised methods are illustrated below:
[0155] For masked language modeling (MLM), 30% of the sample text lines in the sample image are randomly masked. The position information of the masked sample text lines is used to reconstruct the text embedding features of the masked sample text lines based on the position information of the unmasked sample text lines and their layout information.
[0156] For masked image modeling (MIM), 40% of the sample image blocks in the sample image are randomly masked, and the token encoding (i.e., the initial image visual features) corresponding to the masked sample image blocks is restored based on the content of the unmasked sample image blocks.
[0157] For text-image patch alignment (WPA), each sample text line in a sample image can be aligned to a sample image patch. In order to learn the alignment relationship between text and image patches, the image-text alignment sub-model can predict whether the sample image patch corresponding to the sample text line is masked, so as to learn the fine-grained alignment features between text and image.
[0158] S609: If the reference model to be trained obtained from iterative training satisfies the convergence condition, the model parameters of the visual feature extraction sub-model are used as the target model parameters.
[0159] In other words, the layout information extraction in this embodiment is based on a pre-trained model, which includes a general model pre-training stage and a document layout detection fine-tuning stage. In the general model pre-training stage, a backbone network of a visual feature extraction sub-model with strong representational capabilities is obtained. Then, in the document layout detection fine-tuning stage, the weight parameters (i.e., target model parameters) of the backbone network of the visual feature extraction sub-model are loaded, achieving a lightweight model with fast response speed and good performance.
[0160] S610: In the document layout detection fine-tuning stage, the initial sampling network model is initialized based on the target model parameters. The initial sampling network model to be trained is used to generate the target sampling network model, which is used to determine the sampling feature map corresponding to the text region image.
[0161] For a detailed description of S610, please refer to the above embodiments, which will not be repeated here.
[0162] In this embodiment, layout information extraction is performed based on a pre-trained model. The pre-trained model includes a general model pre-training stage and a document layout detection fine-tuning stage. In the general model pre-training stage, a backbone network of a visual feature extraction sub-model with strong representational capabilities is obtained. Then, in the document layout detection fine-tuning stage, the weight parameters (i.e., target model parameters) of the backbone network of the visual feature extraction sub-model are loaded, achieving a lightweight model with fast response speed and good performance. The layout information determination method based on the pre-trained model can effectively improve the accuracy of layout information extraction and significantly improve detection performance for ambiguous layout categories. In this embodiment, the device implementing the layout information determination method can be used as an internal device of a text-to-image converter, where the extracted layout information can effectively support downstream recognition tasks. Alternatively, it can be used as a separate layout information determination device to directly extract layout information. Compared to heuristic rule-based methods and deep learning-based general object detection methods, the layout information determination method in this embodiment has higher recall and better robust generalization performance. Compared to methods that use multiple auxiliary information fusion, this method requires fewer model parameters, resulting in a faster model response and effectively improving the efficiency of layout information determination.
[0163] Figure 8 This is a schematic diagram according to the sixth embodiment of the present disclosure.
[0164] like Figure 8 As shown, the layout information determining device 80 includes:
[0165] The first acquisition module 801 is used to acquire a document image, wherein the document image includes: a text region image;
[0166] The second acquisition module 802 is used to acquire a sampled feature map corresponding to the text region image, wherein the sampled feature map includes: at least one feature point; and
[0167] The determination module 803 is used to determine the layout information corresponding to the text region image based on the sampled feature map and at least one feature point.
[0168] In some embodiments of this disclosure, the second acquisition module 802 is specifically used for:
[0169] The document image is input into the target sampling network model, and the sampled feature map corresponding to the text region image is obtained from the output of the target sampling network model;
[0170] Among them, the target sampling network model has learned the mapping relationship between the document image and the sampling feature map corresponding to the text region image in the document image during the document layout detection fine-tuning stage.
[0171] In some embodiments of this disclosure, the number of feature points is multiple; such as Figure 9 As shown, Figure 9 The schematic diagram is based on the seventh embodiment of this disclosure. The layout information determining device 90 includes: a first acquisition module 901, a second acquisition module 902, and a determining module 903, wherein the determining module 903 includes:
[0172] The first determining submodule 9031 is used to determine the layout category and candidate position information of the candidate text box to which each feature point belongs based on the sampled feature map;
[0173] The second determining submodule 9032 is used to determine the positional offset information between the feature point and the candidate text box based on the sampled feature map and the candidate position information of the candidate text box; and
[0174] The third determination submodule 9033 is used to determine the target text box from multiple candidate text boxes based on the position offset information, and to use the layout category of the target text box and the candidate position information together as layout information.
[0175] In some embodiments of this disclosure, a target text region prediction model is connected after the target sampling network model, and the target sampling network model is used to determine the sampling feature map corresponding to the text region image; wherein, the first determining submodule 9031 is specifically used for:
[0176] The sampled feature map is input into the target text region prediction model, and the layout category and candidate position information of the candidate text box to which each feature point belongs are obtained from the output of the target text region prediction model. The target text region prediction model has learned the mapping relationship between each feature point in the sampled feature map, the layout category of the candidate text box to which it belongs, and the candidate position information.
[0177] In some embodiments of this disclosure, a target text location regression model is further connected after the target sampling network model; wherein, the second determination submodule 9032 is specifically used for:
[0178] The sampled feature map and the candidate position information of the candidate text box are input into the target text position regression model, and the position offset information between each feature point output by the target text position regression model and the candidate text box is obtained.
[0179] Among them, the target text position regression model has learned the mapping relationship between each feature point in the sampled feature map, the candidate position information of the candidate text box, and the position offset information between the feature point and the candidate text box.
[0180] In some embodiments of this disclosure, the device 90 further includes:
[0181] The first training module 904 is used to train the target sampling network model based on the following method:
[0182] In the document layout detection and fine-tuning stage, sample images are acquired, including: sample region images, which have corresponding labeled feature maps;
[0183] Obtain the sampling network model to be trained;
[0184] The sample image is input into the sampling network model to be trained, and the predicted feature map output by the sampling network model corresponding to the sample region image in the sample image is obtained; and
[0185] The sampling network model to be trained is iteratively trained based on the labeled feature map and the predicted feature map until it is determined that the sampling network model obtained by iterative training meets the convergence condition. The sampling network model obtained by iterative training is then used as the target sampling network model.
[0186] In some embodiments of this disclosure, the first training module 904 is specifically used for:
[0187] Obtain the initial sampling network model;
[0188] Obtain the target model parameters, which are determined based on sample images during the pre-training phase of the general model; and
[0189] Configure the initial sampling network model according to the target model parameters to obtain the sampling network model to be trained.
[0190] In some embodiments of this disclosure, the device 90 further includes:
[0191] The second training module 905 is used to determine the target model parameters during the general model pre-training phase based on the following method:
[0192] In the pre-training phase of the general model, a reference model to be trained is obtained;
[0193] From the sample image, determine the sample text line information, where the sample image has corresponding labeled text line information;
[0194] The sample image is segmented to obtain multiple sample image blocks, where the labeled text line information has corresponding labeled image block information;
[0195] The reference model to be trained is iteratively trained based on sample text line information, multiple sample image patches, labeled text line information, and labeled image patch information; and
[0196] The training continues until the reference model obtained through iterative training meets the convergence condition, at which point the model parameters of the reference model obtained through iterative training are used as the target model parameters.
[0197] In some embodiments of this disclosure, the reference model to be trained includes: a language recognition sub-model, a visual feature extraction sub-model, and an image-text alignment sub-model;
[0198] The second training module 905 is specifically used for:
[0199] The sample text line information is input into the language recognition sub-model, and the text embedding features output by the language recognition sub-model are obtained;
[0200] The sample image patch is input into the visual feature extraction sub-model, and the initial image visual features output by the visual feature extraction sub-model are obtained.
[0201] Process the text embedding features to obtain the target text features, and process the initial image visual features to obtain the target image visual features;
[0202] Aligning target text features and target image visual features based on an image-text alignment sub-model yields target image-text alignment features; and
[0203] The reference model to be trained is iteratively trained based on the loss values between the target image-text alignment features, the labeled text line information, and the labeled image patch information, until the reference model to be trained obtained by iterative training satisfies the convergence condition.
[0204] In some embodiments of this disclosure, the second training module 905 is further configured to:
[0205] Based on the information in the sample text line, determine the position encoding information of the sample text line;
[0206] The text embedding features and positional encoding information are concatenated to obtain the target text features.
[0207] In some embodiments of this disclosure, the second training module 905 is further configured to:
[0208] Determine the location encoding information corresponding to the sample image block;
[0209] The visual features of the initial image and the location encoding information are concatenated to obtain the visual features of the target image.
[0210] In some embodiments of this disclosure, the second training module 905 is further configured to:
[0211] If the reference model to be trained obtained from iterative training satisfies the convergence condition, the model parameters of the visual feature extraction sub-model are used as the target model parameters.
[0212] It should be noted that the foregoing explanation of the method for determining layout information also applies to the layout information determining device of this embodiment, and will not be repeated here.
[0213] In this embodiment, by acquiring a document image, which includes a text region image, and acquiring a sampling feature map corresponding to the text region image, wherein the sampling feature map includes at least one feature point, and determining the layout information corresponding to the text region image based on the sampling feature map and at least one feature point, the accuracy of the layout information determination can be effectively improved, and the detection effect can be significantly improved for ambiguous layout information.
[0214] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0215] Figure 10 A schematic block diagram of an example electronic device that can be used to implement the layout information determination method of embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0216] like Figure 10As shown, device 1000 includes a computing unit 1001, which can perform various appropriate actions and processes according to a computer program stored in read-only memory (ROM) 1002 or a computer program loaded into random access memory (RAM) 1003 from storage unit 1008. The RAM 1003 may also store various programs and data required for the operation of device 1000. The computing unit 1001, ROM 1002, and RAM 1003 are interconnected via bus 1004. Input / output (I / O) interface 1005 is also connected to bus 1004.
[0217] Multiple components in device 1000 are connected to I / O interface 1005, including: input unit 1006, such as keyboard, mouse, etc.; output unit 1007, such as various types of monitors, speakers, etc.; storage unit 1008, such as disk, optical disk, etc.; and communication unit 1009, such as network card, modem, wireless transceiver, etc. Communication unit 1009 allows device 1000 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0218] The computing unit 1001 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1001 performs the various methods and processes described above, such as the layout information determination method. For example, in some embodiments, the layout information determination method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 1008. In some embodiments, part or all of the computer program may be loaded and / or installed on device 1000 via ROM 1002 and / or communication unit 1009. When the computer program is loaded into RAM 1003 and executed by the computing unit 1001, one or more steps of the layout information determination method described above may be performed. Alternatively, in other embodiments, the computing unit 1001 may be configured to perform the layout information determination method by any other suitable means (e.g., by means of firmware).
[0219] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0220] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0221] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0222] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0223] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), the Internet, and blockchain networks.
[0224] Computer systems can include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. A server can be a cloud server, also known as a cloud computing server or cloud host, a hosting product within the cloud computing service ecosystem, addressing the shortcomings of traditional physical hosts and VPS (Virtual Private Server, or simply "VPS") services, such as high management difficulty and weak business scalability. Servers can also be servers for distributed systems or servers incorporating blockchain technology.
[0225] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0226] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A method for determining layout information, the method comprising: Acquire a document image, wherein the document image includes: a text region image; The document image is input into a target sampling network model, and a sampling feature map corresponding to the text region image is obtained from the output of the target sampling network model. The sampling feature map includes at least one feature point; the target sampling network model has learned the mapping relationship between the document image and the sampling feature map corresponding to the text region image in the document image during the document layout detection fine-tuning stage; and Based on the sampled feature map and the at least one feature point, determine the layout information corresponding to the text region image; The number of feature points is multiple; The step of determining the layout information corresponding to the text region image based on the sampled feature map and the at least one feature point includes: Based on the sampled feature map, determine the layout category and candidate position information of the candidate text box to which each feature point belongs; Based on the sampled feature map and the candidate position information of the candidate text boxes, determine the positional offset information between the feature points and the candidate text boxes; and Based on the position offset information, a target text box is determined from among the candidate text boxes, and The layout category and candidate position information of the target text box are used together as the layout information; A target text region prediction model is connected after the target sampling network model, wherein the target sampling network model is used to determine the sampling feature map corresponding to the text region image; The step of determining the layout category and candidate position information of the candidate text box to which each feature point belongs based on the sampled feature map includes: The sampled feature map is input into the target text region prediction model, and the layout category and candidate position information of the candidate text box to which each feature point belongs are obtained from the output of the target text region prediction model; The target text region prediction model has learned the mapping relationship between each feature point in the sampled feature map, the layout category of the candidate text box to which it belongs, and the candidate position information. A target text location regression model is also connected after the target sampling network model; The step of determining the positional offset information between the feature point and the candidate text box based on the sampled feature map and the candidate position information of the candidate text box includes: The sampled feature map and the candidate position information of the candidate text box are input into the target text position regression model, and the position offset information between each feature point output by the target text position regression model and the candidate text box is obtained; The target text position regression model has learned the mapping relationship between each feature point in the sampled feature map, the candidate position information of the candidate text box, and the position offset information between the feature point and the candidate text box.
2. The method according to claim 1, wherein, The target sampling network model is trained in the following way: During the document layout detection and fine-tuning stage, sample images are acquired, wherein the sample images include: sample region images, and the sample region images have corresponding labeled feature maps; Obtain the sampling network model to be trained; The sample image is input into the sampling network model to be trained, and the predicted feature map output by the sampling network model corresponding to the sample region image in the sample image is obtained; and The sampling network model to be trained is iteratively trained based on the labeled feature map and the predicted feature map until the sampling network model obtained by iterative training satisfies the convergence condition, and the sampling network model obtained by iterative training is used as the target sampling network model.
3. The method according to claim 2, wherein, The process of obtaining the sampled network model to be trained includes: Obtain the initial sampling network model; Obtain the target model parameters, wherein the target model parameters are determined based on the sample images during the general model pre-training phase; and Configure the initial sampling network model according to the target model parameters to obtain the sampling network model to be trained.
4. The method according to claim 3, wherein, During the pre-training phase of the general model, the parameters of the target model are determined based on the following method: During the pre-training phase of the general model, a reference model to be trained is obtained; From the sample image, determine the sample text line information, wherein the sample image has corresponding labeled text line information; The sample image is segmented to obtain multiple sample image blocks, wherein the labeled text line information has corresponding labeled image block information; The reference model to be trained is iteratively trained based on the sample text line information, the multiple sample image patches, the labeled text line information, and the labeled image patch information; and The training continues until the reference model obtained through iterative training meets the convergence condition, at which point the model parameters of the reference model obtained through iterative training are used as the target model parameters.
5. The method according to claim 4, wherein, The reference model to be trained includes: a language recognition sub-model, a visual feature extraction sub-model, and an image-text alignment sub-model; The step of iteratively training the reference model to be trained based on the sample text line information, the multiple sample image blocks, the labeled text line information, and the labeled image block information until the reference model to be trained obtained through iterative training satisfies the convergence condition includes: The sample text line information is input into the language recognition sub-model, and the text embedding features output by the language recognition sub-model are obtained; The sample image block is input into the visual feature extraction sub-model, and the initial image visual features output by the visual feature extraction sub-model are obtained. The text embedding features are processed to obtain the target text features, and the initial image visual features are processed to obtain the target image visual features; Based on the image-text alignment sub-model, the target text features and the target image visual features are aligned to obtain the target image-text alignment features; and The reference model to be trained is iteratively trained based on the loss value between the target image-text alignment features, the labeled text line information, and the labeled image block information until the reference model to be trained obtained by iterative training satisfies the convergence condition.
6. The method according to claim 5, wherein, The process of processing the text embedding features to obtain the target text features includes: Based on the sample text line information, determine the position encoding information of the sample text line; The target text features are obtained by concatenating the text embedding features and the position encoding information.
7. The method according to claim 5, wherein, The process of processing the initial image visual features to obtain the target image visual features includes: Determine the location encoding information corresponding to the sample image block; The initial image visual features and the location encoding information are concatenated to obtain the target image visual features.
8. The method according to any one of claims 5-7, wherein, The step of using the model parameters of the reference model to be trained obtained from the iterative training as the target model parameters includes: If the reference model to be trained obtained from iterative training satisfies the convergence condition, the model parameters of the visual feature extraction sub-model are used as the target model parameters.
9. A layout information determination device, the device comprising: The first acquisition module is used to acquire a document image, wherein the document image includes: a text region image; The second acquisition module is used to input the document image into a target sampling network model and obtain a sampling feature map output by the target sampling network model corresponding to the text region image, wherein the sampling feature map includes: at least one feature point; the target sampling network model has learned the mapping relationship between the document image and the sampling feature map corresponding to the text region image in the document image during the document layout detection fine-tuning stage; and The determining module is used to determine the layout information corresponding to the text region image based on the sampled feature map and the at least one feature point; The number of the feature points is multiple; The determining module includes: The first determining submodule is used to determine the layout category and candidate position information of the candidate text box to which each feature point belongs, based on the sampled feature map; The second determining submodule is used to determine the positional offset information between the feature point and the candidate text box based on the sampled feature map and the candidate position information of the candidate text box; and The third determining submodule is used to determine the target text box from multiple candidate text boxes based on the position offset information, and to use the layout category and candidate position information of the target text box together as the layout information; A target text region prediction model is connected after the target sampling network model, wherein the target sampling network model is used to determine the sampling feature map corresponding to the text region image; Specifically, the first determining submodule is used for: The sampled feature map is input into the target text region prediction model, and the layout category and candidate position information of the candidate text box to which each feature point belongs are obtained from the output of the target text region prediction model; The target text region prediction model has learned the mapping relationship between each feature point in the sampled feature map, the layout category of the candidate text box to which it belongs, and the candidate position information. A target text location regression model is also connected after the target sampling network model; The second determining submodule is specifically used for: The sampled feature map and the candidate position information of the candidate text box are input into the target text position regression model, and the position offset information between each feature point output by the target text position regression model and the candidate text box is obtained; The target text position regression model has learned the mapping relationship between each feature point in the sampled feature map, the candidate position information of the candidate text box, and the position offset information between the feature point and the candidate text box.
10. The apparatus according to claim 9, further comprising: The first training module is used to train the target sampling network model based on the following method: During the document layout detection and fine-tuning stage, sample images are acquired, wherein the sample images include: sample region images, and the sample region images have corresponding labeled feature maps; Obtain the sampling network model to be trained; The sample image is input into the sampling network model to be trained, and the predicted feature map output by the sampling network model corresponding to the sample region image in the sample image is obtained; and The sampling network model to be trained is iteratively trained based on the labeled feature map and the predicted feature map until the sampling network model obtained by iterative training satisfies the convergence condition, and the sampling network model obtained by iterative training is used as the target sampling network model.
11. The apparatus according to claim 10, wherein, The first training module is specifically used for: Obtain the initial sampling network model; Obtain the target model parameters, wherein the target model parameters are determined based on the sample images during the general model pre-training phase; and Configure the initial sampling network model according to the target model parameters to obtain the sampling network model to be trained.
12. The apparatus of claim 11, further comprising: The second training module is used to determine the parameters of the target model during the pre-training phase of the general model based on the following method: During the pre-training phase of the general model, a reference model to be trained is obtained; From the sample image, determine the sample text line information, wherein the sample image has corresponding labeled text line information; The sample image is segmented to obtain multiple sample image blocks, wherein the labeled text line information has corresponding labeled image block information; The reference model to be trained is iteratively trained based on the sample text line information, the multiple sample image patches, the labeled text line information, and the labeled image patch information; and The training continues until the reference model obtained through iterative training meets the convergence condition, at which point the model parameters of the reference model obtained through iterative training are used as the target model parameters.
13. The apparatus according to claim 12, wherein, The reference model to be trained includes: a language recognition sub-model, a visual feature extraction sub-model, and an image-text alignment sub-model; The second training module is specifically used for: The sample text line information is input into the language recognition sub-model, and the text embedding features output by the language recognition sub-model are obtained; The sample image block is input into the visual feature extraction sub-model, and the initial image visual features output by the visual feature extraction sub-model are obtained. The text embedding features are processed to obtain the target text features, and the initial image visual features are processed to obtain the target image visual features; Based on the image-text alignment sub-model, the target text features and the target image visual features are aligned to obtain the target image-text alignment features; and The reference model to be trained is iteratively trained based on the loss value between the target image-text alignment features, the labeled text line information, and the labeled image block information until the reference model to be trained obtained by iterative training satisfies the convergence condition.
14. The apparatus according to claim 13, wherein, The second training module is also used for: Based on the sample text line information, determine the position encoding information of the sample text line; The target text features are obtained by concatenating the text embedding features and the position encoding information.
15. The apparatus according to claim 13, wherein, The second training module is also used for: Determine the location encoding information corresponding to the sample image block; The initial image visual features and the location encoding information are concatenated to obtain the target image visual features.
16. The apparatus according to any one of claims 13-15, wherein, The second training module is further used for: If the reference model to be trained obtained from iterative training satisfies the convergence condition, the model parameters of the visual feature extraction sub-model are used as the target model parameters.
17. An electronic device comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-8.
18. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-8.
19. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-8.
Citation Information
Patent Citations
OCR-based image analysis method, system and device, and medium
CN111539412A
Text image layout analysis method and device, electronic equipment and storage medium
CN115063820A