Method and device for processing document image and electronic equipment
By using OCR technology to capture text features and position information in document image processing and multimodal pre-training combined with image features, the problem that document images cannot be parsed into fine-grained digitized encoding in the prior art is solved, and efficient extraction and utilization of document images are achieved.
Patent Information
- Application Number
- CN202510197685.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-05-23
AI Technical Summary
The document images formed by existing digital archives after scanning cannot be further parsed into fine-grained digitized encoding such as text, tables and illustrations, resulting in inefficient processing and inability to achieve effective information retrieval and utilization.
The text features and position information in the document are captured through OCR and other technologies, combined with the file image features, and multimodal pre-training is used to perform multimodal pre-training to realize structured information extraction of document images.
The ability to extract multi-dimensional information on document images is significantly improved, so that the structured information obtained through this model can be used more effectively, and the processing efficiency and information retrieval ability of document image files are improved.
Smart Images

Figure CN120032386A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of text processing, and specifically, embodiments of the present application relate to a method, device, and electronic device for processing document images. Background Art
[0002] With the rapid development of information technology, digital archives are increasingly used in various fields. However, existing digital archives will form PDF and other image format files (i.e., document images) after scanning, and these document images cannot be further parsed into fine-grained digital codes such as text, tables, and illustrations, which leads to low processing efficiency of digital archives and inability to achieve effective information retrieval and utilization.
[0003] That is to say, the existing technology will scan digital archives to obtain image format files or PDF files (i.e., document images). Since these files cannot be further parsed into fine-grained digital codes such as text, tables, and illustrations, it is impossible to effectively utilize the information in these document images. Summary of the invention
[0004] The purpose of the embodiments of the present application is to provide a method, device and electronic device for processing document images, which uses OCR and other technologies to capture text features and position information in documents and combines document image features to achieve multimodal pre-training, providing support for downstream tasks such as typesetting detection and semantic extraction. Compared with other document image processing models in the same field, the target image file processing model recorded in the embodiments of the present application has better performance and effectiveness.
[0005] In a first aspect, an embodiment of the present application provides a method for processing a document image, the method comprising: acquiring a document image; using the document image as an input of a target image file processing model, and obtaining structured information through the target image file processing model, wherein the target image file processing model comprises: a target text embedding module, a target spatial information embedding module, a target image feature embedding module and a target BERT module, the target spatial information embedding module is used to mine the position information of the elements in the document image, the target image feature embedding module is used to mine the image features of the elements in the document image, the target BERT module is configured to receive the feature information extracted by the target text embedding module, the target spatial information embedding module and the target image feature embedding module and output the structured information, the elements include visual elements and / or language elements, the visual elements include: tables or icons, and the language elements include text.
[0006] The embodiment of the present application significantly improves the ability to extract multi-dimensional information on document images by constructing a network model of a target BERT module that can input text information, location information, and image features, so that structured information corresponding to the document image can be obtained through the model, which facilitates the subsequent further use of the structured information and improves the utilization rate of document image files.
[0007] In some embodiments, the obtaining of structured information through the target image file processing model includes: obtaining text in the document image through the target text embedding module; obtaining position information of the text on the document image through the target spatial information embedding module; and obtaining image information corresponding to the text through the target image feature embedding module.
[0008] Some embodiments of the present application obtain the position information of each text extracted from the document image through the target spatial information embedding module, so that the target BERT module can further combine the position information for feature extraction so that the extracted information is more accurate.
[0009] In some embodiments, the position information is represented by two-dimensional coordinates, and the two-dimensional coordinates are coordinate values of the corresponding text in a two-dimensional coordinate system constructed according to the document image.
[0010] Some embodiments of the present application provide a method for representing position information, which facilitates quantification of position attribute information of each text extracted from a document image.
[0011] In some embodiments, the text includes a first text, and the position information of the first text is represented as (x0, y0, x1, y1), wherein (x0, y0) is the coordinate value of the upper left corner of the bounding box, and (x1, y1) represents the coordinate value of the lower right corner of the bounding box, and the bounding box is used to mark the position of the first text on the document image.
[0012] Some embodiments of the present application provide a text position representation method that uses the coordinates of the upper left corner and the lower right corner of the rectangular box where the text is located to represent the position of the text on the document image, thereby better quantifying the position information of each text.
[0013] In some embodiments, the obtaining of structured information through the target image file processing model further includes: obtaining image information corresponding to the text through the target image feature embedding module, wherein the target image feature embedding module is at least configured to divide the document image into regions corresponding to each text and obtain image features of the regions.
[0014] Some embodiments of the present application respectively extract text block image features of each text obtained from the document image through a target image feature embedding module, so that the information input into the target BERT module (referred to as the BERT model during the training process) includes the text, the position corresponding to the text, and the image block features corresponding to the text, thereby improving the mining of all text semantics and typesetting formats on the document image through this information, and facilitating further processing of this information.
[0015] In some embodiments, the obtaining of image information corresponding to the text through the target image feature embedding module includes: scanning the document image through optical character recognition (OCR) to identify each word included in the document image, and obtaining a bounding box of each word; dividing the document image into multiple small blocks according to the bounding box to obtain multiple document images, wherein each document image corresponds to a word; using the coordinate values corresponding to the bounding box and the block document images corresponding to each bounding box as inputs to a fast region-based convolutional network Faster R-CNN model; using the fast region-based convolutional network Faster R-CNN model to process the input information, and using the output of the region of interest (ROI) pooling layer of the fast region-based convolutional network Faster R-CNN model as the block image region feature corresponding to the corresponding block document image; and embedding the block image region feature as the image corresponding to the corresponding word to obtain the image information.
[0016] Some embodiments of the present application structurally adjust Fast R-CNN, remove the classification part of the fully connected layer, retain the feature extraction module, and use the output of the ROI pooling layer as the extracted image features of each text block. On the one hand, this improves the data processing speed, and on the other hand, it can align the image features and text features of the document.
[0017] In some embodiments, the acquiring of image information corresponding to the text through the target image feature embedding module further includes: using the fast region-based convolutional network Faster R-CNN model to take the entire scanned image corresponding to the document image as a region of interest ROI to generate a corresponding embedding.
[0018] Some embodiments of the present application use global average pooling to generate embedding by taking the entire document image as the region of interest (ROI) to facilitate downstream tasks that require [CLS] tags, and then input the extracted image features ImageFeatures and text embeddings into the target image file processing model for subsequent processing.
[0019] In some embodiments, before obtaining structured information through the target image file processing model, the method further includes: completing pre-training of the image file processing model by setting two tasks, wherein the two tasks include a multi-label document classification task and a masked text prediction task, and the masked text prediction task is configured to predict masked tags based on contextual text information and corresponding position information, and the image file processing model includes: a text embedding module, a spatial information embedding module, an image feature embedding module and a BERT model; and obtaining the target image file processing model by fine-tuning the pre-trained model through a target downstream task, wherein the target downstream task includes typesetting detection and text recognition.
[0020] In some embodiments of the present application, the model can better understand the layout and structure of the document through the spatial embedding based on the tag in the pre-training stage, so that when predicting the mask tag, it not only relies on the contextual text information, but also combines the corresponding spatial information. This helps to improve the performance of the model in document image understanding tasks, because the document layout and text content are closely related in many cases. In this way, the model learns a semantic representation that combines text information and spatial information, which is crucial for understanding the structure of the document.
[0021] In some embodiments, the loss function corresponding to the masked text prediction task is:
[0022] in, Used to characterize the prediction under set conditions The probability of , xi represents any text in the text sequence (i.e., a sequence of multiple texts extracted from the document image), ESpace is a spatial embedding used to represent the position of a word in a document, and n represents the total number of texts included in the text sequence.
[0023] Compared to the Masked Language Model (MLM) in the BERT model, whose goal is to predict the randomly masked words in the input sequence, the embodiments of the present application further introduce spatial embedding ESpace (i.e., position information) in the CoordinateLM model (i.e., the target image file processing model or the image file processing model) to capture the position of a word (as an example of text) in a document.
[0024] In some embodiments, the target image file processing model is obtained by fine-tuning the pre-trained model through the target downstream task, including: for the typesetting detection, the output of the pre-trained model is connected with the entire sample document image embedding, and prediction is performed through the fully connected layer FC; for the text recognition, the output of the pre-trained model is serialized and annotated through semantic tags and links to achieve structured extraction of digital document content.
[0025] In the process of training the image file processing model to obtain the target image file processing model, the embodiment of the present application can learn deep document representation through its own structural design and task-driven method without a large amount of manually annotated data. Under the framework of self-supervised learning, the CoordinateLM model (i.e., the target image file processing model) is trained using a large-scale unlabeled scanned document image dataset, thereby avoiding costly manual annotation.
[0026] In a second aspect, an embodiment of the present application provides a device for processing a document image, the device comprising: a document image acquisition module, configured to acquire a document image; a target file acquisition module, configured to use the document image as an input of a target image file processing model, and obtain structured information through the target image file processing model, wherein the target image file processing model comprises: a target text embedding module, a target spatial information embedding module, a target image feature embedding module and a target BERT module, the target spatial information embedding module is used to mine the position information of the elements in the document image, the target image feature embedding module is used to mine the image features of the elements in the document image, the target BERT module is configured to receive the feature information extracted by the target text embedding module, the target spatial information embedding module and the target image feature embedding module and output the structured information, the elements include visual elements and / or language elements, the visual elements include: tables or icons, and the language elements include text.
[0027] In a third aspect, some embodiments of the present application provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, can implement the method described in any one of the embodiments included in the first aspect.
[0028] In a fourth aspect, some embodiments of the present application provide an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executed, implements a method as described in any one of the embodiments included in the first aspect.
[0029] In a fifth aspect, some embodiments of the present application provide a computer program product, wherein the computer program product includes a computer program, and when the computer program is executed by a processor, the method described in any one of the embodiments included in the first aspect is implemented. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments of the present application will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.
[0031] Figure 1 One of the flowcharts of the method for processing a document image provided in an embodiment of the present application; Figure 2 A second flowchart of a method for processing a document image provided in an embodiment of the present application; Figure 3 One of the process schematic diagrams of the training image file processing model provided in the embodiment of the present application; Figure 4 The second process diagram of the training image file processing model provided in the embodiment of the present application; Figure 5 A model architecture diagram of an image feature embedding module provided in an embodiment of the present application; Figure 6 A block diagram of the device for processing document images provided in an embodiment of the present application; Figure 7 A schematic diagram of the composition of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0032] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application.
[0033] It should be noted that similar reference numerals and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. At the same time, in the description of this application, the terms "first", "second", etc. are only used to distinguish the description and cannot be understood as indicating or implying relative importance.
[0034] Archive digitization refers to the process of automatically extracting information, understanding content, and analyzing data from various forms of original documents. Automatic reading and analysis of large-scale commercial documents has a vital impact on improving corporate productivity and creating economic benefits. Due to the diversity of original archive formats, uneven quality of scanned images, and the complexity of template structures, typesetting detection and semantic extraction for digital archives (i.e., document images) is still a very challenging task. In the field of digital document processing, the academic community has made corresponding progress by combining statistical learning and artificial intelligence technology.
[0035] However, the inventors of this application found that while the research on related technologies has made technological progress, the following problems still exist: (1) Dependence on manually labeled samples: As the forms of business documents tend to be more diverse, the traditional training method that relies on manually labeled samples is very inefficient. Therefore, it is feasible to explore large-scale unlabeled samples for model training. (2) The correlation between text content and spatial information: Existing models usually use pre-trained CV models or NLP models, ignoring other potential semantic information in the file, such as the location information of the text or the file layout. Exploring the joint training of text content and spatial information is a key issue to improve the accuracy of semantic extraction.
[0036] At least to solve the above-mentioned problems, the embodiments of the present application provide a method for processing document images, which can be used for typesetting detection and semantic extraction in the process of document digitization. Different from the related art, the embodiments of the present application add spatial embedding of the relative position of the text (i.e., position information embedding) and document image feature embedding in the input stage to capture the connection between text semantics. The embodiments of the present application also explore the importance of shielding the loss of the visual language model to model training, so as to enhance the level of joint training of text information and position information. It has been found through actual verification that the target image file processing model of the present application is superior to the existing processing models in terms of performance and effectiveness in multiple task scenarios such as typesetting detection, semantic information extraction, and document image classification.
[0037] Please see Figure 1 , Figure 1 A method for processing a document image is provided for an embodiment of the present application, and the method exemplarily includes: S110, acquiring a document image.
[0038] In some embodiments of the present application, S110 exemplarily includes reading an image corresponding to the document to be processed (i.e., a document image) from a storage unit. For example, the document image is a PDF or other image format file obtained by scanning files of various formats.
[0039] S120, using the document image as an input of a target image file processing model, and obtaining structured information through the target image file processing model.
[0040] It should be noted that if Figure 2 As shown, the target image file processing model described in S120 includes: a target text embedding module 210, a target spatial information embedding module 220, a target image feature embedding module 230 and a target BERT module 250. Figure 2 The target text embedding module 210 is used to mine the text information of the elements in the document image, the target spatial information embedding module 220 is used to mine the position information of the elements in the document image, the target image feature embedding module 230 is used to mine the image features of the corresponding elements of the document image, and the target BERT module 250 is configured to receive the feature information extracted by the target text embedding module 210, the target spatial information embedding module 220 and the target image feature embedding module 230, and output the structured information. The elements include visual elements and / or language elements, the visual elements include: tables or icons, and the language elements include text.
[0041] For example, in some embodiments of the present application, Figure 2 As shown, the input of the target text embedding module 210 is the document image and the output is the text feature of the element, the input of the target spatial information embedding module 220 is the document image and the output is the position information, and the input of the target image feature embedding module 230 is the document image and the output is the text image feature. In some embodiments of the present application, the target image file processing model 200 also exemplarily includes a target feature fusion module 240, which is at least used to align text features, position information and text image features, and provide the processed data to the target BERT module 250.
[0042] It is not difficult to understand that the embodiments of the present application significantly improve the ability to extract multi-dimensional information from document images by constructing a network model of a target BERT module that can input text information, location information, and image features, so that the files output by the model are editable or searchable.
[0043] It should be noted that Figure 2 The target image file processing model 200 is based on Figure 3 The pre-training and fine-tuning process is obtained. Figure 3 The process of pre-training and fine-tuning is explained exemplarily.
[0044] In order to obtain the target image file processing model (or referred to as the CoordinateLM model) through the training model, some embodiments of the present application set two model tasks, namely masked vision-language and multi-label document classification, to pre-train it, so as to learn deep-level features. Unlabeled data is fully utilized in the form of self-supervision to reduce the traditional manual annotation cost. Other embodiments of the present application also use fine-tuning technology to fine-tune the pre-trained model. The fine-tuning process sets two downstream tasks, namely typesetting detection and text recognition classification, so that a small amount of manually labeled samples can be used for model fine-tuning. In some embodiments of the present application, for typesetting detection, the output of the pre-trained model is connected to the embedding of the entire document image, and prediction is performed through FC. In some embodiments of the present application, for text recognition, the output of the pre-trained model is serialized and annotated through semantic tags and links to achieve structured extraction of digital document content.
[0045] Understandably, Figure 3 Image file processing model and Figure 2 The target image file processing model architecture is the same as Figure 2 The difference is Figure 3 The training process includes a pre-training module and a fine-tuning module, where the fine-tuning module is used to fine-tune the parameters of the trained model.
[0046] like Figure 3 As shown, the image file processing model includes: a text embedding module 110, a spatial information embedding module 120, an image feature embedding module 130, a feature fusion module 140, a pre-training module 150 and a fine-tuning module 160. Figure 2 The difference is that the input Figure 3 That is, in some embodiments of the present application, before obtaining structured information through the target image file processing model in step S220, the method further includes the following method for training the model: The pre-training of the image file processing model is completed by setting two tasks, wherein the two tasks include a multi-label document classification task and a masked text prediction task, the masked text prediction task is configured to predict masked tags based on contextual text information and corresponding position information, and the image file processing model includes: a text embedding module, a spatial information embedding module, an image feature embedding module and a BERT model.
[0047] That is to say, through Figure 3 The pre-training module 150 sets two tasks to complete the text embedding module, the spatial information embedding module, the image feature embedding module and the BERT model ( Figure 3Not shown, as an example, the model is included in the pre-training module 150) pre-training. In some embodiments of the present application, the fine-tuning module 160 can also be used to fine-tune the model obtained after the pre-training of the pre-trained model, so that the final Figure 2 The target image format file processing module shown, wherein the model in training is also referred to as the image format file processing module. In some embodiments of the present application, Figure 3 The model also includes a feature fusion module 140. After training, the model corresponds to Figure 2 The target feature fusion module 240 is shown.
[0048] It should be noted that the above two tasks include a multi-label document classification task and a masked text prediction task, and the masked text prediction task is configured to predict masked marks based on contextual text information and corresponding position information; the target image file processing model is obtained by fine-tuning the pre-trained model through the target downstream task, wherein the target downstream task includes typesetting detection and text recognition.
[0049] In some embodiments of the present application, the model can better understand the layout and structure of the document through the spatial embedding (i.e., position information) marked in the pre-training stage, so that when predicting the mask tag, it not only relies on the contextual text information, but also can combine the corresponding position information. This helps to improve the performance of the model in the document image understanding task, because the document layout and text content are closely related in many cases. In this way, the model of the embodiment of the present application learns a semantic representation that combines text information and position information, which is crucial for understanding the structure of the document. Compared with the Masked Language Model (MLM) in the BERT model, the goal is to predict the words that are randomly masked in the input sequence. The embodiment of the present application further introduces the spatial embedding ESpace (i.e., position information) in the CoordinateLM model (i.e., the target image file processing model or the image file processing model) to capture the position of the word (as an example of text) in the document.
[0050] In some embodiments of the present application, the target image file processing model is obtained by fine-tuning the pre-trained model through the target downstream task, including: for the typesetting detection, connecting the output of the pre-trained model with the entire sample document image embedding, and predicting through the fully connected layer FC (Fully Connected layer); for the text recognition, serializing and annotating the output of the pre-trained model through semantic tags and links to achieve structured extraction of digital document content.
[0051] It should be noted that the typesetting detection task aims to correctly classify each document image. Unlike existing image-based methods, the embodiments of the present application not only include image representations, but also combine text and spatial information. In order to fine-tune the image file processing model CoordinateLM on this task, the output of the image file processing model CoordinateLM is connected to the entire image embedding, and then a softmax layer is used for category prediction.
[0052] It is not difficult to understand that the embodiments of the present application can learn deep document representations through their own structural design and task-driven methods without requiring a large amount of manually annotated data in the process of training the image file processing model to obtain the target image file processing model. Under the framework of self-supervised learning, the CoordinateLM model (i.e., the target image file processing model) is trained using a large-scale unlabeled scanned document image dataset, thereby avoiding costly manual annotation.
[0053] Next, we combine the loss function Figure 3 The pre-training process performed by the pre-training module 150 is described below.
[0054] The BERT model is an attention-based bidirectional language modeling method that shows effective knowledge transfer effects in self-supervised tasks with large-scale training data. The BERT architecture is a multi-layer bidirectional Transformer encoder that accepts a sequence of tokens and stacks multiple layers to produce the final representation. Specifically, given a set of tokens processed using WordPiece, the input embedding is calculated by adding the corresponding word embedding (token), position embedding (position), and segment embedding (segment), and a contextual representation with an adaptive attention mechanism is generated through a multi-layer bidirectional Transformer. BERT uses two target tasks during pre-training to learn language representations: masked language modeling (MLM) and next sentence prediction (NSP), where MLM randomly masks some input tokens and the goal is to restore these masked tokens; NSP is a binary classification task that takes a pair of sentences as input and classifies whether they are two consecutive sentences.
[0055] As described above, some embodiments of the present application are Figure 2The pre-training stage of the model of the architecture also adopts two task-driven methods: Masked Visual-Language Model (MVLM, corresponding to the masked text prediction task) and Multi-label Document Classification (MDC). The Masked Visual-Language Model (MVLM) method of the embodiment of the present application is inspired by the Masked Language Model (MLM) in the related BERT model. Unlike the related art, some embodiments of the present application randomly mask some tags (such as token or segment) in the input sequence in MVLM (Masked Vision-Language Model), while retaining the position information embedding corresponding to these tags, so even if the text content of these tags is masked, their corresponding position information is still input into the model as known information. The purpose of this method is to use the position information of text tags in the document to enhance the model's learning of text information. For example, in some embodiments of the present application, spatial embedding represents the position of each text on a document page (usually represented by the coordinates of the upper left corner and the lower right corner, such as (x0, y0, x1, y1)). Even if the text content is masked, the model can still use this known position information to predict the masked text content.
[0056] The model provided by some embodiments of the present application can input tag-based spatial embedding, so that the model can better understand the layout and structure of the document, so that when predicting mask tags, it not only relies on contextual text information, but also can combine the spatial information of the corresponding text. This helps to improve the performance of the model in document image understanding tasks, because document layout and text content are closely related in many cases. In this way, the image file processing model of some embodiments of the present application learns a semantic representation that combines text information and spatial information to more accurately mine the structure of the document. During the pre-training process, some embodiments of the present application expose the model under training to a large number of scanned document images. By predicting masked tags, the pre-trained model gradually grasps the relationship between the visual and language modalities of the document.
[0057] That is to say, compared with the MLM target in the BERT model of the related art, which is to predict the randomly masked words in the input sequence, the pre-training stage of the embodiment of the present application further introduces spatial embedding ESpace (i.e., position information) to capture the position of a word (as an example of text, the text can also be a character in Chinese) in the document.
[0058] For example, in some embodiments of the present application, given an input text sequenceX ={ , ,..., } (where n is the sequence length, and each text in the text sequence is the text extracted from the document image), randomly masking a certain proportion of the text input tokens x i For the special [MASK] tag, the output of the model is to predict the probability distribution of these masked tags. In some embodiments of the present application, the loss function of MVLM (i.e., the loss function corresponding to the masked text prediction task) can be expressed as:
[0059] in, In the given For all inputs except The probability, that is, Used to characterize the Prediction for all inputs except The probability of x i Represents any text in a text sequence (i.e., a sequence of texts (e.g., words or Chinese characters) in the sample documents of this training). E Space is a spatial embedding used to represent the position of a text (eg, a word or a Chinese character, etc.) in a document, and n represents the total number of texts included in the text sequence.
[0060] In the multi-label document classification (MDC) task of some embodiments of the present application, the pre-training of the model is guided by utilizing the metadata of the document, such as multi-label information related to the document. Labels include the type, subject, source, etc. of the document. By introducing these labels, the model learns the classification information of the document while learning the representation of the document content. During the pre-training process, the model needs to learn a comprehensive representation that can capture both text, image information and category information based on the input document image and the corresponding label. The above tasks require the model to be able to recognize and distinguish documents of different fields and types, thereby providing powerful document-level representation capabilities for subsequent fine-grained tasks. For each document image D, there is a corresponding set of labels associated with it. Y ={ , ,..., },in, m is the total number of labels. The loss function of MDC adopts the strategy of multi-label classification, such as binary cross entropy loss, which can be expressed as:
[0061] here is the sigmoid function, which is a linear combination of the model output Converted to probability, is the true value of label j. For multi-label classification, Can be 0 or 1.
[0062] It should be noted that the embodiments of the present application use a large-scale unlabeled scanned document image dataset, such as the IIT-CDIP test dataset, when implementing the above two tasks. This dataset provides rich text and coordinate information. By performing self-supervised learning on this data, the trained model can be trained without a large amount of manually labeled data through its own structural design (for example, Figure 2 or Figure 3 Model architecture) and task-driven methods are used to learn deep document representation. Under the framework of self-supervised learning, the trained model is trained using a large-scale unlabeled scanned document image dataset, thereby avoiding costly manual annotation. It is precisely because of the CoordinateLM provided in the embodiment of the present application that this technical purpose is achieved. CoordinateLM is a pre-trained model that adds spatial location features and image features on the basis of BERT.
[0063] In some embodiments of the present application, during the pre-training process, the model minimizes the above loss function and The total loss function of pre-training can be expressed as:
[0064] Among them, λ is the weight parameter used to balance the two tasks.
[0065] The following further illustrates the process of fine-tuning the pre-trained model. Figure 3 The fine-tuning module 160 executes, that is, fine-tuning the CoordinateLM using the fine-tuning technology.
[0066] In some embodiments of the present application, a series of steps can be taken in the process of fine-tuning the CoordinateLM model to ensure that the model can achieve optimal performance on a specific document image understanding task. For example, as described above, in some embodiments of the present application, two tasks are selected for fine-tuning: typesetting detection and text recognition classification. Corresponding strategies and methods are adopted for each task to achieve the best fine-tuning effect.
[0067] In some embodiments of the present application, the typesetting detection task aims to correctly classify each document image. Unlike existing image-based methods, the embodiments of the present application not only include image representations, but also combine text and spatial information. In order to fine-tune the fine-tuned model (i.e., the model obtained after pre-training the CoordinateLM model) on this task, the output of the CoordinateLM model is connected with the entire text as an image embedding, and then a softmax layer is used for category prediction. For example, in some embodiments of the present application, the model is fine-tuned for 30 cycles with a batch size of 40 and a learning rate of 2e-5. An end-to-end training strategy is adopted in the fine-tuning process, that is, the model is updated on each task-specific dataset (such as Figure 2 This strategy allows the model to leverage the knowledge it has learned during the pre-training phase and adjust it for specific tasks.
[0068] In some embodiments of the present application, for the text recognition task, the model goal is to extract and structure the text content in the scanned document image, and the text recognition task exemplarily includes two subtasks: semantic tagging and semantic linking. The semantic tagging task is to identify and classify each entity in the document. By treating semantic tagging as a sequence labeling problem, the CoordinateLM is fine-tuned to adapt to this task. The final representation of the model predicts the label of each tag through a linear layer and a softmax layer. The model is trained for 100 cycles with a batch size of 16 and a learning rate of 5e-5, thereby achieving performance improvements in various tasks.
[0069] like Figure 4 As shown, the image file processing model of some embodiments of the present application can be applied in the field of document image understanding. The model may be called the CoordinateLM model. The embodiments of the present application improve the understanding of scanned document images by jointly modeling text and position information. The CoordinateLM model is an extension of the BERT model architecture, adding two embedding mechanisms: spatial embedding (for example, using the coordinate information of the element to represent the position to obtain spatial embedding) and image embedding (for example, using the image features of the element as image embedding) to capture the position information and visual feature information of the text in the document. Therefore, the multimodal information of the CoordinateLM model of the embodiments of the present application are text feature embedding TextEmbedding, spatial feature embedding SpaceEmbedding, and image feature embedding ImageFeatures.
[0070] In some embodiments of the present application, Figure 4The text feature embedding (i.e., TextEmbedding) is a text embedding obtained through the standard BERT model. For example, BERT uses a bidirectional Transformer encoder to learn context-dependent word vectors from the input text. The input text sequence is first tokenized, and then each word is converted into a word embedding, and then the context-dependent representation is generated through the BERT model.
[0071] In some embodiments of the present application, spatial feature embedding (i.e., SpaceEmbedding) is related to the relative position of the words input to the BERT model in the text, and spatial embedding is intended to model the relative position of the text in the document image. In order to represent the position of elements in a scanned document image, the embodiment of the present application regards the document image as a two-dimensional coordinate system, and the element or text position can be represented by (x0, y0, x1, y1), where (x0, y0) corresponds to the position of the upper left corner of the bounding box where the element is located on the document image, and (x1, y1) represents the position of the lower right corner of the bounding box where the element is located on the document image, that is, the maximum horizontal distance and maximum vertical distance of the bounding box in the two-dimensional coordinate system. This setting supports the model to capture the position information of elements (e.g., text) and enhance the model's structured understanding of the document. Specifically expressed as:
[0072] in, and is the two-dimensional coordinate of a word (as an example of text in a language element) in a document, Embedding and Embedding is the learned embedding vector. The spatial embedding for each word (an example of a text included as a language element) can be represented as: .
[0073] like Figure 4 As shown, the target image file processing model of some embodiments of the present application exemplarily includes: Fast Regional Convolutional Neural Network Fast R-CNN (corresponding to Figure 3 Image feature embedding module), PDF parser based on text recognition OCRPDF Parser (corresponding to Figure 3 The spatial information embedding module and the text embedding module are at least used to realize text feature and text spatial feature extraction), BERT model and the fully connected layer FC-Layer, where the FC-Layer is used to map the parameters learned by the model to specific categories or make predictions, such as subsequent recognition or classification tasks.
[0074] Figure 4The input of Fast R-CNN is the document image, and the output is the image feature embedding vector (used to characterize the document image features, Figure 4 exemplified as each text image block), Figure 4 The input of the OCR PDF Parser module is the document image and the output includes the layout position embedding (the upper left corner coordinates and the lower right corner coordinates of the bounding box where the corresponding text is located, i.e. Figure 4 Layout Position Embedding (x 1 ,y 1 ) and Layout Position Embedding (x 0 ,y 0 )), text relative position embedding Position Embedding and text feature embedding Text Embedding. These outputs are jointly input into the BERT model to obtain the text embedding feature vector Text Embedding. Then, this information is fused with the image feature embedding vector Image Embedding (as a joint intrusion Co-Embedding) and input into the FC-Layer module. The FC-Layer module is used to output multiple tasks Tasks, which include: Layout detection, semantics understanding, and semantic extraction.
[0075] The models obtained after the pre-training and fine-tuning phase of the CoordinateLM model in some embodiments of the present application can realize the segmentation of digital archive typesetting and the formatted output of written text. This step is the core link of converting the deep learning ability of the model into practical application, aiming to realize the efficient conversion of document images to structured information. For example, the CoordinateLM model is first used to analyze the digital archive (i.e., document image) as a whole to identify the key visual elements and language elements in the document, including text blocks, tables, charts, and images. Through an in-depth understanding of the position and semantic content of these elements, the model can determine their relative position and hierarchical relationship in the document. Next, typesetting segmentation is performed, which involves decomposing the document image into independent and operable components, such as identifying structures such as titles, paragraphs, lists, and tables from a continuous text stream. This step usually requires the application of image processing and natural language processing techniques to ensure the accuracy and reliability of segmentation. This is followed by the formatted output of written text. At this stage, the identified text information is converted into a standardized and formatted format, such as HTML, XML, or JSON. It is not only easy to read and understand, but also supports further automated processing and analysis, such as searching, indexing, and data mining.
[0076] For example, in some embodiments of the present application, the model outputs a file that converts the text recognized in the scanned image into a uniform font and size, adds appropriate titles and tags, and generates clickable links and buttons. Ultimately, through this step, the original, unstructured document image is converted into a digital archive that can be efficiently used by computers and users, which not only greatly improves the efficiency of document processing, but also opens up new possibilities for the dissemination and utilization of knowledge.
[0077] The following is an example of using Figure 2 The process of obtaining structured information from the target image file processing model.
[0078] For example, in some embodiments of the present application, the structured information is obtained through the target image file processing model, including: obtaining the text in the document image (for example, words or text in the document, etc.) through the target text embedding module; obtaining the position information of the text on the document image through the target spatial information embedding module.
[0079] It is not difficult to understand that some embodiments of the present application obtain the position information of each text extracted from the document image through the target spatial information embedding module, so that the target BERT module can further combine the position information for feature extraction so that the extracted information is more accurate.
[0080] In some embodiments of the present application, the position information is represented by two-dimensional coordinates, and the two-dimensional coordinates are the coordinate values of the corresponding text in the two-dimensional coordinate system constructed according to the document image. For example, in some embodiments of the present application, the text includes a first text, and the position information of the first text is represented by (x0, y0, x1, y1), where (x0, y0) is the coordinate value of the upper left corner of the bounding box, and (x1, y1) represents the coordinate value of the lower right corner of the bounding box, and the bounding box is used to mark the position of the first text on the document image.
[0081] It is not difficult to understand that some embodiments of the present application provide a method for representing position information, which is convenient for quantifying the position attribute information of each text extracted from a document image. Some embodiments of the present application provide a method for representing the position of the text on the document image by using the coordinates of the upper left corner and the lower right corner of the rectangular box where the text is located, which better quantifies the position information of each text.
[0082] For example, in some embodiments of the present application, S220 obtains structured information through the target image file processing model, and also includes: obtaining image information corresponding to the text through the target image feature embedding module, wherein the target image feature embedding module is at least configured to divide the document image into regions corresponding to each text and obtain image features of the regions.
[0083] It can be understood that some embodiments of the present application respectively extract text block image features of each text obtained from the document image through the target image feature embedding module, so that the information input into the target BERT module (referred to as the BERT model during the training process) includes the text, the position corresponding to the text, and the image block features corresponding to the text, thereby improving the mining of all text semantics and typesetting formats on the document image through this information, and facilitating further processing of this information.
[0084] In some embodiments of the present application, the method of obtaining the image information corresponding to the text through the target image feature embedding module includes: scanning the document image through optical character recognition (OCR) to identify each word included in the document image, and obtaining a bounding box of each word, that is, the OCR technology scans the document image and outputs the text information therein; dividing the document image into multiple small blocks according to the bounding box to obtain multiple document images, wherein each document image corresponds to a word; using the coordinate values corresponding to the bounding box and the block document images corresponding to each bounding box (for example, the image features of the image blocks corresponding to the words) as inputs to a Faster R-CNN model based on a fast region-based convolutional network; using the Faster R-CNN model to process the input information, and using the output of the region of interest (ROI) pooling layer of the Faster R-CNN model as the block image region features corresponding to the corresponding block document image; and embedding the block image region features as the image corresponding to the corresponding word to obtain the image information. That is to say, some embodiments of the present application use the bounding box of a word obtained by OCR (an example of a text included in a language element) as input, and the improved Faster R-CNN in the embodiments of the present application is used to extract the image block features corresponding to each word, so that each word not only has text features (for example, word embedding) but also contains image features (features extracted from Faster R-CNN). This method can fuse the information of images and text and improve the model's ability to understand the content of the document.
[0085] It is not difficult to understand that some embodiments of the present application structurally adjust Fast R-CNN, remove the classification part of the fully connected layer, retain the feature extraction module, and use the output of the ROI pooling layer as the extracted image features of each text block. On the one hand, the data processing speed is improved, and on the other hand, the image features and text features of the document can be aligning.
[0086] For example, in some embodiments of the present application, the region of interest ROI pooling layer is used to map the candidate region to a feature map, the feature map is obtained by the feature extraction module of the fast region-based convolutional network Faster R-CNN model, the candidate region is a potential text region identified by the fast region-based convolutional network Faster R-CNN model, and the potential text region includes the bounding box of the corresponding text. In other words, the candidate region is obtained after FastRCNN scanning, and the region includes but is not limited to the region where the text information is located, that is, the range is wider, and it is the region where the model "believes" that there may be text.
[0087] It can be understood that some embodiments of the present application use the ROI pooling layer to complete the candidate region and feature map (i.e., the feature map of the text image extracted corresponding to each text block), which can complete the correspondence between text and image features, and then use the text image features as the input of the target BERT module.
[0088] For example, in some embodiments of the present application, obtaining the image information corresponding to the text through the target image feature embedding module also includes: using the Faster R-CNN model to take the entire scanned image corresponding to the document image as a region of interest ROI to generate a corresponding embedding.
[0089] Some embodiments of the present application use global average pooling to generate embedding by taking the entire document image as the region of interest (ROI) to facilitate downstream tasks that require [CLS] tags, and then input the extracted features ImageFeatures and text embeddings into the target image file processing model for subsequent processing.
[0090] Combine the following Figure 5 An improved Faster R-CNN model provided in some embodiments of the present application is exemplarily described, where the model is used to implement the functions of an image feature embedding module (or a target image feature embedding module, both of which have the same structure).
[0091] Image features are image information corresponding to text information. In order to utilize the image features of a document and align them with the text features, some embodiments of the present application add an image embedding layer (i.e., the improved Faster R-CNN model of the present application, which is used to extract the image features of the image blocks in the area where each word or text is located) to represent the image features of the document. For example, some embodiments of the present application use Figure 5 The improved Fast R-CNN model shown divides the document image into ROI regions, each of which corresponds to a word (as an example of text extracted from the document image).
[0092] The Fast R-CNN provided by the embodiment of the related art is configured to perform the following steps: In the first step, a candidate region is generated by selective search technology, which includes but is not limited to the target text or target word region. The target text or target word is each text or word located on the document image, and does not specifically refer to a certain type of text or word on the document image.
[0093] In the second step, the feature extraction layer CNN of Fast R-CNN is used to extract features of the input image to obtain a feature map. The feature map represents the image information (intermediate state). Subsequent pooling and other layers are required to finally obtain the image features.
[0094] The third step is to use the ROI pooling layer to map the candidate region to the feature map. This is because the candidate region is still a matrix at this time and needs to be converted into a vector through pooling.
[0095] In the fourth step, fully connected layers (fc6 and fc7) are used to further process and classify the features.
[0096] Since the Fast R-CNN model provided by the related art cannot achieve the technical purpose of this application, some embodiments of this application make structural adjustments to Fast R-CNN, remove the classification part of the fully connected layer, retain the feature extraction module, and use the output of the ROI pooling layer as the image embedding. For example, in some embodiments of this application, for the [CLS] tag, the entire document image is used as the region of interest (ROI) and global average pooling is used to generate an embedding to facilitate downstream tasks that require the [CLS] tag. The extracted features ImageFeatures and text embedding are then input together into the CoordinateLM model for subsequent processing.
[0097] In some embodiments of the present application, by integrating multiple embeddings, CoordinateLM can more accurately interpret scanned document images and improve the accuracy of information extraction and document understanding. The final output of the model is jointly calculated by the following formula:
[0098] Figure 5 The improved Faster R-CNN model includes: 13 convolutional layers (i.e. 13 conv layers), thirteen relu layers, four pooling layers (i.e. 4 pooling layers), feature mapping layers (i.e. Figure 5 Feature Map layer), regenerate layer (i.e. Figure 4 Reshape layer), Softmax layer Proposal layer and region of interest ROI pooling layer (i.e. Figure 4 ROIPooling layer), wherein, in some embodiments of the present application, the output of the ROI pooling layer is used as the image feature of each text block image extracted, and the input Figure 5 is the image of each text block obtained from the document image (i.e. Figure 5 imaginepieces).
[0099] Figure 5 The input, i.e., multiple text image blocks, are obtained by, for example, obtaining the bounding box of each word (as an example of text) in the document through OCR (Optical Character Recognition); the word bounding box in the OCR result is used to divide the image into multiple small blocks, each of which corresponds to a word. Then, the Faster R-CNN model is used to process these small blocks (i.e. Figure 5 The image blocks are input), and corresponding image region features are generated. These image region features are used as image embeddings corresponding to the words. In some embodiments of the present application, for the special [CLS] tag, the Faster R-CNN model is used to take the entire scanned document image as the region of interest (ROI) to generate the corresponding embedding. This embedding can be used for downstream tasks that require [CLS] tag representation, such as classification tasks.
[0100] Some embodiments of the present application utilize Faster R-CNN to generate feature representations of different regions in an image. By inputting image blocks, Faster R-CNN can extract high-level features of these regions. For example, in some embodiments of the present application, by taking the word bounding box generated by OCR as input, Faster R-CNN extracts the image block features corresponding to each word. In this way, each word not only has text features (e.g., word embedding), but also contains image features (features extracted from Faster R-CNN). This method can fuse the information of images and text and enhance the model's ability to understand the content of the document. For the entire document image, Faster R-CNN processes the entire image to generate features, which are used for [CLS] tags. In this way, in subsequent tasks such as classification, the [CLS] tag can contain global information for the entire document.
[0101] like Figure 6 As shown, some embodiments of the present application provide a device for processing document images. It should be understood that the device is similar to the above-mentioned Figure 1 The method embodiment corresponds to the method embodiment and can execute each step involved in the above method embodiment. The specific functions of the device can refer to the description above. To avoid repetition, the detailed description is appropriately omitted here. The device includes at least one software function module that can be stored in the memory in the form of software or firmware or solidified in the operating system of the device. The device for processing document images includes: a document image acquisition module 510 and a target file acquisition module 520.
[0102] The document image acquisition module is configured to acquire a document image.
[0103] The target file acquisition module is configured to use the document image as an input of a target image file processing model, and obtain structured information through the target image file processing model, wherein the target image file processing model includes: a target text embedding module, a target spatial information embedding module, a target image feature embedding module and a target BERT module, the target text embedding module is used to mine the text information of the elements in the document image, the target spatial information embedding module is used to mine the position information of the elements in the document image, the target image feature embedding module is used to mine the image features of the elements in the document image, the target BERT module is configured to receive the feature information extracted by the target text embedding module, the target spatial information embedding module and the target image feature embedding module, and output the structured information, the elements include visual elements and / or language elements, the visual elements include: tables or icons, and the language elements include text.
[0104] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working process of the device described above can refer to Figure 1 The corresponding process in the method will not be described in detail here.
[0105] Some embodiments of the present application provide a computer-readable storage medium having a computer program stored thereon. When the program is executed by a processor, the method described in any one of the embodiments of the method for processing a document image can be implemented.
[0106] Some embodiments of the present application provide a computer program product, which includes a computer program. When the computer program is executed by a processor, the method described in any one of the embodiments of the method for processing a document image is implemented.
[0107] like Figure 7 As shown, some embodiments of the present application provide an electronic device 600, including a memory 610, a processor 620, and a computer program stored in the memory 610 and executable on the processor 620, wherein the processor 620 reads the program through a bus 630 and implements a method as described in any one of the embodiments of the method for processing a document image as described above when executing the program.
[0108] Processor 620 can process digital signals and can include various computing structures, such as complex instruction set computer structure, reduced instruction set computer structure, or a structure that implements a combination of multiple instruction sets. In some examples, processor 620 can be a microprocessor.
[0109] The memory 610 may be used to store instructions executed by the processor 620 or data related to the execution of instructions. These instructions and / or data may include code to implement some or all functions of one or more modules described in the embodiments of the present application. The processor 620 of the present embodiment may be used to execute the instructions in the memory 610 to implement Figure 1 The memory 610 includes a dynamic random access memory, a static random access memory, a flash memory, an optical memory or other memory known to those skilled in the art.
[0110] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely schematic. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the devices, methods and computer program products according to multiple embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of a code, and the module, a program segment or a part of a code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart can be implemented with a dedicated hardware-based system that performs a specified function or action, or can be implemented with a combination of dedicated hardware and computer instructions.
[0111] In addition, the functional modules in the various embodiments of the present application may be integrated together to form an independent part, or each module may exist separately, or two or more modules may be integrated to form an independent part.
[0112] If the function is implemented in the form of a software function module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, and other media that can store program codes.
[0113] The above description is only an embodiment of the present application and is not intended to limit the scope of protection of the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application should be included in the scope of protection of the present application. It should be noted that similar reference numerals and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further defined and explained in the subsequent drawings.
[0114] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
[0115] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.
Claims
1. A method for processing a document image, characterized in that: The method comprises: Acquire document images; Using the document image as input of a target image file processing model, and obtaining structured information through the target image file processing model; in, The target image file processing model includes: a target text embedding module, a target spatial information embedding module, a target image feature embedding module and a target BERT module; The target space information embedding module is used to mine the position information of the elements in the document image, and the target image feature embedding module is used to mine the image features of the elements in the document image; The target BERT module is configured to receive the information extracted by the target text embedding module, the target spatial information embedding module and the target image feature embedding module and output the structured information; The elements include visual elements and / or language elements. The visual elements include: tables or icons, and the language elements include texts.
2. The method according to claim 1, characterized in that The obtaining of structured information through the target image file processing model includes: Acquire the text in the document image through the target text embedding module; Acquiring the position information of the text on the document image through the target space information embedding module; The image information corresponding to the text is acquired through the target image feature embedding module.
3. The method according to claim 2, characterized in that The position information is represented by two-dimensional coordinates, and the two-dimensional coordinates are coordinate values of the corresponding text in a two-dimensional coordinate system constructed according to the document image.
4. The method according to claim 3, characterized in that The text includes a first text, and the position information of the first text is represented by (x0, y0, x1, y1), wherein (x0, y0) is the coordinate value of the upper left corner of the bounding box, and (x1, y1) represents the coordinate value of the lower right corner of the bounding box, and the bounding box is used to mark the position of the first text on the document image.
5. The method according to claim 4, characterized in that The acquiring the image information corresponding to the text through the target image feature embedding module includes: Scanning the document image by optical character recognition (OCR) to identify each word included in the document image and obtain a bounding box of each word; Segmenting the document image into a plurality of small blocks according to the boundary box to obtain a plurality of document images, wherein each document image corresponds to a word; Using the coordinate values corresponding to the bounding boxes and the block document images corresponding to the bounding boxes as inputs to a fast region-based convolutional network Faster R-CNN model; Processing input information using the Faster R-CNN model based on the fast region-based convolutional network, and using the output of the region-of-interest (ROI) pooling layer of the Faster R-CNN model based on the fast region-based convolutional network as a block image region feature corresponding to the corresponding block document image; The block image region feature is embedded as an image corresponding to a corresponding word to obtain the image information.
6. The method according to claim 4, characterized in that The acquiring of image information corresponding to the text by the target image feature embedding module further includes: The Faster R-CNN model based on the fast region convolution network is used to take the entire scanned image corresponding to the document image as the region of interest (ROI) and generate the corresponding embedding.
7. The method according to claim 6, characterized in that Before obtaining the structured information through the target image file processing model, the method further includes: Pre-training the image file processing model is completed by setting two tasks, wherein the two tasks include a multi-label document classification task and a masked text prediction task, wherein the masked text prediction task is configured to predict a masked tag according to contextual text information and corresponding position information, and the image file processing model includes: a text embedding module, a spatial information embedding module, an image feature embedding module, and a BERT model; The target image file processing model is obtained by fine-tuning the pre-trained model through target downstream tasks, wherein the target downstream tasks include typesetting detection and text recognition; for the typesetting detection, the output of the pre-trained model is connected with the entire sample document image embedding, and prediction is performed through the fully connected layer FC; for the text recognition, the output of the pre-trained model is serialized and annotated through semantic tags and links to achieve structured extraction of digital document content.
8. The method according to claim 7, characterized in that The loss function corresponding to the masked text prediction task is: in, Used to characterize the prediction under set conditions The probability of x i Represents any text in the text sequence, E Space is a spatial embedding used to represent the position of text in a document, and n represents the total number of texts included in the text sequence.
9. A device for processing document images, characterized in that: The device comprises: A document image acquisition module, configured to acquire a document image; The target file acquisition module is configured to use the document image as an input of a target image file processing model, and obtain structured information through the target image file processing model, wherein the target image file processing model includes: a target text embedding module, a target spatial information embedding module, a target image feature embedding module and a target BERT module, the target spatial information embedding module is used to mine the position information of the elements in the document image, the target image feature embedding module is used to mine the image features of the elements in the document image, the target BERT module is configured to receive the information extracted by the target text embedding module, the target spatial information embedding module and the target image feature embedding module and output the structured information, the elements include visual elements and / or language elements, the visual elements include: tables or icons, and the language elements include text.
10. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method according to any one of claims 1 to 8 when executing the computer program.
Citation Information
Cited By
Multi-modal document content cross-platform analysis system
CN120726658A
Multi-modal document analysis method and device and computer readable storage medium
CN121170830A
Multi-modal document parsing method, device and computer readable storage medium
CN121170830B