Document digitization method and device, equipment, storage medium and program product
By detecting layout elements in document image files and extracting multimodal features, the problem of low text recognition accuracy in existing technologies has been solved, achieving high-precision document digitization and improving the accuracy of document content and user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- UC MOBILE CHINA CO LTD
- Filing Date
- 2025-12-31
- Publication Date
- 2026-04-24
AI Technical Summary
Existing technologies have low accuracy in recognizing various types of text, such as artistic fonts, rare characters, and formulas. This results in errors and omissions in the converted documents, affecting subsequent editing and collaboration efficiency.
By detecting layout elements in document image files, extracting target elements and their context images, using a pre-trained text recognition model for multimodal feature extraction, and combining visual and semantic features, a digital document of the document image file is generated.
It improves the recognition accuracy of various text types, including printed text, handwritten text, and formulas, enhances the robustness of text recognition and the accuracy of document layout reproduction, reduces the number of times users need to modify digital documents, and improves the user experience.
Smart Images

Figure CN121921791A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of document conversion technology, and in particular to a document digitization method, apparatus, device, storage medium, and program product. Background Technology
[0002] With the widespread adoption of mobile internet and smart terminals, document digitization has become a core user need across various scenarios and at high frequency. Whether it's students converting class notes and exam papers into editable formats for review and organization, professionals digitizing paper reports and contracts for collaborative editing and efficient management, or researchers extracting charts and formulas from papers for further processing, users urgently need a solution that can quickly convert paper documents or image-based documents into editable, standard formats.
[0003] However, the existing OCR (Optical Character Recognition) technology used in restoring the layout has low recognition accuracy when dealing with various types of text such as artistic fonts, rare characters, formulas, and handwritten fonts. This results in problems such as typos and omissions in the converted documents, which not only increases the additional cost of manual proofreading, but also seriously affects the efficiency of subsequent editing and collaboration, and cannot meet users' in-depth needs for document digitization.
[0004] Therefore, there is an urgent need to provide a document digitization processing solution that takes into account the recognition of multiple types of text and the restoration of layout, so as to improve the accuracy of the content in digitized documents. Summary of the Invention
[0005] This application provides a document digitization method, apparatus, device, storage medium, and program product. By utilizing the visual and semantic features of the image containing text elements in the original image and its context image, text recognition is achieved, improving recognition accuracy and thus improving the accuracy of document image file content restoration.
[0006] In a first aspect, embodiments of this application provide a document digitization method, comprising: detecting layout elements in a document image file; for a target element in the layout elements, cropping an image of the target element in the document image file to obtain a target element image, and cropping a context image of the target element; wherein the target element is an element that needs to be text-recognized; inputting the target element image and the context image into a pre-trained text recognition model, and performing multimodal feature extraction on the target element image and the context image of the target element through the text recognition model to obtain element visual features, context visual features, element semantic features, and context semantic features; and obtaining the text recognition result of the target element based on the element visual features, context visual features, element semantic features, and context semantic features; and generating a digitized document of the document image file based on the text recognition result of the detected target element and the layout attributes of each detected layout element.
[0007] Secondly, embodiments of this application provide a document digitization device, comprising: an element detection module for detecting layout elements in a document image file; an image cropping module for cropping an image of a target element in the document image file to obtain a target element image and a context image of the target element; wherein the target element is an element for which text recognition is required; a text recognition module for inputting the target element image and the context image into a pre-trained text recognition model, and extracting multimodal features from the target element image and the context image of the target element through the text recognition model to obtain element visual features, context visual features, element semantic features, and context semantic features; and obtaining the text recognition result of the target element based on the element visual features, context visual features, element semantic features, and context semantic features; and a document digitization module for generating a digitized document of the document image file based on the text recognition result of the detected target element and the layout attributes of each detected layout element.
[0008] Thirdly, embodiments of this application provide an electronic device, including a memory and a processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory, causing the processor to perform the method and / or various implementations provided in the first aspect of this application.
[0009] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the method and / or various implementation methods provided in the first aspect of this application.
[0010] Fifthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the method and / or various implementation methods provided in the first aspect of this application.
[0011] The document digitization method, apparatus, device, storage medium, and program products provided in this application, for user-provided document image files, such as document images and PDFs, achieve the detection of layout elements in the file through an element detection step; for target elements containing text, the multimodal features of the image of the target element captured in the original file and its context image, including visual and semantic features, are used to recognize the text in the target element. Text recognition through multimodal features improves the accuracy of recognizing various types of text, such as printed text, handwritten text, formulas, and artistic fonts, and overcomes the interference of image background on text recognition; by introducing contextual semantics and visual features, the recognition accuracy of short words and handwritten text is improved, enhancing the robustness of text recognition; using the text recognition results and the layout attributes of each element obtained during element detection, an editable digital document is generated, improving the fidelity of the digital document content. Simultaneously, the high-precision text recognition results improve the fidelity of the document layout, reducing the number of times users need to modify the digital document and improving the user experience. Attached Figure Description
[0012] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0013] Figure 1 An application scenario diagram provided for an embodiment of this application;
[0014] Figure 2 A flowchart illustrating a document digitization method provided in an embodiment of this application;
[0015] Figure 3 This is a schematic diagram of the structure of a layout element detection model provided in an embodiment of this application;
[0016] Figure 4 This is a schematic diagram of the structure of the text recognition model provided in the embodiments of this application;
[0017] Figure 5 A flowchart illustrating another document digitization method provided in this application embodiment;
[0018] Figure 6 A schematic diagram of the structure of the document digitization device provided in the embodiments of this application;
[0019] Figure 7This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.
[0020] The accompanying drawings have illustrated specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to specific embodiments. Detailed Implementation
[0021] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0022] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with relevant laws, regulations and standards, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0023] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.
[0024] With the popularization of mobile internet and the widespread use of smart devices, it has become a frequent need for users to take pictures or scan paper documents (such as work documents, study notes, book pages, test papers, contracts, etc.) with their smartphones and convert them into editable digital formats (such as Word, Excel, etc.).
[0025] For example, students need to digitize their class notes for review, professionals need to convert paper reports into electronic documents for collaborative editing, researchers need to extract charts and formulas from papers into editable content, and corporate finance personnel need to convert invoices, bills, and other documents into editable formats for statistical accounting.
[0026] For example, Figure 1 An application scenario diagram provided for an embodiment of this application, such as Figure 1As shown, users use user terminals 10, such as smartphones, tablets, scanning devices, etc., to photograph or scan paper documents, generating image or PDF document image carriers, and then uploading them to a document processing platform such as... Figure 1 The document processing server 20 in the document processing platform uses technologies such as object detection and OCR recognition to identify various elements and extract their basic attributes. Subsequently, based on the coordinates, styles, and hierarchical relationships between elements, the document processing platform determines the positional relationships and arrangement of each element. Finally, according to the layout of the original document, the platform restructures each element in a structured manner to obtain editable digital format documents such as Word and Excel.
[0027] In practical applications, user-uploaded images or PDF documents often contain diverse font types, such as handwritten notes in student notes and lesson plans, markings on exam papers, complex formulas in research papers, and artistic or anti-counterfeiting fonts in contracts. Existing OCR technologies are mostly designed for standardized fonts like printed text. Their accuracy is generally insufficient for non-standard fonts and diverse content, easily leading to character misjudgments, omissions, and misaligned formula elements, resulting in missing or incorrect content. These errors not only distort the core content of digitized documents but also directly disrupt the logical connections between elements, affecting the accuracy of subsequent positional judgments and layout restoration. Ultimately, the converted editable document contains typos, omissions, incorrect formulas, and incomplete content, requiring users to spend considerable time manually proofreading and correcting, severely reducing the efficiency of document digitization and the user experience.
[0028] Based on this, this application provides a document digitization method that, when recognizing text in a document image carrier, significantly improves the recognition accuracy of diverse content such as non-standard fonts, complex formulas, and special symbols by coordinating visual and semantic multimodal features and deeply integrating contextual information.
[0029] Figure 2 This is a flowchart illustrating a document digitization method provided in an embodiment of this application. This document digitization method can be executed by any device or module with corresponding data processing capabilities, such as a server or user terminal. Figure 2 As shown, the document digitization method includes:
[0030] Step S201: Detect layout elements in the document image file.
[0031] The document image file can be in image or PDF format. For example, it can be a document image obtained by taking a picture of a paper document, or a PDF file uploaded by the user. Layout elements are various visual components in the document image file that constitute the document structure and image reading logic, including titles, paragraphs, tables, illustrations, stamps, etc.
[0032] For example, users can upload pre-stored document images or PDF files through the software interface of the document processing software installed on their user terminals to achieve the upload of document image files, or they can take images of documents through the relevant controls in the software interface and upload them to obtain document image files in image format.
[0033] User-uploaded document image files are transmitted to the server via network communication, and the server executes the method provided in this application to convert the document image files into digital documents.
[0034] After receiving the document image file uploaded by the user, the server can first preprocess the document image file, such as noise reduction, pagination, and standardization. Then, it can perform element detection on the preprocessed document image file to obtain the layout elements and their typesetting attributes.
[0035] Taking a PDF file containing multiple pages as an example, during preprocessing, the PDF file can be paginated and each page can be converted into an image in a preset format. The converted image can also be processed by subject detection, rotation, correction and other methods.
[0036] Layout element detection can be performed using a deep learning model, which can automatically identify and extract various layout elements in the input image, such as paragraphs, illustrations, seals, tables, etc. It can also determine the relative position and arrangement of each layout element based on the layout attributes of the identified layout elements.
[0037] For example, layout elements in document image files can be detected based on finely tuned general object detection models, such as Faster R-CNN (Faster Region-based Convolutional Neural Network) or YOLO (You Only Look Once) model, or multiple detection models can be used, with each model responsible for detecting layout elements of the corresponding type.
[0038] Figure 3 This is a schematic diagram of a layout element detection model provided in an embodiment of this application. The layout element detection model is used to detect layout elements in document image files, such as... Figure 3As shown, the input image for the layout element detection model is an image obtained through preprocessing of a document image file. The layout element detection model includes a backbone network, a neck network, a head network, and a post-processing module. The backbone network extracts visual features from the input image, such as multi-scale visual features. The neck network fuses features from different levels to achieve multi-scale feature complementarity. The head network predicts the bounding boxes (bboxes), categories, and confidence scores of the layout elements contained within the fused features. The post-processing module optimizes the output of the head network by filtering the detection boxes (i.e., bounding boxes) through Non-Maximum Suppression (NMS) and thresholding. NMS filters overlapping detection boxes; if multiple detection boxes correspond to the same layout element, the detection box with the highest confidence score is retained. Thresholding filtering removes detection boxes with confidence scores below a threshold, reducing false detections. The output of the post-processing module is the output of the layout element detection model, yielding the detection boxes, categories, and confidence scores of each layout element in the input image.
[0039] Step S202: For the target element in the layout elements, extract the image of the target element in the document image file to obtain the target element image and the context image of the target element.
[0040] The target element is the element that needs to be recognized, that is, the target element contains the text that needs to be recognized.
[0041] Optional target elements include headings, paragraphs, stamps, and tables, and may also include flowcharts.
[0042] For each detected target element, based on the location information of the target element obtained during detection, such as the bounding box, the image of the target element is extracted from the image where the target element is located, thus obtaining the target element image.
[0043] The context image of a target element is the region within a preset range surrounding the target element in the image containing the target element; it is the image of the context element of the target element. The context image of the target element can be extracted from the image containing the target element by expanding the bounding box of the target element and then using the expanded bounding box as a basis.
[0044] The bounding box of a target element can be expanded according to a set ratio or a set number of pixels. Alternatively, based on the type of the target element and its context elements, expansion parameters in each direction can be determined, and the bounding box of the target element can be expanded according to the expansion parameters in each direction.
[0045] Taking a line of text within a paragraph as the target element as an example, the context elements can be the text preceding and following the target element. If there is no preceding or following text, the preceding or following layout element adjacent to the target element can be used as the context element for that target element.
[0046] Step S203: Input the target element image and the context image of the target element into the pre-trained text recognition model. The text recognition model performs multimodal feature extraction on the target element image and the context image of the target element to obtain element visual features, context visual features, element semantic features and context semantic features. Based on the element visual features, context visual features, element semantic features and context semantic features, the text recognition result of the target element is obtained.
[0047] Among them, element visual features are visual features extracted from the target element image, and contextual visual features are visual features extracted from the context image of the target element. Visual features refer to quantifiable information extracted from an image that can describe its content, structure, or semantics, such as character shape features and texture features. Element semantic features are the semantic features corresponding to the target element. Contextual semantic features are the semantic features corresponding to the context elements of the target element.
[0048] Element visual features and contextual visual features can be extracted through a visual feature extraction layer. This visual feature extraction layer can use the Swin Transformer model, which captures local visual features of the input image (target element image or context image) based on the window self-attention mechanism, such as the thickness of character strokes and the shape of radicals, as well as global layout features such as the arrangement direction of text lines and the spatial relationship between characters, to obtain element visual features and contextual visual features.
[0049] Semantic features, including element semantic features and contextual semantic features, are essentially the logical relationships between characters. Candidate text sequences can be identified first using corresponding visual features—element visual features and contextual visual features. Then, a pre-trained language model maps the characters in the candidate text sequences to a low-dimensional dense vector space and mines the semantics of the resulting vectors to obtain semantic features. Semantic features include character-level semantics, sequence-level semantics, and scene-level semantics. Character-level semantics represents the meaning of a single character, sequence-level semantics represents the contextual dependencies between characters, and scene-level semantics represents the type of scene to which the sequence belongs, the scene-specific formatting rules, terminology constraints, and other information.
[0050] Optionally, multimodal feature extraction is performed on the target element image and the context image of the target element to obtain element visual features, context visual features, element semantic features, and context semantic features. This includes: extracting visual features from the target element image and the context image, incorporating positional encoding, and performing serialization encoding to obtain element visual features and context visual features; generating element embedding features and context embedding features based on the element visual features and context visual features; and mining the semantic features of the element embedding features and context embedding features through a self-attention mechanism to obtain element semantic features and context semantic features.
[0051] Incorporating positional encoding enables the model to perceive the spatial positional information of different visual units (such as pixel blocks, feature map grids, etc.), thus compensating for the lack of positional information in the extracted visual features.
[0052] Taking the Transformer-type model for visual feature extraction as an example, during visual feature extraction, the Transformer-type model divides the input image (target element image or target element context image) into fixed-size pixel patches, such as 16 pixels × 16 pixels. Each pixel patch is flattened into a one-dimensional vector and transformed into a feature vector matching the dimension of the Transformer-type model through linear projection, thus achieving the initial extraction of visual features. In the positional encoding stage, a two-dimensional coordinate system is first established, and each pixel patch corresponds to a unique spatial index in this coordinate system. Based on the spatial index of the pixel patch, a positional code is generated using learnable positional encoding, sine / cosine positional encoding, etc., and the positional code is integrated into the extracted visual features.
[0053] In addition to Transformer-type models, CNN models can also be used to extract visual features, obtain feature maps, and flatten the feature maps into H×W feature vectors. During positional encoding, positional codes for each feature vector are generated and concatenated with the corresponding feature vectors to achieve the integration of positional codes.
[0054] After incorporating positional encoding, serialization encoding is required to adapt to the input format of subsequent steps. The core of serialization encoding is to convert two-dimensional features (visual features incorporating positional encoding) into a one-dimensional sequence.
[0055] After obtaining the element visual features and context visual features of the one-dimensional sequence, embedding processing is used to obtain their respective embedding features, namely element embedding features and context embedding features. Then, a self-attention mechanism can be used to capture the dependencies between any two elements in the sequence, analyze the semantic associations within the input embedding features, and further mine cross-image semantic associations between element embedding features and context embedding features, thereby extracting element semantic features and context semantic features more comprehensively and accurately.
[0056] The serialized element embedding feature sequence and the context embedding feature sequence are concatenated to form a unified global feature sequence. The global feature sequence is then passed through three different linear projection layers to generate a query vector Q, a key vector K, and a value vector V. Attention weights are obtained through multi-head self-attention calculation. The attention weights are then weighted and summed with the value vector V to obtain a new feature vector. This new feature vector is then passed through a feedforward neural network, residual connections, and layer normalization to obtain a complete feature sequence that integrates global semantics. From this sequence, element semantic features and context semantic features are extracted.
[0057] Optionally, element embedding features and context embedding features are generated based on element visual features and context visual features, including: generating initial candidate text sequences corresponding to target elements and initial candidate text sequences corresponding to context elements based on element visual features and context visual features; mapping each discrete character in the initial candidate text sequences corresponding to target elements and initial candidate text sequences corresponding to context elements to a low-dimensional dense word embedding vector, and supplementing the temporal position information of each word embedding vector through text position encoding to obtain element embedding features and context embedding features.
[0058] Element embedding features are the embedding features obtained by word embedding and text position encoding of the initial candidate text sequence corresponding to the target element. Context embedding features are the embedding features obtained by word embedding and text position encoding of the initial candidate text sequence corresponding to the context elements.
[0059] Visual features can be converted into discrete character sequences using visual-text mapping models, such as cross-modal mapping layers, to obtain the corresponding initial candidate text sequences. Subsequently, word embedding processing is performed on the initial candidate text sequences, mapping each discrete character in the sequence into a low-dimensional dense word embedding vector through a pre-trained word embedding matrix. To compensate for the lack of temporal positional information in the word embedding vectors, text positional encoding, such as sine / cosine encoding or learnable positional encoding, is introduced to supplement the temporal positional information of each word embedding vector within the text sequence. Finally, element embedding features and context embedding features that fuse character semantics and temporal positional information are obtained.
[0060] After obtaining features that include both semantic and visual modalities, namely element visual features, contextual visual features, element semantic features, and contextual semantic features, feature fusion is performed on the multimodal features. The fused features are then used to obtain the text recognition results of the target element or the target element image.
[0061] In addition to the recognized text sequence, the text recognition results may also include some auxiliary information, such as semantic association labels and the confidence level of each character in the text sequence.
[0062] During feature fusion, visual features can be enhanced based on masks representing character regions in visual features (including element visual features and contextual visual features).
[0063] An encoder can be used to serialize the fused features (or multimodal fusion features). A multi-head self-attention mechanism is used to enhance the global feature interactions and temporal correlations of characters, mapping two-dimensional visual features to one-dimensional serialized encoded features. Then, based on these one-dimensional serialized encoded features, a beam search decoding strategy is used to generate candidate text sequences corresponding to the target element character by character. By filtering and verifying the candidate text sequences, the final text recognition result of the target element can be obtained.
[0064] Optionally, based on element visual features, contextual visual features, element semantic features, and contextual semantic features, the text recognition result of the target element is obtained, including: using a cross-attention mechanism to fuse the element visual features, contextual visual features, element semantic features, and contextual semantic features to obtain multimodal fusion features; generating a target candidate text sequence based on the multimodal fusion features; and performing low-confidence screening and semantic verification on the target candidate text sequence to obtain the text recognition result.
[0065] The cross-attention mechanism uses element semantic features and contextual semantic features as queries, and element visual features and contextual visual features as keys and values. By calculating the attention weight of semantic features on visual features, it fully integrates spatial location information at the visual level with association information at the semantic level, and finally generates multimodal fusion features that combine visual representation and semantic association.
[0066] Using this multimodal fusion feature as input, a target candidate text sequence is generated through a decoding layer such as a Transformer decoder. The target candidate text sequence is a preliminary prediction of the text content of the target element.
[0067] Low confidence filtering is used to remove characters or segments in the sequence whose confidence is lower than a preset threshold, so as to ensure that the retained text content has sufficient predictive reliability; semantic verification combines contextual semantic features to verify whether the semantic logic of the target candidate text sequence conforms to the association rules between the target element and the context element. After these two filtering and verification steps, an accurate and reliable text recognition result is finally obtained.
[0068] Optional, Figure 4 This is a schematic diagram of the structure of the text recognition model provided in the embodiments of this application, such as... Figure 4 As shown, the text recognition model includes a visual feature extraction module, a word embedding layer, a multimodal fusion module, a decoder, and a text recognition task head.
[0069] The steps of extracting multimodal features from the target element image and its context image using a text recognition model to obtain element visual features, context visual features, element semantic features, and context semantic features; and obtaining the text recognition result of the target element based on the element visual features, context visual features, element semantic features, and context semantic features, further include:
[0070] The visual feature extraction module extracts visual features from the input target element image and context image, incorporates positional encoding, and performs serialization encoding to obtain element visual features and context visual features.
[0071] The decoder generates an initial candidate text sequence corresponding to the target element and an initial candidate text sequence corresponding to the context element based on the element's visual features and the context's visual features.
[0072] Through the word embedding layer, each discrete character in the initial candidate text sequence corresponding to the target element and the initial candidate text sequence corresponding to the context element is mapped into a low-dimensional dense word embedding vector. The temporal position information of each word embedding vector is supplemented by text position encoding to obtain element embedding features and context embedding features.
[0073] The semantic features of element embedding features and context embedding features are mined through the multimodal fusion module using an attention mechanism to obtain element semantic features and context semantic features; and the element visual features, context visual features, element semantic features and context semantic features are fused through an attention mechanism to obtain multimodal fused features.
[0074] The decoder generates a sequence of target candidate text elements based on multimodal fusion features; and
[0075] By using the text recognition task head, the target candidate text sequence is subjected to low-confidence screening and semantic verification to obtain the text recognition result.
[0076] See also Figure 4The visual feature extraction module comprises a backbone network, a positional encoding module, and a Transformer encoder. The backbone network extracts visual features from the input image (including the target element image and its context image), extracting visual morphological information such as character shape, texture, and layout to obtain initial visual features. The positional encoding module uses the pixel coordinate system of the input image to obtain the two-dimensional spatial positional information corresponding to the initial visual features, and fuses it into the initial visual features after positional encoding to obtain initial visual features with positional encoding. The Transformer encoder performs serialization encoding on the initial visual features with positional encoding, strengthening the spatial association and feature interaction between characters through a self-attention mechanism, and outputting the corresponding visual features, namely element visual features and contextual visual features.
[0077] For example, the backbone network can adopt the Swing Transformer model.
[0078] The decoder can be a Transformer decoder, which uses the autoregressive sequence generation mechanism provided by the self-attention layer and the feedforward network to map the semantic features of elements and context semantic features to the character probability distribution at each position, filter high-probability characters and associate them according to the temporal sequence corresponding to the spatial position encoding, and generate the initial candidate text sequence corresponding to the target element and the initial candidate text sequence corresponding to the context element.
[0079] The token embedding layer maps each discrete character in the initial candidate text sequence to a low-dimensional dense token embedding vector, preserving the inherent semantic attributes of the characters. This yields token embedding vectors, and a temporal index is assigned to each token embedding vector. The text temporal position code is generated through sine, cosine, and trigonometric function encoding to supplement the temporal position information of the characters in the sequence. The token embedding vectors are then fused with the corresponding text temporal position codes to obtain the embedding features, namely element embedding features and context embedding features.
[0080] The multimodal feature fusion module, based on a self-attention mechanism, performs semantic interaction on the input embedded features to mine character-level semantics, sequence-level semantics, and scene-level semantics, obtaining semantic features, namely element semantic features and context semantic features. Based on the built-in cross-attention layer, it performs weighted fusion on element visual features, context visual features, element semantic features, and context semantic features to obtain multimodal fusion features. The decoder strengthens the temporal correlation between characters in the multimodal fusion features through the self-attention layer, and then maps it to a more accurate character probability distribution through a feedforward network. Through an optimized autoregressive sequence generation mechanism, it filters the best candidates character by character to obtain a more accurate candidate text sequence for the target element, namely the target candidate text sequence.
[0081] Each character in the target candidate text sequence corresponds to a prediction probability, which is the confidence level output by the model when generating that character, i.e., the character-level confidence level; the confidence level of each target candidate text sequence is the average of the confidence levels of each character in it.
[0082] The text recognition task head is used to eliminate target candidate text sequences with confidence scores below a preset threshold, and to perform semantic verification on each retained target candidate text sequence, such as correcting typos, grammatical errors, and resolving semantic contradictions. If the semantic verification still fails after the correction process, the target candidate text sequence is eliminated, and the retained target candidate text sequence after semantic verification is taken as the text recognition result.
[0083] Step S204: Based on the text recognition results of the detected target elements and the layout attributes of the detected layout elements, a digital document of the document image file is generated.
[0084] The layout attributes can include position, layout, style, orientation, and other attributes.
[0085] Based on the detected positional attributes of each layout element, the positional relationships between layout elements can be determined. Based on these relationships, a hierarchical framework for the document page is constructed, generating structured data. Based on the layout attributes of the layout elements, format parameters are mapped to each node in the structured data, and the corresponding text recognition results are bound to the target elements. Through the document generation engine, elements with bound format parameters and text content are rendered into a blank document according to the hierarchy and order of the structured data, resulting in a digital document image file.
[0086] For example, the digital document can be a Word document, an Excel document, or other easily accessible document.
[0087] The document generation engine renders elements with bound format parameters and text content into a blank document according to the hierarchy and order in the structured data. It can also make layout adjustments based on the layout attributes of each layout element, such as font, font size, color, and alignment, to ensure that the layout document is highly consistent with the original image.
[0088] The document generation engine renders elements with bound formatting parameters and text content into a blank document according to the hierarchy and order of the structured data. Layout reconstruction is then performed. During layout reconstruction, the system activates a rules engine, automatically adjusting layout parameters based on layout analysis results to ensure the output format is compatible with common office software.
[0089] For structured elements such as tables and lists, multimodal fusion algorithms, combined with image recognition and natural language processing techniques, can be used to extract row relationships, column relationships, hierarchical structure, and data attribute information from the structured elements. For example, for tables, the data in each column is extracted to determine the connections and hierarchical relationships between columns, which is used to restore the structural information of the structured elements when generating digital documents.
[0090] The text recognition results of the detected target elements and the layout attributes of each detected layout element can be input into a pre-trained layout restoration model. The document digitization model determines the logical relationships between each layout element and generates a document logical structure tree. Based on the document logical structure tree output by the layout restoration model, a digital document, such as a Word document, can be generated.
[0091] The document digitization method provided in this embodiment, for user-provided document image files, such as document images and PDFs, detects layout elements in the file through an element detection step; for target elements containing text, it utilizes multimodal features, including visual and semantic features, of the image of the target element captured in the original file and its context image to recognize the text in the target element. This multimodal text recognition improves the accuracy of recognizing various types of text, such as printed text, handwritten text, formulas, and artistic fonts, and overcomes the interference of image background on text recognition; by introducing contextual semantics and visual features, it improves the recognition accuracy of short words and handwritten text, enhancing the robustness of text recognition; and by using the text recognition results and the layout attributes of each element obtained during element detection, it generates an editable digital document, improving the fidelity of the digital document content. Simultaneously, the high-precision text recognition results improve the document layout fidelity, reducing the number of times users need to modify the digital document and enhancing the user experience.
[0092] Figure 5 This is a flowchart illustrating another document digitization method provided in this application embodiment. Figure 2 Based on the examples, the document digitization method is described in detail, such as... Figure 5 As shown, this document digitization method may specifically include the following steps:
[0093] Step S501: Obtain the document image file uploaded by the user.
[0094] Step S502: Detect layout elements in the document image file.
[0095] Optionally, before detecting layout elements in the document image file, the method further includes: preprocessing the document image file to detect layout elements in the preprocessed document image file; the preprocessing includes at least one of the following: subject detection, rotation, correction, and pagination.
[0096] Among them, subject detection is used to identify the parts of document image files that need to be digitally reconstructed, filtering out impurities other than the subject.
[0097] For example, when a user takes a picture of a paper document, there may be irrelevant objects in the background such as a desktop, other books, or hands, or there may be black borders or shadows at the edges of the image. Subject detection can segment out the main area containing only the valid content of the document, remove irrelevant impurities, and avoid impurities interfering with the subsequent layout element detection.
[0098] Rotation is used to correct orientation deviations in document image files, ensuring that the reading direction of the document content aligns with the standard layout orientation. For example, if a user tilts their hand while taking a picture, causing the image to rotate at a certain angle, image rotation can automatically rotate the image to a horizontal orientation (or a vertical orientation that conforms to conventional reading habits) based on the text line direction and page border features in the document. This improves the accuracy of subsequent layout element position coordinate detection and text orientation recognition.
[0099] Image correction is used to correct geometric distortions in document image files, such as image distortion caused by paper wrinkles, perspective effects, or image quality degradation caused by insufficient light or blurry shooting. Through image correction, operations such as perspective correction, distortion removal, and sharpening enhancement can be performed to restore the document's borders, text lines, and table lines to normal, and ensure that the shape, position, layout, and other attributes of layout elements are not affected by distortion.
[0100] Pagination is used to split multi-page document images into individual, independent document images. For example, if a user takes a series of photos of a multi-page paper document into a single long image, or merges multiple pages into one image file during scanning, the system can automatically identify and split each page based on features such as page edges, page spacing, and header / footer positions, generating single-page images that correspond one-to-one with the page numbers of the original paper document. This avoids misalignment of layout elements caused by confusion between multiple pages.
[0101] Step S503: For the target element in the layout elements, extract the image of the target element in the document image file to obtain the target element image and the context image of the target element.
[0102] Target elements include headings, paragraphs, stamps, tables, and flowcharts.
[0103] When the target element to be detected is a flowchart, the flowchart can be split, and each sub-graph after splitting can be used as a target element image.
[0104] When the target element to be detected is a seal, the seal can be separated from the image in which it is located, and the separated seal image can be used as a target element image.
[0105] Step S504: Preprocess the target element image and the context image respectively.
[0106] Preprocessing can include grayscale conversion, noise reduction, binarization, and size normalization.
[0107] Step S505: Identify the character regions in the preprocessed target element image and the preprocessed context image respectively, and generate a first mask matrix and a second mask matrix based on the character regions.
[0108] The character region is the area where the characters are located in the image. It can be the area covered by all characters or the area where each character is located. The first mask matrix is used to represent the character region in the preprocessed target element image, and the second mask matrix is used to represent the character region in the preprocessed context image.
[0109] After preprocessing, each pixel in the image can be traversed, and based on the pixel's grayscale value, it can be determined whether the pixel belongs to a character or the background; based on the pixels that belong to characters, the character region in the image can be determined.
[0110] It can also perform connected component detection on the pixels corresponding to characters to obtain each connected region; based on the physical features of the characters, such as size and aspect ratio, invalid connected regions are filtered out; based on the center distance and boundary distance of adjacent connected regions, it is determined whether to merge two adjacent connected regions to obtain the region of each character in the image. Based on the region of each character in the image, a corresponding mask matrix is generated, namely the first mask matrix and the second mask matrix. In the mask matrix, all element values within the character region range are 1, and the rest are 0.
[0111] Step S506: Based on the first mask matrix and the second mask matrix, image enhancement is performed on the preprocessed target element image and the preprocessed context image, respectively.
[0112] Based on the first mask matrix, enhancement operations can be performed on the character regions in the preprocessed target element image to improve the contrast between the character regions and the background. Alternatively, Laplacian sharpening can be performed on the character regions based on the first mask matrix, and the sharpening kernel can be convolved with the pixels of the character regions to enhance the gray-scale transitions at the stroke edges, filling gaps while making the stroke edges more distinct.
[0113] Enhancement operations can be performed on the character regions in the preprocessed context image based on the second mask matrix to improve the contrast between the character regions and the background. Laplacian sharpening can also be performed on the character regions based on the second mask matrix, and the sharpening kernel can be convolved with the pixels of the character regions to enhance the gray-scale transitions at the stroke edges, filling gaps while making the stroke edges more distinct.
[0114] For example, the pixel value of a pixel with a value of 1 in the mask matrix can be increased or decreased to improve the contrast between the character area and the background.
[0115] Step S507: Input the enhanced context image and the enhanced target element image into the text recognition model. The text recognition model extracts multimodal features from the enhanced target element image and the enhanced context image respectively. Based on the extracted multimodal features, the text recognition result of the target element is obtained.
[0116] The extracted multimodal features include element visual features, contextual visual features, element semantic features, and contextual semantic features.
[0117] Step S508: Perform layout understanding on the document image file to obtain the layout parameters of the document image file.
[0118] Layout understanding is used to determine the macro-layout framework of document image files, and to clarify the functional divisions and structural rules of pages.
[0119] Layout parameters include, but are not limited to, page orientation, page size, margins, number of columns, column parameters, and text orientation.
[0120] Step S509: Input the layout parameters of the document image file, the layout attributes of each detected page element, and the text recognition results of each target element into the pre-trained layout restoration model to obtain the document logical structure tree of the document image file.
[0121] The document logical structure tree is a structured tree data representing the hierarchical logical relationships between layout elements. Within the document logical structure tree, nodes represent logical relationships such as inclusion, subordination, and parallelism between layout elements, with the root node corresponding to the entire document image file.
[0122] The input to the layout restoration model can also include document images, which can be images of photographed paper documents uploaded by the user, or images obtained after preprocessing the document image files uploaded by the user.
[0123] The layout restoration model comprises an input layer, an image encoder, a text encoder, a multimodal feature fusion module, and a prediction layer. The input layer maps the input layout parameters and typography attributes to a predefined feature space to obtain attribute features. It also performs word segmentation and embedding on the input text recognition results, converting them into language feature vectors and establishing preliminary relationships between features across different dimensions. The image encoder encodes the input document image, the text encoder encodes the text recognition results of the target elements, and the multimodal feature fusion module fuses the attribute features from the image encoder, text encoder, and input layer. The prediction layer predicts the logical relationships between layout elements based on the multimodal features output by the multimodal feature fusion module, such as parallel, inclusive, and irrelevant relationships. It then performs hierarchical partitioning and sorting of layout elements, constructing a document logical structure tree based on the logical relationships, hierarchical levels, and order of the layout elements.
[0124] The image encoder in the layout restoration model can be a Discrete Variational Auto Encoder (DUE). Specifically, it divides the input document image into blocks of fixed size, converts each block into a continuous feature vector, and then quantizes it into discrete visual symbol vectors using a codebook. Specifically, it first calculates the similarity between the continuous feature vector and the prototype feature vector in the codebook, finds the best-matching prototype feature vector, and uses the index (integer) of this prototype feature vector as the visual symbol vector of the continuous feature vector. Deep encoding is then performed on the visual symbol vector to extract visual features such as layout structure, element shape, and spatial relationships. Deep encoding can be implemented using a Transformer encoder or convolutional layers.
[0125] The text encoder is used to capture the contextual dependencies within the text based on the self-attention mechanism, and to enhance the semantic features of the text through a feedforward neural network to obtain language feature vectors.
[0126] The multimodal feature fusion module can use BEiT-3's Multiway-Transformer.
[0127] The pre-training tasks for the typography restoration model include: MLM (Masked Language Modeling), MIM (Masked Image Modeling), and WPA (Word-Patch Alignment).
[0128] The MLM task involves randomly occluding some tokens in a text and then using a model to predict the original content corresponding to the occluded tokens based on the contextual information of the occluded tokens.
[0129] The MIM task involves segmenting an image into fixed-size patches, with an image encoder extracting features from each patch to generate discrete visual tokens as visual labels for that patch. During pre-training, a certain proportion of patches are randomly occluded, and the model predicts the visual tokens of the occluded patches.
[0130] The WAP task involves randomly covering parts of an image, and the model predicts which tokens in the input text correspond to the covered patches.
[0131] Step S510: Based on the document logical structure tree of the document image file, generate an editable, standard-format digital document.
[0132] Based on the hierarchy, order, and format of each layout element in the document's logical structure tree, the document generation engine is invoked to render the corresponding layout elements, resulting in a digital document that highly reproduces the document image file.
[0133] In this embodiment, image enhancement is achieved through a mask matrix obtained by character region recognition, enabling the text recognition model to perform text recognition based on the enhanced image. This effectively improves the problems of blurry and low recognizability of text in low-quality images, significantly enhancing the recognition accuracy and robustness of the text recognition model. Layout understanding accurately extracts macroscopic layout parameters of the document, laying a structured foundation for subsequent processing and avoiding errors in element location and classification. Secondly, the layout restoration model integrates layout parameters, layout attributes, and text recognition results, generating a logically clear document logical structure tree through multimodal information linkage, ensuring accurate restoration of the hierarchical relationships, spatial order, and functional attributes of document elements. Finally, an editable standard format document is generated based on the document logical structure tree, fully preserving the original document's layout, format style, and content information, while achieving efficient conversion from unstructured images to structured editable documents. This significantly reduces the cost of manual reformatting, improves the automation and accuracy of document digitization processing, and ensures the universality and editing flexibility of the output document.
[0134] In order to overcome the problems of high cost, long cycle and limited sample coverage of manual annotation training samples during the training of the aforementioned text recognition model, this application also provides a method for automatic synthesis and annotation of training samples.
[0135] Optionally, the method further includes: collecting multiple real training samples; integrating, segmenting, and normalizing the multi-source corpus to obtain a corpus; generating multiple initial training samples based on the matching results of each corpus in the corpus with the fonts in a preset font library; performing at least one of the following operations on at least some of the initial training samples to obtain synthetic training samples, and training the text recognition model with a training set composed of multiple real training samples, multiple initial training samples, and synthetic training samples: adjusting the background color of the initial training samples based on preset foreground and background color difference constraints; performing deformation processing on the text in the initial training samples based on a deformation simulation model; the deformation simulation model is constructed by extracting the deformation features of the text in multiple real training samples; injecting noise into the initial training samples; blurring the initial training samples; adding background texture to the initial training samples based on a background texture generation formula; the background texture generation formula is constructed based on the features of the background texture in multiple real training samples.
[0136] Real training samples are either manually labeled samples or automatically labeled samples that have been manually reviewed. They can include document image files uploaded by users at historical times, such as images of class notes, exam papers, screenshots of paper formulas, PDF documents, etc.
[0137] Multi-source corpora are comprehensive collections of corpora built to support text recognition tasks, originating from multiple dimensions and covering diverse linguistic needs. Examples include multilingual lexicons, multilingual dictionary example sentences, and topic corpora. Topic corpora are corpora related to specific topics, such as academic papers and contracts.
[0138] By using a unified data format, multi-source corpora can be deduplicated, filtered, and removed to form a unified set of original corpora with distinct content. Furthermore, based on the grammatical rules of different languages, the corpora of different languages are segmented into words, breaking down continuous text into discrete lexical units (tokens).
[0139] The multi-source corpus contains formulas. The normalization operation of the formulas is to unify the formulas of different formats and forms into a standardized representation, thereby eliminating format differences and standardizing the structural form.
[0140] After obtaining the corpus through the aforementioned processing, for each piece of corpus, based on the matching result between the corpus and the font in the preset font library, the initial training sample of the corpus is obtained by rendering with the matched font.
[0141] By sequentially configuring the initial training samples with at least one of the following settings, the training samples are augmented to obtain synthetic training samples: setting font, setting color, setting deformation, adding or injecting noise, blurring, and adding background texture.
[0142] Among them, setting the font allows you to randomly select a font from Song, Kai, Hei, Times New Roman, and other fonts; setting the color is used to change the color of the font in the corpus.
[0143] Deformation settings are used to geometrically adjust the font shape in the corpus to simulate handwriting or the character distortion effect caused by paper deformation.
[0144] Adding or injecting noise can simulate text interference effects in real-world scenarios. You can randomly select one or more of the following noise types to add: salt and pepper noise, Gaussian noise, text edge noise, and random noise blocks.
[0145] Blur processing is used to simulate the text blurring effect caused by out-of-focus shooting, insufficient scanning resolution, ink diffusion, paper wear, etc. in real-world scenarios. Specifically, Gaussian blur, motion blur, mean blur, bilateral blur, etc. can be randomly selected. The degree of blur can be controlled by adjusting the size and intensity of the blur kernel to achieve blur processing for different corpora.
[0146] Adding background textures is used to add backgrounds with certain texture features to the corpus.
[0147] To improve the similarity between the synthesized training samples and the real training samples, during training sample augmentation, the background color of the initial training samples can be adjusted based on a preset foreground-background color difference constraint. This preset constraint is used to limit the minimum color difference between the adjusted background color and the foreground color. The preset foreground-background color difference constraint can be determined based on a large number of statistically analyzed real training samples.
[0148] To improve the efficiency, versatility, and realism of deformation processing, a deformation simulation model can be used when deforming text in the initial sample. This model can be constructed using deformation features from a large number of pre-extracted real training samples, such as handwritten text samples. Deformation processing includes, but is not limited to, stretching, tilting, bending, local twisting, and rotational deformation.
[0149] An initial training sample can be processed multiple times to generate multiple synthetic training samples.
[0150] To improve the closeness of the background texture to the simulation of the background texture in the real training samples, a background texture can be added to the initial training samples after the above processing by using a preset background texture generation formula to obtain synthetic training samples.
[0151] Based on statistical models, deep learning algorithms, and other methods, the background texture in multiple real training samples can be analyzed to obtain the background texture generation formula.
[0152] When adding a background texture, you need to use a character area mask matrix to clearly divide the character area and the background area, ensuring that the background texture is only superimposed on the background area.
[0153] The native background textures can be extracted from multiple real training samples to obtain a set of background texture templates. A background texture is randomly selected from the set of background texture templates. First, the selected background texture is spatially aligned with the initial training sample according to the character region mask matrix. Then, the background texture is added by weighting and fusing it pixel by pixel through the alpha channel.
[0154] After pixel-by-pixel weighted fusion through the Alpha channel, you can use Boson editing to make the background texture and text edges transition naturally, and you can also perform style transfer to make the background texture more in line with the real scene.
[0155] Furthermore, when synthesizing training samples, the corresponding synthesis parameters can be selected based on the font of the initial training samples to configure the initial training samples in terms of color, background color, deformation, background texture, etc.
[0156] Optionally, formula normalization is performed on the multi-source corpus, including: uniformly replacing the same symbol in different formulas of the multi-source corpus with a preset standardized symbol to obtain the standard expression of each formula; merging formulas with the same semantics but different expression forms based on the standard expression; and removing the modifying expressions in the standard expression.
[0157] In this context, the decorative symbols used in formulas have no substantial impact on the formula's semantics, calculation logic, or numerical results; they are only used for the formula's display format, visual presentation, or scene annotation. Examples include "\left" or "\right" in LaTeX for adjusting bracket size, and "\color{}" for setting colors.
[0158] To address formula recognition, all 18 mathematical fonts supported by LaTeX were collected, and a unified Unicode representation was established for each symbol in these 18 fonts. For example, the triangle symbol was uniformly represented as / triangle. Thus, for formulas in multi-source corpora, the symbols in the formulas were replaced with the pre-defined standardized symbols, achieving a unified representation across different mathematical fonts.
[0159] Furthermore, mathematical formulas often contain synonyms with different expressions. Therefore, it's necessary to reduce redundant formula expressions by merging synonyms. For example, 1 / 2, \frac12, \frac{1}{2}, \dfrac{1}{2}, and {1\over2} can all be uniformly represented as \frac{1}{2}.
[0160] In the standard expression of some formulas, there may be redundant curly braces, such as \angle{{A}}. In order to make the standard expression of formulas more concise and consistent, it is necessary to automatically remove the redundant braces in the standard expression, for example, change \angle{{A}} to \angle{A}.
[0161] Corresponding to the document digitization method provided in the foregoing embodiments, this application also provides a document digitization device. Figure 6 A schematic diagram of the structure of the document digitization device provided in the embodiments of this application is shown below. Figure 6 As shown, the document digitization device includes: an element detection module for detecting layout elements in a document image file; an image cropping module for cropping an image of a target element from the layout elements in the document image file to obtain a target element image and a context image of the target element; wherein the target element is an element that needs to be text-recognized; a text recognition module for inputting the target element image and the context image into a pre-trained text recognition model, and extracting multimodal features from the target element image and the context image of the target element to obtain element visual features, context visual features, element semantic features, and context semantic features; and obtaining the text recognition result of the target element based on the element visual features, context visual features, element semantic features, and context semantic features; and a document digitization module for generating a digitized document of the document image file based on the text recognition result of the detected target element and the layout attributes of each detected layout element.
[0162] In one possible implementation, the text recognition module is specifically used for: inputting a target element image and a context image into a pre-trained text recognition model; extracting visual features from the target element image and the context image through the text recognition model, incorporating positional encoding, and performing serialization encoding to obtain element visual features and context visual features; generating element embedding features and context embedding features based on the element visual features and context visual features; mining semantic features of the element embedding features and context embedding features through a self-attention mechanism to obtain element semantic features and context semantic features; and obtaining the text recognition result of the target element based on the element visual features, context visual features, element semantic features, and context semantic features.
[0163] In one possible implementation, the text recognition module, when generating embedding features, specifically performs the following steps: based on element visual features and context visual features, initial candidate text sequences corresponding to the target element and context elements; mapping each discrete character in the initial candidate text sequences corresponding to the target element and context elements to a low-dimensional dense word embedding vector, and supplementing the temporal position information of each word embedding vector through text position encoding to obtain element embedding features and context embedding features.
[0164] In one possible implementation, when the text recognition module obtains the text recognition result of the target element based on the element's visual features, contextual visual features, element's semantic features, and contextual semantic features, it specifically performs the following: it fuses the element's visual features, contextual visual features, element's semantic features, and contextual semantic features through a cross-attention mechanism to obtain multimodal fusion features; it generates a target candidate text sequence based on the multimodal fusion features; and it performs low-confidence screening and semantic verification on the target candidate text sequence to obtain the text recognition result.
[0165] In one possible implementation, the text recognition model includes: a visual feature extraction module, a word embedding layer, a multimodal fusion module, a decoder, and a text recognition task head; the text recognition module is specifically used to: extract visual features from the input target element image and context image via the visual feature extraction module, incorporate positional encoding, and perform serialization encoding to obtain element visual features and context visual features; generate an initial candidate text sequence corresponding to the target element and an initial candidate text sequence corresponding to the context element based on the element visual features and context visual features via the decoder; and map each discrete character in the initial candidate text sequence corresponding to the target element and the initial candidate text sequence corresponding to the context element to the word embedding layer. The algorithm generates low-dimensional, dense word embedding vectors and supplements each word embedding vector with temporal position information through text position encoding to obtain element embedding features and context embedding features. A multimodal fusion module then mines the semantic features of the element embedding features and context embedding features using a self-attention mechanism to obtain element semantic features and context semantic features. A cross-attention mechanism is then used to fuse the element visual features, context visual features, element semantic features, and context semantic features to obtain multimodal fusion features. Based on these multimodal fusion features, a decoder generates a target candidate text sequence for the target element. Finally, a text recognition task head performs low-confidence filtering and semantic verification on the target candidate text sequence to obtain the text recognition result.
[0166] In one possible implementation, the device further includes an image enhancement module, configured to: preprocess the target element image and the context image respectively before inputting them into a pre-trained text recognition model; identify character regions in the preprocessed target element image and the preprocessed context image respectively, and generate a first mask matrix and a second mask matrix based on the character regions; and perform image enhancement on the preprocessed target element image and the preprocessed context image respectively based on the first mask matrix and the second mask matrix, so as to input the enhanced context image and the enhanced target element image into the text recognition model.
[0167] In one possible implementation, the device further includes a training sample augmentation module, configured to: collect multiple real training samples; integrate, segment, and normalize the multi-source corpus to obtain a corpus; generate multiple initial training samples based on the matching results of each corpus in the corpus with the fonts in a preset font library; perform at least one of the following operations on at least some of the initial training samples to obtain synthetic training samples, and train the text recognition model with a training set composed of multiple real training samples, multiple initial training samples, and synthetic training samples: adjust the background color of the initial training samples based on a preset foreground-background color difference constraint; perform deformation processing on the text in the initial training samples based on a deformation simulation model; the deformation simulation model is constructed by extracting the deformation features of the text in multiple real training samples; inject noise into the initial training samples; blur the initial training samples; and add a background texture to the initial training samples based on a background texture generation formula; the background texture generation formula is constructed based on the features of the background texture in multiple real training samples.
[0168] In one possible implementation, when the training sample augmentation module performs formula normalization on the multi-source corpus, it specifically performs the following: for the same symbol in different formulas of the multi-source corpus, it uniformly replaces it with a preset standardized symbol to obtain the standard expression of each formula; based on the standard expression, it merges formulas with the same semantics but different expression forms; and it removes the modifying expressions in the standard expression.
[0169] In one possible implementation, the document digitization module is specifically used for: performing layout understanding on the document image file to obtain the layout parameters of the document image file; inputting the layout parameters of the document image file, the layout attributes of each detected layout element, and the text recognition results of each target element into a pre-trained layout restoration model to obtain the document logical structure tree of the document image file; and generating an editable, standard-format digital document based on the document logical structure tree of the document image file.
[0170] In one possible implementation, the apparatus further includes a document preprocessing module for: preprocessing the document image file before detecting layout elements in the document image file, so as to detect layout elements in the preprocessed document image file; the preprocessing includes at least one of the following: subject detection, rotation, correction, and pagination.
[0171] The document digitization device provided in this embodiment can execute the document digitization method provided in any of the above embodiments. Its implementation principle and technical effect are similar, and will not be described in detail here.
[0172] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 7 As shown, the electronic device 70 provided in this embodiment includes at least one processor 701 and a memory 702. Optionally, the electronic device 70 further includes a communication component 703. The processor 701, memory 702, and communication component 703 are connected via a bus 704.
[0173] In a specific implementation, at least one processor 701 executes computer execution instructions stored in memory 702, causing at least one processor 701 to perform the above-described method.
[0174] The specific implementation process of processor 701 can be found in the above method embodiments, and its implementation principle and technical effect are similar. It will not be repeated here.
[0175] In the above embodiments, it should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this invention can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor.
[0176] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage device.
[0177] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, the buses shown in the accompanying drawings are not limited to a single bus or a single type of bus.
[0178] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.
[0179] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the above-described method.
[0180] The aforementioned readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The readable storage medium can be any available medium accessible to a general-purpose or special-purpose computer.
[0181] An exemplary readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can reside in an Application Specific Integrated Circuit (ASIC). Alternatively, the processor and the readable storage medium can exist as discrete components in the device.
[0182] The division of units is merely a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or units, and may be electrical, mechanical, or other forms.
[0183] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.
[0184] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0185] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0186] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.
[0187] Finally, it should be noted that other embodiments of the invention will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This invention is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein, and is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope.
Claims
1. A document digitization method, characterized in that, include: Detect layout elements in document image files; For the target element in the layout elements, the image of the target element in the document image file is extracted to obtain the target element image, and the context image of the target element is also extracted; wherein, the target element is the element that needs to be text-recognized. The target element image and the context image are input into a pre-trained text recognition model. The text recognition model performs multimodal feature extraction on the target element image and the context image of the target element, respectively, to obtain element visual features, context visual features, element semantic features, and context semantic features. Based on the element visual features, context visual features, element semantic features, and context semantic features, the text recognition result of the target element is obtained. Based on the text recognition results of the detected target elements and the layout attributes of each of the detected layout elements, a digital document of the document image file is generated.
2. The method according to claim 1, characterized in that, The step of performing multimodal feature extraction on the target element image and the context image of the target element to obtain element visual features, context visual features, element semantic features, and context semantic features includes: Visual features are extracted from the target element image and the context image, positional encoding is incorporated, and serialization encoding is performed to obtain the element visual features and the context visual features. Based on the element visual features and the context visual features, generate element embedding features and context embedding features; The semantic features of the element embedding features and context embedding features are obtained by mining the semantic features of the element embedding features and context embedding features through a self-attention mechanism.
3. The method according to claim 2, characterized in that, The step of generating element embedding features and context embedding features based on the element visual features and the context visual features includes: Based on the visual features of the element and the visual features of the context, an initial candidate text sequence corresponding to the target element and an initial candidate text sequence corresponding to the context element are generated. Each discrete character in the initial candidate text sequence corresponding to the target element and the initial candidate text sequence corresponding to the context element is mapped to a low-dimensional dense word embedding vector, and the temporal position information of each word embedding vector is supplemented by text position encoding to obtain the element embedding feature and the context embedding feature.
4. The method according to claim 3, characterized in that, The text recognition result of the target element based on the element's visual features, contextual visual features, element semantic features, and contextual semantic features includes: By using a cross-attention mechanism, the element visual features, contextual visual features, element semantic features, and contextual semantic features are fused to obtain multimodal fusion features; Based on the multimodal fusion features, a target candidate text sequence for the target element is generated; The target candidate text sequence is subjected to low-confidence screening and semantic verification to obtain the text recognition result.
5. The method according to claim 1, characterized in that, Before inputting the target element image into the context image into a pre-trained text recognition model, the method further includes: The target element image and the context image are preprocessed separately. Character regions in the preprocessed target element image and the preprocessed context image are identified respectively, and a first mask matrix and a second mask matrix are generated based on the character regions. Based on the first mask matrix and the second mask matrix, image enhancement is performed on the preprocessed target element image and the preprocessed context image, respectively, so that the enhanced context image and the enhanced target element image are input into the text recognition model.
6. The method according to claim 1, characterized in that, The method further includes: Collect multiple real training samples; The corpus is obtained by integrating, segmenting, and normalizing multi-source corpora using formulas. Based on the matching results of each corpus in the corpus with the fonts in the preset font library, multiple initial training samples are generated. Perform at least one of the following operations on at least a portion of the initial training samples to obtain synthetic training samples, and train the text recognition model on a training set composed of the plurality of real training samples, the plurality of initial training samples, and the synthetic training samples: Based on a preset foreground-background color difference constraint, the background color of the initial training sample is adjusted; Based on the deformation simulation model, deformation processing is performed on the text in the initial training samples; the deformation simulation model is constructed by extracting the deformation features of the text in the multiple real training samples; Noise is injected into the initial training samples; The initial training samples are then blurred. Background texture is added to the initial training samples based on the background texture generation formula; the background texture generation formula is constructed based on the features of the background texture in the multiple real training samples.
7. The method according to claim 6, characterized in that, Formula normalization is performed on multi-source corpora, including: For the same symbol in different formulas of the multi-source corpus, a preset standardized symbol is used to uniformly replace it to obtain the standard expression of each formula; Based on the standard expression, formulas with the same semantics but different forms of expression are merged; Remove the modifiers from the standard expression.
8. The method according to claim 1, characterized in that, The process of generating a digital document from the document image file based on the text recognition results of the detected target elements and the layout attributes of each detected layout element includes: Perform layout understanding on the document image file to obtain the layout parameters of the document image file; The layout parameters of the document image file, the layout attributes of each of the detected layout elements, and the text recognition results of each of the target elements are input into a pre-trained layout restoration model to obtain the document logical structure tree of the document image file. Based on the document logical structure tree of the document image file, an editable, standard-format digital document is generated.
9. The method according to any one of claims 1-8, characterized in that, Before detecting layout elements in the document image file, the method further includes: The document image file is preprocessed to detect layout elements in the preprocessed document image file; the preprocessing includes at least one of the following: subject detection, rotation, correction, and pagination.
10. The method according to any one of claims 1-8, characterized in that, The target elements include headings, paragraphs, stamps, and tables.
11. A document digitization device, characterized in that, include: The element detection module is used to detect layout elements in document image files; The image cropping module is used to crop an image of a target element in the document image file to obtain a target element image, and to crop a context image of the target element; wherein the target element is an element that needs to be text-recognized. The text recognition module is used to input the target element image and the context image into a pre-trained text recognition model, and to extract multimodal features from the target element image and the context image of the target element through the text recognition model to obtain element visual features, context visual features, element semantic features and context semantic features; and to obtain the text recognition result of the target element based on the element visual features, context visual features, element semantic features and context semantic features. The document digitization module is used to generate a digitized document of the document image file based on the text recognition results of the detected target elements and the layout attributes of the detected layout elements.
12. An electronic device, characterized in that, include: Memory, processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory, causing the processor to perform the method as described in any one of claims 1-10.
13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1-10.
14. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method according to any one of claims 1-10.