Document data processing method and device based on large model
By using self-attention and cross-attention mechanisms to extract text and image features in PDF document analysis, the system complexity of PDF document analysis, insufficient formula analysis capabilities and difficult table analysis are solved in the prior art, and a more efficient and accurate PDF document analysis effect is achieved.
Patent Information
- Application Number
- CN202510128695.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-27
- Publication Date
- 2025-05-06
AI Technical Summary
The existing PDF document analysis methods have system complexity, traditional OCR methods lack formula analysis capabilities and table analysis difficulties, resulting in instability in the parsing effect.
By obtaining at least one PDF element of a variety of different types of elements identified from the PDF file, feature extraction of the to be processed text and images based on the self-attention and cross-attention mechanism, text-image cross-attention features are generated, and the analysis result of the PDF file is finally determined.
This method can reduce the error consequences caused by hallucinations while improving the completeness and accuracy of PDF document parsing, and supports unified analysis of text, formulas and tables.
Smart Images

Figure CN119940343A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, in particular to big models and computer vision technology, and specifically to a document data processing method, device, electronic device, computer-readable storage medium and computer program product based on big models. Background Art
[0002] Artificial intelligence is a discipline that studies how to use computers to simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, planning, etc.). It includes both hardware-level and software-level technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing; artificial intelligence software technologies mainly include computer vision technology, speech recognition technology, natural language processing technology, as well as machine learning / deep learning, big data processing technology, knowledge graph technology, and other major directions.
[0003] Portable Document Format (PDF) document parsing is a common task in computer vision. Its core purpose is to extract and process content from PDF documents through algorithms, including text, images, tables, formulas, etc. This technology has important application value in education and teaching, multilingual processing, and auxiliary large language model (LLM) training.
[0004] The methods described in this section are not necessarily methods that have been previously conceived or employed. Unless otherwise indicated, it should not be assumed that any method described in this section is considered to be prior art simply because it is included in this section. Similarly, unless otherwise indicated, the issues mentioned in this section should not be considered to have been recognized in any prior art. Summary of the invention
[0005] The present disclosure provides a document data processing method, device, electronic device, computer-readable storage medium and computer program product based on a large model.
[0006] According to one aspect of the present disclosure, a data processing method is provided, comprising: acquiring at least one PDF element from a plurality of different types of elements identified from a portable document format PDF file; determining an image to be processed and a text to be processed based on the identified PDF elements, wherein the image to be processed includes an image of the at least one identified PDF element, and the text to be processed includes text identified from the image to be processed; performing feature extraction on the text to be processed based on a self-attention mechanism to obtain self-attention features of the text to be processed; performing feature extraction on the self-attention features of the text to be processed and the image features of the image to be processed based on a cross-attention mechanism to obtain text-image cross-attention features for the PDF file; and determining a parsing result of the PDF file based at least on the cross-attention features.
[0007] According to another aspect of the present disclosure, a data processing device is provided, including: an acquisition unit, configured to acquire at least one PDF element from a plurality of different types of elements identified from a portable document format PDF file; a text and image determination unit, configured to determine an image to be processed and a text to be processed based on the identified PDF element, wherein the image to be processed includes an image of at least one identified PDF element, and the text to be processed includes text identified from the image to be processed; a text feature extraction unit, configured to perform feature extraction on the text to be processed based on a self-attention mechanism to obtain a self-attention feature of the text to be processed; a cross-attention feature extraction unit, configured to perform feature extraction on the self-attention feature of the text to be processed and the image feature of the image to be processed based on a cross-attention mechanism to obtain a text-image cross-attention feature for the PDF file; and a parsing result determination unit, configured to determine a parsing result of the PDF file based at least on the cross-attention feature.
[0008] According to another aspect of the present disclosure, an electronic device is also provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to an embodiment of the present disclosure.
[0009] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is further provided, wherein the computer instructions are used to enable the computer to execute the method according to the embodiment of the present disclosure.
[0010] According to another aspect of the present disclosure, a computer program product is provided, including a computer program, wherein the computer program implements the method described in the embodiment of the present disclosure when executed by a processor.
[0011] According to one or more embodiments of the present disclosure, by processing the image features corresponding to the images of different elements identified from the PDF file and the text features of the PDF based on the cross-attention mechanism, the accurate information of the visual input can be referred to during the autoregressive processing of the text features, thereby reducing the erroneous consequences caused by the hallucination phenomenon.
[0012] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The accompanying drawings exemplarily illustrate the embodiments and constitute a part of the specification, and together with the text description of the specification, are used to explain the exemplary implementation of the embodiments. The embodiments shown are for illustrative purposes only and do not limit the scope of the claims. In all drawings, the same reference numerals refer to similar but not necessarily identical elements.
[0014] Figure 1 A schematic diagram showing an exemplary system in which the various methods described herein may be implemented according to an embodiment of the present disclosure;
[0015] Figure 2 A flow chart of a data processing method according to an embodiment of the present disclosure is shown;
[0016] Figure 3 An exemplary process of a data processing method according to an embodiment of the present disclosure is shown;
[0017] Figure 4 An exemplary structure of a multimodal model according to an embodiment of the present disclosure is shown;
[0018] Figure 5 An exemplary block diagram of a data processing device according to an embodiment of the present disclosure is shown;
[0019] Figure 6 A block diagram of an exemplary electronic device that can be used to implement an embodiment of the present disclosure. DETAILED DESCRIPTION
[0020] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for the sake of clarity and conciseness, the description of well-known functions and structures is omitted in the following description.
[0021] In the present disclosure, unless otherwise specified, the use of the terms "first", "second", etc. to describe various elements is not intended to limit the positional relationship, temporal relationship, or importance relationship of these elements, and such terms are only used to distinguish one element from another element. In some examples, the first element and the second element may refer to the same instance of the element, and in some cases, based on the description of the context, they may also refer to different instances.
[0022] The terms used in the description of various examples in this disclosure are only for the purpose of describing specific examples and are not intended to be limiting. Unless the context clearly indicates otherwise, if the number of elements is not specifically limited, the element can be one or more. In addition, the term "and / or" used in this disclosure covers any one of the listed items and all possible combinations.
[0023] The embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.
[0024] Figure 1 FIG. 1 is a schematic diagram of an exemplary system 100 in which various methods and apparatuses described herein may be implemented according to an embodiment of the present disclosure. Figure 1 , the system 100 includes one or more client devices 101, 102, 103, 104, 105, and 106, a server 120, and one or more communication networks 110 coupling the one or more client devices to the server 120. The client devices 101, 102, 103, 104, 105, and 106 may be configured to execute one or more applications.
[0025] In an embodiment of the present disclosure, the server 120 may run one or more services or software applications that enable execution of the method for parsing a PDF according to an embodiment of the present disclosure.
[0026] In some embodiments, server 120 may also provide other services or software applications, which may include non-virtualized environments and virtualized environments. In some embodiments, these services may be provided as web-based services or cloud services, such as provided to users of client devices 101, 102, 103, 104, 105, and / or 106 under a software as a service (SaaS) model.
[0027] exist Figure 1 In the configuration shown, the server 120 may include one or more components that implement the functions performed by the server 120. These components may include software components, hardware components, or a combination thereof that can be executed by one or more processors. Users operating client devices 101, 102, 103, 104, 105, and / or 106 may in turn utilize one or more client applications to interact with the server 120 to utilize the services provided by these components. It should be understood that a variety of different system configurations are possible, which may differ from the system 100. Therefore, Figure 1 is one example of a system for implementing the various methods described herein and is not intended to be limiting.
[0028] The user may use client devices 101, 102, 103, 104, 105 and / or 106 to obtain information input by the user or output results to the user. The client device may provide an interface that enables the user of the client device to interact with the client device. The client device may also output information to the user via the interface. Figure 1 Only six client devices are depicted, but one skilled in the art will appreciate that the present disclosure may support any number of client devices.
[0029] Client devices 101, 102, 103, 104, 105 and / or 106 may include various types of computer devices, such as portable handheld devices, general-purpose computers (such as personal computers and laptop computers), workstation computers, wearable devices, smart screen devices, self-service terminal devices, service robots, game systems, thin clients, various messaging devices, sensors or other sensing devices, etc. These computer devices may run various types and versions of software applications and operating systems, such as MICROSOFT Windows, APPLE iOS, UNIX-like operating systems, Linux or Linux-like operating systems (such as GOOGLE Chrome OS); or include various mobile operating systems, such as MICROSOFT Windows Mobile OS, iOS, Windows Phone, Android. Portable handheld devices may include cellular phones, smart phones, tablet computers, personal digital assistants (PDAs), etc. Wearable devices may include head-mounted displays (such as smart glasses) and other devices. Game systems may include various handheld game devices, Internet-enabled game devices, etc. Client devices are capable of executing various different applications, such as various Internet-related applications, communication applications (such as email applications), short message service (SMS) applications, and may use various communication protocols.
[0030] The network 110 may be any type of network known to those skilled in the art that may support data communications using any of a variety of available protocols, including but not limited to TCP / IP, SNA, IPX, etc. By way of example only, the one or more networks 110 may be a local area network (LAN), an Ethernet-based network, a token ring, a wide area network (WAN), the Internet, a virtual network, a virtual private network (VPN), an intranet, an extranet, a blockchain network, a public switched telephone network (PSTN), an infrared network, a wireless network (e.g., Bluetooth, WIFI), and / or any combination of these and / or other networks.
[0031] Server 120 may include one or more general purpose computers, dedicated server computers (e.g., PC (personal computer) servers, UNIX servers, mid-range servers), blade servers, mainframe computers, server clusters, or any other suitable arrangement and / or combination. Server 120 may include one or more virtual machines running virtual operating systems, or other computing architectures involving virtualization (e.g., one or more flexible pools of logical storage devices that may be virtualized to maintain a server's virtual storage device). In various embodiments, server 120 may run one or more services or software applications that provide the functionality described below.
[0032] The computing units in the server 120 may run one or more operating systems including any of the above operating systems and any commercially available server operating systems. The server 120 may also run any of a variety of additional server applications and / or middle-tier applications, including HTTP servers, FTP servers, CGI servers, JAVA servers, database servers, etc.
[0033] In some implementations, server 120 may include one or more applications to analyze and consolidate data feeds and / or event updates received from users of client devices 101, 102, 103, 104, 105, and / or 106. Server 120 may also include one or more applications to display data feeds and / or real-time events via one or more display devices of client devices 101, 102, 103, 104, 105, and / or 106.
[0034] In some embodiments, the server 120 may be a server of a distributed system, or a server combined with a blockchain. The server 120 may also be a cloud server, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology. A cloud server is a host product in a cloud computing service system to solve the defects of difficult management and weak business scalability in traditional physical hosts and virtual private servers (VPS) services.
[0035] The system 100 may also include one or more databases 130. In some embodiments, these databases may be used to store data and other information. For example, one or more of the databases 130 may be used to store information such as audio files and video files. The databases 130 may reside in various locations. For example, the database used by the server 120 may be local to the server 120, or may be remote from the server 120 and may communicate with the server 120 via a network-based or dedicated connection. The databases 130 may be of different types. In some embodiments, the databases used by the server 120 may be, for example, relational databases. One or more of these databases may store, update, and retrieve data to and from the databases in response to commands.
[0036] In some embodiments, one or more of the databases 130 may also be used by applications to store application data. The databases used by the applications may be different types of databases, such as a key-value store, an object store, or a conventional store backed by a file system.
[0037] Figure 1 The system 100 may be configured and operated in various ways to enable the application of various methods and apparatuses described in the present disclosure.
[0038] At present, mainstream PDF document parsing methods usually rely on object detection models. First, various elements in the document (such as text, tables, formulas, etc.) are detected, and then dedicated models are used to parse different elements. However, this solution has significant limitations.
[0039] (1) System complexity: Multiple analytical models are required to work together, which leads to high engineering complexity of the overall system and increases development and maintenance costs.
[0040] (2) Limitations of traditional OCR methods: Although traditional OCR parsing methods (such as PaddleOCR, RapidOCR, and EasyOCR) perform well in plain text parsing, they lack the ability to parse formulas.
[0041] (3) Difficulties in table parsing: Due to the complexity of table structure and the diversity of styles, conventional table parsing models (such as SLANet and StrucTexT) are difficult to generalize well to various table styles, resulting in unstable parsing results.
[0042] These problems have become the main bottleneck in PDF document parsing tasks, and more efficient and robust technical solutions are urgently needed to break through the current technical barriers.
[0043] As multimodal technology has become more powerful in OCR tasks, the direct extraction of PDF document content through multimodal models (such as InternVL2, Qwen2VL, MiniCPM-V) has attracted widespread attention in recent years. These multimodal models rely on powerful model architectures and massive amounts of high-quality image and text training data, and not only excel in text parsing, but also show significant advantages in formula and table parsing.
[0044] However, most existing multimodal models are designed for general scenarios and are not trained for PDF document parsing. Therefore, if they are not optimized specifically (such as for specific prompts), random answers or hallucinations may occur.
[0045] In addition, single-step parsing models such as Nougat and GOT-OCR provide another end-to-end solution for PDF document parsing. However, due to the limitations of image input methods or model structures, problems such as incomplete parsing or missing details may occur.
[0046] In order to improve the effect of PDF parsing, the present disclosure provides a new data processing method for parsing PDF documents.
[0047] Figure 2 An exemplary flow chart of a data processing method according to an embodiment of the present disclosure is shown.
[0048] In step S202, at least one PDF element among a plurality of different types of elements identified from a portable document format PDF file is obtained.
[0049] In step S204, an image to be processed and text to be processed are determined based on the identified PDF elements, wherein the image to be processed includes an image of at least one identified PDF element, and the text to be processed includes text recognized from the image to be processed.
[0050] In step S206, feature extraction is performed on the text to be processed based on the self-attention mechanism to obtain self-attention features of the text to be processed.
[0051] In step S208, feature extraction is performed on the self-attention features of the text to be processed and the image features of the image to be processed based on the cross-attention mechanism to obtain text-image cross-attention features for the PDF file.
[0052] In step S210, a parsing result of the PDF file is determined based at least on the cross-attention feature.
[0053] By utilizing the data processing method provided by the present invention, by processing the image features corresponding to the images of different elements identified from the PDF file and the text features of the PDF based on the cross-attention mechanism, the accurate information of the visual input can be referred to during the autoregressive processing of the text features, thereby reducing the erroneous consequences caused by the hallucination phenomenon.
[0054] The principles of the present disclosure will be described in detail below.
[0055] In step S202, at least one PDF element among a plurality of different types of elements identified from a portable document format PDF file is obtained.
[0056] The PDF file may be processed by means of layout analysis to determine at least one PDF element included in the PDF file. In some embodiments, the different types of elements included in the PDF file may include at least one of the following: text, image, table, formula, etc. In some examples, the identified PDF elements may be further refined into: page header, image title, image, text title, text, table title, table, footnote, etc.
[0057] The layout analysis can be performed by performing a target detection task of a predetermined type of element on a PDF document. For example, the page can be partitioned according to the content type of the document (such as title, text, formula, table, image, etc.). In the example, with the support of massive annotated data and synthetic data, the SOTA (state-of-the-art) model can usually achieve a higher mAP (mean Average Precision) index, thereby providing a good analysis basis for subsequent PDF parsing. It is understandable that without departing from the principles of the present disclosure, any suitable method can be used to perform layout analysis on a PDF file.
[0058] In step S204, an image to be processed and text to be processed are determined based on the identified PDF elements, wherein the image to be processed includes an image of at least one identified PDF element, and the text to be processed includes text identified from the image to be processed. In some examples, the image to be processed may contain only the content of a single PDF element. In other examples, the image to be processed may contain the content of multiple PDF elements, for example, including a portion of the identified PDF elements or all of the PDF elements.
[0059] Through the layout analysis performed in step S202, the PDF page can be partitioned according to the different content types included in the PDF document, and images corresponding to the PDF elements of each type can be obtained. Furthermore, text recognition can be performed from the identified PDF elements to obtain the to-be-processed text of the PDF document included in the identified PDF elements.
[0060] In step S206, feature extraction is performed on the text to be processed based on the self-attention mechanism to obtain self-attention features of the text to be processed.
[0061] In some embodiments, the text to be processed can be processed using a deep learning model based on a self-attention mechanism to obtain the self-attention features of the text to be processed. For example, a Transformer network (or any other deep learning network capable of extracting text features) including at least one self-attention layer can be used to process the embedding vector of the text to be processed to obtain the self-attention features of the text to be processed. An exemplary deep learning model for extracting text features can be a Qwen model.
[0062] In step S208, feature extraction is performed on the self-attention features of the text to be processed and the image features of the image to be processed based on the cross-attention mechanism to obtain text-image cross-attention features for the PDF file.
[0063] Some self-attention layers in the deep learning model used to process the text to be processed can be replaced with cross-attention layers, or additional cross-attention layers can be inserted into the self-attention layers in the deep learning model, and the self-attention features of the text to be processed and the image features of the image to be processed can be processed by the cross-attention layers to obtain text-image cross-attention features. Using the above method, the accurate information of the visual input can be referenced in the process of feature extraction of the text, thereby reducing the erroneous consequences caused by the hallucination phenomenon.
[0064] In the example, four self-attention layers can be used to extract features from the embedding vector of the text to be processed in sequence to obtain the self-attention features of the text to be processed. Then, a cross-attention layer can be used to extract features from the self-attention features of the text to be processed and the image features of the image to be processed to obtain text-image cross-attention features. The structure of four self-attention layers connected to one cross-attention layer can be repeated to improve the accuracy of feature extraction.
[0065] Taking Qwen2.5-0.5B as an example, the model originally contains 24 decoder layers (DecoderLayer) of the self-attention structure. In the exemplary architecture of the present disclosure, the original structure of the model can be modified to add 6 DecoderLayer layers of the cross-attention structure, constituting a total of 30 DecoderLayer layers, of which the 3rd, 8th, 13th, 18th, 23rd, and 28th layers are cross-attention structures. This design ensures that the model can efficiently utilize visual information and reduce the probability of hallucinations during the generation process.
[0066] The DecoderLayer structure code of the exemplary cross-attention structure is as follows:
[0067]
[0068] The following will describe a method for acquiring image features of the image to be processed used in step S208.
[0069] The images to be processed corresponding to different types of PDF elements may have different sizes. The image processing process supporting dynamic resolution input may be implemented by dividing the image to be processed into a plurality of image units (tokens) of predetermined sizes, thereby retaining the information in the original image to the greatest extent.
[0070] In some embodiments, the image features of the image to be processed may be determined by dividing the image to be processed into a plurality of image units according to a predetermined size, and performing feature extraction on an image vector formed by the plurality of image units to obtain the image features of the image to be processed.
[0071] Before the image is input into the model, the length and width of the image can be fine-tuned to ensure that its size is divisible by the predetermined image unit size (patch_size). For example, the size can be adjusted so that the image is divisible by patch_size. Then the image is divided into multiple image units according to patch_size. For example, when patch_size=14, for an image with an original resolution of 166x29, the image size is first fine-tuned to 168×28, and then the image is divided into 24 image units, and the image vector formed by the 24 image units obtained by the division can be input into a deep learning model for feature extraction of visual information for processing, wherein each element in the image vector corresponds to a divided image unit. In the example, the adjusted image size can make the image divisible by n*patch_size, where n is the downsampling multiple used when compressing the image in the subsequent processing process. It can be understood that, without departing from the principles of the present disclosure, other methods can also be used to divide the image to be processed, so that images of different sizes to be processed can be converted into a combination of a series of image units based on the same method. Processed images of different input sizes can be converted into combinations of different numbers of image units, thereby avoiding information loss caused by forcibly adjusting the input image size.
[0072] In some embodiments, the image information of the divided image units can be encoded by two-dimensional rotation position encoding. In the related art, one-dimensional absolute position encoding is often used to encode image information. However, for images whose inherent attribute is two-dimensional, one-dimensional encoding will lose part of the position information. In the embodiments of the present disclosure, encoding the image information by two-dimensional rotation encoding can retain the two-dimensional position information in the image, and thus improve the accuracy of PDF file parsing. An exemplary two-dimensional rotation encoding can be 2D-RoPE. Those skilled in the art can select any other suitable two-dimensional rotation encoding means according to actual conditions.
[0073] In some embodiments, feature extraction can be performed on an image vector based on a self-attention mechanism to obtain self-attention features of multiple image units. For example, a deep learning model based on a self-attention mechanism can be used to extract features from an image vector formed by multiple image units. Before feature extraction is performed on the image vector, the image vector can be linearly transformed. In some examples, the linearly transformed image vector can better characterize the image information. An exemplary deep learning model for extracting image features can be SigLIP-SO400M. Without departing from the principles of the present disclosure, other suitable deep learning models can also be used to extract image features.
[0074] In some embodiments, before extracting the self-attention features of the text to be processed and the image features of the image to be processed based on the cross-attention mechanism, the self-attention features of multiple image units can be compressed to align the dimensions of the image features of the image to be processed and the self-attention features of the text to be processed. In the example, the Pixel Shuffle technology in InternVL2 can be used to compress the image features, and an exemplary compression ratio can be 0.5, that is, downsampling by 2 times. Taking the example of the aforementioned image vector length of 25 as an example, after Pixel Shuffle, the vector length is compressed to 6, thereby significantly reducing the subsequent calculation amount.
[0075] In step S210, a parsing result of the PDF file is determined based at least on the cross-attention feature.
[0076] The parsing result may be determined based on the output of the deep learning model used to process the text to be processed.
[0077] In an example, the parsing result may be in an editable Latex format. In another example, the parsing result may also be in an HTML format or any other suitable format. The parsing result may be edited to restore the layout of the PDF document, and functions such as extraction, processing, and analysis of the parsed PDF document content may be implemented.
[0078] In some embodiments, the method 200 can be implemented using a trained large model, wherein the large model can include a deep learning model for processing the text to be processed as a text processing backbone network, and includes a deep learning model for processing the image to be processed as an image processing backbone network. In some examples, prompt information can be input to the large model to indicate the form of the parsing result, for example, in Latex format. The above method can be used to guide the model to focus on the parsing task in the format specified by the prompt information, thereby reducing the learning complexity of the model.
[0079] The results of PDF parsing using the embodiments of the present disclosure can be used in the following scenarios:
[0080] ■ Document content extraction
[0081] ●Extract editable text from PDF files for easy indexing, analysis or archiving.
[0082] ●Parse table data into Latex or HTML format to facilitate subsequent data analysis and processing.
[0083] ■Data processing and analysis
[0084] ●Parse PDF files such as contracts and financial reports and extract key fields for quick reference or automated processing.
[0085] ●Extract charts, data and references from academic papers for research or literature management.
[0086] ■Education and teaching
[0087] ●Digitize the contents of textbooks, handouts or examination papers and integrate them into the online learning system.
[0088] ●Extract key information, charts or examples from teaching PDFs to assist with lesson preparation and classroom teaching.
[0089] ■Multi-language processing
[0090] ●Extract multilingual texts from PDF documents and translate them automatically.
[0091] ●Supports generation of Markdown documents or PDF files in the target language for archiving and publishing of multilingual content.
[0092] ■Tool support
[0093] ●As an open source tool, it provides basic PDF parsing functions to meet general document processing needs.
[0094] ●Supports parsing of PDF files in scanned or image form, solving the problem of content extraction from non-digital documents.
[0095] Using the above method, the embodiment of the present disclosure provides a lightweight multimodal model for PDF document parsing. The model can achieve unified parsing of text, formulas and tables through pre-set prompt information Prompt, so there is no need to rely on multiple models to parse different types of PDF elements. In addition, in terms of model architecture, by adaptively adjusting and dividing the input original image, the method of the embodiment of the present disclosure supports the original resolution input of the image, and can adaptively obtain more accurate visual information. Furthermore, by introducing the cross-attention mechanism of text information and visual information, the hallucination problem that occurs during the parsing process can be reduced, thereby significantly improving the completeness and accuracy of the parsing. The model and method for PDF parsing according to the embodiment of the present disclosure show high efficiency, flexibility and reliability in different scenarios, providing comprehensive technical support for various PDF document parsing tasks.
[0096] Figure 3 An exemplary process of a data processing method according to an embodiment of the present disclosure is shown.
[0097] like Figure 3As shown, the input PDF file 301 can be analyzed for layout, and multiple PDF elements can be obtained, including title 302, text 303, image title 304, table title 305, formula 306, table 307, and others 308. The images and texts corresponding to the PDF elements 302 to 307 can be input into the multimodal model 309. The multimodal model 309 can be a large model that has been trained and is trained to perform combined Figure 2 The described data processing method is used to obtain the parsing result 310 for PDF layout recovery. Other elements 308 may include information such as headers and footers. In an embodiment of the present disclosure, parsing other elements 308 may be abandoned.
[0098] Figure 4 FIG. 1 shows an exemplary structure of a multimodal model according to an embodiment of the present disclosure. Figure 4 Multimodal model implementation described Figure 3 Multimodal models in 309.
[0099] like Figure 4 As shown, the multimodal model may include an image processing backbone network 410 , a text processing backbone network 420 , and a connector 430 connecting the image processing backbone network 410 and the text processing backbone network 420 .
[0100] The image processing backbone network 410 can be implemented by a model structure such as SigLIP-SO400M to obtain stronger visual feature expression capabilities. The text processing backbone network can be implemented by a model structure such as Qwen2.5, which is suitable for the text parsing requirements in OCR tasks and avoids the redundancy of complex reasoning and understanding functions of the large language model LLM.
[0101] like Figure 4 As shown, the image processing backbone network 410 may include a linear transformation layer 411 and a first hidden layer 412. The linear transformation layer 411 may perform a linear transformation on the input image to be processed. Before the input linear transformation layer 411, the image processing backbone network 410 may be combined with the first hidden layer 412. Figure 2 The described method preprocesses the input image, such as dividing the input image into a plurality of image units of a predetermined size to achieve image input with dynamic resolution. The first hidden layer 412 may include a plurality of connected self-attention layers and a feedforward network. The first hidden layer 412 may be used to further extract the image features output by the linear transformation layer 411 to obtain the self-attention features of the input image.
[0102] The text processing backbone network 420 may include an embedding layer 421 and a second hidden layer 422. Among them, the embedding layer 421 can vectorize the input text and obtain the embedding vector of the input text. The second hidden layer 422 can be used to extract features from the embedding vector of the input text. Among them, the second hidden layer 422 may include a self-attention layer, a cross-attention layer, and a feedforward network. The self-attention layer in the second hidden layer 422 can be used to obtain the self-attention features of the input text, and the cross-attention layer can be used to jointly process the self-attention features of the text and the image features from the image processing backbone network 410 to consider visual information when generating text features. Connector 430 can be used to align the dimensions of the image features output by the image processing backbone network 410 and the self-attention features of the text. As mentioned above, the connector 430 can be implemented using Pixel Shuffle technology to improve the processing efficiency of the model by compressing the image features.
[0103] Figure 5 An exemplary block diagram of a data processing device according to an embodiment of the present disclosure is shown. Figure 5 The data processing device 500 shown in FIG. Figure 2 A data processing method 200 is described.
[0104] like Figure 5 As shown, the device 500 may include an acquisition unit 510, a text and image determination unit 520, a text feature extraction unit 530, a cross-attention feature extraction unit 540 and a parsing result determination unit 550.
[0105] The acquisition unit 510 may be configured to acquire at least one PDF element from among a plurality of different types of elements identified from a portable document format PDF file.
[0106] The text and image determination unit 520 may be configured to determine an image to be processed and text to be processed based on the identified PDF elements, wherein the image to be processed includes an image of at least one identified PDF element, and the text to be processed includes text recognized from the image to be processed.
[0107] The text feature extraction unit 530 may be configured to perform feature extraction on the text to be processed based on a self-attention mechanism to obtain a self-attention feature of the text to be processed.
[0108] The cross-attention feature extraction unit 540 may be configured to perform feature extraction on the self-attention features of the text to be processed and the image features of the image to be processed based on the cross-attention mechanism to obtain text-image cross-attention features for the PDF file.
[0109] The parsing result determining unit 550 may be configured to determine a parsing result of the PDF file based at least on the cross-attention feature.
[0110] In some embodiments, the image features of the image to be processed are determined by: dividing the image to be processed into a plurality of image units according to a predetermined size; and performing feature extraction on an image vector formed by the plurality of image units to obtain the image features of the image to be processed.
[0111] In some embodiments, the image elements are encoded using two-dimensional rotational position encoding.
[0112] In some embodiments, performing feature extraction on an image vector formed by a plurality of image units to obtain image features of an image to be processed includes: performing feature extraction on the image vector based on a self-attention mechanism to obtain self-attention features of a plurality of image units.
[0113] In some embodiments, before performing feature extraction on the image vector, a linear transformation is performed on the image vector.
[0114] In some embodiments, feature extraction of self-attention features of the text to be processed and image features of the image to be processed based on the cross-attention mechanism includes: using four self-attention layers to sequentially extract features of the embedding vector of the text to be processed to obtain the self-attention features of the text to be processed; using cross-attention layers to extract features of the self-attention features of the text to be processed and image features of the image to be processed to obtain text-image cross-attention features.
[0115] In some embodiments, the device 500 may also include a compression unit, which is configured to compress the self-attention features of multiple image units before extracting the self-attention features of the text to be processed and the image features of the image to be processed based on the cross-attention mechanism to align the dimensions of the image features of the image to be processed and the self-attention features of the text to be processed.
[0116] In some embodiments, the parsed results are in Latex format.
[0117] In some embodiments, the parsing result is in Latex format by inputting prompt information to the large model executing the data processing method.
[0118] In some embodiments, the multiple types of elements include: text, image, table, formula.
[0119] It should be understood that Figure 5 The various modules or units of the apparatus 500 shown in FIG. 5 can be used in conjunction with the reference Figure 2The steps in the method 200 described above correspond to each other. Therefore, the operations, features and advantages described above for the method 200 are also applicable to the device 500 and the modules and units included therein. For the sake of brevity, some operations, features and advantages are not repeated here.
[0120] Although specific functionality is discussed above with reference to specific modules, it should be noted that the functionality of the various units discussed herein may be separated into multiple units, and / or at least some functionality of multiple units may be combined into a single unit.
[0121] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0122] According to an embodiment of the present disclosure, an electronic device is also provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to the embodiment of the present disclosure.
[0123] According to an embodiment of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is further provided, wherein the computer instructions are used to enable the computer to execute the method according to the embodiment of the present disclosure.
[0124] According to an embodiment of the present disclosure, a computer program product is further provided, including a computer program, wherein the computer program implements the method according to the embodiment of the present disclosure when executed by a processor.
[0125] refer to Figure 6 , a block diagram of an electronic device 600 that can be used as a server or client of the present disclosure will now be described, which is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workbenches, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.
[0126] like Figure 6As shown, the electronic device 600 includes a computing unit 601, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the electronic device 600 can also be stored. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0127] Multiple components in the electronic device 600 are connected to the I / O interface 605, including: an input unit 606, an output unit 607, a storage unit 608, and a communication unit 609. The input unit 606 can be any type of device that can input information to the electronic device 600. The input unit 606 can receive input digital or character information and generate key signal input related to user settings and / or function control of the electronic device, and can include but is not limited to a mouse, a keyboard, a touch screen, a track pad, a track ball, a joystick, a microphone, and / or a remote controller. The output unit 607 can be any type of device that can present information, and can include but is not limited to a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 608 can include but is not limited to a disk, an optical disk. The communication unit 609 allows the electronic device 600 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks, and can include but is not limited to a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset, such as a Bluetooth device, an 802.11 device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.
[0128] The computing unit 601 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 601 performs the various methods and processes described above, such as method 200. For example, in some embodiments, the method 200 may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 608. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the computing unit 601, one or more steps of the method 200 described above may be performed. Alternatively, in other embodiments, the computing unit 601 may be configured to perform the method 200 in any other appropriate manner (e.g., by means of firmware).
[0129] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0130] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.
[0131] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0132] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0133] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), the Internet, and a blockchain network.
[0134] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server combined with a blockchain.
[0135] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.
[0136] Although the embodiments or examples of the present disclosure have been described with reference to the accompanying drawings, it should be understood that the above-mentioned methods, systems and devices are merely exemplary embodiments or examples, and the scope of the present invention is not limited by these embodiments or examples, but only by the claims after authorization and their equivalent scope. Various elements in the embodiments or examples can be omitted or replaced by their equivalent elements. In addition, each step can be performed in an order different from that described in the present disclosure. Further, the various elements in the embodiments or examples can be combined in various ways. It is important that with the evolution of technology, many elements described herein can be replaced by equivalent elements that appear after the present disclosure.
Claims
1. A data processing method, comprising: Retrieve at least one PDF element from among a plurality of different types of elements identified from a portable document format PDF file; Determining an image to be processed and text to be processed based on the identified PDF elements, wherein the image to be processed includes an image of at least one identified PDF element, and the text to be processed includes text identified from the image to be processed; Performing feature extraction on the text to be processed based on the self-attention mechanism to obtain the self-attention features of the text to be processed; Extracting the self-attention features of the text to be processed and the image features of the image to be processed based on a cross-attention mechanism to obtain text-image cross-attention features for the PDF file; and A parsing result of the PDF file is determined based at least on the cross-attention feature.
2. The data processing method according to claim 1, wherein: The image features of the image to be processed are determined in the following manner: Dividing the image to be processed into a plurality of image units according to a predetermined size; Feature extraction is performed on the image vector formed by the multiple image units to obtain image features of the image to be processed.
3. The data processing method according to claim 2, wherein: The image unit is encoded using two-dimensional rotational position encoding.
4. The data processing method according to claim 2, wherein: Extracting features from the image vectors formed by the plurality of image units to obtain image features of the image to be processed includes: Feature extraction is performed on the image vector based on a self-attention mechanism to obtain self-attention features of the multiple image units.
5. The data processing method according to claim 4, wherein: Before performing feature extraction on the image vector, a linear transformation is performed on the image vector.
6. The data processing method according to claim 1, wherein: Extracting features of the self-attention features of the text to be processed and the image features of the image to be processed based on the cross attention mechanism includes: Using four self-attention layers to sequentially extract features from the embedding vector of the text to be processed to obtain self-attention features of the text to be processed; A cross-attention layer is used to extract the self-attention features of the text to be processed and the image features of the image to be processed to obtain the text-image cross-attention features.
7. The data processing method according to claim 1, further comprising: Before extracting the self-attention features of the text to be processed and the image features of the image to be processed based on the cross-attention mechanism, the self-attention features of the multiple image units are compressed to align the dimensions of the image features of the image to be processed and the self-attention features of the text to be processed.
8. The data processing method according to claim 1, wherein: The parsing result is in Latex format.
9. The data processing method according to claim 8, wherein: The data processing method is executed by a large model, and prompt information is input into the large model to indicate that the parsing result has the Latex format.
10. The data processing method according to claim 1, wherein: The multiple types of elements include: text, image, table, and formula.
11. A data processing device, comprising: An acquisition unit configured to acquire at least one PDF element from among a plurality of different types of elements identified from a portable document format PDF file; a text and image determination unit configured to determine an image to be processed and a text to be processed based on the identified PDF elements, wherein the image to be processed includes an image of at least one identified PDF element, and the text to be processed includes text recognized from the image to be processed; A text feature extraction unit is configured to perform feature extraction on the text to be processed based on a self-attention mechanism to obtain a self-attention feature of the text to be processed; a cross-attention feature extraction unit configured to extract features of the self-attention features of the to-be-processed text and the image features of the to-be-processed image based on a cross-attention mechanism to obtain a text-image cross-attention feature for the PDF file; and A parsing result determining unit is configured to determine a parsing result of the PDF file based at least on the cross-attention feature.
12. An electronic device comprising: at least one processor; as well as a memory communicatively coupled to the at least one processor; in The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 10.
13. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-10.
14. A computer program product comprising a computer program, wherein: When the computer program is executed by a processor, the method according to any one of claims 1 to 10 is implemented.
Citation Information
Cited By
Artificial intelligence text analysis and extraction and key point source positioning method
CN120180252A