An intelligent parsing method for scientific paper PDFs based on multi-task learning

By performing layout analysis and multi-task learning on the PDF of scientific and technological papers, the problem of traditional methods being unable to identify complex typesetting elements is solved, efficient text structured recognition and layout analysis is achieved, and document parsing effect is improved.

CN119832580BActive Publication Date: 2025-07-22DOCUMENT & INFORMATION CENT OF CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411914709.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-24
Publication Date
2025-07-22
Estimated Expiration
2044-12-24

AI Technical Summary

Technical Problem

Existing text extraction methods cannot accurately identify complex typesetting elements such as formulas and charts in PDFs of scientific and technological papers, and the layout of different types of contents varies greatly, resulting in poor results in structured recognition and layout analysis.

Method used

Using a multi-task learning method, the layout analysis of the PDF of scientific and technological papers is carried out, and the positions of the formulas between pictures, tables and rows are identified and masked. Combined with OCR processing and deep multi-task analysis model, text position and content feature extraction are carried out to realize reading order, text structure function and citation extraction.

Benefits of technology

It improves the accuracy of OCR processing, enhances the accuracy of text structure function recognition, improves the speed and accuracy of reading order recognition, and improves the overall analytical efficiency and performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119832580B_ABST
    Figure CN119832580B_ABST
Patent Text Reader

Abstract

The present invention provides an intelligent parsing method for scientific and technological paper PDFs based on multi-task learning, which relates to the technical field of paper content mining and includes: masking the positions of pictures, tables and in-line formulas to obtain a masked picture of the scientific and technological paper page; performing OCR processing to obtain the text content and text positions; performing preprocessing to obtain text position features and text content features; and processing through a deep multi-task parsing model to simultaneously obtain the reading order recognition result, the text structure function recognition result, the paragraph detection result and the citation extraction recognition result. The present invention solves the technical problem that traditional text extraction methods usually cannot accurately identify complex layout elements such as formulas and charts when processing PDF formats, and the layout differences of different types of content on the page are relatively large. It is difficult for traditional methods to achieve accurate structured recognition and precise layout analysis, resulting in poor document parsing effects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of thesis content mining, and particularly to an intelligent parsing method for scientific and technological thesis PDFs based on multi-task learning. Background Art

[0002] Currently, the object of scientific and technological thesis content mining has gradually shifted from bibliographic information to full-text information. The fine-grained parsing of full-text content is crucial for many downstream tasks such as domain knowledge base construction and innovation evaluation.

[0003] The parsing of scientific and technological thesis PDF content is to extract the text content in the PDF file and convert it into a text storage form such as JSON. It is necessary to identify the reading order of the page text, convert it into a one-dimensional linear structure, extract the pictures, tables, and formulas in the page, and also identify the structural functions of the text, such as titles, abstracts, main texts, captions, references, etc.

[0004] Existing general PDF parsing methods usually only focus on restoring the reading order of the text in the PDF, removing content such as headers and footers, and cannot meet the requirements of content structuring for academic PDF parsing. Among the existing dedicated academic PDF parsing methods, for example, the GROBID project has good parsing effects on English journal PDFs, but poor parsing effects on Chinese academic PDFs, and there are differences in the content layout between Chinese and English thesis PDF files. Summary of the Invention

[0005] The present application provides an intelligent parsing method for scientific and technological thesis PDFs based on multi-task learning, aiming to solve the technical problems that traditional text extraction methods usually cannot accurately identify complex layout elements such as formulas and charts when processing PDF formats, and the layout differences of different types of content in the page are large, making it difficult for traditional methods to achieve accurate structural recognition and precise layout analysis, resulting in poor document parsing effects.

[0006] An intelligent parsing method for scientific and technological thesis PDFs based on multi-task learning disclosed in the present application includes: performing layout analysis on a scientific and technological thesis PDF file to obtain pictures and picture positions, tables and table positions, in-line formulas and in-line formula positions, masking the picture positions, the table positions, and the in-line formula positions to obtain a masked scientific and technological thesis page picture; performing OCR processing on the masked scientific and technological thesis page picture to obtain text content and text positions; preprocessing the text content and the text positions to obtain text position features and text content features; and processing the text position features and the text content features through a deep multi-task parsing model to simultaneously obtain a reading order recognition result, a text structure function recognition result, a paragraph detection result, and a citation extraction recognition result.

[0007] One or more technical solutions provided in this application have at least the following technical effects or advantages:

[0008] By performing layout analysis on scientific and technological paper PDF files, the pictures and their positions, tables and their positions, in-line formulas and their positions are accurately obtained and masked. This preprocessing step reduces the interference of images, tables, and in-line formulas on the text parsing process, thereby improving the accuracy of subsequent OCR processing and text content parsing; by combining the semantic features of the text with various position features, the layout information of the PDF document is fully utilized, improving the accuracy of text structure function recognition, especially the recognition accuracy of picture captions and table captions; by using the method of generating sorting scores through a regression function for reading order detection, compared with traditional pairwise sorting methods, the inference speed and accuracy of the text reading order recognition task are improved; by adopting a multi-task learning framework, the constructed parsing model can simultaneously perform four tasks: reading order detection, text function recognition, paragraph boundary recognition, and reference extraction, which not only improves the overall parsing efficiency but also promotes each other among the tasks, thereby improving the overall parsing performance.

[0009] The above description is only an overview of the technical solutions of this application. In order to be able to understand the technical means of this application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features, and advantages of this application more obvious and understandable, the specific embodiments of this application are specifically given below. Brief Description of the Drawings

[0010] Figure 1 This is a schematic flowchart of a method for intelligent parsing of scientific and technological paper PDFs based on multi-task learning provided by an embodiment of this application.

[0011] Figure 2 This is a schematic flowchart of a method for obtaining text position features and text content features in a method for intelligent parsing of scientific and technological paper PDFs based on multi-task learning provided by an embodiment of this application.

[0012] Figure 3 This is a schematic flowchart of a method for obtaining masked pictures of scientific and technological paper pages in a method for intelligent parsing of scientific and technological paper PDFs based on multi-task learning provided by an embodiment of this application. Detailed Embodiments

[0013] By providing a method for intelligent parsing of scientific and technological paper PDFs based on multi-task learning in an embodiment of this application, the technical problem that traditional text extraction methods usually cannot accurately identify complex layout elements such as formulas and charts when processing PDF formats, and the layout differences of different types of content on the page are large, and it is difficult for traditional methods to achieve accurate structured recognition and precise layout analysis, resulting in poor document parsing effects, is solved.

[0014] After introducing the basic principle of this application, various non-restrictive implementation manners of this application will be specifically introduced below in conjunction with the accompanying drawings of the specification.

[0015] As Figure 1 shown, an embodiment of this application provides a method for intelligent parsing of scientific and technological paper PDFs based on multi-task learning. The method includes:

[0016] Step S100: Perform layout analysis on the scientific and technological paper PDF file to obtain pictures and their positions, tables and their positions, in-line formulas and their positions, and mask the positions of the pictures, the positions of the tables, and the positions of the in-line formulas to obtain a masked scientific and technological paper page picture;

[0017] Use a PDF-to-image tool to convert each page in the scientific and technological paper PDF file into an image, generating a first scientific and technological paper page picture. Whether the PDF file is a scanned type or a text-selectable type, it is uniformly converted into a pure image format first. This is because when directly using existing PDF parsing libraries (such as PDFminer) to process text-selectable PDF files, the extracted text content is prone to garbled characters, and pictures and formulas are not fully extracted. When encountering a scanned PDF file, it is even impossible to process. Use a layout analysis model to analyze the converted page image, identify the positions of three types of elements therein, including pictures, tables, and in-line formulas. According to the obtained positions of the pictures, tables, and in-line formulas, crop the image, extract the regions containing these three types of elements, save them as separate images, and use a preset color, such as gray or white, to cover the regions of the pictures and tables in the cropped image region to obtain a masked scientific and technological paper page picture. At this time, the image only retains the text and formula information.

[0018] Step S200: Perform OCR processing on the masked scientific and technological paper page picture to obtain text content and text positions;

[0019] Use an OCR model to process the masked scientific and technological paper page picture. The OCR model extracts the text content by analyzing the characters and text arrangement in the image. Specifically, the OCR model converts the characters in the image into machine-readable text content through character recognition and provides the corresponding coordinate positions, that is, text box positions, for each text region, providing a basis for further text analysis.

[0020] Step S300: Preprocess the text content and text positions to obtain text position features and text content features;

[0021] Preprocess the text content and text position to obtain text character features, 1D position features, 2D position features, and page relative position features; preprocess special elements such as pictures, tables, and inline formulas obtained from layout analysis to obtain special element token markers and special element position features. Merge the extracted text character features and special element token markers to form a comprehensive content vector feature, which contains the content information of the text and special elements, and the model will use these vectors for subsequent parsing tasks.

[0022] In addition to the content vector feature, the position feature is also very important. The spatial distribution of text and formulas in the page is very helpful for understanding the structure of the document. Therefore, various extracted position features, including the 1D and 2D positions of the text and the spatial features of the formulas, are merged into a content position feature for subsequent model processing.

[0023] Step S400: Process the text position feature and the text content feature through a deep multi-task parsing model to obtain the reading order recognition result, text structure function recognition result, paragraph detection result, and reference extraction recognition result at the same time.

[0024] The deep multi-task parsing model includes a reading order recognition task channel, a paragraph detection task channel, a structure function recognition task channel, and a reference extraction task channel. Each channel is responsible for processing a certain aspect of the document. Each task channel will output a specific result according to the content vector feature and the position feature. Among them, each task channel has its own independent loss function, and is trained and optimized according to this loss function. The key to multi-task learning is to improve the generalization ability of the model by sharing part of the network structure and using the relationship between multiple tasks.

[0025] According to the reading order recognition task channel, output the reading order recognition result, that is, the sorting score or label of each text line. The model will output the correct order of the text lines according to these labels; according to the paragraph detection task channel, output the paragraph detection result, which is a binary classification result (1 or 0), where 1 indicates that the text line is the start of a paragraph, and 0 indicates otherwise; according to the structure function recognition task channel, output the text structure recognition result, indicating that the task recognizes different structural functions of the text, such as titles, texts, formulas, references, etc.; according to the reference extraction task channel, output the reference extraction recognition result, that is, the reference information in the text, such as cited literature, formula citations, chart citations, etc.

[0026] Furthermore, as Figure 2 shown, preprocess the text content and text position to obtain text position features and text content features, including:

[0027] Step S310: Preprocess the text content and text positions to obtain text character features, 1D position features, 2D position features, and page relative position features, and tokenize the text for the pictures, the tables, and the in-line formulas;

[0028] Step S320: Add the text character features and the text tokens into the text content features;

[0029] Step S330: Add the 1D position features, the 2D position features, and the page relative position features into the text position features.

[0030] Extract character information from each line of text output by OCR, and perform encoding processing on the text line content to obtain the vector representation of each character, that is, the text character features.

[0031] The 1D position feature refers to the position of the text line in the PDF page. The position of each line of text can be regarded as the order of the line on the page, that is, the arrangement order from top to bottom. Encode the text line content starting from 0 at the beginning, generate the relative position of the text in the page, and usually assign a numerical value to each text line in order, arranged from top to bottom, representing its serial number in the page.

[0032] The 2D position feature refers to the spatial position information of the text line. Considering its horizontal and vertical coordinates on the page, encode the position of each line of text (such as the coordinates of the beginning and end of the line). Specifically, normalize the upper left and lower right coordinates of the text line position, and also normalize the width and height of the page to a preset interval. For example, map the coordinate values to between 0 and 1, or map them to a fixed interval, such as [1, 1024].

[0033] The page relative position feature refers to the overall position of the text line relative to the page. This feature takes into account the position of the text line in the page and the position of the page in the entire document, calculates the ratio of the page number where the text line is located to the total number of pages in the entire document, and scales this ratio to a preset interval.

[0034] For elements such as pictures, tables, and in-line formulas, they need to be specially identified during the document parsing process. These elements do not directly contain processable text content, but are represented in the form of special tokens. These tokens act as placeholders in subsequent processing stages to help the model identify and distinguish these non-text elements.

[0035] Combine the text character features and the text tokens to form a complete text content feature. This feature contains the semantic features of the text itself (such as text embeddings) and the identification features of images, tables, etc., which is convenient for subsequent processing.

[0036] Combine the 1D position feature, 2D position feature, and page relative position feature to form a complete text position feature, which helps the model more accurately understand the spatial distribution and sequential relationship of text lines on the page.

[0037] Furthermore, as Figure 3 shown, obtain the masked scientific and technological paper page image, including:

[0038] Step S110: Receive the scientific and technological paper PDF file and convert it into an initial scientific and technological paper page image;

[0039] Step S120: Analyze the initial scientific and technological paper page image through a layout analysis model to obtain the image position, table position, and in-line formula position;

[0040] Step S130: Crop the initial scientific and technological paper page image according to the image position, the table position, and the in-line formula position to obtain images, tables, and in-line formulas;

[0041] Step S140: Mask the initial scientific and technological paper page image with a preset color according to the image position, the table position, and the in-line formula position to obtain the masked scientific and technological paper page image.

[0042] Receive and load the scientific and technological paper PDF file to be parsed. At this time, the file is still in a mixed format of text and images and may contain content such as text, formulas, images, tables, etc. Use software tools or scripts to convert each page into a high-quality image to obtain the initial scientific and technological paper page image. This step uniformly converts the PDF file into an image format for subsequent processing.

[0043] Layout analysis aims to identify and extract various elements from a document image, such as text, images, tables, headings, etc., and understand their spatial and logical relationships. This technology uses a variety of methods, including rule-based, machine learning-based, and deep learning-based, to segment the document image into different semantic regions and classify and extract relationships between these regions. Layout analysis has a wide range of applications, playing a crucial role from document digitization and information extraction to document understanding and robotic process automation, providing a basis for efficient processing and understanding of document information.

[0044] The layout analysis model is an algorithm used to extract and identify different elements from a page image, aiming to divide the page into multiple regions and classify them according to the content of these regions. For example, by using deep learning methods such as convolutional neural networks (CNNs) or region-based convolutional neural networks (R-CNNs), elements in the image can be efficiently and accurately extracted and classified. After model analysis, the positions of the picture, table, and the formula are output, and these positions are usually represented in the form of rectangular boxes. Among them, pictures refer to images, charts, illustrations, etc. in the page, and these pictures usually have specific identifiers such as image resolution, color mode, etc.; tables refer to table elements in the page, and these tables usually contain information such as data, row and column structures.

[0045] According to the positions of the picture, table, and in-line formula, the initial scientific research paper page image is cropped. For the cropping of the formula, precise cropping can be carried out according to the rectangular area where the formula is located to ensure that only the formula content is retained. Similarly, for the table and picture, cropping is also carried out according to their positions to avoid the influence of other unnecessary regions on subsequent processing. Through cropping, regions containing formulas, pictures, or tables are obtained.

[0046] In order to prevent the cropped area from affecting OCR processing, a preset color is selected to cover the cropped area, such as gray, white, or other background colors. The purpose of doing this is to clear the unwanted image areas and only retain the text areas, thereby avoiding interfering with OCR recognition. After the covering is completed, a covered scientific research paper page image is generated, which contains the covered areas, while the text part remains unchanged.

[0047] Furthermore, the text content and text positions are preprocessed to obtain text character features, 1D position features, 2D position features, and page relative position features. Tokenizing the text for the pictures, the tables, and the in-line formulas includes:

[0048] Step S311: According to the text content and the text positions, obtain the text line content and text line positions;

[0049] Step S312: Encode the text line content to obtain text character features;

[0050] Step S313: Starting from 0 at the beginning, encode according to the text line content to obtain the 1D position features;

[0051] Step S314: Perform 2D position encoding on the text line positions to obtain the 2D position features, including:

[0052] Step S3141: Normalize the upper-left corner coordinates and lower-right corner coordinates of the text line position, as well as the width and height of the page to which the text line position belongs, into a preset interval to obtain the 2D position feature;

[0053] Step S315: Encode the page position of the text line position to obtain the page relative position feature, including:

[0054] Step S3151: Calculate the ratio of the page number of the text line to the total number of pages of the text line position, and scale the ratio to a preset interval to obtain the page relative position feature;

[0055] Step S316: Identify the text tokens for the picture, the table, and the in-line formula.

[0056] Use an OCR model to perform text detection on the PDF page to obtain all text line information on the page, including the content and position of each line. Extract the character information in the OCR output by text line to form independent text line content. For example, different paragraphs and headings on a page will be divided into multiple lines, and the content of each line is a continuous character sequence; obtain the position information of the text line in the page. The position information of each line includes the upper-left corner coordinates and the lower-right corner coordinates. The text line position can represent the spatial range of the text in the page and provide a basis for subsequent encoding of the position information.

[0057] Convert the characters of the text content into feature vectors. First, tokenize the string of each text line. Tokenization can be performed using WordPiece or SentencePiece. During tokenization, it is necessary to ensure that each number is separated separately, which is particularly important for subsequent reference extraction tasks. The features of each character can be represented by character embeddings. The obtained text character features can help the model understand the semantic information of the text.

[0058] The 1D position feature is mainly used to represent the relative position of each character in its text line. For each line of text, starting from the first character of the line, encode it in the order of the characters within the line, and assign a unique number (starting from 0) to each character. This number represents the relative position of the character in the text line. For example, the 1D position feature of the first character in a line of text is 0, the second character is 1, and so on. The 1D position feature helps the model understand the order and structure of the characters within the text line.

[0059] The 2D position feature is used to represent the two-dimensional spatial position of a text line on a page. For each text line, the coordinates of the upper left corner and the lower right corner of its bounding box on the page are extracted. These coordinates reflect the horizontal and vertical positions of the text line on the page. To eliminate the influence of page size differences on the features, first, the page size is scaled to a fixed layout interval, such as [1200, 1800], so that one side of the page is the same as the length or width of the interval. Then, the scaling ratio is calculated, and based on this scaling ratio value, the coordinate values of all text lines, including width and height, are normalized and scaled to a preset interval [1200, 1800]. The 2D position feature helps the model understand the specific position of the text line on the page. For example, whether the text line is in the upper left corner or the lower right corner of the page, whether it is adjacent to other text lines, and the relative distance between them.

[0060] The page relative position feature aims to represent the relative position of a text line in the entire PDF document. Specifically, for each text line, we first obtain the page number where the text line is located, and then calculate the ratio of this page number to the total number of pages in the document. For example, if a text line is on page 3 of a document with a total of 10 pages, then the page relative position feature value of this text line is 3 / 10 = 0.3. This ratio can be regarded as a kind of normalized position information, which maps the position of the text line in the document to an interval of [0, 1]. The closer the value is to 0, the closer the text line is to the beginning of the document; the closer the value is to 1, the closer the text line is to the end of the document.

[0061] The functions of the page relative position feature are as follows: 1) Provide global context information. The page relative position feature provides the position information of the text line in the entire document, which helps the model understand the context of the text. For example, the title usually appears at the beginning of the document, while the references usually appear at the end of the document. 2) Distinguish different types of content. Different types of content usually have certain regularities in their positions in the document. For example, the abstract usually appears at the beginning of the document, while the conclusion usually appears at the end of the document. The page relative position feature can help the model distinguish these different types of content.

[0062] In addition to the construction of text element features, three types of special elements, namely images, tables, and in-line formulas, also need to be taken into consideration. By introducing special element tokens and their position information, the model can distinguish different types of content and understand the spatial relationships between them and the text. For example, the character token 21000 represents a table, 21001 represents an image on the page, and 21002 represents an in-line formula on the page. The information of special elements is crucial for the reading order recognition task. The presence of multi-column tables and images on the page can change the original normal text reading order. After merging and splicing the token marks of special elements with text character marks, a comprehensive content vector feature is formed, which includes both the semantic information of the text and the structural information of the formula.

[0063] Combine different types of position features such as the 1D position feature, the 2D position feature, the page relative position feature, and the formula position feature into a unified content position feature. The splicing method can also be used to merge these position features. Through these features, the model can learn the spatial layout information of the text and the formula on the page, including their relative positions, arrangement patterns, and whether they are in the same paragraph, etc.

[0064] Furthermore, the deep multi-task parsing model includes:

[0065] A shared encoder for extracting the representation of input features; a reading order recognition task channel that constructs a regression function based on a multi-layer fully connected network to predict the sorting score of text lines; a text structure function recognition task channel that constructs a multi-classifier based on a multi-layer fully connected network to predict the structural function type of text lines; a paragraph detection task channel that constructs a binary classifier based on a multi-layer fully connected network to predict whether a text line is the starting sentence of a paragraph; a reference extraction task channel that is constructed based on a sequence annotation head to predict the reference positions in the text.

[0066] Multi-task learning is a machine learning method aimed at improving the generalization ability of the model by simultaneously learning multiple related tasks. Different from training models separately in single-task learning, multi-task learning trains a single model to complete multiple tasks. Specifically, the model uses a pre-trained causal language model as the shared text representation encoder. In this embodiment, the parameters of the pre-trained LayoutLMv3 are selected as the initialization parameters of the model. The shared encoder is the core component of the model, which is used to uniformly encode the input multi-modal features and extract global and local representations. The purpose of the shared encoder is to share representations between different tasks to enhance the collaborative learning ability between tasks and reduce redundant calculations.

[0067] The reading order recognition task channel is used to predict the reading order of text. Reading order recognition refers to arranging the text in a PDF page into a one-dimensional linear sequence according to the normal human reading order. The reading order recognition task channel uses the semantic features and positional features of the text to infer its correct arrangement order. Usually, by analyzing the up, down, left, and right positional relationships of each text block on the page, the model can determine which text should appear first and which should appear later. The model will establish relationships between the text blocks on the page according to these rules and output the reading order recognition result, which is an ordered text sequence that rearranges the text blocks on the PDF page into a linear sequence to ensure that the text conforms to the order of human reading literature. For the captions of pictures and tables, we distinguish their reading order from that of the main text. We believe that images and tables generally appear as floating elements in the main text. Therefore, the model predicts the reading order of the main text and caption text separately. The reading order of caption text includes the captions of pictures and tables. When performing segmentation, the results of paragraph detection and structure function recognition will be used together. If a text line is recognized as starting a new paragraph, it means that the text line is the caption text of the next element.

[0068] Structure function recognition refers to classifying each line of text into the following 10 categories: table caption, picture caption, main text paragraph, title (including main title and subtitle), abstract, author name, keyword, institution, reference, other auxiliary information (including header, footer, copyright information, etc.). The structure function recognition channel uses deep learning models such as BERT, LSTM, CNN, etc., mainly using semantic features and assisted by positional features to determine which of the above 10 categories the text line belongs to. For each input text line, the model will output a category label to obtain the text structure recognition result, indicating which part of the document the text belongs to, such as title, main text paragraph, table caption, etc.

[0069] Paragraph detection refers to correctly breaking text into lines, where a title belongs to a separate line, a paragraph belongs to a separate line, a reference belongs to a separate line, the abstract content belongs to a separate line, and the keyword content belongs to a separate line. The paragraph detection model infers which text belongs to the same paragraph and which text should be divided into different paragraphs based on the spatial distance between text blocks, line spacing, and layout structure. Usually, there are obvious intervals between paragraphs, and there are also certain semantic differences between the title and the body text. The model distinguishes the title and the body text, the body text and the references, etc. according to the position features. For example, usually the title is relatively short, while the body text is in a regular layout, and the references are usually located at the bottom of the page and separated from the body text. After analysis, the paragraph detection result is output, which is a tokenized result indicating whether each text block belongs to a paragraph. The paragraph detection task channel can not only identify the body text paragraphs, but also distinguish different content types such as titles, paragraphs, and references.

[0070] For the reading order recognition task, the paragraph recognition task, and the input of the multi-layer fully connected neural network corresponding to the structural function recognition task is the first character of each line of text. The first character can be a letter, a number, or a separately added CLS character. The first character integrates all the information of the text line and can complete the reading order recognition, paragraph detection, and output of the structural function recognition task based on the features of this first character. While in the citation extraction task channel, the encoded layer feature representation of all the characters of the complete text line is input, not just the first character. The model uses all the characters of the text line for the classification prediction of citation marking.

[0071] Citation extraction refers to the recognition of the citation position in the body text of the literature PDF. For example, in the text "Zhang San et al. [1] proposed...", [1] is the citation mark. In the text "Maning(2018) on the chemical properties of deep crustal fluids and the carriers of these fluids", Maning(2018) is the citation mark. The citation extraction task channel uses a deep learning model to process the input features. The model analyzes the structure of the text and identifies possible citation marks such as "[1]", "author, year", etc. The citation marks usually have a fixed format, so the model extracts the citation information by identifying these formats and outputs the citation extraction recognition result, which is the recognized citation information.

[0072] Furthermore, the deep multi-task parsing model is obtained by training with a multi-task loss function, and the multi-task loss function is:

[0073] Total loss = w1*Loss1 + w2*Loss2 + w3*Loss3 + w4*Loss4

[0074] where Total lossCharacterize the overall loss value of multi-task parsing. w1, w2, w3, and w4 represent the weights of each loss value. Loss1 represents the ListMLE loss function of the reading order recognition task. Loss2 represents the cross-entropy loss function of the text structure function recognition task. Loss3 represents the cross-entropy loss function of the paragraph detection task. Loss4 represents the cross-entropy loss function of the reference extraction task.

[0075] Specifically, the formula of the multi-task loss function is as follows:

[0076] Total loss = w1 * Loss1 + w2 * Loss2 + w3 * Loss3 + w4 * Loss4;

[0077] Among them, Loss2, Loss3, and Loss4 are cross-entropy loss functions, and Loss1 is the ListMLE loss function. The ListMLE loss function is a loss function used for ranking learning. It directly optimizes the likelihood probability of the entire sorted list. This method is based on the Plackett-Luce model, which treats each sorted list as a permutation and tries to maximize the probability of the observed permutation. In practical applications, ListMLE minimizes the loss by calculating the negative log-likelihood. Its goal is to make the ranking generated by the model as close as possible to the ideal ranking. Compared with pointwise and pairwise methods, ListMLE takes into account the relative order relationship of the entire document sequence, so it usually can obtain better ranking results.

[0078] This method trains the constructed parsing model using the multi-task learning method, optimizing and learning the 4 tasks simultaneously. During the model training process, for the setting of the four loss weights w1, w2, w3, and w4, multiple methods are adopted, including manually setting fixed weights, DWA, and GradNorm methods to adaptively adjust and learn the weights. According to this multi-task parsing loss function, the initial reading order recognition task channel, the initial paragraph detection task channel, the initial structure function recognition task channel, and the initial reference extraction task channel are trained to obtain the deep multi-task parsing model.

[0079] Furthermore, the output of the reading order recognition task channel is the sorting score of text lines. Among them, the reading order recognition task includes the paragraph text sorting task and the caption text sorting task. The losses of the paragraph text sorting task and the caption text sorting task are calculated separately. The loss of the reading order recognition task is the sum of the paragraph text sorting task loss and the caption text sorting task loss.

[0080] The main output of the reading order recognition task channel is the sorting score for each line of text. This sorting score is a continuous numerical value that represents the reading order of the text lines on the page. The higher the score, the higher the priority. Among them, the reading order recognition task includes the paragraph text sorting task and the caption text sorting task. The paragraph text sorting task is responsible for sorting the text lines of the main text paragraphs, with the goal of ensuring that the text lines within the paragraphs are arranged in the correct contextual order, thus restoring the natural reading order. The caption text sorting task is responsible for sorting the captions of pictures and tables, with the goal of ensuring that the caption text is correctly aligned with the associated picture or table and sorted in the correct reading order.

[0081] The paragraph text sorting task and the caption text sorting task each use the ListMLE sorting loss function to directly optimize the probability of the sorted list. By maximizing the ideal sorting, it ensures that the sorting result of the paragraph text is close to the actual reading order. The total loss of the final reading order recognition task is the sum of the loss of the paragraph text sorting task and the loss of the caption text sorting task. This way of calculating the loss independently ensures that the model can optimize the sorting quality of the paragraph text and the caption text separately, while also enabling joint optimization through the total loss.

[0082] Furthermore, the output of the text structure and function recognition task channel is the structure and function type of the text line, and the predicted output includes at least 10 categories. The 10 categories include: table caption, picture caption, paragraph text, title, abstract, author name of the paper, keywords, institution, reference, and unclassified. Among them, the title includes the Chinese title of the paper and chapter titles, and the unclassified includes content from other articles, English title, English abstract, header, and footer.

[0083] The goal of the text structure and function recognition task channel is to predict the structure and function category to which each line of text belongs, and the output is the specific function type label for each line of text, with a total of 10 categories. Among them, the table caption represents the explanatory text related to the table, usually appearing immediately after the table; the picture caption represents the explanatory text related to the picture, usually appearing immediately after the picture; the paragraph text is the ordinary paragraph in the main text, usually the main content of the article; the title includes the main title and subtitle, such as the chapter title of the section; the abstract is the abstract part of the paper, usually after the title and author information, used to summarize the content of the paper; the author name of the paper is the name information of the paper author, usually after the title or abstract; the keywords refer to the keyword part of the paper, usually used to mark the core content or field of the research; the institution refers to the institution information of the author, usually after the author name; the reference refers to the list of references cited in the paper, usually at the end of the article; the unclassified includes other text that does not belong to the above categories, specifically including content from other articles, English title, English abstract, header, and footer, etc.

[0084] Furthermore, the output of the paragraph detection task channel is 1 or 0. When it is 1, it indicates that the text is the starting sentence of a paragraph. When it is 0, it indicates that the text line is not the starting sentence of a paragraph.

[0085] The paragraph detection task channel uses a multi-layer fully connected network to construct a binary classifier for discriminating the paragraph structure of the text. After obtaining the representations of each character in the encoding layer, only the first character of each text line is input into the binary classifier. The output result of the classifier being 0 represents that the text line does not start a new paragraph, and the result being 1 represents that the text line is the starting sentence of a paragraph and will start a new paragraph.

[0086] Furthermore, after obtaining the text character features, it further includes:

[0087] Step S460: Obtain the first and last characters of the text line according to the text character features;

[0088] Step S470: Input the first and last characters of the text line into the reading order recognition task channel, the paragraph detection task channel, and the text structure function recognition task channel with the text character features;

[0089] Step S480: Input the complete text character features into the reference extraction task channel with the text character features.

[0090] For each line of text, the first character of the line is extracted according to the character features to obtain the first and last characters of the text line. The first and last characters can be letters, numbers, or other symbols, depending on the format of the text. The first and last characters provide information about the reading order. Especially in documents with multi-column or multi-paragraph layouts, the first and last characters help determine the starting position of the text block.

[0091] The first and last characters of the text line are input into the reading order recognition task channel, the paragraph detection task channel, and the text structure function recognition task channel. Among them, in the reading order recognition task channel, the first and last characters help determine the correct order of the text. The first and last characters provide the starting identifier of the text line, helping the model infer the relative order of each text line on the page. In the paragraph detection task channel, the first and last characters help identify the starting sentence of the paragraph. Usually, the first line of a paragraph has a unique layout (such as a first-line indent). The extraction of the first and last characters can help the model identify the paragraph separation. In the text structure function recognition task channel, the first and last characters help the model understand the structure of the text. For example, a title usually starts with a specific character (such as a capital letter), and the body paragraphs usually have different layouts. The first and last characters can provide these clues. Each task channel combines the first and last characters of the text line with other text features for task-specific processing.

[0092] Input the complete text character features in the reference extraction task channel. In the reference extraction task channel, the input is the complete text character features, not just the first and last characters. The model uses the text content extracted by OCR (including the whole sentence or paragraph) and combines the character features of the text to identify the reference markers.

[0093] In summary, the technical effects of a scientific paper PDF intelligent parsing method based on multi-task learning provided by the embodiments of the present application are as follows:

[0094] By performing layout analysis on the scientific paper PDF file, accurately obtain the pictures and their positions, tables and their positions, in-line formulas and their positions, and mask them. This preprocessing step reduces the interference of images, tables, and in-line formulas on the text parsing process, thereby improving the accuracy of subsequent OCR processing and text content parsing; by combining the semantic features of the text with various position features, make full use of the layout information of the PDF document, improve the accuracy of text structure function recognition, especially the recognition accuracy of picture captions and table captions; by using the regression function to generate the sorting score for reading order detection, compared with the traditional pairwise sorting method, improve the inference speed and accuracy of the text reading order recognition task; by adopting the multi-task learning framework, the constructed parsing model can simultaneously perform four tasks: reading order detection, text function recognition, paragraph boundary recognition, and reference extraction, which not only improves the overall parsing efficiency, but also promotes each other among the tasks, thereby improving the overall parsing performance.

[0095] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for intelligent parsing of scientific paper PDFs based on multi-task learning, characterized in that The method includes: Performing layout analysis on a scientific and technological paper PDF file to obtain images and their positions, tables and their positions, in-line formulas and their positions, masking the positions of the images, the positions of the tables, and the positions of the in-line formulas to obtain a masked scientific and technological paper page image; Performing OCR processing on the masked scientific and technological paper page image to obtain text content and text positions; Preprocessing the text content and text positions to obtain text position features and text content features; Processing the text position features and the text content features through a deep multi-task parsing model to simultaneously obtain a reading order recognition result, a text structure function recognition result, a paragraph detection result, and a citation extraction recognition result; Preprocessing the text content and text positions to obtain text position features and text content features, including: Preprocessing the text content and text positions to obtain text character features, 1D position features, 2D position features, and page relative position features, and labeling text tokens for the images, the tables, and the in-line formulas; Adding the text character features and the text tokens into the text content features; Adding the 1D position features, the 2D position features, and the page relative position features into the text position features; Preprocessing the text content and text positions to obtain text character features, 1D position features, 2D position features, and page relative position features, and labeling text tokens for the images, the tables, and the in-line formulas, including: Obtaining text line content and text line positions according to the text content and the text positions; Encoding the text line content to obtain text character features; Encoding from the first position starting from 0 according to the text line content to obtain the 1D position features; Performing 2D position encoding on the text line positions to obtain the 2D position features, including: Normalizing the upper left coordinate and the lower right coordinate of the text line position, as well as the width and height of the page to which the text line position belongs, to a preset interval to obtain the 2D position features; Encoding the page position of the text line position to obtain the page relative position features, including: Calculating the ratio of the text line page number of the text line position to the total number of pages, and scaling the ratio to a preset interval to obtain the page relative position features; Labeling the images, the tables, and the in-line formulas as special text tokens.

2. The intelligent PDF parsing method for scientific papers based on multi-task learning according to claim 1, characterized in that Performing layout analysis on a scientific and technological paper PDF file to obtain a masked scientific and technological paper page image, including: Receiving a scientific and technological paper PDF file and converting it into an initial scientific and technological paper page image; Analyzing the initial scientific and technological paper page image through a layout analysis model to obtain image positions, table positions, and in-line formula positions; Cropping the initial scientific and technological paper page image according to the image positions, the table positions, and the in-line formula positions to obtain images, tables, and in-line formulas; According to the positions of the pictures, the tables, and the in-line formulas, use a preset color to cover the pictures on the initial scientific and technological paper page to obtain a covered scientific and technological paper page picture.

3. A method for intelligent parsing of scientific paper PDFs based on multi-task learning as claimed in claim 1, characterized in that The deep multi-task parsing model includes: A shared encoder for extracting the representation of the input features; A reading order recognition task channel that constructs a regression function based on a multi-layer fully connected network to predict the sorting scores of text lines; A text structure function recognition task channel that constructs a multi-classifier based on a multi-layer fully connected network to predict the structural function types of text lines; A paragraph detection task channel that constructs a binary classifier based on a multi-layer fully connected network to predict whether a text line is the starting sentence of a paragraph; A citation extraction task channel that is constructed based on a sequence annotation head to predict the citation positions in the text.

4. The intelligent PDF parsing method for scientific papers based on multi-task learning according to claim 3, wherein, The deep multi-task parsing model is obtained by training with a multi-task loss function, and the multi-task loss function is: Total loss = w1 * Loss1 + w2 * Loss2 + w3 * Loss3 + w4 * Loss4 Among them, Total loss represents the overall loss value of multi-task parsing. w1, w2, w3, and w4 represent the weights of each loss value. Loss1 represents the ListMLE loss function of the reading order recognition task. Loss2 represents the cross-entropy loss function of the text structure function recognition task. Loss3 represents the cross-entropy loss function of the paragraph detection task. Loss4 represents the cross-entropy loss function of the citation extraction task.

5. The intelligent PDF parsing method for scientific papers based on multi-task learning according to claim 3, wherein The output of the reading order recognition task channel is the sorting score of the text line. Among them, the reading order recognition task includes a paragraph text sorting task and a caption text sorting task. The losses of the paragraph text sorting task and the caption text sorting task are calculated separately, and the loss of the reading order recognition task is the sum of the paragraph text sorting task loss and the caption text sorting task.

6. The intelligent PDF parsing method for scientific papers based on multi-task learning according to claim 3, characterized in that The output of the text structure function recognition task channel is the structural function type of the text line, and the predicted output includes at least 10 categories. The 10 categories include: table caption, picture caption, paragraph text, title, abstract, paper author name, keyword, institution, reference, and unclassified. Among them, the title includes the Chinese version of the paper title and chapter titles, and the unclassified includes content from other articles, English version title, English version abstract, header, and footer.

7. The intelligent PDF parsing method for scientific papers based on multi-task learning according to claim 3, characterized in that The output of the paragraph detection task channel is 1 or 0. When it is 1, it indicates that the text line is the starting sentence of a paragraph, and when it is 0, it indicates that the text line is not the starting sentence of a paragraph.

8. The intelligent PDF parsing method for scientific papers based on multi-task learning according to claim 1, characterized in that After obtaining the text character features, it further includes: According to the text character features, obtain the first and last characters of the text line; The first and last characters of the text line are input into the reading order recognition task channel, the paragraph detection task channel, and the text structure function recognition task channel with the text character features; The complete text character features are input into the citation extraction task channel with the text character features.

Citation Information

Patent Citations

  • Document image intelligent analysis and processing method based on multi-modal information

    CN117173730A

  • KR20240114180A