PDF content positioning method based on large model

Through the PDF content positioning technology based on large models, the problems of low efficiency and poor accuracy of large PDF files are solved, efficient and accurate information extraction is achieved, and the digital transformation of enterprises is supported.

CN120144752APending Publication Date: 2025-06-13NANJING WANDE INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510303115.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently process large PDF files, resulting in low information extraction efficiency, poor accuracy, and easy introduction of human errors.

Method used

Using PDF content positioning technology based on big models, sample PDF files are collected and marked through the preliminary data preparation stage, corpus and positioning model libraries are built, and pre-trained language models are used for deep learning vector processing, vector query and similarity evaluation are realized to quickly locate and extract key information.

Benefits of technology

It improves the retrieval efficiency and accuracy of specific content in PDF documents, reduces manual intervention, reduces error rate, improves the overall quality of data processing, and supports the digital transformation and intelligent upgrade of enterprises.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144752A_ABST
    Figure CN120144752A_ABST
Patent Text Reader

Abstract

The invention discloses a PDF content positioning method based on a large model, and the method is characterized in that the method comprises the following steps: 1, carrying out the early-stage data preparation; 2, extracting paragraphs of the target PDF file; and step 3, processing paragraphs of the target PDF file. According to the method, the retrieval efficiency and accuracy of the specific content in the PDF document can be improved, human resources can be greatly liberated, the working efficiency is improved, the error rate is reduced, and powerful technical support is provided for digital transformation and intelligent upgrading of enterprises. In addition, due to the introduction of the closed-loop training step, the system can be continuously self-optimized to adapt to continuously changing business requirements and data characteristics.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a large model-based PDF content location technology, which is mainly applied to text processing and information retrieval, especially in scenarios where efficient content location of large PDF documents is required. Background Art

[0002] In the digital age, enterprises are faced with the need to process a large number of PDF files, which contain rich business data and information. However, the process of manually processing and extracting information from large PDF files is not only time-consuming and laborious, but also prone to errors due to human factors, affecting work efficiency and accuracy. Especially in business processes, when it is necessary to accurately identify and extract key fields and information from a large amount of unstructured data, this challenge becomes even more obvious.

[0003] With the increase in business volume and the expansion of file size, traditional manual processing methods are no longer able to meet the needs of enterprises. When manually processing large PDF documents, not only a large amount of time is spent on searching and comparing information, but also due to the complexity of the file content, it is difficult to ensure that every piece of information can be accurately identified and processed. In addition, the inevitable omissions and errors in manual operations may lead to the omission of important information or incorrect entry, thus affecting the accuracy of business decisions and processes.

[0004] In order to improve work efficiency and reduce errors, enterprises need an automated technology to assist or replace manual processing of large PDF files. This technology needs to be able to quickly locate key paragraphs in the file, identify the information that needs to be processed, and at the same time filter out irrelevant "junk paragraphs" to reduce interference in the processing process. In this way, the efficiency and accuracy of information extraction can be greatly improved, manual intervention can be reduced, costs can be lowered, and the overall quality of data processing can be enhanced. Summary of the Invention

[0005] The object of the present invention is to provide a method that can effectively process large PDF files and accurately and quickly locate key information.

[0006] To achieve the above object, the technical solution of the present invention discloses a large model-based PDF content location method, which is characterized by including the following steps:

[0007] Step 1, preliminary data preparation, specifically including:

[0008] Step 101, define the target business and the business fields of the target business, where the business fields represent the key information that needs to be particularly concerned about and extracted when processing PDF files for the target business;

[0009] Step 102, prepare and collect sample PDF files:

[0010] For each business field f in the target business, collect a set of representative sample PDF files to form a sample set Pdf = {pdf1, pdf2,..., pdfn};

[0011] Step 103: Extract paragraphs and tables:

[0012] Extract the paragraphs and tables in each sample PDF file to obtain a paragraph set Paragraph = {paragraph1, paragraph2,..., paragraphn} and a table set Table = {table1, table2,..., table n};

[0013] Step 104: Annotation and corpus construction:

[0014] Manually process each sample PDF file, mark the specific location and value of the business field f in the sample PDF file to form a processed corpus entity Corpus. After the annotation is completed, store these processed corpus entities Corpus in the corpus;

[0015] Step 105: Vector calculation and model library construction:

[0016] Extract the corpus entity set from the corpus obtained in Step 104, perform vectorization processing on the paragraphs or tables in each processed corpus entity Corpus in the set, calculate its vector value VectorValue, and store the vector value VectorValue together with other information of the corresponding processed corpus entity Corpus in the positioning model library to form a positioning model entity;

[0017] Step 2: Extract the paragraphs of the target PDF file, which specifically includes the following steps:

[0018] Step 201: Obtain the target PDF file, and use a professional PDF parsing tool to perform document structure analysis to identify the paragraphs and tables in the document;

[0019] Step 202: For each identified paragraph and table, perform text extraction to obtain a paragraph set TargetParagraph = {targetParagraph1, targetParagraph2,..., targetParagraphn} and a table set TargetTable = {targetTable1, targetTable2,..., targetTable n};

[0020] Step 203: Merge the extracted paragraph and table texts to form a complete PDF text set TargetText = {targetText1, targetText2,..., targetTextn};

[0021] Step 3: Paragraph processing of the target PDF file, which specifically includes the following steps:

[0022] Step 301: For each text paragraph in the PDF text set TargetText, apply a pre-trained large language model for deep learning vectorization processing to generate a corresponding vector set TargetVectorValue = {targeVectorValue1, targeVectorValue2,..., targeVectorValuem};

[0023] Step 302: Initialize the traversal index i of the field set Field = {field1, field2,..., fieldn} to 1;

[0024] Step 303: Extract the currently processed field fieldi from the field set Field through the index i;

[0025] Step 304: Initialize the traversal index j of the PDF text set TargetText to 1;

[0026] Step 305: Extract the currently processed paragraph TargetTextj from the PDF text set TargetText through the index j;

[0027] Step 306: For the current field fieldi, perform a vector query operation and set the maximum number of results returned by the vector library to x for subsequent similarity evaluation.

[0028] Step 307: After querying the vector library, accumulate the similarity scores of the obtained x results, calculate the average similarity score count(score) / x, associate this score with the corresponding paragraph TargetTextj to form an entity TargetTextAndScore, and store it in the set TargetTextAndScoreArrayim;

[0029] Step 308: Based on the index j, determine whether the TargetText set has been traversed. If not, increment j and continue traversing. If completed, proceed to the next step;

[0030] Step 309: Sort the set TargetTextAndScoreArrayim to ensure that paragraphs with higher similarity scores are ranked first;

[0031] Step 3010: Based on the index i, determine whether the field set Field has been traversed completely. If not, increment i and continue traversing. If completed, proceed to the next step;

[0032] Step 3011: Construct a two-dimensional set TargetTextAndScoreArraynm[][], which contains the similarity score information of all fields and corresponding paragraphs;

[0033] Step 3012: Traverse the two-dimensional set TargetTextAndScoreArraynm[][]. For each field Fieldx, extract the corresponding subset TargetTextAndScoreArrayxm[] and filter out paragraphs with high similarity according to the set threshold. Among them, for paragraphs that do not reach the threshold, decide whether to retain them according to the sorting result of top y;

[0034] Step 3013: The subset TargetTextAndScoreArrayxm[] of each field Fieldx contains all located paragraphs that meet the conditions.

[0035] Preferably, in step 101, after defining the business fields, obtain the business field set Field = {f1, f2,..., fn}.

[0036] Preferably, in step 102, all sample PDF files cover different business scenarios and file formats.

[0037] Preferably, in step 102, when collecting sample PDF files, pay attention to the copyright and privacy issues of the files to ensure the legality and compliance of all sample files.

[0038] Preferably, in step 103, use natural language processing techniques and document parsing tools to extract paragraphs and tables.

[0039] Preferably, in step 104, each processed corpus entity Corpus contains: file ID, paragraph paragraph or table table, business ID, business field f, and the value annotated by the annotator from the paragraph.

[0040] Preferably, in step 105, use a pre-trained language model to convert the text into points in a high-dimensional vector space.

[0041] Preferably, in step 105, the positioning model entity includes a file ID, a paragraph or a table, a service ID, a service field f, the value annotated by the annotator from the paragraph, and the vector value VectorValue of the paragraph or the table.

[0042] Preferably, after step 3, the following steps are further included:

[0043] Step 4, closed-loop training:

[0044] Starting from the preliminary data preparation stage of step 1, re-examine and optimize the dataset construction process to improve the quality and effect of model training.

[0045] The present invention can improve the retrieval efficiency and accuracy of specific content in PDF documents, greatly liberate human resources, improve work efficiency, reduce error rates, and provide strong technical support for the digital transformation and intelligent upgrading of enterprises. In addition, the introduction of the closed-loop training step enables the system to continuously self-optimize and adapt to changing business requirements and data characteristics. Description of the Drawings

[0046] Figure 1 Schematically shows the process of the pdf positioning content solution based on a large model. Detailed Embodiments

[0047] The following further elaborates the present invention in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by the appended claims of this application.

[0048] A method for accurately positioning PDF disclosed in an embodiment of the present invention includes the following steps:

[0049] The first step, preliminary data preparation, includes the following steps:

[0050] Step 1: Define service B and its service fields

[0051] In this step, first, it is necessary to clarify the specific service scope and requirements of service B, and then determine the set of service fields Field = {f1, f2,..., fn} that need to be processed in service B. These fields represent the key information that service B needs to pay special attention to and extract when processing PDF files. For example, if service B is financial auditing, then the service fields may include "company name", "financial statements", "date", etc.

[0052] Step 2: Collection and preparation of sample PDF files

[0053] For each business field f in business B, collect a set of representative sample PDF files to form a sample set Pdf = {pdf1, pdf2,..., pdfn}. These sample files should cover different business scenarios and file formats to ensure that the trained model has good generalization ability. When collecting sample files, it is also necessary to pay attention to copyright and privacy issues of the files to ensure the legality and compliance of all sample files.

[0054] Step 3: Paragraph and Table Extraction

[0055] Conduct in-depth text analysis on each sample PDF file, and use natural language processing (NLP) techniques and document parsing tools to extract paragraphs Paragraph = {paragraph1, paragraph2,..., paragraphn} and tables Table = {table1, table2,..., table n} in the document. The key to this step lies in accurately identifying the boundaries of paragraphs and tables, as well as correctly handling complex situations such as multi-page and nested tables. The extracted paragraphs and tables will serve as the basis for subsequent annotation and model training.

[0056] Step 4: Annotation and Corpus Construction

[0057] Have professional annotators conduct meticulous manual processing on each sample PDF file, mark the specific positions and values of the business field f in the document, and form processed corpus entities Corpus. Each Corpus entity contains the following information: file ID, paragraph (or table), business ID, business field f, and the value annotated by the annotator from the paragraph. After annotation, store these Corpus entities in the corpus to provide high-quality training data for subsequent model training.

[0058] Step 5: Vector Calculation and Model Library Construction

[0059] Extract the Corpus set from the corpus, vectorize the paragraphs (or tables) in each Corpus entity, and calculate their vector values VectorValue. This step usually requires using pre-trained language models (such as BERT, GPT, etc.) to convert the text into points in a high-dimensional vector space. The calculated vector values VectorValue will be stored in the positioning model library together with other information of the Corpus entity to form a positioning model entity, including file ID, paragraph (table), business ID, business field f, the values annotated by the annotator from the paragraph, the vector value VectorValue of the paragraph (table), etc. The positioning model library will serve as the core resource for subsequent PDF content positioning and processing.

[0060] Through the above detailed data preparation steps, we can build a high-quality corpus and positioning model library, providing a solid data foundation for subsequent PDF content positioning and processing. This can not only improve the accuracy of positioning and processing, but also shorten the time for model training and optimization, ultimately achieving efficient and accurate PDF content processing.

[0061] Step 2: Extract paragraphs from the target PDF file, including the following steps:

[0062] Step 1: Obtain the target PDF file and use a professional PDF parsing tool to analyze the document structure to identify paragraphs and tables in the document. Through this analysis, we can distinguish different parts of the text content, including titles, main texts, footnotes, etc., and store them classified.

[0063] Step 2: For each identified paragraph and table, perform text extraction to obtain the paragraph set TargetParagraph = {targetParagraph1, targetParagraph2,..., targetParagraphn} and the table set TargetTable = {targetTable1, targetTable2,..., targetTable n}. This step requires accurately capturing the boundaries of paragraphs and tables and maintaining the integrity and accuracy of the content.

[0064] Step 3: Merge the extracted paragraph and table texts to form a complete PDF text set TargetText = {targetText1, targetText2,..., targetTextn}. This merging process needs to ensure that the order and structure of the text are consistent with the original PDF file so that subsequent vector calculations and similarity evaluations can accurately reflect the content of the original document.

[0065] Step 3: Paragraph Processing of the Target PDF File

[0066] Step 1: For each text paragraph in the set TargetText, apply a pre-trained large language model for deep learning vectorization processing to generate a corresponding vector set TargetVectorValue = {targeVectorValue1, targeVectorValue2,..., targeVectorValuem}. This step ensures that the complex semantic information of the text is effectively represented in the vector space.

[0067] Step 2: Initialize the traversal index i of the field set Field = {field1, field2,..., fieldn} to 1 to ensure that each business field can be processed in sequence in the subsequent steps.

[0068] Step 3: Extract the currently processed field fieldi from the Field set through the index i.

[0069] Step 4: Initialize the traversal index j of the TargetText set to 1 to prepare for processing each paragraph in the set.

[0070] Step 5: Extract the currently processed paragraph TargetTextj from the TargetText set through the index j.

[0071] Step 6: For the current field fieldi, perform a vector query operation and set the maximum number of results returned by the vector library to x for subsequent similarity evaluation.

[0072] Step 7: After querying the vector library, accumulate the similarity scores of the obtained x results, calculate the average similarity score count(score) / x, associate this score with the corresponding paragraph TargetTextj to form an entity TargetTextAndScore, and store it in the set TargetTextAndScoreArrayim.

[0073] Step 8: Determine whether the index j has traversed the entire TargetText set. If not, increment j and continue traversing; if so, proceed to the next step.

[0074] Step 9: Sort the set TargetTextAndScoreArrayim to ensure that paragraphs with higher similarity scores are ranked first.

[0075] Step 10: Determine whether the index i has traversed the entire Field set. If not, increment i and continue traversing; if so, proceed to the next step.

[0076] Step 11: After the above loop, finally construct a two-dimensional set TargetTextAndScoreArraynm[][], which contains the similarity score information for all fields and corresponding paragraphs.

[0077] Step 12: Traverse the two-dimensional set TargetTextAndScoreArraynm[][]. For each field Fieldx, extract the corresponding subset TargetTextAndScoreArrayxm[], and filter out paragraphs with high similarity according to a set threshold (e.g., 90%). For paragraphs that do not reach the threshold, decide whether to retain them according to the sorting result of the top y.

[0078] Step 13: Finally, the subset TargetTextAndScoreArrayxm[] for each field Fieldx will contain all the located paragraphs that meet the conditions, providing accurate data support for subsequent business processing.

[0079] Fourth step, closed-loop training, including the following steps:

[0080] Starting from the preliminary data preparation stage of the first step, re-examine and optimize the dataset construction process to improve the quality and effect of model training.

[0081] The above technical solution can not only improve the accuracy of PDF content location, but also greatly enhance the processing efficiency, reduce manual intervention, and ultimately achieve automated PDF content processing.

[0082] For example, it is necessary to quickly locate specific contract terms, such as "confidentiality agreement" or "liability for breach of contract" clauses, from a large number of legal documents for review, comparison, or extraction. These contract terms are usually contained in PDF format files, and the number of files is huge. Manual search is time-consuming and error-prone. Then, the above technical solution is used for processing, which specifically includes the following steps:

[0083] Step 1: Preprocessing and feature extraction: Extract the text content from the PDF file and convert it into a processable text format. Preprocess the extracted text, including word segmentation, removal of stop words, punctuation marks, etc. Extract features of each paragraph, such as word frequency, location information, semantic features, etc.

[0084] Step 2, Deep learning model construction: Select a pre-trained model suitable for text classification and entity recognition, such as RoBERTa. Fine-tune the model so that it can recognize specific terms in the contract, such as "confidentiality agreement", "liability for breach of contract", etc.

[0085] Step 3, Paragraph vector calculation: Use the fine-tuned deep learning model to calculate the vector representation of each contract paragraph.

[0086] Step 4, Build a vector database: Store the vector representation of each paragraph in the vector database and associate it with information such as contract ID, clause type, clause content, paragraph location, etc.

[0087] Step 5, New file content location: When a new contract file needs to be processed, extract all paragraphs in the file and calculate the vector representation of each paragraph. Query the top 10 similar paragraphs that match a specific clause type through the vector database.

[0088] Step 6, Similarity calculation: For each paragraph, calculate the similarity score between it and the top 10 similar paragraphs and calculate the average value. If the average similarity exceeds a preset threshold (e.g., 90%), then the paragraph is considered a target clause; otherwise, the top 10 similar paragraphs are used as candidate paragraphs.

[0089] Step 7, Result output: Output the identified target clause paragraphs and provide them to legal professionals for further review or automatically extract them into the contract management system.

[0090] Step 8, Model iteration and optimization: According to the feedback and review results of legal professionals, continuously adjust and optimize the model parameters and similarity calculation methods to improve the accuracy and efficiency of clause location.

[0091] Through the above implementation methods, legal professionals can quickly locate specific clauses in the contract, greatly improving work efficiency, reducing repetitive labor, and reducing legal risks caused by manual search errors.

Claims

1. A PDF content positioning method based on a large model, characterized in that: The following steps are involved: Step 1: Preliminary data preparation, including: Step 101: define a target business and a business field of the target business, wherein the business field represents key information that the target business needs to pay special attention to and extract when processing a PDF file; Step 102: Prepare and collect sample PDF files: For each business field f in the target business, collect a set of representative sample PDF files to form a sample set Pdf = {pdf1, pdf2, ..., pdfn}; Step 103: Extract paragraphs and tables: Extract the paragraphs and tables in each sample PDF file to obtain a paragraph set Paragraph = {paragraph1, paragraph2, ..., paragraphn} and a table set Table = {table1, table2, ..., table n}; Step 104: Annotation and corpus construction: Manually process each sample PDF file, mark the specific position and value of the business field f in the sample PDF file, and form the processed corpus entity Corpus. After the marking is completed, these processed corpus entities Corpus are stored in the corpus; Step 105: Vector calculation and model library construction: Extract a corpus entity set from the corpus obtained in step 104, perform vectorization processing on the paragraphs or tables in each processed corpus entity Corpus in the set, calculate its vector value VectorValue, and store the vector value VectorValue together with other information of the corresponding processed corpus entity Corpus into the positioning model library to form a positioning model entity; Step 2: Extract the target PDF file paragraphs, including the following steps: Step 201: Obtain the target PDF file and use a professional PDF parsing tool to perform document structure analysis to identify paragraphs and tables in the document; Step 202: For each identified paragraph and table, perform text extraction to obtain a paragraph set TargetParagraph = {targetParagraph1, targetParagraph2, ..., targetParagraphn} and a table set TargetTable = {targetTable1, targetTable2, ..., targetTable n}; Step 203: Merge the extracted paragraph and table texts to form a complete PDF text set TargetText={targetText1, targetText2, ..., targetTextn}; Step 3: Process the target PDF file paragraphs, including the following steps: Step 301: for each text paragraph in the PDF text set TargetText, a pre-trained large language model is applied to perform deep learning vectorization processing to generate a corresponding vector set TargetVectorValue={targetVectorValue1,targetVectorValue2,...,targetVectorValuem}; Step 302, initialize the traversal index i of the field set Field={field1, field2, ..., fieldn} to 1; Step 303: extract the current field fieldi to be processed from the field set Field by using the index i; Step 304, initialize the traversal index j of the PDF text set TargetText to 1; Step 305, extracting the current paragraph TargetTextj to be processed from the PDF text set TargetText by using the index j; Step 306: Perform a vector query operation on the current field fieldi, and set the maximum number of results returned by the vector library to x, so as to perform subsequent similarity evaluation. Step 307: After querying the vector library, the similarity scores of the obtained x results are accumulated, and the average similarity score count(score) / x is calculated. The score is associated with the corresponding paragraph TargetTextj to form an entity TargetTextAndScore, and is stored in the collection TargetTextAndScoreArrayim. Step 308: Based on the index j, determine whether the TargetText collection has been traversed. If not, j is incremented and the traversal continues. If it is completed, proceed to the next step. Step 309, sorting the set TargetTextAndScoreArrayim to ensure that the paragraphs with higher similarity scores are arranged at the front; Step 3010: Based on the index i, determine whether the field set Field has been traversed. If not, i is incremented and the traversal continues. If it is completed, proceed to the next step. Step 3011, construct a two-dimensional set TargetTextAndScoreArraynm[][], which contains similarity score information of all fields and corresponding paragraphs; Step 3012: traverse the two-dimensional set TargetTextAndScoreArraynm[][], extract the corresponding sub-set TargetTextAndScoreArrayxm[] for each field Fieldx, and filter out paragraphs with high similarity according to the set threshold. For paragraphs that do not reach the threshold, decide whether to retain them according to the sorting result of top y; Step 3013: The sub-collection TargetTextAndScoreArrayxm[] of each field Fieldx contains all located paragraphs that meet the conditions.

2. A PDF content positioning method based on a large model as claimed in claim 1, characterized in that: In step 101, after defining the business field, a business field set Field = {f1, f2, ..., fn} is obtained.

3. A PDF content positioning method based on a large model as claimed in claim 1, characterized in that: In step 102, all sample PDF files cover different business scenarios and file formats.

4. A PDF content positioning method based on a large model as claimed in claim 1, characterized in that: In step 102, when collecting sample PDF files, pay attention to the copyright and privacy issues of the files to ensure the legality and compliance of all sample files.

5. A PDF content positioning method based on a large model as claimed in claim 1, characterized in that: In step 103, paragraphs and tables are extracted using natural language processing technology and document parsing tools.

6. A PDF content positioning method based on a large model as claimed in claim 1, characterized in that: In step 104, each processed corpus entity Corpus includes: a document ID, a paragraph or a table, a business ID, a business field f, and a value annotated by the annotator from the paragraph.

7. A PDF content positioning method based on a large model as claimed in claim 1, characterized in that: In step 105, the text is converted into points in a high-dimensional vector space using a pre-trained language model.

8. A PDF content positioning method based on a large model as claimed in claim 1, characterized in that: In step 105, the positioning model entity includes a document ID, a paragraph or a table, a business ID, a business field f, a value annotated by annotators from a paragraph, and a vector value VectorValue of the paragraph or the table.

9. A PDF content positioning method based on a large model as claimed in claim 1, characterized in that: After step 3, also include: Step 4: Closed-loop training: Starting from the preliminary data preparation stage in step 1, review and optimize the dataset construction process to improve the quality and effect of model training.