Document intelligent classification information extraction method and device

By constructing semantic vector search services and using multiple pre-trained models, the efficiency and accuracy of diversified electronic document information extraction is solved, and intelligent document classification and information extraction is realized without preset templates, which improves the efficiency and accuracy of petroleum document processing.

CN120372004APending Publication Date: 2025-07-25DAQING OILFIELD CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410082773.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-19
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

When facing diversified electronic documents, the prior art requires preset fixed templates for information extraction, resulting in large workload and low extraction accuracy, which cannot meet the needs of rapid extraction of different document types.

Method used

The semantic vector search service is constructed using UIE, RocetQA, Chinese_OCR and ChatGLM3-6B models. By identifying the document format and converting it into a predetermined format, a classification directory is constructed, and key information is extracted using Prompt prompt words, combining multiple models to realize intelligent classification and information extraction.

Benefits of technology

It realizes intelligent classification and key information extraction without preset fixed templates, improves work efficiency, reduces labor costs, and ensures the accuracy of information extraction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120372004A_ABST
    Figure CN120372004A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of petroleum document processing, in particular to a document intelligent classification information extraction method and device. The method comprises the following steps: converting a to-be-extracted document into a predetermined extraction format; constructing a classification directory of document categories and setting corresponding key information; the method comprises the following steps: constructing a semantic vector retrieval service on the basis of a UIE (Unified Identity Element) model, a RooctQA model, a ChineseOCR model and a ChatGLM3-6B model; and generating Prompt cue words according to the retrieved document paragraphs, and inputting the Prompt cue words into the ChatGLM3-6B model to extract corresponding key information. The problems that due to the fact that document formats are diversified, the workload of information extraction through a preset fixed template is large, the purpose of improving quality and efficiency cannot be achieved, the same method cannot meet the requirement for rapid extraction of information of all electronic document types, and the extraction accuracy is not high are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of oil document processing, and particularly to a method and device for intelligent classification information extraction of documents. Background Art

[0002] Existing electronic document structured information extraction solutions only extract key information for a certain type of document, such as word documents or pdf format documents. There are mainly two methods: one is to preset a fixed template, classify according to the template style attributes, and then use different key information extraction methods for different types of paragraphs. The other is to identify paragraphs through neural networks and then extract them through preset attribute format rules.

[0003] Existing methods rely on preset attribute format rules, require manual rule updates, limit the flexibility and adaptability of the methods, and at the same time require a large amount of manually labeled data, with high costs. In addition, neural network recognition has limited understanding of the context of documents, thus affecting the accuracy of information extraction. Summary of the Invention

[0004] The present invention provides a method and device for intelligent classification information extraction of documents to solve the problems that due to the diverse document formats, the workload of information extraction through preset fixed templates is extremely large, the purpose of improving quality and efficiency cannot be achieved, the same method cannot meet the rapid extraction of information for all types of electronic documents, and the extraction accuracy is not high.

[0005] According to one aspect of the present invention, a method for intelligent classification information extraction of documents is provided, including: Obtain the document to be extracted, determine whether the format of the document is a predetermined extraction format. If not, convert the document into a predetermined extraction format; Construct a classification directory for the document categories and set the key information corresponding to each document category in the directory; Based on the document and the classification directory, construct a semantic vector retrieval service based on the UIE, RocetQA, Chinese_OCR, and ChatGLM3-6B models, and use the semantic vector retrieval service to retrieve the document paragraphs and / or table contents corresponding to the input extraction entity retrieval statement; Generate a Prompt prompt word according to the document paragraph and input it into the ChatGLM3-6B model. The ChatGLM3-6B model extracts the corresponding key information from the document paragraph according to the Prompt prompt word, and summarizes the key information and / or table contents and outputs them in a predetermined output format.

[0006] Preferably, the predetermined extraction format is: pdf file format.

[0007] Preferably, the method for constructing the classification directory of the document category includes: Input all the documents into the UIE model, locate the cover page of each document in the predetermined extraction format, and extract the keywords of the cover page; Construct a document classification directory using Python office automation based on the keywords of the cover page.

[0008] Preferably, the method for constructing a semantic vector retrieval service based on the UIE, RocetQA, Chinese_OCR, and ChatGLM3-6B models according to the document and the classification directory includes: Extract the text in the document through the Chinese_OCR model; Convert the extracted text and the classification directory into corresponding text vectors through the RocetQA model and store them in Elasticsearch; Input the pictures of all the pages where the tables in the document are located into the UIE model; Establish the connection between the ChatGLM3-6B model and the RocetQA and UIE models to complete the construction of the semantic vector retrieval service.

[0009] Preferably, the method for inputting the pictures of all the pages where the tables in the document are located into the UIE model includes: Identify the pages where the tables are located in the document through a layout analysis algorithm; Input the screenshots of the pages where the tables are located into the UIE model.

[0010] Preferably, the method for retrieving the document paragraph corresponding to the input extraction entity retrieval statement using the semantic vector retrieval service includes: Input the extraction entity retrieval statement into the RocetQA model, and the RocetQA model converts the extraction entity retrieval statement into a corresponding text vector; Retrieve the corresponding text vector paragraph in the Elasticsearch according to the text vector corresponding to the extraction entity retrieval statement; Extract the text content corresponding to the text vector paragraph, and this text content is the document paragraph corresponding to the extraction entity retrieval statement.

[0011] Preferably, the method for retrieving the table content corresponding to the input extraction entity retrieval statement using the semantic vector retrieval service includes: The table content includes the first extracted table content and the second extracted table content; Among them, the method for retrieving the first extracted table content corresponding to the input extraction entity retrieval statement is: Input the extracted entity retrieval statement into the RocetQA model, and the RocetQA model converts the extracted entity retrieval statement into a corresponding text vector; According to the text vector corresponding to the extracted entity retrieval statement, retrieve the corresponding table text vector in the Elasticsearch; Extract the table text content corresponding to the table text vector, and the table text content is the first extracted table content corresponding to the extracted entity retrieval statement; Among them, the method for retrieving the second extracted table content corresponding to the input extracted entity retrieval statement is: Input the extracted entity retrieval statement into the UIE model, and the UIE model extracts the text content in the table picture corresponding to the extracted entity retrieval statement inside it, and the text content in the table picture is the second extracted table content corresponding to the extracted entity retrieval statement.

[0012] Preferably, the method for generating the Prompt prompt word according to the document paragraph includes: Combine the document paragraph with the key information corresponding to the document category to which the document paragraph belongs and a preset prompt word template to form the Prompt prompt word.

[0013] According to one aspect of the present invention, there is provided a document intelligent classification information extraction device, including: A format conversion unit, configured to obtain a document to be extracted, determine whether the format of the document is a predetermined extraction format, and if not, convert the document into the predetermined extraction format; A directory construction unit, configured to construct a classification directory of the document category and set the key information corresponding to each document category in the directory; A retrieval unit, configured to construct a semantic vector retrieval service based on the UIE, RocetQA, Chinese_OCR, and ChatGLM3-6B models according to the document and the classification directory, and use the semantic vector retrieval service to retrieve the document paragraph and / or table content corresponding to the input extracted entity retrieval statement; An information extraction unit, configured to generate a Prompt prompt word according to the document paragraph and input it into the ChatGLM3-6B model. The ChatGLM3-6B model extracts the corresponding key information in the document paragraph according to the Prompt prompt word, and summarizes the key information and / or table content and outputs it in a predetermined output format.

[0014] The present invention has at least the following beneficial effects: The present invention provides a method and apparatus for intelligent classification and information extraction of documents. Based on the characteristics of multiple open-source pre-trained language models, a semantic vector retrieval service is constructed to achieve intelligent classification of a large number of petroleum electronic documents and automatic extraction of key information, replacing manual work. Without the need for a preset fixed template, the work efficiency is greatly improved, the labor cost is reduced, and the accuracy of the extracted information is ensured at the same time. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] The accompanying drawings are incorporated herein and form a part of this specification. These drawings illustrate embodiments consistent with the present invention and, together with the specification, are used to explain the technical solutions of the present invention.

[0016] Figure 1 FIG. shows a flowchart of a method for intelligent classification and information extraction of documents according to an embodiment of the present invention; Figure 2 FIG. shows a flowchart of document text information extraction according to an embodiment of the present invention; Figure 3 FIG. shows a flowchart of document table information extraction according to an embodiment of the present invention. DETAILED DESCRIPTION

[0017] Various exemplary embodiments, features, and aspects of the present invention will be described in detail below with reference to the accompanying drawings. The same reference numerals in the drawings denote elements having the same or similar functions. Although various aspects of the embodiments are shown in the drawings, the drawings are not necessarily drawn to scale unless otherwise specified.

[0018] The term "exemplary" used herein means "serving as an example, embodiment, or illustration". Any embodiment described as "exemplary" herein is not necessarily to be construed as superior to or better than other embodiments.

[0019] As used herein, the term "and / or" merely describes an association relationship between associated objects and indicates that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the term "at least one" as used herein means any one of a plurality or any combination of at least two of a plurality. For example, including at least one of A, B, and C may represent including any one or more elements selected from the set consisting of A, B, and C.

[0020] In addition, for a better description of the present invention, numerous specific details are given in the following detailed embodiments. Those skilled in the art should understand that the present invention can be implemented without some specific details. In some instances, methods, means, elements, and circuits well-known to those skilled in the art are not described in detail so as to highlight the gist of the present invention.

[0021] Figure 1A flowchart showing a method for extracting document intelligent classification information according to an embodiment of the present invention; Figure 2 A flowchart showing the extraction of document text information according to an embodiment of the present invention; Figure 3 A flowchart showing the extraction of document table information according to an embodiment of the present invention. As Figures 1-3 As shown, a method for extracting document intelligent classification information includes: Step S01: Obtain the document to be extracted, and determine whether the format of the document is a predetermined extraction format. If not, convert the document into a predetermined extraction format; Step S02: Construct a classification directory for the document categories, and set the key information corresponding to each document category in the directory; Step S03: Based on the document and the classification directory, construct a semantic vector retrieval service using the UIE, RocetQA, Chinese_OCR, and ChatGLM3-6B models, as well as the Elasticsearch vector database. Use the semantic vector retrieval service to retrieve the document paragraphs and / or table content corresponding to the input extraction entity retrieval statement; Step S04: Generate a Prompt prompt word according to the document paragraph, and input it into the ChatGLM3-6B model. The ChatGLM3-6B model extracts the corresponding key information from the document paragraph according to the Prompt prompt word, and summarizes the key information and / or table content and outputs it in a predetermined output format.

[0022] A method for extracting document intelligent classification information provided by an embodiment of the present invention specifically includes the following steps: Step S01: Obtain the document to be extracted, and determine whether the format of the document is a predetermined extraction format. If not, convert the document into a predetermined extraction format.

[0023] In the present invention, the predetermined extraction format is: pdf file format.

[0024] In an embodiment of the present invention, the document with information to be extracted obtained may contain various file formats, such as word, etc. For the convenience of subsequent parsing, all non-pdf format documents need to be uniformly converted into pdf documents to facilitate subsequent unified processing. Document format conversion can be achieved by using existing file format converters.

[0025] In the target oilfield cloud computing system environment, multiple pre-trained large language models such as the Baidu UIE model, Baidu RocetQA (Baidu's open-source "Rocket Q&A" information retrieval model), Tsinghua ChatGLM3-6B (Tsinghua Zhipu Qingyan's third-generation open-source model), and Baidu Chinese_OCR (Baidu's Chinese character recognition model) are deployed offline for subsequent intelligent classification and key information extraction of oil documents. Among them, the Baidu UIE model is built using the Python language and the Baidu PaddlePaddle deep learning framework.

[0026] Step S02: Construct a classification directory for the document categories and set the key information corresponding to each document category in the directory.

[0027] In the present invention, the method for constructing the classification directory of the documents includes: inputting the documents into the UIE model, locating the cover page of each document in the predetermined extraction format, and extracting the keywords of the cover page; and constructing the document classification directory by using Python office automation according to the keywords of the cover page.

[0028] In an embodiment of the present invention, a table of contents is created by identifying the keywords on the first page of the document converted to pdf, including key information such as report category (document category), report date, etc. Report categories include logging report, completion report, drilling well history, drilling core description, drilling geological design, comprehensive logging record, drilling geological design, and so on.

[0029] Take the first page or the first two pages of the document as a picture and input it into the Baidu UIE model. The Baidu UIE model extracts the text information on the input picture and automatically identifies the document category. Specifically, the identified document category can be obtained through the intelligent question and answer of the Baidu UIE model. For example, when asking the model: "What is the document category?", the Baidu UIE model will automatically give the identified document category, and the table of contents can be created according to the document category.

[0030] According to the document category identified by the Baidu UIE model, use Python office automation technology to automatically construct a document classification directory, and let the program automatically loop through all the pdf electronic documents to be extracted, and classify them according to the report category, such as logging report, geological summary report, etc., and finally achieve intelligent classification.

[0031] Step S03: Based on the UIE, RocetQA, Chinese_OCR, and ChatGLM3-6B models, and the Elasticsearch vector database, construct a semantic vector retrieval service according to the document and the classification directory, and use the semantic vector retrieval service to retrieve the document paragraphs and / or table contents corresponding to the input extraction entity retrieval statement.

[0032] In the present invention, the method for constructing a semantic vector retrieval service based on the document and the classification directory, the UIE, RocetQA, Chinese_OCR, and ChatGLM3-6B models, and the Elasticsearch vector database includes: extracting the text in the document through the Chinese_OCR model; converting the extracted text and the classification directory into corresponding text vectors through the RocetQA model and storing them in the Elasticsearch vector database; inputting the pictures of all pages where the tables in the document are located into the UIE model; establishing a connection between the ChatGLM3-6B model and the RocetQA and UIE models to complete the construction of the semantic vector retrieval service.

[0033] In the present invention, the method of inputting the pictures of all pages where the tables in the document are located into the UIE model includes: identifying the pages where the tables are located in the document through a layout analysis algorithm; inputting the screenshots of the pages where the tables are located into the UIE model.

[0034] In an embodiment of the present invention, constructing a semantic vector retrieval service includes constructing text information and constructing table information.

[0035] Constructing text information requires constructing a semantic vector retrieval service based on the Baidu RocetQA model and the Baidu Chinese_OCR model in combination with Elasticsearch (vector database).

[0036] Specifically, it includes: first, extracting all text information in the pdf format document by using the open-source Baidu Chinese_OCR model. Then, converting the extracted text and the corresponding classification directory information into their corresponding text vectors through the Baidu RocetQA model and storing them in Elasticsearch (vector database).

[0037] Among them, for the table information content in the document, two processing methods are adopted at the same time. One is to extract the table text through the Baidu Chinese_OCR model and store it in Elasticsearch, and the other is to input the picture of the table into Baidu UIE.

[0038] Setting the key information corresponding to each document category in the Tsinghua ChatGLM3-6B model, and the Tsinghua ChatGLM3-6B model is used to extract the corresponding key information from the extraction results of the Baidu RocetQA model and the Baidu UIE model.

[0039] The Baidu Chinese_OCR model is an optical character recognition (OCR) model specifically for Chinese. It is designed to solve the problem of recognizing Chinese text in various backgrounds. Chinese text recognition presents specific challenges due to its unique character set and writing style. The Chinese_OCR model can accurately recognize Chinese characters from images. It is widely used in fields such as document digitization, information extraction, and automated form processing, especially in Chinese environments. The Baidu RocketQA model is a question-answering system dedicated to improving the performance of question-answering tasks. It uses advanced deep learning techniques to understand and answer users' questions. This model can quickly and accurately find the questions and answers most relevant to the user's query, thus providing high-quality question-answering services in various fields such as customer service, education, and search engines. The Baidu UIE model UIE (Universal Information Extraction) is a general information extraction model. It aims to automatically identify and extract key information from text, such as entities, relationships, and events. The UIE model can process various types of text data, such as news, social media posts, and scientific articles, and extract useful information from them. This is very valuable for applications such as data analysis, content management systems, and intelligent assistants. The ChatGLM3-6B model is a large language model specifically designed for dialogue systems. It is trained on a large corpus of text to understand and generate natural and fluent conversations. The model can generate high-quality conversations in various topics and scenarios. It plays a role in multiple fields such as chatbots, virtual assistants, and automated customer service systems, providing a smooth user experience.

[0040] In the present invention, the method for retrieving the document paragraph corresponding to the input extraction entity retrieval statement by using the semantic vector retrieval service includes: inputting the extraction entity retrieval statement into the RocetQA model, and the RocetQA model converts the extraction entity retrieval statement into a corresponding text vector; according to the text vector corresponding to the extraction entity retrieval statement, retrieving the corresponding text vector paragraph in the Elasticsearch; extracting the text content corresponding to the text vector paragraph, and this text content is the document paragraph corresponding to the extraction entity retrieval statement.

[0041] In an embodiment of the present invention, as Figure 2As shown in the figure, the specific retrieval process is as follows: Input the keywords during retrieval, that is, extract the entity retrieval statement; convert the input extracted entity retrieval statement into its corresponding text vector (retrieval sentence to vector) through the Baidu RocetQA model; perform a relevance match between the text vector corresponding to the extracted entity retrieval statement and the text vector paragraphs stored in Elasticsearch, and screen out several matching results (candidate results), and sort these several matching results in descending order of relevance degree (candidate sorting); select the result with the highest relevance degree or the first several candidate sorting results in the candidate sorting as the final text vector paragraph positioning result (result positioning), and the actual text corresponding to this positioning result is the document paragraph corresponding to the finally retrieved extracted entity retrieval statement. For example, if the extracted entity retrieval statement is "single well information", then perform a relevance match between the input text "single well information" and the corresponding text vectors in Elasticsearch through the Baidu RocetQA model, find several results with the highest relevance, and extract the text content of the paragraphs where each result is located, so as to obtain the document paragraph content corresponding to "single well information".

[0042] In the present invention, the method for retrieving the table content corresponding to the input extracted entity retrieval statement by using the semantic vector retrieval service includes: The table content includes the first extracted table content and the second extracted table content; Among them, the method for retrieving the first extracted table content corresponding to the input extracted entity retrieval statement is: Input the extracted entity retrieval statement into the RocetQA model, and the RocetQA model converts the extracted entity retrieval statement into the corresponding text vector; According to the text vector corresponding to the extracted entity retrieval statement, retrieve its corresponding table text vector in the Elasticsearch; Extract the table text content corresponding to the table text vector, and this table text content is the first extracted table content corresponding to the extracted entity retrieval statement; Among them, the method for retrieving the second extracted table content corresponding to the input extracted entity retrieval statement is: Input the extracted entity retrieval statement into the UIE model, and the UIE model extracts the text content in the table picture corresponding to the extracted entity retrieval statement inside it, and this text content in the table picture is the second extracted table content corresponding to the extracted entity retrieval statement.

[0043] In the embodiment of the present invention, as Figure 3 shown, the table data in the document will be extracted in two ways, one is through vector semantic retrieval, and the other is through the Baidu UIE model for extraction.

[0044] The process of extracting table content through vector semantic retrieval is the same as that of document paragraph extraction, that is, the input extraction entity retrieval statement is converted into its corresponding text vector through the Baidu RocetQA model; the text vector corresponding to the extraction entity retrieval statement is correlated with the table text vector paragraphs stored in Elasticsearch for matching, and several matching results are screened out. These several matching results are sorted in descending order according to the degree of correlation; the candidate with the highest degree of correlation or the first several candidate sorting results in the candidate sorting are selected as the final table text vector positioning result, and the actual text content corresponding to this positioning result is the first extracted table content corresponding to the finally retrieved extraction entity retrieval statement.

[0045] The process of extracting table content through the Baidu UIE model is as follows: The extraction entity retrieval statement is input into the Baidu UIE model, and through the intelligent question-answering function of the Baidu UIE model, the corresponding table text information is extracted from the table pictures stored in the UIE model.

[0046] When extracting table information using only the Baidu UIE or semantic vector retrieval alone, their accuracy may not meet the requirements. For example, the Baidu UIE model is relatively accurate in extracting parts with a relatively fixed format in the table, but it is difficult to extract complex tables. It is necessary to try table data annotation to improve the extraction accuracy. Since there may be blanks in the table, the information will be automatically filled forward during extraction, which may ultimately lead to incorrect extraction of data with blanks by the UIE model; semantic vector retrieval extracts after converting table data into text, and there may also be errors during the conversion, resulting in inaccurate final extraction results. By using both the Baidu UIE and semantic vector retrieval methods to jointly extract table data, they can complement and verify each other. If the table content extracted by one method is incorrect, the result extracted by the other method can be used as the standard.

[0047] Step S04: Generate a Prompt prompt word according to the document paragraph and input it into the ChatGLM3-6B model. The ChatGLM3-6B model extracts the corresponding key information from the document paragraph according to the Prompt prompt word, and summarizes the key information and / or table content and outputs it in a predetermined output format.

[0048] In the present invention, the method for generating the Prompt prompt word according to the document paragraph includes: combining the document paragraph with the key information corresponding to the document category to which the document paragraph belongs and a preset prompt word template to form the Prompt prompt word.

[0049] In an embodiment of the present invention, the preset prompt word template is: Please help me perform an information extraction task. I want to extract key information from the following text: "Target text fragment...". The key information to be extracted includes: ["Well name", "Well type",...].

[0050] Among them, the target text fragment is the document paragraph corresponding to the extraction entity retrieval statement retrieved through the semantic vector retrieval service, and / or the table content retrieved through the semantic vector retrieval service and Baidu UIE. The key information is the set key information corresponding to the document category where the retrieved document paragraph and / or table content is located.

[0051] For example, if the input extraction entity retrieval statement is: Single well information, after retrieving the document paragraph where the single well information is located through the semantic vector retrieval service, combine the document paragraph with the preset prompt word template and the key information corresponding to the document paragraph, such as well name, well type, etc., to form a prompt prompt word.

[0052] The prompt prompt word is dynamically formed. Input the dynamically formed prompt prompt word into the Tsinghua ChatGLM3-6B model. The Tsinghua ChatGLM3-6B model will extract the text related to the key information from the document paragraph and the table content, and summarize it into an Excel table and output it.

[0053] In an embodiment of the present invention, if an extraction entity retrieval statement is "What is the single well overview of Well No. 11?", input the extraction entity retrieval statement into the Baidu RocetQA model. After the Baidu RocetQA model converts the extraction entity retrieval statement into a corresponding text vector, perform a correlation match with the text vector paragraphs stored in Elasticsearch, screen out several matching results, and select the one with the highest degree of correlation as the document paragraph corresponding to the retrieved extraction entity retrieval statement; combine the retrieved document paragraph, the preset prompt word template, and the key information corresponding to the document paragraph to form a prompt prompt word, specifically: Please help me perform an information extraction task. I want to extract key information from the following text: "Well No. 11 is located 10.5 km northwest of Bayan Gun, Xin Barag Left Banner, Hulun Buir City, Inner Mongolia Autonomous Region. It is a wildcat well on Structure No. Dongba-20 in the Dongbayan Gun Structural Belt of the Huhehu Sag in the Huhehu Depression of the Hailar Basin. Its ordinate is 5329509.6 m and its abscissa is 20630241.9 m. The drilling purpose is...". The key information to be extracted includes: ["Well name", "Abscissa", "Ordinate"]; input the formed prompt prompt word into the Tsinghua ChatGLM3-6B model, and the model will output a table containing key information such as Well name: Well No. 11, Abscissa: 20630241.9 m, Ordinate: 5329509.6 m.

[0054] It can be understood that, without violating the principle logic, the above-mentioned method embodiments mentioned in the present invention can be combined with each other to form combined embodiments. Due to space limitations, the present invention will not elaborate further.

[0055] The execution subject of the document intelligent classification information extraction method can be a document intelligent classification information extraction device. For example, the document intelligent classification information extraction method can be executed by a terminal device, a server, or other processing devices. Among them, the terminal device can be a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc. In some possible implementation manners, the document intelligent classification information extraction method can be implemented by a processor invoking computer-readable instructions stored in a memory.

[0056] Those skilled in the art can understand that in the above method of the specific implementation manner, the writing order of each step does not mean a strict execution order and does not constitute any limitation to the implementation process. The specific execution order of each step should be determined according to its function and possible internal logic.

[0057] The present invention also provides a document intelligent classification information extraction device, including: a format conversion unit, configured to obtain a document to be extracted, determine whether the format of the document is a predetermined extraction format, and if not, convert the document into the predetermined extraction format; a directory construction unit, configured to construct a classification directory of the document categories and set the key information corresponding to each document category in the directory; a retrieval unit, configured to construct a semantic vector retrieval service based on the UIE, RocetQA, Chinese_OCR, and ChatGLM3-6B models according to the document and the classification directory, and use the semantic vector retrieval service to retrieve the document paragraphs and / or table content corresponding to the input extraction entity retrieval statement; an information extraction unit, configured to generate a Prompt prompt word according to the document paragraph and input it into the ChatGLM3-6B model. The ChatGLM3-6B model extracts the corresponding key information from the document paragraph according to the Prompt prompt word, and summarizes the key information and the table content and outputs them in a predetermined output format.

[0058] In some embodiments, the functions or modules and units included in the device provided by the embodiments of the present invention can be used to execute the methods described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be elaborated here.

[0059] The present invention applies offline deployment and uses a variety of open-source pre-trained large language models to achieve intelligent classification of petroleum electronic documents based on the Baidu UIE model, defines the entity names to be extracted, i.e., key information, according to the classification results. For each type of electronic report, key information extraction is performed on the text based on Baidu RocetQA and Tsinghua ChatGLM, and at the same time, key information is automatically extracted from the tables based on Baidu UIE. Finally, an Excel table is output according to the entity names to complete the key information extraction of all electronic documents. The present invention belongs to the forefront field of information technology artificial intelligence. Through the application and verification of actual data from thousands of documents, good results have been achieved.

[0060] The present invention fully applies artificial intelligence, cloud computing, and big data technologies, deploys pre-trained language models offline, and realizes intelligent classification of a large number of petroleum electronic documents to replace manual work and automatic extraction of key information. After actual extraction tests on the logging reports of an oilfield, the accuracy rate of key information extraction is over 80%, providing new methods and ideas for data quality control and governance, providing technical support for the digital and intelligent development of enterprises, and providing new methods and ideas for the enterprise data quality control and governance work.

[0061] The embodiments of the present invention have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art in the technical field without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, practical applications, or improvements to technologies in the market, or to enable other ordinary skill in the art in the technical field to understand the disclosed embodiments.

Claims

1. A method for intelligent classification information extraction of documents, characterized in that Including: Obtain the document to be extracted, and determine whether the format of the document is the predetermined extraction format. If not, convert the document into the predetermined extraction format; Construct a classification directory for the document categories, and set the key information corresponding to each document category in the directory; Based on the document and the classification directory, construct a semantic vector retrieval service based on the UIE, RocetQA, Chinese_OCR, and ChatGLM3-6B models. Use the semantic vector retrieval service to retrieve the document paragraphs and / or table content corresponding to the input extraction entity retrieval statement; Generate a Prompt prompt word according to the document paragraph, and input it into the ChatGLM3-6B model. The ChatGLM3-6B model extracts the corresponding key information in the document paragraph according to the Prompt prompt word, and summarizes the key information and / or table content and outputs it in the predetermined output format.

2. The method for intelligent classification information extraction of documents according to claim 1, characterized in that: The predetermined extraction format is: pdf file format.

3. The method for extracting document intelligent classification information according to claim 1, wherein The method for constructing the classification directory of the document categories includes: Input all the documents into the UIE model, locate the cover page of each document in the predetermined extraction format, and extract the keywords of the cover page; Construct a document classification directory by using Python office automation according to the keywords of the cover page.

4. The method for extracting document intelligent classification information according to any one of claims 1-3, characterized in that, The method for constructing a semantic vector retrieval service based on the UIE, RocetQA, Chinese_OCR, and ChatGLM3-6B models according to the document and the classification directory includes: Extract the text in the document through the Chinese_OCR model; Convert the extracted text and the classification directory into corresponding text vectors through the RocetQA model and store them in Elasticsearch; Input the pictures of all the pages where the tables in the document are located into the UIE model; Establish the connection between the ChatGLM3-6B model and the RocetQA and UIE models to complete the construction of the semantic vector retrieval service.

5. The method for extracting document intelligent classification information according to claim 4, characterized in that, The method for inputting the pictures of all the pages where the tables in the document are located into the UIE model includes: Identify the pages where the tables are located in the document through the layout analysis algorithm; Input the screenshots of the pages where the tables are located into the UIE model.

6. The method for extracting document intelligent classification information according to claim 1, characterized in that The method for using the semantic vector retrieval service to retrieve the document paragraphs corresponding to the input extraction entity retrieval statement includes: Input the extraction entity retrieval statement into the RocetQA model, and the RocetQA model converts the extraction entity retrieval statement into a corresponding text vector; Retrieve the corresponding text vector paragraph in the Elasticsearch according to the text vector corresponding to the extraction entity retrieval statement; Extract the text content corresponding to the text vector paragraph, and this text content is the document paragraph corresponding to the extraction entity retrieval statement.

7. The method for extracting document intelligent classification information according to claim 1, wherein The method for using the semantic vector retrieval service to retrieve the table content corresponding to the input extraction entity retrieval statement includes: The table content includes the first extracted table content and the second extracted table content; Among them, the method for retrieving the first extracted table content corresponding to the extraction entity retrieval statement input for retrieval is as follows: Input the extraction entity retrieval statement into the RocetQA model, and the RocetQA model converts the extraction entity retrieval statement into a corresponding text vector; According to the text vector corresponding to the extraction entity retrieval statement, retrieve the corresponding table text vector in the Elasticsearch; Extract the table text content corresponding to the table text vector, and this table text content is the first extracted table content corresponding to the extraction entity retrieval statement; Among them, the method for retrieving the second extracted table content corresponding to the extraction entity retrieval statement input for retrieval is as follows: Input the extraction entity retrieval statement into the UIE model, and the UIE model extracts the text content in the table picture corresponding to the extraction entity retrieval statement inside it, and this text content in the table picture is the second extracted table content corresponding to the extraction entity retrieval statement.

8. The method for extracting document intelligent classification information according to any one of claims 1-7, characterized in that, The method for generating a Prompt prompt word according to the document paragraph includes: Combine the document paragraph with the key information corresponding to the document category to which the document paragraph belongs and a preset prompt word template to form the Prompt prompt word.

9. An intelligent document classification information extraction device, characterized in that, It includes: A format conversion unit, which is used to obtain the document to be extracted, judge whether the format of the document is the predetermined extraction format, and if not, convert the document into the predetermined extraction format; A directory construction unit, which is used to construct a classification directory of the document category and set the key information corresponding to each document category in the directory; A retrieval unit, which is used to construct a semantic vector retrieval service based on the UIE, RocetQA, Chinese_OCR, and ChatGLM3-6B models according to the document and the classification directory, and use the semantic vector retrieval service to retrieve the document paragraph and / or table content corresponding to the input extraction entity retrieval statement; An information extraction unit, which is used to generate a Prompt prompt word according to the document paragraph and input it into the ChatGLM3-6B model. The ChatGLM3-6B model extracts the corresponding key information in the document paragraph according to the Prompt prompt word, and summarizes the key information and / or table content and outputs it in a predetermined output format.