Document Information Extraction Method, Device, Storage Medium and Electronic Device

Through the combination of the pre-trained factor extraction model and the large language model, the problems of low accuracy and lack of information in the prior art are solved, and efficient and accurate extraction of official document information is achieved, and the efficiency and accuracy of official document management are improved.

CN119180282BActive Publication Date: 2025-07-22NAVAL UNIV OF ENG PLA
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202411299782.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-18
Publication Date
2025-07-22
Estimated Expiration
2044-09-18

AI Technical Summary

Technical Problem

The official document information extraction method based on the BERT derivative model in the prior art has problems such as low accuracy and lack of information, especially when facing multiple extraction results and noisy data, it is difficult to ensure the accuracy and completeness of the results.

Method used

The pre-trained feature extraction model obtains official document elements and their corresponding official document paragraphs, and generates prompt words to indicate that the large language model optimizes official document paragraphs, and trains the model in combination with multi-task fine-tuning and weight adjustment to ensure the accuracy and completeness of official document information.

Benefits of technology

It realizes the accurate and complete extraction of official document information from pending official documents, improves the efficiency and accuracy of official document management, and reduces the need for manual intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119180282B_ABST
    Figure CN119180282B_ABST
Patent Text Reader

Abstract

The present application provides a method, apparatus, storage medium, and electronic device for extracting official document information. Among them, the electronic device obtains the official document to be processed; processes the official document to be processed through a pre-trained element extraction model to obtain the official document elements in the official document to be processed and the official document paragraphs corresponding to the official document elements; generates a prompt for instructing a preset large language model to optimize the official document paragraphs according to the official document elements and the official document paragraphs corresponding to the official document elements; and sends the prompt to the preset large language model for processing to obtain the optimized official document paragraphs. In this way, since the element extraction model has been trained specifically and is good at extracting relatively accurate official document elements, and the large language model is used to optimize the official document paragraphs corresponding to the official document elements, relatively accurate and complete official document information can be extracted from the official document to be processed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of natural language processing, and in particular, to a method, apparatus, storage medium, and electronic device for extracting official document information. Background Art

[0002] An official document, that is, an official business document, refers to a formal document formed by government agencies, enterprises, institutions, social organizations, etc. in administrative management activities or daily work for the purposes of conveying decisions, instructions, notices, reports, requests, approvals, etc. With the continuous expansion of business, the number of official documents accumulated within the unit has increased sharply, and the types and contents have also tended to be diversified and complicated. At present, the traditional electronic document management method has shown deficiencies. Its sorting and classification work completely rely on manual operations, and the label classification of documents is greatly affected by personal subjective judgments, which not only reduces the query efficiency but also easily causes errors. In this situation, it is particularly urgent and important to introduce artificial intelligence technology to improve the management level of unit official documents, realize the automatic extraction of key elements of official documents, and construct a systematic official document relationship and network circulation context. Through intelligent management, not only can the work efficiency be improved, but also the accuracy and standardization of official document management can be ensured, thereby promoting the effective utilization of internal information resources and knowledge inheritance within the unit.

[0003] It should be noted that official document element extraction includes two main links: document parsing and element extraction. In the document parsing stage, it is mainly to identify the full text content of official documents of different types and formats. In the element extraction stage, at present, the problem is mainly transformed into a sequence labeling problem. Through deep learning technology, semantic representation of the text can better obtain semantic information, and each word in the text is classified with multiple labels in combination with context features to predict the category and position of the element to which it belongs. However, the current extraction method based on the BERT-derived model has problems of low accuracy and information loss. Summary of the Invention

[0004] In order to overcome at least one deficiency in the prior art, the present application provides a method, apparatus, storage medium, and electronic device for extracting official document information, specifically including:

[0005] In a first aspect, the present application provides a method for extracting official document information, the method including:

[0006] Obtain an official document to be processed;

[0007] Process the official document to be processed through a pre-trained element extraction model to obtain the official document elements in the official document to be processed and the official document paragraphs corresponding to the official document elements;

[0008] Generate a prompt for instructing a pre-trained large language model to optimize the official document paragraph based on the official document elements and the official document paragraph corresponding to the official document elements;

[0009] Send the prompt to the pre-trained large language model for processing to obtain an optimized official document paragraph.

[0010] Combined with the optional implementation manners of the first aspect, the generating a prompt for instructing a pre-trained large language model to optimize the official document paragraph based on the official document elements and the official document paragraph corresponding to the official document elements includes:

[0011] Obtain a pre-written prompt template;

[0012] Fill the official document elements and the official document paragraph corresponding to the official document elements into the reserved positions in the prompt template to obtain the prompt.

[0013] Combined with the optional implementation manners of the first aspect, the method further includes a training method for the element extraction model, and the training method includes:

[0014] Obtain an official document sample and a training label of the official document sample, where the training label includes sample official document elements in the official document sample and a sample official document paragraph corresponding to the sample official document elements;

[0015] Input the official document sample into a model to be trained for processing to obtain a first prediction result of the sample official document elements and a second prediction result of the sample official document paragraph;

[0016] Obtain a first model loss of the model to be trained according to the first prediction result of the sample official document elements;

[0017] Obtain a second model loss of the model to be trained according to the second prediction result of the sample official document paragraph;

[0018] Obtain a comprehensive loss of the model to be trained according to the first model loss and the second model loss;

[0019] If the model to be trained does not meet the preset convergence condition, update the model to be trained according to the comprehensive loss, and return to the step of inputting the official document sample into the model to be trained for processing to obtain a first prediction result of the sample official document elements and a second prediction result of the sample official document paragraph, and execute until the preset convergence condition is met, and then obtain the element extraction model.

[0020] Combined with the optional implementation manners of the first aspect, the obtaining a comprehensive loss of the model to be trained according to the first model loss and the second model loss includes:

[0021] Weight the first model loss and the second model loss according to their respective weights to obtain the comprehensive loss of the model to be trained.

[0022] Combined with the optional implementation manner of the first aspect, the weight of the first model loss decreases as the number of iterations increases, and the weight of the second model loss increases as the number of iterations increases.

[0023] Combined with the optional implementation manner of the first aspect, the expressions of the weight α1 of the first model loss and the weight α2 of the second model loss are respectively:

[0024]

[0025] In the formula, step represents the current number of iterations, and total_step represents the total number of iterations.

[0026] Combined with the optional implementation manner of the first aspect, the sample official document paragraph includes at least one of text content, table content, and image content.

[0027] In a second aspect, the present application further provides an official document information extraction device, and the device includes:

[0028] An official document acquisition module, configured to acquire an official document to be processed;

[0029] An information extraction module, configured to process the official document to be processed through a pre-trained element extraction model to obtain the official document elements and the official document paragraphs corresponding to the official document elements in the official document to be processed;

[0030] An information optimization module, configured to generate a prompt word for instructing a preset large language model to optimize the official document paragraph according to the official document elements and the official document paragraphs corresponding to the official document elements; send the prompt word to the preset large language model for processing to obtain an optimized official document paragraph.

[0031] Combined with the optional implementation manner of the second aspect, the information optimization module is further specifically configured to:

[0032] Obtain a pre-written prompt word template;

[0033] Fill the official document elements and the official document paragraphs corresponding to the official document elements into the reserved positions in the prompt word template to obtain the prompt word.

[0034] Combined with the optional implementation manner of the second aspect, the device further includes a model training module, and the model training module is configured to:

[0035] Obtain a document sample and training labels for the document sample, where the training labels include sample document elements in the document sample and sample document paragraphs corresponding to the sample document elements;

[0036] Input the document sample into the model to be trained for processing, and obtain a first prediction result for the sample document elements and a second prediction result for the sample document paragraphs;

[0037] Based on the first prediction result of the sample document elements, obtain a first model loss for the model to be trained;

[0038] Based on the second prediction result of the sample document paragraphs, obtain a second model loss for the model to be trained;

[0039] Based on the first model loss and the second model loss, obtain a comprehensive loss for the model to be trained;

[0040] If the model to be trained does not meet the preset convergence condition, update the model to be trained according to the comprehensive loss, and return to execute the step of inputting the document sample into the model to be trained for processing to obtain a first prediction result for the sample document elements and a second prediction result for the sample document paragraphs, until the preset convergence condition is met, and then obtain the element extraction model.

[0041] Combined with the optional implementation manners of the second aspect, the model training module is further specifically configured to:

[0042] Perform weighting according to the respective weights of the first model loss and the second model loss to obtain a comprehensive loss for the model to be trained.

[0043] Combined with the optional implementation manners of the second aspect, the weight of the first model loss decreases as the number of iterations increases, and the weight of the second model loss increases as the number of iterations increases.

[0044] Combined with the optional implementation manners of the second aspect, the respective expressions of the weight α1 of the first model loss and the weight α2 of the second model loss are:

[0045]

[0046] In the formula, step represents the current number of iterations, and total_step represents the total number of iterations.

[0047] Combined with the optional implementation manners of the second aspect, the sample document paragraphs include at least one of text content, table content, and image content.

[0048] In a third aspect, the present application also provides a storage medium storing a computer program which, when executed by a processor, implements the official document information extraction method described above.

[0049] In a fourth aspect, the present application also provides an electronic device, which includes a processor and a memory. The memory stores a computer program which, when executed by the processor, implements the official document information extraction method.

[0050] Compared with the prior art, the present application has the following beneficial effects:

[0051] The present application provides an official document information extraction method, device, storage medium and electronic device. Among them, the electronic device obtains an official document to be processed; processes the official document to be processed through a pre-trained element extraction model to obtain the official document elements in the official document to be processed and the official document paragraphs corresponding to the official document elements; generates a prompt word for instructing a preset large language model to optimize the official document paragraphs according to the official document elements and the official document paragraphs corresponding to the official document elements; and sends the prompt word to the preset large language model for processing to obtain an optimized official document paragraph. In this way, since the element extraction model is trained in a targeted manner and is good at extracting relatively accurate official document elements, and the large language model is used to optimize the official document paragraphs corresponding to the official document elements, relatively accurate and complete official document information can be extracted from the official document to be processed. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required in the embodiments. It should be understood that the following drawings only show some embodiments of the present application, and thus should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0053] Figure 1 It is a schematic flow chart of the official document information extraction method provided by the embodiment of the present application;

[0054] Figure 2 It is a schematic overview diagram of the official document information extraction method provided by the embodiment of the present application;

[0055] Figure 3 It is a schematic structural diagram of the official document information extraction device provided by the embodiment of the present application;

[0056] Figure 4 It is a schematic structural diagram of the electronic device provided by the embodiment of the present application.

[0057] Icons: 11 - Document acquisition module; 12 - Information extraction module; 13 - Information optimization module; 21 - Memory; 22 - Processor; 23 - Communication unit; 24 - System bus. Detailed implementation manners

[0058] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some, but not all, of the embodiments of the present application. Components of the embodiments of the present application generally described and illustrated in the accompanying drawings herein can be arranged and designed in a variety of different configurations.

[0059] Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the claimed present application, but merely represents selected embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts fall within the scope of protection of the present application.

[0060] It should be noted that like reference numerals and letters denote like items in the following drawings. Therefore, once an item is defined in one drawing, it does not require further definition and explanation in subsequent drawings.

[0061] In the description of the present application, it should be noted that the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be construed as indicating or implying relative importance. In addition, the terms "include", "comprise", or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device including a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or device. Without further limitation, an element defined by the phrase "including one..." does not exclude the presence of additional identical elements in the process, method, article, or device including the element.

[0062] Based on the above statement, as introduced in the background art, the current extraction method based on the BERT-derived model has problems of low accuracy and information loss.

[0063] Exemplarily, in the related art, the full text content of a document is used as training data. After the full text elements are marked by experts, mainstream BERT-derived models, such as RoBERTa, ELECTRA, and MacBERT models, are used for target fine-tuning training. And the extracted model obtained by fine-tuning is deployed online for users to extract the element information in the full text. However, it is found in the practical process that there are two deficiencies in the current processing flow:

[0064] 1. If an element appears multiple times in the official document content, it will bring multiple extraction results, and post-processing customization development is required to ensure the accuracy of the results;

[0065] 2. The extraction information of elements is broken due to noise data and fuzzy boundaries, and post-processing positioning and completion of the extraction results are required to make up for the information loss.

[0066] Based on the discovery of the above technical problems, the inventor has proposed the following technical solutions through creative labor to solve or improve the above problems. It should be noted that the defects existing in the above solutions in the prior art are the results obtained by the inventor after practice and careful research. Therefore, the process of discovering the above problems and the solutions proposed in the embodiments of the present application below for the above problems should be the contributions made by the inventor to the present application during the invention creation process, and should not be understood as the technical content known to those skilled in the art.

[0067] In view of the above problems, this embodiment provides an official document information extraction method applied to an electronic device. In this method, the electronic device obtains the official document to be processed; processes the official document to be processed through a pre-trained element extraction model to obtain the official document elements and the official document paragraphs corresponding to the official document elements in the official document to be processed; generates a prompt for instructing a preset large language model to optimize the official document paragraphs according to the official document elements and the official document paragraphs corresponding to the official document elements; and sends the prompt to the preset large language model for processing to obtain the optimized official document paragraphs. In this way, since the element extraction model is trained specifically and is good at extracting relatively accurate official document elements, and the large language model is used to optimize the official document paragraphs corresponding to the official document elements, relatively accurate and complete official document information can be extracted from the official document to be processed.

[0068] It should be noted that the electronic device implementing this method can be, but is not limited to, a mobile terminal, a tablet computer, a laptop computer, a desktop computer, a server, etc. with sufficient computing power. When it is a server, the server can be a single server or a server group. The server group can be centralized or distributed (for example, the server can be a distributed system). In some embodiments, the server can be local or remote relative to the user terminal. In some embodiments, the server can be implemented on a cloud platform; only as an example, the cloud platform can include a private cloud, a public cloud, a hybrid cloud, a community cloud, a distributed cloud, an inter-cloud, a multi-cloud, etc., or any combination thereof. In some embodiments, the server can be implemented on an electronic device with one or more components.

[0069] To make the solution provided in this embodiment clearer, the following takes a server as an electronic device and elaborates on each step of the method in combination with Figure 1 The following elaborates on each step of the method in detail. However, it should be understood that the operations in the flowchart may not be implemented in sequence, and steps without a logical context relationship can be reversed or implemented simultaneously. In addition, those skilled in the art can add one or more other operations to the flowchart or remove one or more operations from the flowchart under the guidance of the content of this application. As Figure 1 shown, the method includes:

[0070] S1, obtain the official document to be processed.

[0071] S2, process the official document to be processed through a pre-trained element extraction model to obtain the document elements in the official document to be processed and the document paragraphs corresponding to the document elements.

[0072] For this element extraction model, this embodiment also provides a training method for the element extraction model. The following elaborates on this training method in detail:

[0073] Step 1, obtain the official document sample and the training label of the official document sample. The training label includes the sample document elements in the official document sample and the sample document paragraphs corresponding to the sample document elements.

[0074] It should be understood that this embodiment does not limit the number of official document samples. It can collect multiple historical official document documents, identify the document elements in the documents through techniques such as layout analysis, and aggregate the paragraph content in the documents to form labels in the format of {"title": [AAA], "paragraph": [BBB], "table": [CCC], "picture": [DDD]}; then, analyze and sample different official document theme types to make the data volume ratio of each theme type appropriate; finally, divide the historical official document documents marked with labels into official document samples including a training set and a test set.

[0075] For the above labels, "title", "paragraph", "table", and "picture" in the labels can all be regarded as document elements, and the text in "[]" is the document paragraph corresponding to the document element in text form. This document paragraph can be either the original text in the official document sample or an appropriate summary of the original text content. It should be noted that considering that official documents not only have ordinary text content but also tables and pictures, etc. Therefore, the sample document paragraphs corresponding to the sample document elements include at least one of text content, table content, and image content.

[0076] Exemplarily, to make the above tags easier to understand, the above tags are described exemplarily below. Official documents usually contain document elements such as "first-level, second-level or third-level headings", "document number", "title", "issuing agency", "issuing agency logo", "main recipient agency", "paragraph", etc. Then the tags of the official document sample can be expressed as {"document number": [...], "title": [...], "issuing agency": [...], "issuing agency logo": [...]} . Among them, the content filled in "[...]" corresponds to the descriptive text of the document element. For example, the content filled in "[...]" in "issuing agency": [...] is the name of the issuing agency.

[0077] Based on the introduction of the official document sample in the above embodiment, the training method further includes:

[0078] Step 2, input the official document sample into the model to be trained for processing, and obtain the first prediction result of the sample official document element and the second prediction result of the sample official document paragraph.

[0079] In this embodiment, a relatively mature neural network model in the field of natural processing can be selected as the model to be trained. For example, in this embodiment, Chinese-BERT-WWM is selected as the model to be trained. Chinese-BERT-WWM (Chinese BERT with Whole Word Masking) is a pre-trained model based on the BERT (Bidirectional Encoder Representations from Transformers) model, which is specially designed for Chinese natural language processing tasks. Compared with the original BERT model, Chinese-BERT-WWM adopts the Whole Word Masking technology, better considers the characteristics of the Chinese language, and improves the performance of the model in Chinese natural language processing tasks. And, in order to adapt to the task requirements of this embodiment, the structural parameters of the model are adjusted on the basis of maintaining the original structure of Chinese-BERT-WWM. Specifically, for the BERT core in Chinese-BERT-WWM, its number of network layers is 12, the hidden layer size is 768, the number of attention heads is 12, the learning rate set is 3e-5, the number of fine-tuning iterations is 24, the weight decay set is 0.01, and the gradient accumulation step number set is 1.

[0080] The model to be trained predicts the sample official document element and the corresponding sample official document paragraph, and updates the model parameters according to the difference between the prediction result and the tag, so that the model can complete the expected element extraction task. Therefore, in this embodiment, the training method further includes:

[0081] Step 3: Obtain the first model loss of the model to be trained based on the first prediction result of the sample official document elements.

[0082] Step 4: Obtain the second model loss of the model to be trained based on the second prediction result of the sample official document paragraphs.

[0083] Step 5: Obtain the comprehensive loss of the model to be trained based on the first model loss and the second model loss.

[0084] It can be understood that in this embodiment, the model is trained to simultaneously extract official document elements and the corresponding official document paragraphs by means of multi-task fine-tuning, and the model to be trained is fine-tuned by combining the model losses of both. As an alternative implementation, the comprehensive loss of the model to be trained is obtained by weighting the weights of the first model loss and the second model loss respectively.

[0085] Regarding the weights of the above two model losses, the more conventional solution in the art often pre-sets weights of a fixed size. However, in this embodiment, in order to enable the trained extraction model to have good effects in both element extraction and official document paragraph extraction, a new weight adjustment method is proposed. Specifically, the weight of the first model loss decreases as the number of iterations increases, and the weight of the second model loss increases as the number of iterations increases. Among them, the loss function of the model to be trained can be expressed as:

[0086] Loss = α1×loss1 + α2×loss2;

[0087] In the formula, Loss represents the comprehensive loss, loss1 represents the first model loss, loss2 represents the second model loss, α1 represents the weight of the first model loss, and α2 represents the weight of the second model loss.

[0088] Since the weight of the first model loss decreases as the number of iterations increases, and the weight of the second model loss increases as the number of iterations increases, it can be understood that in the early stage of model training, element extraction is the main focus; after the extraction accuracy of the official document elements gradually stabilizes, it is necessary to further conduct more refined learning on the official document paragraphs to which the official document elements belong. That is, throughout the training process, α1 gradually decays and α2 gradually increases. In this embodiment, the decay process of α1 can be controlled by the following expression:

[0089]

[0090] Correspondingly, the increase process of α2 can be controlled by the following expression:

[0091]

[0092] In the formula, step represents the current iteration number, and total_step represents the total number of iterations.

[0093] Step 6: Determine whether the model to be trained meets the preset convergence condition. If not, return to Step 2 for execution until the preset convergence condition is met, and then obtain the element extraction model.

[0094] In this way, by iterating the above Steps 2 to 6 until the model to be trained meets the preset convergence condition, the model to be trained at this time is used as the element extraction model to extract document elements and the corresponding document paragraphs from the document to be processed. For the convenience of users, the extraction model can be deployed to the server so that users can access it online.

[0095] Based on the above implementation, to extract document elements and the corresponding document paragraphs from the document to be processed, continue to refer to Figure 1 , the document information extraction method provided in this embodiment further includes:

[0096] S3: Generate a prompt for instructing a preset large language model to optimize the document paragraph according to the document element and the corresponding document paragraph.

[0097] In an optional implementation, the server obtains a pre-written prompt template; fills the document element and the corresponding document paragraph into the reserved positions in the prompt template to obtain the prompt.

[0098] Exemplarily, for the document elements extracted by the element extraction model, obtain the overall document paragraph corresponding to the document element, and return it in the following format: {"Document Element A": [XXX], "Document Element B": [AAAAAAXXXAAAAA],...}. For example, {"First-level Heading": [Implementation Opinions on Further Strengthening Work Safety], "Second-level Heading": [Implementing and Implementing the Work Safety Responsibility System], "First Line of Paragraph": [To further strengthen work safety, ensure the safety of people's lives and property, in accordance with relevant laws and regulations, and in light of the actual situation in our city, the following implementation opinions on strengthening work safety are hereby put forward],...}.

[0099] Taking the "First Line of Paragraph" in the above example: [To further strengthen work safety, ensure the safety of people's lives and property, in accordance with relevant laws and regulations, and in light of the actual situation in our city, the following implementation opinions on strengthening work safety are hereby put forward] as an example, filling it into the reserved position in the template, the following prompt can be obtained:

[0100] As a tool for optimizing and completing factor extraction results, please help me optimize the extraction results from the following extracted official document elements and their corresponding official document paragraph texts. Content other than paragraphs cannot be used. The paragraph content is: [In order to further strengthen safety production work and ensure the safety of people’s lives and property, in accordance with relevant laws and regulations and combined with the actual situation of our city, the following implementation opinions are proposed to strengthen safety production work]. The results are returned in the form of a dictionary {"first line of paragraph": [optimized factor results]}.

[0101] Based on the prompt words assembled in the embodiment, they are sent to the preset language model to optimize the official document paragraphs, mainly including completing and polishing the extracted official document paragraphs. Figure 1 The document information extraction method provided in this embodiment also includes:

[0102] S4, sending the prompt words to the preset large language model for processing to obtain an optimized document paragraph.

[0103] In this embodiment, a currently mature large language model can be selected as the above-mentioned preset large language model. For example, the preset large language model can be, but is not limited to, ChatGLM 6B model, ChatGPT model, Claude model, etc.

[0104] The above embodiment introduces the implementation process of the document information extraction method. To make the above embodiment easier to understand, the above embodiment is summarized as follows: Figure 2 Multiple links shown.

[0105] Based on the same inventive concept as the document information extraction method provided in this embodiment, this embodiment also provides a document information extraction device, which includes at least one software function module that can be stored in a memory or fixed in an electronic device in the form of software. The processor in the electronic device is used to execute the executable module stored in the memory. For example, the software function module and computer program included in the device. Please refer to Figure 3 , functionally speaking, the device may include:

[0106] The document acquisition module 11 is used to acquire the documents to be processed;

[0107] The information extraction module 12 is used to process the official document to be processed by using a pre-trained element extraction model to obtain the official document elements in the official document to be processed and the official document paragraphs corresponding to the official document elements;

[0108] The information optimization module 13 is used to generate prompt words for instructing a preset large language model to optimize the official document paragraph according to the official document elements and the official document paragraphs corresponding to the official document elements; and send the prompt words to the preset large language model for processing to obtain the optimized official document paragraph.

[0109] In this embodiment, the official document acquisition module 11 is used to implement Figure 1 step S1 in Figure 1 ; the information extraction module 12 is used to implement step S2 in

[0110] ; and the information optimization module 13 is used to implement steps S3 and S4 in

[0111] . Therefore, for the detailed descriptions of the above modules, reference can be made to the specific implementation manners of the corresponding steps. Of course, for the above modules, in view of having the same inventive concept as the official document information extraction method provided in this embodiment, the above modules can also be used to implement other steps or sub-steps of this method, which will not be elaborated herein.

[0112] In addition, in each embodiment of the present application, the functional modules can be integrated together to form an independent part, or each module can exist alone, or two or more modules can be integrated to form an independent part.

[0113] It should also be understood that if the above implementation manners are implemented in the form of software function modules and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application.

[0112] Therefore, this embodiment also provides a storage medium that stores a computer program. When the computer program is executed by a processor, it implements the official document information extraction method provided in this embodiment. Among them, the storage medium can be various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc.

[0113] An electronic device for implementing the official document information extraction method provided in this embodiment. As Figure 4 shown, the electronic device may include a processor 22 and a memory 21. And the memory 21 stores a computer program. The processor reads and executes the computer program corresponding to the above implementation manner in the memory 21 to implement the official document information extraction method provided in this embodiment.

[0114] Continue to refer to Figure 4, the electronic device further includes a communication unit 23. Each of the memory 21, the processor 22, and the communication unit 23 is directly or indirectly electrically connected to each other through a system bus 24 to achieve data transmission or interaction.

[0115] Among them, the memory 21 can be an information recording device based on any electronic, magnetic, optical or other physical principles for recording execution instructions, data, etc. In some embodiments, the memory 21 can be, but is not limited to, a volatile memory, a non-volatile memory, a storage drive, etc.

[0116] In some embodiments, the volatile memory can be a random access memory (RAM); in some embodiments, the non-volatile memory can be a read only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a flash memory, etc.; in some embodiments, the storage drive can be a disk drive, a solid state drive, any type of storage disk (such as an optical disk, a DVD, etc.), or a similar storage medium, or a combination thereof, etc.

[0117] The communication unit 23 is used to transmit and receive data via a network. In some embodiments, the network may include a wired network, a wireless network, an optical fiber network, a telecommunication network, an intranet, the Internet, a local area network (LAN), a wide area network (WAN), a wireless local area network (WLAN), a metropolitan area network (MAN), a wide area network (WAN), a public switched telephone network (PSTN), a Bluetooth network, a ZigBee network, or a near field communication (NFC) network, etc., or any combination thereof. In some embodiments, the network may include one or more network access points. For example, the network may include a wired or wireless network access point, such as a base station and / or a network switching node, and one or more components of the service request processing system may be connected to the network through the access point to exchange data and / or information.

[0118] The processor 22 may be an integrated circuit chip with signal processing capabilities, and the processor may include one or more processing cores (e.g., a single-core processor or a multi-core processor). By way of example only, the above-mentioned processor may include a central processing unit (CPU), an application specific integrated circuit (ASIC), an application specific instruction-set processor (ASIP), a graphics processing unit (GPU), a physics processing unit (PPU), a digital signal processor (DSP), a field programmable gate array (FPGA), a programmable logic device (PLD), a controller, a microcontroller unit, a reduced instruction set computing (RISC), or a microprocessor, etc., or any combination thereof.

[0119] It can be understood that Figure 4The structure shown is only illustrative. The electronic device may also have more or fewer components than Figure 4 shown, or have a different configuration from Figure 4 shown. Figure 4 Each of the components shown may be implemented by hardware, software, or a combination thereof.

[0120] It should be understood that the devices and methods disclosed in the above embodiments may also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or actions, or may be implemented by a combination of dedicated hardware and computer instructions.

[0121] As described above, these are only various embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed in the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A method for extracting official document information, characterized in that, The method includes: Obtain the official document to be processed; Process the official document to be processed through a pre-trained element extraction model to obtain the official document elements in the official document to be processed and the official document paragraphs corresponding to the official document elements; Generate a prompt for instructing a preset large language model to optimize the official document paragraph according to the official document elements and the official document paragraphs corresponding to the official document elements; Send the prompt to the preset large language model for processing to obtain an optimized official document paragraph; The method further includes a training method for the element extraction model, and the training method includes: Obtain an official document sample and a training label for the official document sample, where the training label includes the sample official document elements in the official document sample and the sample official document paragraphs corresponding to the sample official document elements; Input the official document sample into the model to be trained for processing to obtain a first prediction result of the sample official document elements and a second prediction result of the sample official document paragraphs; Obtain a first model loss loss1 of the model to be trained according to the first prediction result of the sample official document elements; obtain a second model loss loss2 of the model to be trained according to the second prediction result of the sample official document paragraphs; perform weighting according to the weights of the first model loss and the second model loss respectively to obtain a comprehensive loss Loss = α1loss1 + α2loss2 of the model to be trained; where the expressions of the weights α1 of the first model loss and α2 of the second model loss are respectively: In the formula, step represents the current iteration number, and total_step represents the total number of iterations; If the model to be trained does not meet the preset convergence condition, update the model to be trained according to the comprehensive loss, and return to the step of inputting the official document sample into the model to be trained for processing to obtain a first prediction result of the sample official document elements and a second prediction result of the sample official document paragraphs, and execute until the preset convergence condition is met to obtain the element extraction model.

2. The official document information extraction method according to claim 1, characterized in that, The generating a prompt for instructing a preset large language model to optimize the official document paragraph according to the official document elements and the official document paragraphs corresponding to the official document elements includes: Obtain a pre-written prompt template; Fill the official document elements and the official document paragraphs corresponding to the official document elements into the reserved positions in the prompt template to obtain the prompt.

3. The official document information extraction method according to claim 1, wherein The sample official document paragraphs include at least one of text content, table content, and image content.

4. An official document information extraction device, characterized in that, The device includes: An official document acquisition module for obtaining the official document to be processed; An information extraction module for processing the official document to be processed through a pre-trained element extraction model to obtain the official document elements in the official document to be processed and the official document paragraphs corresponding to the official document elements; An information optimization module for generating a prompt for instructing a preset large language model to optimize the official document paragraph according to the official document elements and the official document paragraphs corresponding to the official document elements; sending the prompt to the preset large language model for processing to obtain an optimized official document paragraph; The device further includes a model training module, and the model training module is configured to: Obtain a document sample and a training label of the document sample, where the training label includes sample document elements in the document sample and sample document paragraphs corresponding to the sample document elements; Input the document sample into the model to be trained for processing, and obtain a first prediction result of the sample document elements and a second prediction result of the sample document paragraphs; According to the first prediction result of the sample document elements, obtain a first model loss loss1 of the model to be trained; according to the second prediction result of the sample document paragraphs, obtain a second model loss loss2 of the model to be trained; perform weighting according to the weights of the first model loss and the second model loss respectively, and obtain a comprehensive loss Loss = α1loss1 + α2loss2 of the model to be trained; where the expressions of the weight α1 of the first model loss and the weight α2 of the second model loss are respectively: In the formula, step represents the current iteration number, and total_step represents the total number of iterations; If the model to be trained does not meet the preset convergence condition, update the model to be trained according to the comprehensive loss, and return to execute the step of inputting the document sample into the model to be trained for processing to obtain the first prediction result of the sample document elements and the second prediction result of the sample document paragraphs until the preset convergence condition is met, and then obtain the element extraction model.

5. A storage medium, characterized in that, The storage medium stores a computer program, and when the computer program is executed by a processor, it implements the document information extraction method according to any one of claims 1-3.

6. An electronic device, characterized in that, The electronic device includes a processor and a memory, the memory stores a computer program, and when the computer program is executed by the processor, it implements the document information extraction method according to any one of claims 1-3.

Citation Information

Patent Citations

  • Document domain element extraction method and device

    CN114282518A

  • Key information extraction model training method and device and storage medium thereof

    CN115587184A

  • Scoring method and device, equipment and storage medium

    CN116227500A

  • Method for generating article content and electronic equipment

    CN117875274A