A Performance Evaluation Method for Document Understanding Models

By summarizing the four basic capabilities required for document understanding models and building corresponding performance evaluation benchmarks and methods, the problem that the existing technology cannot effectively evaluate document understanding models is solved, and a more comprehensive and accurate model evaluation is achieved.

CN116340465BActive Publication Date: 2025-06-13INST OF SOFTWARE - CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310391444.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-13
Publication Date
2025-06-13
Estimated Expiration
2043-04-13

AI Technical Summary

Technical Problem

Existing natural language processing performance evaluation benchmarks are mainly aimed at short texts, and it is not possible to effectively and comprehensively evaluate the performance of document understanding models, especially when dealing with long documents and complex structures.

Method used

A performance evaluation benchmark and method for document understanding model is proposed. By summarizing the four basic abilities required for document understanding: document classification, structural analysis, information extraction and translator capabilities, and building performance evaluation benchmarks and methods based on these abilities.

Benefits of technology

This method can more comprehensively evaluate the model's document understanding ability, reveal the advantages and shortcomings of the model, and promote the research and improvement of the document understanding model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116340465B_ABST
    Figure CN116340465B_ABST
Patent Text Reader

Abstract

The present invention discloses a performance evaluation method for a document understanding model, and its steps include: 1) constructing a performance evaluation benchmark; 2) processing document data according to the benchmark to obtain a data set for testing different performances; 3) implementing a text classification model with the document understanding model to be tested as the backbone, training and testing on the document classification data set to obtain the document classification ability value of the model; 4) implementing a sequence annotation model with the document understanding model as the backbone, training and testing on the document structure analysis data set to obtain the document structure analysis ability value of the model; 5) implementing a question-answering model with the document understanding model as the backbone, training and testing on the document information extraction data set to obtain the document information extraction ability value of the model; 6) implementing a generation model with the document understanding model as the backbone, training and testing on the document transcribing data set to obtain the document transcribing ability value of the model; 7) obtaining the evaluation result of the model according to the above results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of natural language processing, and specifically, to a method for evaluating the performance of a document understanding model. Background Art

[0002] Documents are the basic units of natural language organization and widely exist in real-world text data, such as news reports, academic papers, legal contracts, government reports, etc. Document understanding aims to automatically extract, interpret, and process the required information from documents, which has received wide attention as the basis for many downstream tasks, driving the rapid emergence of many new document understanding models. How to effectively and comprehensively evaluate the document understanding capabilities of these models has become an urgent problem to be solved.

[0003] In the field of natural language processing, existing performance evaluation benchmarks are mainly constructed for short texts. The most famous benchmark GLUE and its improved version SuperGLUE evaluate the natural language understanding capabilities of models through the performance of models on sentence or sentence pair tasks. However, a document is not a simple aggregation of isolated sentences or paragraphs. First, the length of a document is usually much longer than that of a sentence or paragraph, and due to the limitation of running memory, it cannot be completely input into a neural network model. Second, a document usually has a potentially complex structure, and understanding the associations and hierarchical relationships of basic elements in the document is a key part of correct document understanding. Finally, the information in a document is usually scattered throughout the full text, and document understanding requires more context than short text understanding. Considering the unique challenges faced by document understanding, existing short text performance evaluation benchmarks cannot effectively distinguish the document understanding capabilities of models.

[0004] Therefore, it is necessary to design a document understanding model evaluation benchmark and method to effectively and comprehensively evaluate the document understanding capabilities of models, so as to explore the current situation, ability boundaries, and possible defects of document understanding models, and promote the development of the field of natural language processing in document understanding. Summary of the Invention

[0005] To overcome the deficiency that existing benchmarks cannot effectively and comprehensively evaluate the document understanding capabilities of models, the present invention proposes a performance evaluation benchmark and method for a document understanding model, and the specific content includes: 1. Summarization of the basic capabilities required for document understanding; 2. Construction of a performance evaluation benchmark based on the above division of document understanding capabilities; 3. A performance evaluation method based on the above performance evaluation benchmark. Constructing a performance evaluation benchmark based on the division of document understanding capabilities can more comprehensively evaluate the document understanding capabilities of models, and at the same time can show the advantages and weaknesses of the document understanding capabilities of models, thereby providing a favorable basis for the research and improvement of document understanding models.

[0006] In the first aspect, the present invention summarizes four basic capabilities required for document understanding, including:

[0007] 1. Document classification ability, which evaluates whether the document understanding model can completely understand the semantics of the document, such as the theme or stance of the document;

[0008] 2. Document structure analysis ability, which evaluates whether the document understanding model can analyze and utilize the potential structural information of the document, such as discourse structure or argumentative structure;

[0009] 3. Document information extraction ability, which evaluates whether the document understanding model can identify and aggregate the relevant information scattered throughout the full text, such as long-distance co-reference;

[0010] 4. Document transcribing ability, which evaluates whether the document understanding model can capture and transcribe the key information in the document, such as the document abstract.

[0011] In the second aspect, the present invention provides a method for constructing a performance evaluation benchmark for a document understanding model, including:

[0012] 1. Collect the document data of the document understanding task to obtain multiple data sets; wherein, the documents in the data sets need to be natural texts, with a length exceeding 512 characters, and the key information is scattered throughout the full text;

[0013] 2. Manually aggregate the collected data sets into 4 categories according to the basic capabilities required for document understanding in the first aspect above;

[0014] 3. Create corresponding tasks and unify the data format according to each category of document set; specifically, create a classification task for the data set used to test the document classification ability, create a sentence-level sequence annotation task for the data set used to test the document structure analysis ability, create a multi-answer question-and-answer task for the data set used to test the document information extraction ability, and create a text generation task for the data set used to test the document transcribing ability.

[0015] In the third aspect, the present invention provides a performance evaluation method for a document understanding model, including:

[0016] 1. Implement a text classification model with the document understanding model to be tested as the backbone as the detection model, train and test the detection model on the document classification data set, and obtain the evaluation result of the document classification ability of the document understanding model to be tested;

[0017] Specifically, the text classification model is composed of two layers of multi-layer perceptron layers connected after the document understanding model to be tested; 2. Implement a sequence annotation model with the document understanding model to be tested as the backbone as the detection model, and in the document structure analysis

[0018] analysis data set, train and test the detection model to obtain the document structure analysis ability of the document understanding model to be tested

[0019] Evaluation results; specifically, the sequence labeling model consists of a document understanding model to be detected followed by a conditional random field; 3. Implement a question-answering model with the document understanding model to be tested as the backbone as the detection model, train and test the detection model on the document information extraction dataset, and obtain the evaluation results of the document information extraction ability of the document understanding model to be tested; specifically, the question-answering model consists of a document understanding model to be detected, a softmax layer that predicts the probability of a word in the document as the start or end of an answer, an answer number prediction component, and a selection component;

[0020] 4. Implement a generation model with the document understanding model to be tested as the backbone as the detection model, train and test the detection model on the document transcription dataset, and obtain the evaluation results of the document transcription ability of the document understanding model to be tested; specifically, the generation model consists of a document understanding model to be detected followed by a decoder composed of 12 layers of Transformer Decoder;

[0021] 5. Calculate the arithmetic mean of the evaluation results of the above different abilities to obtain the comprehensive document understanding performance evaluation results of the test model.

[0022] Furthermore, the method for evaluating the document classification ability of the model is as follows:

[0023] 1. For the document d to be classified in the document classification dataset, construct the input: X = {[CLS]d[SEP]}; where the [CLS] and [SEP] symbols are the start symbol and the separator symbol respectively, used to help the model understand the start and end of the input sequence;

[0024] 2. Input the constructed sequence X into the document understanding model in the text classification model;

[0025] 3. Obtain the top-level start symbol [CLS] word vector as the representation vector H of the document d to be classified

[0026] ; [CLS] ;

[0027] 4. Input the representation vector H of the document d to be classified [CLS] into the multi-layer perceptron layer for classification to obtain the category of the document d to be classified;

[0028] 5. Synthesize the classification results of all documents in the document classification dataset to obtain the final classification accuracy.

[0029] Furthermore, the method for evaluating the document structure analysis ability of the model is as follows:

[0030] 1. For the document d = {s 1 , s 2 , …, s n}, first insert the start symbol [CLS] before each sentence to construct the input: X = {[CLS]s 1 [CLS]s 2 …[CLS]s n}}; s n is the nth sentence in document d;

[0031] 2. Input the constructed sequence X into the document understanding model to be tested;

[0032] 3. Obtain the top-level start symbol [CLS] word vector as the representation vector H of the following sentences [CLS] ;

[0033] 4. Input the vector representation H of the sentence [CLS] into the conditional random field to obtain the label y corresponding to each sentence = CRF(H [CLS] ); 5. Synthesize the classification results of all documents in the document structure analysis dataset to obtain the final document structure analysis performance.

[0034] Furthermore, the method for evaluating the model's document information extraction ability is as follows:

[0035] 1. Given the question q and the document d in the document information extraction dataset, concatenate the two using the model's general concatenation symbol to construct the input: X = {[CLS]q[SEP]d[SEP]};

[0036] 2. Input the constructed sequence X into the document understanding model to be tested;

[0037] 3. Obtain the vector representation H of each word w i in X i ;

[0038] 4. Dot-multiply the vector representation H of each word i with the start vector S and pass through the first softmax layer to obtain the probability that each word w i is the start boundary of the answer: where H i is the vector representation of the ith word w i in document d, H j is the vector representation of the jth word in document d, the start vector S is a newly introduced auxiliary variable, and its initial value is randomly initialized and then trained and updated on the dataset; where the objective function is to maximize the log-likelihood sum of the correct start and end positions of the answer in the given document d, and the start vector S is updated according to the gradient descent method; SH i is the product result of vector S and H i , is the exponential operation on the vector product result;

[0039] 5. Dot the vector representation H of each word i with the end vector E, and pass it through the second softmax layer to obtain the probability of each word w i as the end boundary of the answer: where the end vector E is a newly introduced auxiliary variable, whose initial value is randomly initialized and then trained and updated on the dataset; where the objective function is to maximize the log-likelihood sum of the correct start and end positions of the answer in the given document d, and the end vector E is updated according to the gradient descent method; EH i is the product of the vector E and H i result, is the exponential operation on the vector product result;

[0040] 6. The answer number prediction component calculates the probability of the text segment (k, l) as the answer based on the probability of each word as the start and end boundaries of the answer:

[0041] 7. Obtain the vector representations H Q and H P of the sub-top layer start symbols [CLS] and [SEP] as the representations of the question and the document; the document understanding model to be tested generally has a multi-layer transformer architecture, and there are corresponding word vectors in each layer. The sub-top layer refers to the second-to-last layer transformer architecture of the document understanding model, and the top layer refers to the last layer transformer architecture of the document understanding model.

[0042] 8. Based on the representation vectors of the question and the document, calculate the probability distribution p span = softmax(FFN([H Q ; H P ; H [CLS] )), and select the answer number t with the highest probability as the final answer number;

[0043] 9. The selection component selects the t text segments with the highest probability as the answer according to the determined answer number t;

[0044] 10. Integrate the extraction results of all documents in the document information extraction dataset to obtain the final performance of the document information extraction ability; the performance of the document information extraction ability can be quantified by calculating the overlap degree between the model-predicted answer and the correct answer.

[0045] Furthermore, the method for evaluating the model transcribing ability is:

[0046] 1. Given the document d in the document transcribing dataset, construct the input: X = {[CLS]d[SEP]};

[0047] 2. Input the constructed sequence X into the document understanding model to be tested, and obtain the vector representation H corresponding to the sequence X;

[0048] 3. Input the vector representation H into the decoder composed of 12 concatenated Transformer Decoders to obtain the transcribed text Y; The 12 Transformer Decoders are concatenated in series to obtain the decoder;

[0049] 4. Synthesize the transcription results of all documents in the document transcription dataset to obtain the final performance of the document transcription ability.

[0050] In a fourth aspect, the present invention further provides a computer device, which is characterized by including a memory and a processor. The memory stores a computer program, and the computer program is configured to be executed by the processor. The computer program includes instructions for executing the steps in the above method.

[0051] In a fifth aspect, the present invention further provides a computer-readable storage medium, on which a computer program is stored. The computer program is characterized in that when executed by a processor, it implements the steps of the above method.

[0052] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0053] 1. The performance evaluation method and benchmark proposed by the present invention are for document construction. Compared with the existing short text performance evaluation benchmarks, it can evaluate the document understanding ability of the model more accurately and discriminately, reveal the development status of the document understanding model, and provide a reference for the research in this field;

[0054] 2. The performance evaluation method and benchmark proposed by the present invention are constructed based on the division of document understanding ability. The multi-dimensional evaluation system makes the evaluation of document understanding ability more comprehensive, and can better reveal the advantages and disadvantages of the model, enabling R & D personnel to select a suitable document understanding model according to the task requirements;

[0055] 3. The performance evaluation method and benchmark proposed by the present invention are constructed based on multi-tasks, providing a platform for the research of transfer learning and general models in the field of document understanding. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 This is an example of constructing a performance evaluation benchmark for a document understanding model provided by the present invention.

[0057] Figure 2 This is a schematic flowchart of a performance evaluation method for a document understanding model provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0058] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. The examples given are only for explaining the present invention and are not intended to limit the scope of the present invention.

[0059] Figure 1 This is an example of constructing a performance evaluation benchmark for a document understanding model provided by the present invention. In Figure 1 this, the evaluation of document understanding ability can be carried out from four dimensions: document classification, document structure analysis, document information extraction, and document transcription. The overall evaluation can be obtained by comprehensively considering the above four index dimensions.

[0060] Among them, the document classification ability refers to the ability of the model to capture the overall semantics of the document; specifically, given a sequence s or a sequence pair (s 1 , s 2 ), if the model gives the correct classification l for most inputs, it indicates that the model has good document classification ability. As an example, here we choose the Hyperpartisan and ContractNLI datasets as representatives, corresponding to the classification task and natural language inference task respectively.

[0061] The document structure analysis ability refers to the ability of the model to capture and utilize document structure information; specifically, given a document d = {s 1 , s 2 , …, s n}, if the model can correctly identify the labels t = {t 1 , t 2 , …, t n} corresponding to most sentences, it indicates that the model has good document structure analysis ability. As an example, here we choose the ECOM, RR, and GUM datasets as representatives, corresponding to opinion mining, argument mining, and discourse parsing tasks respectively.

[0062] The document information extraction ability refers to the ability of the model to identify and extract relevant information scattered throughout the text; specifically, given a question q and a document d, if the model can correctly extract the answers a = {a 1 , a 2 , …, a m} to most questions from the document, it indicates that the model has good document information extraction ability. As an example, here we choose the LitBank and NarrativeQA datasets as representatives, corresponding to coreference resolution and question answering tasks respectively.

[0063] The document transcription ability refers to the ability of the model to capture and transcribe key information in the document; specifically, given an input sequence s in , if the model can transcribe a smooth new text sequence s out according to the task requirements, it indicates that the model has good document transcription ability. As an example, here Qasper, SummScreen, and GovReport datasets are selected as representatives, corresponding to the question answering and summarization tasks respectively.

[0064] Based on the above performance evaluation benchmark for the document understanding model, Figure 2 is a schematic flowchart of a performance evaluation method for a document understanding model provided by the present invention. Using the model to be tested as the main model, a text classification model, a sequence annotation model, a question answering model, and a generation model are respectively implemented, trained and tested on the corresponding performance evaluation datasets. Finally, the ability evaluation index scores for document classification, document structure information, document information extraction, and document transcription of the model to be tested can be obtained respectively. Finally, the arithmetic mean of the four ability evaluation index scores is calculated to obtain the comprehensive evaluation of the document understanding ability of the model to be tested.

[0065] Although specific embodiments of the present invention are disclosed for illustrative purposes, which are intended to help understand the content of the present invention and implement it accordingly, those skilled in the art can understand that: without departing from the spirit and scope of the present invention and the appended claims, various substitutions, changes, and modifications are possible. Therefore, the present invention should not be limited to the content disclosed in the best embodiments, and the scope of protection required by the present invention is defined by the scope of the claims.

Claims

1. A method for evaluating the performance of a document understanding model, the steps of which include: 1) Construct a performance evaluation benchmark for the document understanding model, where the performance evaluation benchmark includes document classification ability, document structure analysis ability, document information extraction ability, and document transcription ability; 2) Collect document data for the document understanding task, and process the collected document data according to the performance evaluation benchmark to obtain a data set for testing different performances; including a document classification data set, a document structure analysis data set, a document information extraction data set, and a document transcription data set; 3) Implement a text classification model with the document understanding model to be tested as the backbone, train and test the text classification model on the document classification data set, and obtain the document classification ability value of the document understanding model according to the test results; 4) Implement a sequence annotation model with the document understanding model to be tested as the backbone, train and test the sequence annotation model on the document structure analysis data set, and obtain the document structure analysis ability value of the document understanding model according to the test results; 5) Implement a question-answering model with the document understanding model to be tested as the backbone, train and test the question-answering model on the document information extraction data set, and obtain the document information extraction ability value of the document understanding model according to the test results; 6) Implement a generation model with the document understanding model to be tested as the backbone, train and test the generation model on the document transcription data set, and obtain the document transcription ability value of the document understanding model according to the test results; 7) Obtain the document understanding performance evaluation result of the document understanding model according to the document classification ability value, document structure analysis ability value, document information extraction ability value, and document transcription ability value of the document understanding model obtained above.

2. The method according to claim 1, characterized in that the formats of the respective data sets are unified, and a corresponding task is created for each data set; among them, a classification task is created for the document classification data set, a sequence annotation task at the sentence level is created for the document structure analysis data set, a multi-answer question-answering task is created for the document information extraction data set, and a text generation task is created for the document transcription data set.

3. The method according to claim 1 or 2, characterized in that the text classification model is composed of two multi-layer perceptron layers connected after the document understanding model to be tested; the method for training and testing the text classification model is: 31) For the document d to be classified in the document classification data set, construct an input sequence X = {[CLS]d[SEP]}; 32) Input the constructed sequence X into the document understanding model to be tested; 33) Obtain the word vector of the top-level start symbol [CLS] as the representation vector H of the document d to be classified [CLS] ; 34) Represent the vector H [CLS] Input it into the multi-layer perceptron layer for classification to obtain the category of the document d to be classified; 35) Synthesize the classification results of all documents in the document classification data set to obtain the classification accuracy as the document classification ability value.

4. The method according to claim 1 or 2, characterized in that the sequence annotation model is composed of a conditional random field connected after the document understanding model to be detected; the method for training and testing the sequence annotation model is: 41) For the document d = {s 1 , s 2 , …, s n} in the document structure analysis dataset, insert the start symbol [CLS] before each sentence to construct the input sequence X = {[CLS]s 1 [CLS]s 2 … [CLS]s n}; s n is the nth sentence in the document d; 42) Input the constructed sequence X into the document understanding model to be tested; 43) Obtain the word vector of the top-level start symbol [CLS] as the representation vector H of the following sentence [CLS] ; 44) Input the vector representation H of the sentence [CLS] into the conditional random field to obtain the label y corresponding to each sentence, where y = CRF(H [CLS] ); 45) Combine the classification results of all documents in the comprehensive document structure analysis dataset to obtain the final document structure analysis performance as the document structure analysis ability value.

5. The method according to claim 1 or 2, wherein, the question answering model is composed of a document understanding model to be detected, a softmax layer, an answer number prediction component, and a selection component; the method for training and testing the question answering model is: 51) For a given question q and a document d in the document information extraction dataset, construct an input sequence X = {[CLS]q[SEP]d[SEP]}; 52) Input the constructed sequence X into the document understanding model to be tested; 53) Obtain the vector representation H of each word w in sequence X i ; i ; 54) Dot the vector representation H of each word i with the initial vector S, and pass through the first softmax layer to obtain the probability of each word as the start boundary of the answer; 55) Dot the vector representation H of each word i with an end vector E, and pass through a second softmax layer to obtain the probability of each word being the end boundary of the answer; 56) The answer number prediction component calculates the probability of the corresponding text segment as the answer based on the probability of each word as the start boundary of the answer and the probability of each word as the end boundary of the answer; 57) Obtain the word vector of the second top-level start symbol [CLS] as the representation vector H of the question Q ; Obtain the word vector of the second top-level start symbol [SEP] as the representation vector H of the document P ; 58) The selection component is based on the representation vectors H Q 、H P , calculates the probability distribution p of the number of answers span and determines the number of answers t, and then selects the t text segments with the highest probability as the answer; 59) Combine the extraction results of all documents in the document information extraction dataset to obtain the final document information extraction ability performance as the document information extraction ability value.

6. The method according to claim 1, wherein, the generation model is composed of a document understanding model to be detected followed by a decoder; the method for training and testing the generation model is: 61) For a document d in the document transcription dataset, construct an input sequence X = {[CLS]d[SEP]}; 62) Input the constructed sequence X into the document understanding model to be tested to obtain the vector representation H corresponding to the sequence X; 63) Input the vector representation H into the decoder to obtain the transcribed text Y; 64) Combine the transcription results of all documents in the document transcription dataset to obtain the final document transcription ability performance as the document transcription ability value.

7. The method according to claim 6, wherein, the decoder is composed of 12 layers of Transformer Decoder.

8. A server, wherein, it includes a memory and a processor, the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program includes instructions for executing the steps in any one of claims 1 to 7.

9. A computer-readable storage medium, on which a computer program is stored, wherein, the computer program realizes the steps of the method according to any one of claims 1 to 7 when executed by a processor.

Citation Information

Patent Citations

  • BERT-based machine reading understanding method, apparatus and device, and storage medium

    CN112464641A

  • Multi-engine intelligent question answering system for multi-type knowledge base

    CN115238101A