A method, apparatus and electronic device for generating a title extraction model
The training samples are generated by fusion of text and image features, and the LayoutXLM training model and self-attention mechanism are used to solve the problem of low efficiency and accuracy in title extraction and serial number correction of long and diverse documents in title extraction and serial number correction, achieving efficient and accurate title extraction and serial number correction.
Patent Information
- Application Number
- CN202210413888.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-15
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2042-04-15
AI Technical Summary
In the prior art, documents with long and diverse layouts have low extraction efficiency and low accuracy during title extraction and serial number correction processing, which affects the user experience.
The training samples are generated from the fusion of text and image feature information, and the LayoutXLM training model is used to extract document titles, and feature fusion and semantic entity recognition are performed through self-attention mechanism and conditional random field to generate a title extraction model.
It improves the accuracy and efficiency of document title extraction, realizes accurate extraction of multi-level titles and serial number correction, and improves user experience.
Smart Images

Figure CN114724166B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of data processing technologies, in particular to natural language processing, deep learning, optical character recognition, data processing and other technical fields, and specifically relates to a method, an apparatus, and an electronic device for generating a title extraction model. Background Art
[0002] Document intelligence refers to the process by which a computer automatically reads, understands, and analyzes documents. The popularization of deep learning technologies has greatly promoted the development of the field of document intelligence represented by document information extraction. The extraction of multi-level titles (including titles in the body text format) and title number error correction for PDF documents are widely applied under the requirements of document structuring, abstract extraction, and reducing document error rates.
[0003] Generally, the documents to be processed are characterized by long lengths and diverse formats. However, in the prior art, when extracting titles and performing number error correction on documents with long lengths and diverse formats, the extraction efficiency is low, and the accuracy of the extracted titles is low, which reduces the user experience. Summary of the Invention
[0004] The present disclosure provides a method, an apparatus, and an electronic device for generating a title extraction model.
[0005] According to one aspect of the present disclosure, a method for generating a title extraction model is provided. The method includes: obtaining a document sample, where the document in the document sample is in image format; extracting text feature information from the document in the document sample, and extracting image feature information from the document, where the text feature information represents the text content and text position of the text included in the document sample, and the image feature information represents the document layout of the document included in the document sample; annotating the document sample based on the text feature information to obtain an annotated document sample; performing feature fusion on the annotated document sample and the image feature information to obtain a training sample; and generating a title extraction model based on the training sample, where the title extraction model is used to extract titles from a document to be processed.
[0006] As can be seen from the above, the present disclosure uses a method of fusing semantic features extracted from text and images in a document sample to generate a training sample for training a title extraction model, and uses the training sample to train the title extraction model to extract titles from a document to be processed. The solution provided by the present disclosure achieves the purpose of extracting titles in a document, realizes the effect of improving the extraction accuracy of document titles, and further solves the problem of low extraction accuracy existing in the prior art when extracting titles of documents.
[0007] According to another aspect of the present disclosure, there is provided a generating device for a title extraction model, including: an acquisition module configured to acquire document samples, wherein the documents in the document samples are in image format; a feature extraction module configured to perform text feature extraction on the documents in the document samples to obtain text feature information, and perform image feature extraction on the documents to obtain image feature information, wherein the text feature information characterizes the text content and text positions of the texts included in the document samples, and the image feature information characterizes the document layout of the documents included in the document samples; an annotation module configured to annotate the document samples based on the text feature information to obtain the annotated document samples; a feature fusion module configured to perform feature fusion on the annotated document samples and the image feature information to obtain training samples; a model generation module configured to generate a title extraction model based on the training samples, wherein the title extraction model is used to extract titles from documents to be processed.
[0008] According to another aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the above-mentioned generating method of the title extraction model.
[0009] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the above-mentioned generating method of the title extraction model.
[0010] According to another aspect of the present disclosure, there is provided a computer program product including a computer program, and the computer program realizes the above-mentioned generating method of the title extraction model when executed by a processor.
[0011] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0013] Figure 1 is a flowchart of the generating method of the title extraction model according to the first embodiment of the present disclosure;
[0014] Figure 2 is a schematic diagram of the title level display form according to the first embodiment of the present disclosure;
[0015] Figure 3 is a block diagram of the generating of the title extraction model according to the second embodiment of the present disclosure;
[0016] Figure 4 It is a schematic diagram of a document to be processed according to the third embodiment of the present disclosure;
[0017] Figure 5 It is a schematic diagram of a generating device of a title extraction model according to the third embodiment of the present disclosure;
[0018] Figure 6 It is a block diagram of an electronic device for implementing the generating method of the title extraction model according to the embodiments of the present disclosure. Detailed implementation manners
[0019] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted below.
[0020] It should be noted that in the technical solutions of the present disclosure, the acquisition, storage, and application of user personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0021] In addition, it should be noted that in each of the embodiments provided by the present disclosure, the electronic device can be used as an execution subject to execute the methods provided by each of the embodiments.
[0022] Embodiment 1
[0023] According to an embodiment of the present disclosure, the present disclosure provides a generating method of a title extraction model, wherein, Figure 1 It is a flowchart of the method, and it can be seen from Figure 1 that the method at least includes the following steps:
[0024] Step S102, obtain a document sample, wherein the documents in the document sample are in image format.
[0025] In step S102, the document sample consists of at least one document, and each document consists of at least one page of the document, and each page of the document is in image format. For example, the above-mentioned document can be, but is not limited to, a PDF (Portable Document Format) document.
[0026] It should be noted that in the present disclosure, the documents in the document sample can be documents with long lengths and various formats. For example, in the present disclosure, the document sample can be a sample composed of a prospectus with a length of hundreds of pages and a complex format.
[0027] In addition, it should be noted that the solution provided by the present disclosure can be applied to various scenarios where it is necessary to extract the document title. In different scenarios, the corresponding document samples are also different. For example, in the scenario of corporate debt financing, the above-mentioned document sample can be a sample composed of PDF version bill prospectuses; in the scientific research scenario, the above-mentioned document sample can be a sample composed of PDF version academic papers. That is, in the present disclosure, the document sample can be determined based on different scenarios.
[0028] Step S104: Extract text feature information from the documents in the document sample, and extract image feature information from the documents.
[0029] In step S104, the text feature information characterizes the text content and text position of the text included in the document sample, and the image feature information characterizes the document layout of the documents included in the document sample.
[0030] Optionally, the electronic device can perform a segmentation process on each document in the document sample, and segment the document into multiple pictures in units of pages. Among them, the picture formats of the multiple pictures can be the same or partially the same. The picture formats of the multiple pictures can be PNG format (i.e., bitmap format), and can also be JPG format, JPEG format, PSD format, TIFF format, etc.
[0031] Furthermore, after obtaining the multiple pictures, the electronic device performs text recognition on each picture through OCR (Optical Character Recognition) technology, so as to obtain the above-mentioned text feature information. Among them, the text feature information at least includes text content information and text position information. The text content information at least includes the text content in units of lines and the text content in units of characters in the picture, and the text position information at least includes the coordinate information in units of lines and the coordinate information in units of characters in the picture.
[0032] In addition, during the process of extracting image feature information from the pictures, the electronic device can extract the color information, character size information, text line spacing information, text alignment information, etc. of the text in the pictures, and generate a feature map based on the information extracted above.
[0033] It should be noted that the above-mentioned text position information can be the coordinate information of the two points of the upper left corner and the lower right corner of the line or character relative to the upper left corner of the picture.
[0034] In addition, it should be noted that compared with the prior art solutions that only extract titles through image OCR or only through document text, the present disclosure utilizes the semantic features of both text and image modalities to process document samples, which can avoid the problems of poor title extraction effect and low extraction efficiency existing in the solutions of only extracting titles through image OCR or only through document text, thereby improving the extraction accuracy and efficiency of document titles.
[0035] Furthermore, compared with the prior art in which effective features need to be screened and continuously iterated during the process of extracting title features through machine learning, in the present disclosure, text feature information and image feature information are directly extracted from document samples. This process does not require feature engineering to screen features, simplifies the feature extraction steps, and thus simplifies the generation steps of the title extraction model.
[0036] Step S106: Label the document sample based on the text feature information to obtain the labeled document sample.
[0037] In step S106, the electronic device can label the title of the document sample based on a labeling platform to obtain the labeled document sample, where the labeling platform can be, but is not limited to, the Doccano labeling platform.
[0038] Step S108: Perform feature fusion on the labeled document sample and the image feature information to obtain a training sample.
[0039] In step S108, the electronic device can use a self-attention mechanism based on spatial perception to achieve feature fusion. The self-attention mechanism is a mechanism that focuses on local information, such as image regions in an image. Compared with the self-attention mechanism that usually captures the relationship between tokens based on absolute positions, the self-attention mechanism based on spatial perception can make full use of semantic relative positions and spatial relative positions to calculate attention weights, thereby capturing local invariance in the document layout.
[0040] Step S110: Generate a title extraction model based on the training sample, where the title extraction model is used to extract the title in the document to be processed.
[0041] In step S110, the electronic device can train the LayoutXLM training model based on the training sample to obtain the title extraction model. Since the LayoutXLM training model has the characteristic of cross-language, the title extraction model trained through the solution provided by the present disclosure can also be applied to the extraction of document titles in cross-language scenarios, thereby improving the scalability and applicability of the title extraction model.
[0042] In addition, after obtaining the title extraction model, the user can use the title extraction model to extract the titles in the document to be processed and determine the title level corresponding to each title. In addition, after extracting the titles in the document to be processed and the title level corresponding to each title, the electronic device can also display the titles in a tree form according to the title level. For example, in the form of the title level display shown in Figure 2 to display the title level.
[0043] Based on the solution defined in the above steps S102 to S110, it can be known that the present disclosure adopts a method of fusing semantic features extracted from the text and images in the document sample to generate training samples for training the title extraction model, and uses the training samples to train the title extraction model to extract the titles in the document to be processed.
[0044] It is easy to notice that in the above process, the text feature information and the image feature information are directly extracted from the document sample. This process does not require feature engineering to screen features, simplifies the feature extraction steps, and further simplifies the generation steps of the title extraction model. In addition, the present disclosure uses the semantic features of both text and image modalities to process the document sample, and then obtains the training samples. That is, the present disclosure generates the training samples from a multi-modal perspective, thereby improving the accuracy of the training samples, and further improving the accuracy of the title extraction model for extracting document titles.
[0045] It can be seen that the solution provided by the present disclosure achieves the purpose of extracting the titles in the document, realizes the effect of improving the extraction accuracy of the document titles, and further solves the problem of low extraction accuracy existing in the prior art when extracting the titles of PDF documents.
[0046] Embodiment 2
[0047] According to an embodiment of the present disclosure, the present disclosure also provides a method for generating a title extraction model. In this embodiment, the electronic device generates a title extraction model based on the Figure 3 shown block diagram for generating the title extraction model.
[0048] In an optional embodiment, after obtaining the document sample, the electronic device extracts text features from the documents in the document sample to obtain text feature information.
[0049] Specifically, the electronic device performs segmentation processing on the documents included in the document sample to obtain multiple images corresponding to the documents. Then, text recognition is performed on each of the multiple images to obtain a first document and a second document. Each document has at least one page. The first document at least includes: first text content in units of lines and first position information of at least one line of text in the corresponding document. The second document at least includes: the first text content, the first position information, second text content in units of characters, and second position information of at least one character in the corresponding document. The text feature information at least includes the first text content, the first position information, the second text content, and the second position information.
[0050] Optionally, as Figure 3 shown, the electronic device cuts each sample in the document sample into multiple images by page. For example, PNG images. Then, each image is used to generate the above-mentioned first document through optical character recognition technology (i.e., OCR technology). The first document contains text content in units of lines (i.e., the first text content) and position information of each line of text (i.e., the first position information). After obtaining the above-mentioned first text content, the electronic device also uses a syntax analyzer (e.g., PDF syntax analyzer) to perform syntax analysis on the first text content and uses carriage return characters to connect the text content in units of lines.
[0051] Similar to the first document, the second document contains text content in units of lines and related position information, and also contains text content in units of characters and related position information. The document format of the first document is different from that of the second document. Optionally, the first document is a txt document and the second document is a JSON document. As Figure 3 shown, in the embedding layer of the title extraction model, the first text vector is the vector corresponding to the text content in units of lines, the second text vector is the text vector corresponding to the text content in units of characters, and the first information vector and the second information vector respectively represent the vectors corresponding to the above-mentioned first position information and the second position information.
[0052] It should be noted that from the above content, in the present disclosure, in the process of training the model, not only image data and text data are used, but also text coordinates (i.e., the above-mentioned first position information and second position information) are used, enriching the modal information of the data used in model training. Compared with the existing title extraction models, the solution proposed in the present disclosure can more fully integrate text features and image features, thereby optimizing the title extraction effect of the title extraction model.
[0053] In addition, it should also be noted that as Figure 3As shown, after the electronic device obtains multiple images corresponding to the document sample, it extracts image features based on the ResNet_FPN backbone network integrated in the open-source Detectron2 object detection toolbox to obtain a feature map, and the feature map vector corresponding to the feature map is in the embedding layer of the title extraction model.
[0054] Furthermore, after obtaining the text feature information and the image feature information, the electronic device annotates the document sample based on the text feature information to obtain an annotated document sample. Specifically, the electronic device annotates the title of the first document to obtain the annotated first document, and then obtains the title content in the annotated first document, and annotates the title of the second document based on the title content to obtain the annotated document sample.
[0055] Optionally, the electronic device imports the first document into the Doccano annotation platform for entity annotation, that is, the Doccano annotation platform annotates the title of the first document. Then, the electronic device obtains the entity value (i.e., the title content) corresponding to the first document according to the entity annotation, and locates the relevant information in the second document in units of lines through string matching to implement the annotation of the second document, so as to obtain the annotated document sample. For example, in the first document, "I. Basic Information of the Initiating Institution" is the title. After the electronic device annotates it in the first document, during the annotation of the second document, the electronic device searches for the text content of "I. Basic Information of the Initiating Institution" from the text content of the second document in units of lines, and in the second document, marks the searched text content as the title.
[0056] It should be noted that by annotating the title of the document sample, the trained title extraction model can accurately extract the title of the document to be processed, thereby improving the accuracy of document title extraction.
[0057] Even further, after annotating the document sample, the electronic device performs feature fusion on the annotated document sample and the image feature information to obtain a training sample. Specifically, the electronic device determines the text sequence feature information based on the annotated document sample, and performs feature fusion on the text sequence feature information and the image feature information to obtain a training sample. Among them, the text sequence feature information includes at least one of the following: the label corresponding to the document sample, entity information, and label identification.
[0058] Optionally, as Figure 3As shown, the electronic device uses a self-attention mechanism based on spatial perception for feature fusion at the feature fusion layer. Specifically, the electronic device encapsulates the label, entity information, label identifier, and image feature information corresponding to the document sample at the feature fusion layer, and then performs iterative processing in the data loader to obtain training samples.
[0059] It should be noted that to prevent the problem of information exposure in the bidirectional language model, in this disclosure, after the iterative processing in the data loader, the electronic device can also add an attention mask, Attention_Mask, to improve data security.
[0060] In addition, it should also be noted that in this disclosure, by using a self-attention mechanism based on spatial perception to achieve feature fusion of text sequence feature information and image feature information, it is possible to fully utilize the semantic relative position and spatial relative position to calculate the attention weight, and then capture the local invariance in the document layout. Moreover, this disclosure uses the semantic features of both text and image modalities to process the document sample, and then obtains the training sample, that is, this disclosure generates the training sample from a multi-modal perspective, thereby improving the accuracy of the training sample, and further improving the accuracy of the title extraction model for extracting the document title.
[0061] In addition, before performing feature fusion on the text sequence feature information and the image feature information, the electronic device can also adjust the image size. For example, the image size can be adjusted to 3*224*224 to avoid the problem of increasing the computational complexity of the electronic device due to too large an image size, and also avoid the problem of inaccurate calculation results of the electronic device due to too small an image size.
[0062] In an optional embodiment, in the process of determining the text sequence feature information based on the annotated document sample, the electronic device performs label conversion on the annotated document sample to obtain the label corresponding to the document sample. Then, the index value and label corresponding to the title content in the first annotated document are encapsulated to obtain entity information, and the label is subjected to identification conversion to obtain a label identifier.
[0063] Optionally, the electronic device uses a sequence annotation method to perform label conversion on the title in the annotated document sample to obtain the above-mentioned label. Among them, the BIO (Beginning-Inside-Outside) method can be used for label conversion, where O and B represent first-level headings, and I represents a first-level heading. Then, the electronic device encapsulates the start index and end index of the title content in the first document, as well as the label, to obtain the above-mentioned entity information. Furthermore, the electronic device performs identification conversion on the label. For example, the token is converted into a token identifier, token_id, to obtain the label identifier.
[0064] It should be noted that by determining the text sequence feature information, the annotation of the document sample is realized, and further, the trained title extraction model can accurately extract the title of the document to be processed, improving the accuracy of document title extraction. In addition, the labels generated by data annotation are also indispensable data for the entire supervised training.
[0065] Furthermore, the electronic device can also perform semantic entity recognition on the text sequence feature information to determine the title level of at least one title included in the document sample.
[0066] Optionally, the electronic device can also perform semantic entity recognition on the text sequence feature information in the CRF (Conditional Random Field) layer to determine the title level corresponding to each title.
[0067] It should be noted that by performing semantic recognition on the text sequence feature information, the title level corresponding to each title in the document sample can be determined, which not only ensures the accuracy of title extraction but also the accuracy of the title level.
[0068] In addition, it should also be noted that as Figure 3 shown, after determining the text sequence feature information, the electronic device performs fine-tuning in the fully connected layer and the conditional random field layer. Therefore, compared with the softmax layer, the solution provided by the present disclosure adds a constraint on whether the predicted label is legal, reducing the probability of prediction errors.
[0069] Furthermore, the electronic device can also obtain the label length of the label identifier corresponding to the target document, and when the label length is greater than the preset length, the target document is segmented into multiple sub-documents. The target document is any one of the document samples.
[0070] Optionally, the above preset length can be, but is not limited to, the hyperparameter max_len. In addition, when segmenting the target document, the electronic device can set a random number for the number of segments, and the number of pages corresponding to each sub-document can be the same or different.
[0071] It should be noted that segmenting the target document according to the label length can increase the number of training samples, thus ensuring the training accuracy of the title extraction model.
[0072] Based on the above, taking the aforementioned prospectus as an example, using the solution provided by the present disclosure, 4072 pieces of data were marked. Among them, 3257 pieces were randomly selected as the training set, 408 pieces as the validation set, and 407 pieces as the test set for model training. Before model training, download the LayoutXML-baseBase pre-trained model on the electronic device, specify the path to load the pre-trained model, specify the maximum number of training steps as 1000, save the detection file checkpoint every 500 steps, and the warmup_ratio of the iteration rate can be set to 0.1.
[0073] It should be noted that in the above example, other pre-trained models can also be loaded, for example, the checkpoint model.
[0074] In addition, the effect of the title extraction model is verified on the test set as follows:
[0075] Precision: 92.1% Recall: 98.4% F1-score: 95.1%
[0076] On the same dataset, the effect of the Bert+CRF model is as follows:
[0077] Precision: 91.9% Recall: 93.5% F1-score: 92.7%
[0078] Among them, the F1-score is an index used in statistics to measure the accuracy of a binary classification model, which can be determined by precision and recall.
[0079] It can be seen that the solution provided by the present disclosure is superior to the Bert+CRF model in terms of the index effect of the extraction model.
[0080] Embodiment 3
[0081] According to an embodiment of the present disclosure, the present disclosure also provides a method for generating a title extraction model. In this embodiment, after generating the title extraction model, the electronic device can also detect and / or correct the titles extracted by the title extraction model.
[0082] Specifically, the electronic device first obtains the document to be processed, extracts the title of the document to be processed based on the title extraction model to obtain at least one title corresponding to the document to be processed. Then, determine the index order and title level of at least one title in the document to be processed, and determine the subordinate relationship between at least one title based on the index order and title level to generate a multi-way tree. Finally, detect the title numbers of at least one title based on the nodes corresponding to the multi-way tree to obtain a detection result, where the detection result indicates whether there is an error or omission in the title number.
[0083] It should be noted that during the process of extracting the title of the document to be processed by the title extraction model, not only can the title in the document to be processed be obtained, but also the title level of the document title can be obtained. For example, in the schematic diagram of the document to be processed shown in Figure 4 , the first-level title, second-level title, third-level title, and fourth-level title are shown. The electronic device can determine the parent-child relationship between the titles according to the index order of the titles in the text of the document to be processed and the title level, and generate a multi-way tree as shown in Figure 2 . Then, the electronic device can recursively traverse the child nodes of each node and determine whether the child nodes of each node conform to the incremental order. If the child nodes do not conform to the incremental order, the electronic device determines that there is an error in the title number; if it is determined according to the incremental order that there is a missing title number between two child nodes, the electronic device determines that there is a missing title number and corrects the title number. For example, in Figure 2 , there is a missing title number between "3. Credit Risk" and "5. Liquidity Risk". At this time, the electronic device changes "5. Liquidity Risk" to "4. Liquidity Risk" and modifies the remaining title numbers.
[0084] In addition, it should be noted that the solution provided by the present disclosure can implement error detection and correction of the title level, avoiding the problem of low efficiency of error detection and correction of the title level caused by manual error detection and correction of the title level in the prior art, and improving the efficiency of error detection and correction of the title level.
[0085] Based on the above-mentioned first to third embodiments, the present disclosure provides a multi-modal title extraction and serial number error correction solution for a title extraction model. This solution can respectively fuse the semantic features extracted from the text and the image, and adopt a fine-tuning task similar to named entity recognition (NER) - semantic entity recognition (SER) to achieve the extraction of multi-level titles in the PDF document. Then, the extracted titles are generated into a multi-way tree according to the title level, and recursively determine whether the child nodes of each tree node conform to the incremental order, so as to determine whether the title number is incorrect or missing, and perform error correction.
[0086] Compared with the Bert pre-training model in the field of NLP (Natural Language Processing), the solution provided by the present disclosure not only does not require feature engineering to screen features, but also makes full use of the semantic features of both text and image modalities. In addition, compared with the softmax layer, the present disclosure introduces a Conditional Random Field (CRF) during fine-tuning. Compared with the softmax layer, it adds a constraint on whether the predicted label is legal and reduces the probability of prediction errors. Additionally, the title extraction model uses the LayoutXLM training model, and the LayoutXLM training model has the characteristic of cross-language, which makes the solution proposed by the present disclosure have better scalability and applicability.
[0087] Embodiment 4
[0088] According to an embodiment of the present disclosure, the present disclosure also provides a generating device for a title extraction model, wherein, Figure 5 is a schematic diagram of the device, from Figure 5 it can be seen that the device includes: an acquisition module 501, a feature extraction module 503, a labeling module 505, a feature fusion module 507, and a model generation module 509.
[0089] Among them, the acquisition module 501 is used to acquire a document sample, wherein the document in the document sample is in image format; the feature extraction module 503 is used to perform text feature extraction on the document in the document sample to obtain text feature information, and perform image feature extraction on the document to obtain image feature information, wherein the text feature information characterizes the text content and text position of the text included in the document sample, and the image feature information characterizes the document layout of the document included in the document sample; the labeling module 505 is used to label the document sample based on the text feature information to obtain the labeled document sample; the feature fusion module 507 is used to perform feature fusion on the labeled document sample and the image feature information to obtain a training sample; the model generation module 509 is used to generate a title extraction model based on the training sample, wherein the title extraction model is used to extract the title in the document to be processed.
[0090] It should be noted that the above acquisition module 501, feature extraction module 503, labeling module 505, feature fusion module 507, and model generation module 509 correspond to steps S102 to S110 in the above embodiment. The examples and application scenarios implemented by the five modules and the corresponding steps are the same, but are not limited to the content disclosed in the above embodiment.
[0091] Optionally, the feature extraction module includes: a first segmentation module and a text recognition module. Among them, the first segmentation module is used to segment the document included in the document sample to obtain multiple images corresponding to the document; the text recognition module is used to perform text recognition on at least one of the multiple images to obtain a first document and a second document, where the first document at least includes: first text content in units of lines and first position information of at least one line of text in the corresponding document, and the second document at least includes: the first text content, the first position information, second text content in units of characters, and second position information of at least one character in the corresponding document, and the text feature information at least includes the first text content, the first position information, the second text content, and the second position information.
[0092] Optionally, the annotation module includes: a first annotation module, a first acquisition module, and a second annotation module. Among them, the first annotation module is used to perform title annotation on the first document to obtain the annotated first document; the first acquisition module is used to acquire the title content in the annotated first document; the second annotation module is used to perform title annotation on the second document based on the title content to obtain the annotated document sample.
[0093] Optionally, the feature fusion module includes: a first determination module and a first fusion module. Among them, the first determination module is used to determine text sequence feature information based on the annotated document sample, where the text sequence feature information at least includes one of the following: a label corresponding to the document sample, entity information, and a label identifier; the first fusion module is used to perform feature fusion on the text sequence feature information and the image feature information to obtain a training sample.
[0094] Optionally, the first determination module includes: a first conversion module, a packaging module, and a second conversion module. Among them, the first conversion module is used to perform label conversion on the annotated document sample to obtain the label corresponding to the document sample; the packaging module is used to package the index value and the label corresponding to the title content in the annotated first document to obtain entity information; the second conversion module is used to perform identifier conversion on the label to obtain a label identifier.
[0095] Optionally, the generating device of the title extraction model further includes: an entity recognition module, which is used to perform semantic entity recognition on the text sequence feature information to determine the title level of at least one title included in the document sample.
[0096] Optionally, the generating device of the title extraction model further includes: a second acquisition module and a second segmentation module. Among them, the second acquisition module is used to acquire the label length of the label identifier corresponding to the target document, where the target document is any one of the documents in the document sample; the second segmentation module is used to segment the target document into multiple sub-documents when the label length is greater than a preset length.
[0097] Optionally, the generating device of the title extraction model further includes: a third acquisition module, a title extraction module, a second determination module, a third determination module, and a detection module. Among them, the third acquisition module is used to acquire the document to be processed; the title extraction module is used to extract the title of the document to be processed based on the title extraction model to obtain at least one title corresponding to the document to be processed; the second determination module is used to determine the index order and title level of at least one title in the document to be processed; the third determination module is used to determine the subordinate relationship between at least one title based on the index order and title level to generate a multi-way tree; the detection module is used to detect the title numbers of at least one title based on the nodes corresponding to the multi-way tree to obtain a detection result, where the detection result indicates whether there is an error or omission in the title number.
[0098] Embodiment 5
[0099] According to an embodiment of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium, and a computer program product.
[0100] Figure 6 FIG. shows a schematic block diagram of an exemplary electronic device 600 that can be used to implement the embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processing, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0101] As Figure 6 shown, the device 600 includes a computing unit 601, which can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 602 or the computer program loaded from the storage unit 608 into the random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the device 600 can also be stored. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. The input / output (I / O) interface 605 is also connected to the bus 604.
[0102] Multiple components in device 600 are connected to I / O interface 605, including: an input unit 606, such as a keyboard, a mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a magnetic disk, an optical disc, etc.; and a communication unit 609, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 609 allows device 600 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0103] The computing unit 601 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 executes the various methods and processes described above, such as the method for generating a title extraction model. For example, in some embodiments, the method for generating a title extraction model can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the computing unit 601, one or more steps of the method for generating a title extraction model described above can be executed. Alternatively, in other embodiments, the computing unit 601 can be configured to execute the method for generating a title extraction model by any other suitable means (e.g., by means of firmware).
[0104] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0105] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor or controller, the program codes cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program codes can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0106] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0107] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).
[0108] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.
[0109] A computer system can include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, or a server of a distributed system, or a server incorporating a blockchain.
[0110] It should be understood that various forms of the processes shown above can be used, steps can be reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this is not limited herein.
[0111] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. A method for generating a title extraction model, comprising: Obtaining document samples, wherein the documents in the document samples are in image format; Performing text feature extraction on the documents in the document samples to obtain text feature information, and performing image feature extraction on the documents to obtain image feature information, wherein the text feature information characterizes the text content and text positions of the texts included in the document samples, and the image feature information characterizes the document layout of the documents included in the document samples; Annotating the document samples based on the text feature information to obtain annotated document samples; Performing feature fusion on the annotated document samples and the image feature information to obtain training samples; Generating a title extraction model based on the training samples, wherein the title extraction model is used to extract titles from documents to be processed; Wherein, performing text feature extraction on the documents in the document samples to obtain text feature information further includes: Performing segmentation processing on the documents included in the document samples to obtain multiple images corresponding to the documents; Performing text recognition on at least one of the multiple images to obtain a first document and a second document, wherein the first document at least includes: first text content in units of lines and first position information of at least one line of text in the corresponding document, and the second document at least includes: the first text content, the first position information, second text content in units of characters, and second position information of at least one character in the corresponding document.
2. The method according to claim 1, wherein, The text feature information at least includes the first text content, the first position information, the second text content, and the second position information.
3. The method according to claim 2, wherein, Annotating the document samples based on the text feature information to obtain annotated document samples includes: Performing title annotation on the first document to obtain an annotated first document; Obtaining the title content in the annotated first document; Performing title annotation on the second document based on the title content to obtain the annotated document samples.
4. The method according to claim 3, wherein, Performing feature fusion on the annotated document samples and the image feature information to obtain training samples includes: Determining text sequence feature information based on the annotated document samples, wherein the text sequence feature information at least includes one of the following: labels corresponding to the document samples, entity information, and label identifiers; Performing feature fusion on the text sequence feature information and the image feature information to obtain the training samples.
5. The method according to claim 4, wherein, Determining text sequence feature information based on the annotated document samples includes: Performing label conversion on the annotated document samples to obtain labels corresponding to the document samples; Encapsulating the index value corresponding to the title content in the annotated first document and the labels to obtain entity information; Performing identifier conversion on the labels to obtain label identifiers.
6. According to the method described in claim 4, the method further includes: Performing semantic entity recognition on the text sequence feature information to determine the title levels of at least one title included in the document samples.
7. The method according to claim 5, the method further comprising: Obtaining the label length of the label identifier corresponding to the target document, wherein the target document is any one of the document samples; When the label length is greater than a preset length, splitting the target document into a plurality of sub-documents.
8. The method according to claim 1, the method further comprising: Obtaining the document to be processed; Performing title extraction on the document to be processed based on the title extraction model to obtain at least one title corresponding to the document to be processed; Determining the index order and title level of the at least one title in the document to be processed; Determining the subordinate relationship between the at least one title based on the index order and the title level to generate a multi-way tree; Detecting the title numbers of the at least one title based on the nodes corresponding to the multi-way tree to obtain a detection result, wherein the detection result indicates whether the title numbers are incorrect or missing.
9. A device for generating a title extraction model, comprising: An acquisition module, configured to acquire a document sample, wherein the documents in the document sample are in image format; A feature extraction module, configured to perform text feature extraction on the documents in the document sample to obtain text feature information, and perform image feature extraction on the documents to obtain image feature information, wherein the text feature information characterizes the text content and text position of the text included in the document sample, and the image feature information characterizes the document layout of the documents included in the document sample; A labeling module, configured to label the document sample based on the text feature information to obtain a labeled document sample; A feature fusion module, configured to perform feature fusion on the labeled document sample and the image feature information to obtain a training sample; A model generation module, configured to generate a title extraction model based on the training sample, wherein the title extraction model is used to extract titles in a document to be processed; Wherein, the feature extraction module includes: A first splitting module, configured to split the documents included in the document sample to obtain a plurality of images corresponding to the documents; A text recognition module, configured to perform text recognition on at least one of the plurality of images to obtain a first document and a second document, wherein the first document at least includes: first text content in units of lines and first position information of at least one line of text in the corresponding document, and the second document at least includes: the first text content, the first position information, second text content in units of characters, and second position information of at least one character in the corresponding document.
10. The apparatus according to claim 9, wherein, The text feature information at least includes the first text content, the first position information, the second text content, and the second position information.
11. The device according to claim 10, wherein, The labeling module includes: A first labeling module, configured to perform title labeling on the first document to obtain a labeled first document; A first acquisition module, configured to acquire the title content in the labeled first document; A second annotation module, configured to perform title annotation on the second document based on the title content, so as to obtain the annotated document sample.
12. The apparatus according to claim 11, wherein, The feature fusion module includes: A first determination module, configured to determine text sequence feature information based on the annotated document sample, where the text sequence feature information includes at least one of the following: a label corresponding to the document sample, entity information, and a label identifier; A first fusion module, configured to perform feature fusion on the text sequence feature information and the image feature information to obtain the training sample.
13. The device according to claim 12, wherein, The first determination module includes: A first conversion module, configured to perform label conversion on the annotated document sample to obtain the label corresponding to the document sample; An encapsulation module, configured to encapsulate the index value corresponding to the title content in the first annotated document and the label to obtain entity information; A second conversion module, configured to perform identifier conversion on the label to obtain a label identifier.
14. The apparatus according to claim 12, wherein the apparatus further includes: An entity recognition module, configured to perform semantic entity recognition on the text sequence feature information to determine the title level of at least one title included in the document sample.
15. The apparatus according to claim 13, wherein the apparatus further includes: A second acquisition module, configured to acquire the label length of the label identifier corresponding to the target document, where the target document is any one of the document samples; A second segmentation module, configured to segment the target document into multiple sub-documents when the label length is greater than a preset length.
16. The apparatus according to claim 9, wherein the apparatus further includes: A third acquisition module, configured to acquire the document to be processed; A title extraction module, configured to extract a title from the document to be processed based on the title extraction model to obtain at least one title corresponding to the document to be processed; A second determination module, configured to determine the index order and the title level of the at least one title in the document to be processed; A third determination module, configured to determine the subordination relationship between the at least one title based on the index order and the title level, and generate a multi-way tree; A detection module, configured to detect the title number of the at least one title based on the nodes corresponding to the multi-way tree to obtain a detection result, where the detection result indicates whether the title number is incorrect or missing.
17. An electronic device includes: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method for generating the title extraction model according to any one of claims 1 to 8.
18. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause a computer to execute the method for generating the title extraction model according to any one of claims 1 to 8.
19. A computer program product includes a computer program, and the computer program implements the method for generating the title extraction model according to any one of claims 1 to 8 when executed by a processor.
Citation Information
Patent Citations
Document understanding method and device, electronic equipment and medium
CN113836268A