Pre-training Model Training Method, Device and Electronic Device for Reading Tasks
By layout analysis and training of electronic documents, the reading order between text fragments is determined, and the pre-trained model is trained using preset loss functions, the problem of insufficient accuracy of reading order in reading tasks is solved, and higher prediction accuracy is achieved.
Patent Information
- Application Number
- CN202111350551.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-15
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2041-11-15
AI Technical Summary
The existing pre-trained reading task model has shortcomings in learning the accuracy of reading order and cannot meet human reading order needs.
By analyzing the electronic document, obtaining layout information, determining the first reading order between text fragments, and using preset loss functions to train the pretrained model to learn the correct reading order.
Improve the accuracy of the pre-trained model's prediction in the document reading order, ensuring that the model can read the document in the correct order.
Smart Images

Figure CN114241496B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and more specifically, to artificial intelligence technologies such as natural language processing and deep learning. Background Art
[0002] In related technologies, document pre-training mainly models document data, learns the logical relationships contained within the documents through pre-training, and then fine-tunes on various document analysis tasks and extraction tasks. Pre-trained models usually need to introduce multi-modal information such as text, layout, and images. However, in related technologies, the accuracy of the reading order learned by the pre-trained model for reading tasks is relatively poor and cannot meet the reading order requirements of humans. Summary of the Invention
[0003] This application provides a method, apparatus, and electronic device for training a pre-trained model for reading tasks.
[0004] According to the first aspect of this application, there is provided a method for training a pre-trained model for reading tasks, including:
[0005] Performing layout analysis on an electronic document to obtain layout information of the electronic document;
[0006] Determining a first reading order among text segments in the electronic document according to the layout information;
[0007] Inputting character information of the electronic document into a pre-trained model according to the layout information to predict a second reading order among the text segments;
[0008] Calculating a loss value based on a preset loss function according to the first reading order and the second reading order, and training the pre-trained model based on the loss value.
[0009] According to the second aspect of this application, there is provided a device for training a pre-trained model for reading tasks, including:
[0010] An analysis module for performing layout analysis on an electronic document to obtain layout information of the electronic document;
[0011] A determination module for determining a first reading order among text segments in the electronic document according to the layout information;
[0012] A prediction module for inputting character information of the electronic document into a pre-trained model according to the layout information to predict a second reading order among the text segments;
[0013] A training module, configured to calculate a loss value based on a preset loss function according to the first reading order and the second reading order, and train the pre-trained model based on the loss value.
[0014] According to a third aspect of the present application, there is provided an electronic device, comprising:
[0015] At least one processor; and
[0016] A memory communicatively connected to the at least one processor; wherein,
[0017] The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method according to the first aspect.
[0018] According to a fourth aspect of the present application, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method according to the first aspect.
[0019] According to the technical solution of the present application, the document pre-trained model is trained according to the layout information of the electronic document, so that the pre-trained model learns the implicit relationship of the document reading order. Furthermore, when the pre-trained model is used for document reading order prediction, the accuracy of the prediction result can be improved.
[0020] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The drawings are used to better understand the solution and do not constitute a limitation to the present application. Among them:
[0022] Figure 1 is a schematic diagram according to the first embodiment of the present application;
[0023] Figure 2 is a schematic diagram according to the second embodiment of the present application;
[0024] Figure 3 is a schematic diagram according to the third embodiment of the present application;
[0025] Figure 4 is a schematic diagram of an electronic document proposed in the third embodiment of the present application;
[0026] Figure 5 is a schematic diagram of the document segmentation obtained by recognition proposed in the third embodiment of the present application;
[0027] Figure 6 It is a schematic diagram of the reading order of text fragments in each region proposed in the third embodiment of the present application;
[0028] Figure 7 It is a schematic diagram according to the fourth embodiment of the present application;
[0029] Figure 8 It is a schematic diagram according to the fifth embodiment of the present application;
[0030] Figure 9 It is a schematic diagram according to the sixth embodiment of the present application;
[0031] Figure 10 It is a schematic diagram according to the seventh embodiment of the present application;
[0032] Figure 11 It is a block diagram of an electronic device for implementing the method for training a pre-trained model for a reading task in an embodiment of the present application. Detailed implementation manners
[0033] The following makes an explanation of exemplary embodiments of the present application with reference to the accompanying drawings, including various details of the embodiments of the present application to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present application. Similarly, for clarity and conciseness, the description of well-known functions and structures is omitted below.
[0034] In addition, the term "and / or" herein is merely a description of the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " herein generally represents an "or" relationship between the associated objects before and after.
[0035] In the related art, document pre-training mainly models document data, learns the logical relationships contained in the document through pre-training, and then fine-tunes on various document analysis tasks and extraction tasks. Pre-trained models usually need to introduce multi-modal information such as text, layout, and images. However, in the related art, the accuracy of the reading order learned by the pre-trained model for reading tasks is relatively poor and cannot meet the reading order requirements of humans. Therefore, how to read electronic documents in the correct order has become an urgent problem to be solved.
[0036] Based on the above problems, the present application proposes a method, an apparatus, and an electronic device for training a pre-trained model for a reading task, which can train a document pre-trained model according to the layout information of an electronic document, so that the pre-trained model learns the implicit relationship of the document reading order, and further can improve the accuracy of the prediction result when using the pre-trained model to predict the document reading order.
[0037] Figure 1 It is a schematic diagram according to the first embodiment of the present application. It should be noted that the method for training a pre-trained model for a reading task in the embodiments of the present application can be used in the apparatus for training a pre-trained model for a reading task in the embodiments of the present application, and the apparatus can be configured in an electronic device. As Figure 1 shown, the method for training a pre-trained model for a reading task includes the following steps:
[0038] Step 101, perform layout parsing on the electronic document to obtain the layout information of the electronic document.
[0039] It can be understood that the electronic document can be a text document or a picture document. The text document can directly encode the text information to obtain the electronic document, and the picture document can be obtained through the conversion of scanned documents and pictures. In addition, the electronic document can also be a text and picture mixed document. For example, the document contains both text and tables and pictures.
[0040] As a possible example, performing layout parsing on the electronic document can obtain the character information, the position information of the characters, the segment information, and the position information of the segments of the electronic document. Optionally, a layout parsing tool can be used to perform layout parsing on the electronic document to obtain layout information such as the character information, the position information of the characters, the segment information, and the position information of the segments of the electronic document, so as to provide the correct reading order for document pre-training.
[0041] Step 102, determine the first reading order between each text segment in the electronic document according to the layout information.
[0042] As a possible example, the layout information includes the reading order among various text segments in the electronic document. Assuming that the reading order in the layout information is correct, each text segment is numbered according to the above reading order to obtain the first reading order, that is, the reading order among various text segments in the electronic document. For example, the electronic document includes segment 1, segment 2, and segment 3, and the reading order among segment 1, segment 2, and segment 3 in the layout information is: segment 2, segment 1, and segment 3. Assuming that this reading order is correct, the reading order of segment 1, segment 2, and segment 3 can be marked. For example, the next reading order of segment 1 is segment 3, the next reading order of segment 2 is segment 1, and the next reading order of segment 3 is "empty". "Empty" can be understood as segment 3 being the last unread segment in the electronic document.
[0043] It can be understood that the present application can use this first reading order as a label for training the pre-trained model.
[0044] Step 103: Input the character information of the electronic document into the pre-trained model according to the layout information to predict the second reading order among various text segments.
[0045] As a possible example, the character reading order within each text segment in the electronic document can also be determined according to the layout information. Thus, based on the first reading order and the character reading order within each text segment, the reading order among all characters in the electronic document can be obtained. According to the above reading order among all characters, the character information of the electronic document is input into the pre-trained model, and the pre-trained model can predict the reading order among various text segments according to the character information of the electronic document in the layout information to obtain the second reading order.
[0046] Step 104: Calculate the loss value based on the first reading order and the second reading order according to the preset loss function, and train the pre-trained model based on the loss value.
[0047] As a possible example, by marking the text segments, the first reading order is obtained. By predicting the reading order among text segments through the pre-trained model, the second reading order is obtained. According to the first reading order label and the second reading order, the preset loss function is calculated to obtain the loss value, and the pre-trained model is trained based on the loss value.
[0048] It can be understood that the application can be a native program (nativeApp) installed on the local terminal, or can also be a web program (webApp) of the browser on the local terminal. This embodiment does not limit this.
[0049] As a possible example, the above loss function can be a cross-entropy loss function. Based on the preset cross-entropy loss function, calculate the loss value according to the first reading order and the second reading order, and train the pre-trained model based on the loss value, so that the pre-trained model can learn the correct reading order.
[0050] As a possible example, the trained pre-trained model can be fine-tuned in downstream tasks to support various technologies such as classification, sequence labeling, structure prediction, and sequence generation, and build applications such as abstracting, machine translation, image retrieval, and video annotation. The above downstream tasks can be tasks such as text classification and sequence annotation.
[0051] According to the pre-trained model training method for reading tasks in the embodiments of the present application, by performing layout analysis on the electronic document, the layout information of the electronic document is obtained. According to the layout information, determine the first reading order between each text segment in the electronic document, and input the character information of the electronic document into the pre-trained model according to the layout information. The pre-trained model predicts the second reading order between each text segment. Based on a preset loss function, calculate the loss value according to the first reading order and the second reading order, and train the pre-trained model based on the loss value, so that the pre-trained model can read the document in the correct order according to actual needs.
[0052] To ensure that the reading order of each segment can be determined according to the correlation degree between each segment, optionally, input the character information of the electronic document into the conversion module according to the layout information to obtain the representation information of the characters in the electronic document. According to the representation information of the characters, generate the representation information of each text segment respectively, and input the representation information of each text segment into the self-attention module to obtain the correlation relationship between each text segment. According to the correlation relationship, predict the second reading order between each text segment. Figure 2 It is a schematic diagram according to the second embodiment of the present application. It should be noted that the pre-trained model training method for reading tasks in the embodiments of the present application can be executed by the pre-trained model training device in the embodiments of the present application. In some embodiments of the present application, as Figure 2 shown, the pre-trained model training method for reading tasks includes:
[0053] Step 201, perform layout analysis on the electronic document to obtain the layout information of the electronic document.
[0054] In the embodiments of the present application, step 201 can be implemented in any one of the embodiments of the present application respectively. The embodiments of the present application do not make any limitations on this and will not be elaborated further.
[0055] Step 202, according to the layout information, determine the first reading order between each text segment in the electronic document.
[0056] In the embodiments of the present application, step 202 can be implemented in any one of the various embodiments of the present application. The embodiments of the present application do not limit this and will not elaborate further.
[0057] Step 203, input the character information of the electronic document into the conversion module according to the layout information, and obtain the representation information of the characters in the electronic document.
[0058] It should be noted that in some embodiments of the present application, the pre-trained model may include a conversion module and a self-attention module.
[0059] As a possible example, the representation information of the characters can be the vector representation information of the characters, and the conversion module can be a Transformer conversion module. Optionally, input the character information of the electronic document into the Transformer, and the conversion module Transformer obtains the vector representation information of the characters according to the character information of the electronic document.
[0060] That is to say, taking the character information of the electronic document as the input, using the Transformer as the character encoder to output the character-level feature vector, that is, output the representation information of the characters in the above-mentioned electronic document.
[0061] For example, optical character recognition (OCR) can be performed on the electronic document to obtain text information. By analyzing the text information through an analyzer, the character information in the text, the two-dimensional position information of the characters in the document, the line number information of the characters, the segment information to which the characters belong, and the sorting information of the characters within the segment can be obtained. The information of each segment in the electronic document can also be extracted through a visual encoder to obtain the two-dimensional position information of each segment in the electronic document. Input the information extracted by the above analyzer and visual encoder into the conversion module, so as to obtain the vector representation information of the characters of the character information of the electronic document.
[0062] Step 204, generate the representation information of each text segment according to the representation information of the characters.
[0063] It can be understood that according to the above representation information of the characters, it can be the vector representation information of the characters.
[0064] As a possible example, average the representation information of the characters belonging to the same text segment, so as to generate the representation information of each text segment respectively, that is, by averaging the vectors of the characters, obtain the vector representation information of each text segment.
[0065] Step 205, input the representation information of each text segment into the self-attention module to obtain the correlation relationship between each text segment.
[0066] As a possible example, the vector representation information of each text segment is input into the self-attention module. The self-attention module calculates the weight scores of each text segment relative to other segments according to the vector representation information of each text segment. For example, the weight scores of one segment relative to each of the other segments are calculated, obtaining multiple weight scores. The above calculation of the weight scores is performed for each segment, thereby obtaining the correlation relationships between the respective text segments.
[0067] Step 206: Predict the second reading order between the respective text segments according to the correlation relationships.
[0068] As a possible example, in response to one of the weight scores being the highest according to the calculated weight score values obtained above, it can be predicted that the segment corresponding to the weight score is the next segment of that segment. That is, according to the predicted correlation relationships between the segments, the second reading order between the respective text segments can be obtained.
[0069] Step 207: Calculate a loss value based on a preset loss function according to the first reading order and the second reading order, and train the pre-trained model based on the loss value.
[0070] In the embodiments of the present application, step 207 can be implemented in any one of the various embodiments of the present application. The embodiments of the present application do not make any limitations in this regard and will not be elaborated further.
[0071] According to the pre-trained model training method for reading tasks in the embodiments of the present application, the character information of the electronic document is input into the conversion module according to the layout information, obtaining the representation information of the characters in the electronic document. According to the representation information of the characters, the representation information of each text segment is generated. The representation information of each text segment is input into the self-attention module, obtaining the correlation relationships between the respective text segments. According to the correlation relationships, the second reading order between the respective text segments is predicted, so that the reading order of each segment can be determined according to the degree of correlation between the segments, improving the rationality and accuracy of the reading order of the pre-trained model.
[0072] To ensure that the reading order of the electronic document can be reasonably predicted according to the layout information, optionally, the document image of the electronic document is obtained, the characters in the electronic document are divided according to the document image, obtaining multiple text segments, and the segment information of the multiple text segments and the position information of each of the multiple segments are obtained, thereby obtaining the layout information. Figure 3 It is a schematic diagram according to the third embodiment of the present application. In some embodiments of the present application, as Figure 3 shown, the pre-trained model training method for reading tasks includes:
[0073] Step 301: Obtain the document image of the electronic document.
[0074] As a possible example, the above-mentioned document image can be obtained by scanning a paper document with a scanning device, or can also be a document photo obtained by taking a picture of the document. This embodiment does not make special limitations on this.
[0075] Step 302: Divide the characters in the electronic document according to the document image to obtain multiple text fragments.
[0076] As a possible example, as Figure 4 shown, the document image can be column-divided or block-divided through an image algorithm, so as to obtain the partition information of the electronic document. Optionally, the above-mentioned document image can be converted into a grayscale image, and then the grayscale Figure 2 value can be thresholded to convert it into a black-and-white image. Among them, the background of the document is black, and the characters of the document are white. Optionally, the above-mentioned image algorithm can be the XY Cut algorithm. According to the preset line spacing threshold or column spacing threshold, the position range of column division or the position range of block division is obtained by using the image algorithm, that is, the partition information is obtained.
[0077] As a possible example, as Figure 4 and Figure 5 shown, the document image can be parsed through an OCR method, so as to obtain the text fragments in the document image and the position information of the text fragments.
[0078] As a possible example, as Figure 6 shown, the fragments in the document are subjected to regional division processing according to the partition information to obtain the fragments in each region. The connection order of the fragments in each region is sorted according to the fragment position information of the electronic document to obtain the fragments belonging to the same region.
[0079] Step 303: Obtain the fragment information of multiple text fragments and the position information of each of the multiple fragments to obtain layout information.
[0080] It can be understood that the layout information includes the fragment information of multiple text fragments and the position information of each of the multiple fragments. According to the multiple text fragments obtained by dividing the characters in the electronic document, layout information such as character information, character position information, text fragment information, text fragment position information, regional position information, the reading order between multiple regions, the reading order between multiple text fragments, and the reading order of characters within each text fragment is obtained. Step 304: Determine the first reading order between each text fragment in the electronic document according to the layout information.
[0081] As a possible example, according to the layout information, the reading order among multiple regions is marked. According to each text segment and the segment position information in the electronic document, the region to which each text segment belongs is determined, and at least one text segment belonging to the same region is marked according to the reading order among the segments in the above layout information. According to the reading order among multiple regions and the reading order among text segments within the same region, the first reading order is obtained.
[0082] For example, as Figure 6 shown, optionally, the electronic document is partitioned to obtain 8 regions 601. Figure 6 The numbers "1 - 8" in it represent the reading order among the multiple regions 601 in the above layout information. The 8 regions 601 in the electronic text are marked according to the order among the multiple regions 601 in the above layout information to obtain the reading order labels. Each region 601 can contain at least one text segment 602, and the text segments 602 are marked according to the reading order among the text segments in the layout information. For example, the text segments 602 are marked in the order from left to right and from top to bottom within the region 601 to obtain the reading order among the text segments 602 within the same region 601. According to the reading order among the regions 601 and the reading order among the text segments 602 within each region 601, the first reading order among all the text segments 602 in the entire electronic document can be obtained, and thus the reading order among all the characters in the entire electronic document can be obtained.
[0083] Step 305: Input the character information of the electronic document into the pre-trained model according to the layout information to predict the second reading order among the text segments.
[0084] In the embodiments of the present application, step 305 can be implemented in any one of the embodiments of the present application. The embodiments of the present application do not limit this and will not elaborate further.
[0085] Step 306: Based on a preset loss function, calculate the loss value according to the first reading order and the second reading order, and train the pre-trained model based on the loss value.
[0086] In the embodiments of the present application, step 306 can be implemented in any one of the embodiments of the present application. The embodiments of the present application do not limit this and will not elaborate further.
[0087] According to the pre-trained model training method for reading tasks in the embodiments of the present application, the document image of the electronic document is obtained, the characters in the electronic document are divided according to the document image to obtain multiple text segments, the segment information of the multiple text segments and the position information of each of the multiple segments are obtained, so as to obtain the layout information, and further the reading order of the electronic document can be reasonably predicted according to the layout information, improving the accuracy of the reading order.
[0088] Optionally, to ensure that characters can be segmented based on the correct reading order and thus correct text segments can be obtained, the electronic document is parsed to obtain multiple characters of the electronic document and the position information corresponding to each character, the multiple characters are sorted according to the position information corresponding to each character, and the sorted multiple characters are segmented according to the document image to obtain multiple text segments. Figure 7 It is a schematic diagram according to the fourth embodiment of the present application. It should be noted that the pre-training model training method for the reading task in the embodiments of the present application can be executed by the pre-training model training device in the embodiments of the present application. In some embodiments of the present application, as Figure 7 shown, the pre-training model training method for the reading task includes:
[0089] Step 701, obtain the document image of the electronic document.
[0090] In the embodiments of the present application, step 701 can be implemented in any one of the embodiments of the present application, and the embodiments of the present application do not make any limitations thereto and will not be elaborated further.
[0091] Step 702, parse the electronic document to obtain multiple characters of the electronic document and the position information corresponding to each character.
[0092] As an example of a possible implementation manner, the electronic document can be parsed by OCR to obtain multiple characters of the electronic document and the position information corresponding to each character.
[0093] Step 703, sort the multiple characters according to the position information corresponding to each character.
[0094] According to the position information corresponding to each character, each character can be sorted according to its position in the electronic document. Optionally, the above position information can be the coordinate position of the character in the electronic document, and the characters can be sorted according to the coordinate position of each character.
[0095] Step 704, segment the sorted multiple characters according to the document image to obtain multiple text segments.
[0096] It can be understood that, according to the above document image, partition information is obtained, and the sorted characters in the document are subjected to region segmentation processing according to the partition information to obtain the characters in each region, and thus multiple segments composed of characters can be obtained. It can be understood that since the characters are subjected to region segmentation processing, the character order in the above segments is arranged according to the actual required reading order for sorting the segments according to the layout information.
[0097] Step 705: Obtain the segment information of multiple text segments and the position information of each of the multiple segments to obtain layout information.
[0098] In an embodiment of the present application, step 705 can be implemented in any one of the embodiments of the present application, and the embodiments of the present application do not limit this and will not be elaborated further.
[0099] Step 706: Determine the first reading order among the text segments in the electronic document according to the layout information.
[0100] In an embodiment of the present application, step 706 can be implemented in any one of the embodiments of the present application, and the embodiments of the present application do not limit this and will not be elaborated further.
[0101] Step 707: Input the character information of the electronic document into the pre-trained model according to the layout information, and predict the second reading order among the text segments.
[0102] In an embodiment of the present application, step 707 can be implemented in any one of the embodiments of the present application, and the embodiments of the present application do not limit this and will not be elaborated further.
[0103] Step 708: Calculate a loss value based on a preset loss function according to the first reading order and the second reading order, and train the pre-trained model based on the loss value.
[0104] In an embodiment of the present application, step 708 can be implemented in any one of the embodiments of the present application, and the embodiments of the present application do not limit this and will not be elaborated further.
[0105] According to the pre-trained model training method for reading tasks in the embodiments of the present application, the electronic document is parsed to obtain multiple characters of the electronic document and the position information corresponding to each character. The multiple characters are sorted according to the position information corresponding to each character, and the sorted multiple characters are divided according to the document image to obtain multiple text segments, so that the characters can be divided based on the correct reading order, and then the correct text segments can be obtained.
[0106] To implement the above embodiments, the present application proposes a pre-trained model training device for reading tasks.
[0107] Figure 8 is a schematic diagram according to the fifth embodiment of the present application. As Figure 8 shown, the device includes:
[0108] A parsing module 801, configured to perform layout parsing on the electronic document to obtain the layout information of the electronic document.
[0109] A determination module 802, configured to determine a first reading order among various text segments in an electronic document according to layout information.
[0110] A prediction module 803, configured to input character information of the electronic document into a pre-trained model according to layout information, and predict a second reading order among various text segments.
[0111] A training module 804, configured to calculate a loss value based on a first reading order and a second reading order according to a preset loss function, and train the pre-trained model based on the loss value.
[0112] The pre-trained model training apparatus for a reading task according to an embodiment of the present application obtains layout information of an electronic document by performing layout parsing on the electronic document, determines a first reading order among various text segments in the electronic document according to the layout information, inputs the character information of the electronic document into the pre-trained model according to the layout information, the pre-trained model predicts a second reading order among various text segments, calculates a loss value based on the first reading order and the second reading order according to a preset loss function, and trains the pre-trained model based on the loss value, so that the pre-trained model can read the document in the correct order according to actual requirements.
[0113] Figure 9 It is a schematic diagram according to the sixth embodiment of the present application. As Figure 9 shown, the apparatus includes:
[0114] An analysis module 901, configured to perform layout parsing on the electronic document to obtain layout information of the electronic document.
[0115] A determination module 902, configured to determine a first reading order among various text segments in the electronic document according to the layout information.
[0116] A prediction module 903, configured to input character information of the electronic document into the pre-trained model according to the layout information, and predict a second reading order among various text segments;
[0117] Optionally, the pre-trained model includes a conversion module and a self-attention module; the prediction module 903 includes:
[0118] A first obtaining sub-module 9031, configured to input character information of the electronic document into the conversion module according to the layout information to obtain representation information of characters in the electronic document;
[0119] A generation sub-module 9032, configured to generate representation information of each text segment according to the representation information of characters. Optionally, the generation sub-module includes: a third obtaining sub-module, configured to average the representation information of characters belonging to the same text segment to obtain representation information of each text segment;
[0120] The second acquisition sub-module 9033 is configured to input the representation information of each text segment into the self-attention module to obtain the correlation relationship between each text segment;
[0121] The prediction sub-module 9034 is configured to predict the second reading order between each text segment according to the correlation relationship.
[0122] The training module 904 is configured to calculate a loss value based on a preset loss function according to the first reading order and the second reading order, and train the pre-trained model based on the loss value.
[0123] According to the pre-trained model training device for reading tasks in the embodiments of the present application, the character information of the electronic document is input into the conversion module according to the layout information to obtain the representation information of the characters in the electronic document. According to the representation information of the characters, the representation information of each text segment is generated, and the representation information of each text segment is input into the self-attention module to obtain the correlation relationship between each text segment. According to the correlation relationship, the second reading order between each text segment is predicted, so that the reading order of each segment can be determined according to the degree of correlation between each segment, improving the rationality and accuracy of the reading order of the pre-trained model.
[0124] Figure 10 It is a schematic diagram according to the seventh embodiment of the present application. As Figure 10 shown, the device includes:
[0125] The parsing module 1001 is configured to perform layout parsing on the electronic document to obtain the layout information of the electronic document;
[0126] Among them, the parsing module 1001 includes:
[0127] The fourth acquisition sub-module 10011 is configured to acquire the document image of the electronic document;
[0128] The division sub-module 10012 is configured to divide the characters in the electronic document according to the document image to obtain a plurality of text segments; among them, the division sub-module 10012 includes: the parsing sub-module 10014 is configured to parse the electronic document to obtain a plurality of characters in the electronic document and the position information corresponding to each character; the sorting sub-module 10015 is configured to sort the plurality of characters according to the position information corresponding to each character; the sixth acquisition sub-module 10016 is configured to divide the sorted plurality of characters according to the document image to obtain a plurality of text segments.
[0129] The fifth acquisition sub-module 10013 is configured to acquire the segment information of the plurality of text segments and the position information of each of the plurality of segments to obtain the layout information.
[0130] A determination module 1002, configured to determine a first reading order among various text segments in an electronic document according to layout information.
[0131] A prediction module 1003, configured to input character information of the electronic document into a pre-trained model according to layout information, and predict a second reading order among various text segments.
[0132] A training module 1004, configured to calculate a loss value based on a preset loss function according to the first reading order and the second reading order, and train the pre-trained model based on the loss value.
[0133] The pre-trained model training device for a reading task according to an embodiment of the present application obtains a document image of the electronic document, divides the characters in the electronic document according to the document image to obtain a plurality of text segments, obtains segment information of the plurality of text segments and position information of each of the plurality of segments, so as to obtain layout information, and further can reasonably predict the reading order of the electronic document according to the layout information. In addition, the electronic document is parsed to obtain a plurality of characters of the electronic document and position information corresponding to each character, the plurality of characters are sorted according to the position information corresponding to each character, and the sorted plurality of characters are divided according to the document image to obtain a plurality of text segments, so that the characters can be divided based on the correct reading order, and further the correct text segments can be obtained.
[0134] According to an embodiment of the present application, the present application further provides an electronic device and a readable storage medium.
[0135] As Figure 11 shown, it is a block diagram of an electronic device for a method of training a pre-trained model for a reading task according to an embodiment of the present application. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the present application described and / or claimed herein.
[0136] As Figure 11As shown, the electronic device includes: one or more processors 1101, a memory 1102, and interfaces for connecting the components, including a high-speed interface and a low-speed interface. Each component is interconnected using different buses and can be mounted on a common motherboard or otherwise installed as required. The processor can process instructions executed within the electronic device, including instructions stored in the memory or on the memory for displaying graphical information of a GUI on an external input / output device (such as a display device coupled to the interface). In other embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple electronic devices can be connected, with each device providing part of the necessary operations (such as an array of servers, a set of blade servers, or a multi-processor system). Figure 11 In [example], a single processor 1101 is taken as an example.
[0137] The memory 1102 is the non-transitory computer-readable storage medium provided by the present application. Among them, the memory stores instructions executable by at least one processor, so that the at least one processor executes the method for training a pre-trained model for a reading task provided by the present application. The non-transitory computer-readable storage medium of the present application stores computer instructions, and the computer instructions are used to cause a computer to execute the method for training a pre-trained model for a reading task provided by the present application.
[0138] As a non-transitory computer-readable storage medium, the memory 1102 can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules, such as the program instructions / modules corresponding to the method for training a pre-trained model for a reading task in the embodiments of the present application (for example, Figure 8 the parsing module 801, the determination module 802, the prediction module 803, and the training module 804 shown). By running the non-transitory software programs, instructions, and modules stored in the memory 1102, the processor 1101 executes various functional applications and data processing of the server, that is, implements the method for training a pre-trained model for a reading task in the above method embodiments.
[0139] The memory 1102 may include a program storage area and a data storage area. The program storage area may store an operating system and application programs required for at least one function. The data storage area may store data created according to the use of the electronic device trained with a pre-trained model for reading tasks, etc. In addition, the memory 1102 may include a high-speed random access memory and may also include a non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory 1102 may optionally include a memory remotely provided with respect to the processor 1101, and these remote memories may be connected to the electronic device trained with a pre-trained model for reading tasks through a network. Examples of the above network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0140] The electronic device for training a pre-trained model for reading tasks may further include: an input device 1103 and an output device 1104. The processor 1101, the memory 1102, the input device 1103, and the output device 1104 may be connected through a bus or other means. Figure 11 Taking the connection through the bus as an example.
[0141] The input device 1103 may receive input digital or character information and generate key signal inputs related to the user settings and function controls of the electronic device trained with a pre-trained model for reading tasks, such as input devices like a touch screen, a keypad, a mouse, a trackpad, a touchpad, a pointing stick, one or more mouse buttons, a trackball, a joystick, etc. The output device 1104 may include a display device, an auxiliary lighting device (e.g., an LED), and a haptic feedback device (e.g., a vibration motor), etc. The display device may include but is not limited to a liquid crystal display (LCD), a light-emitting diode (LED) display, and a plasma display. In some embodiments, the display device may be a touch screen.
[0142] Various embodiments of the systems and techniques described herein may be implemented in digital electronic circuit systems, integrated circuit systems, dedicated ASICs (application-specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include: implemented in one or more computer programs that may be executed and / or interpreted on a programmable system including at least one programmable processor, where the programmable processor may be a dedicated or general-purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0143] These computing procedures (also known as programs, software, software applications, or code) include machine instructions for a programmable processor and can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, apparatus, and / or device (e.g., a magnetic disk, an optical disk, a memory, a programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.
[0144] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).
[0145] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0146] A computer system may include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The relationship between the client and the server is generated by computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services ("Virtual Private Server", or simply "VPS").
[0147] It should be understood that various forms of processes shown above can be used, steps can be reordered, added or deleted. For example, the steps described in the present application can be executed in parallel, sequentially or in a different order, as long as the desired results of the technical solutions disclosed in the present application can be achieved, and no limitations are imposed herein.
[0148] The above specific embodiments do not constitute a limitation to the protection scope of the present application. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present application shall be included within the protection scope of the present application.
Claims
1. A pre-training model training method for reading tasks, comprising: Performing layout analysis on an electronic document to obtain layout information of the electronic document; Determining a first reading order among respective text segments in the electronic document according to the layout information; Inputting character information of the electronic document into a pre-training model according to the layout information to predict a second reading order among the respective text segments; Calculating a loss value based on a preset loss function according to the first reading order and the second reading order, and training the pre-training model based on the loss value; Wherein, the pre-training model includes a conversion module and a self-attention module; the inputting character information of the electronic document into the pre-training model according to the layout information to predict a second reading order among the respective text segments includes: Inputting character information of the electronic document into the conversion module according to the layout information to obtain representation information of characters in the electronic document; Generating respective representation information for the respective text segments according to the representation information of characters; Inputting the respective representation information of the respective text segments into the self-attention module to obtain an association relationship among the respective text segments, wherein the attention module calculates a weight of each text segment relative to other segments according to the vector representation information of the respective text segments, so as to obtain the association relationship among the respective text segments; Predicting a second reading order among the respective text segments according to the association relationship.
2. The method according to claim 1, wherein, The generating respective representation information for the respective text segments according to the representation information of characters includes: Calculating an average of the representation information of characters belonging to the same text segment to obtain respective representation information for the respective text segments.
3. The method according to claim 1, wherein The performing layout analysis on an electronic document to obtain layout information of the electronic document includes: Obtaining a document image of the electronic document; Dividing characters in the electronic document according to the document image to obtain a plurality of text segments; Obtaining segment information of the plurality of text segments and position information of the respective segments to obtain the layout information.
4. The method according to claim 3, wherein, The dividing characters in the electronic document according to the document image to obtain a plurality of text segments includes: Parsing the electronic document to obtain a plurality of characters in the electronic document and position information corresponding to each character; Sorting the plurality of characters according to the position information corresponding to each character; Dividing the sorted plurality of characters according to the document image to obtain a plurality of text segments.
5. A pre-training model training device for reading tasks, comprising: An analysis module, configured to perform layout analysis on an electronic document to obtain layout information of the electronic document; A determination module, configured to determine a first reading order among respective text segments in the electronic document according to the layout information; A prediction module, configured to input character information of the electronic document into a pre-training model according to the layout information to predict a second reading order among the respective text segments; A training module, configured to calculate a loss value based on a preset loss function according to the first reading order and the second reading order, and train the pre-trained model based on the loss value; Wherein, the pre-trained model includes a conversion module and a self-attention module; the prediction module includes: A first obtaining sub-module, configured to input the character information of the electronic document into the conversion module according to the layout information to obtain the representation information of the characters in the electronic document; A generating sub-module, configured to generate the representation information of each text segment according to the representation information of the characters; A second obtaining sub-module, configured to input the representation information of each text segment into the self-attention module to obtain the correlation relationship between the text segments, wherein the attention module calculates the weight of each text segment relative to other segments according to the vector representation information of each text segment, so as to obtain the correlation relationship between the text segments; A prediction sub-module, configured to predict the second reading order between the text segments according to the correlation relationship.
6. The device according to claim 5, wherein The generating sub-module includes: A third obtaining sub-module, configured to average the representation information of the characters belonging to the same text segment to obtain the representation information of each text segment.
7. The device according to claim 5, wherein The parsing module includes: A fourth obtaining sub-module, configured to obtain the document image of the electronic document; A dividing sub-module, configured to divide the characters in the electronic document according to the document image to obtain a plurality of text segments; A fifth obtaining sub-module, configured to obtain the segment information of the plurality of text segments and the position information of each of the plurality of segments to obtain the layout information.
8. The device according to claim 7, wherein, The dividing sub-module includes: A parsing sub-module, configured to parse the electronic document to obtain a plurality of characters in the electronic document and the position information corresponding to each character; A sorting sub-module, configured to sort the plurality of characters according to the position information corresponding to each character; A sixth obtaining sub-module, configured to divide the sorted plurality of characters according to the document image to obtain a plurality of text segments.
9. An electronic device, characterized in that, Includes: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1-4.
10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-4.
Citation Information
Patent Citations
Method and apparatus for detecting reading order of document
CN108334805A