A model training, text sorting method and device
Through deep learning model training and text contrast analysis methods, the problem of text order is solved, the sorting efficiency and accuracy are improved, and automated text sorting is realized.
Patent Information
- Application Number
- CN202210106069.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-28
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2042-01-28
AI Technical Summary
The prior art is inefficient and has low accuracy when processing text order incorrectly, and requires manual identification and sorting.
Using a deep learning-based model training method, multiple text pairs are constructed by obtaining sample files and text sequence numbers, and the confidence of text pairs is output using the pre-trained model, and the model parameters are adjusted to determine the correct order of text.
It improves text processing efficiency and the accuracy of sorting results, reduces manual intervention, and realizes automated text sorting.
Smart Images

Figure CN114510920B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, in particular to the fields of deep learning technology and document analysis technology, and specifically relates to a model training method, a text sorting method and an apparatus. Background Art
[0002] In related applications, the text order is often disordered. Currently, the common solution is to manually identify the text order and sort it correctly. This method is inefficient and has low accuracy. Summary of the Invention
[0003] The present disclosure provides a model training method, a text sorting method, an apparatus, a device and a storage medium.
[0004] In a first aspect, the present disclosure provides a model training method, including:
[0005] Obtaining a sample file, a start text and an end text, where the sample file includes at least two sample texts and the serial number of each sample text;
[0006] Constructing a plurality of text pairs, each text pair being constructed by two of the following texts: the start text, the end text, and the at least two sample texts;
[0007] Inputting each of the text pairs into a model, and outputting, by the model, the confidence that the two texts constructing each text pair have a context relationship;
[0008] For each text pair, comparing the confidence with the actual relationship between the two texts constructing the text pair, and adjusting the parameters of the model according to the comparison result.
[0009] In a second aspect, the present disclosure provides a sorting method, including:
[0010] Obtaining a start text, an end text and at least two texts to be sorted;
[0011] Constructing a plurality of first text pairs, each first text pair being constructed by two of the following texts: the start text, the end text, and the at least two texts to be sorted;
[0012] Inputting the plurality of first text pairs into a pre-trained model, and outputting, by the pre-trained model, the confidence that the two texts constructing each first text pair have a context relationship;
[0013] Determining the order of the at least two texts to be sorted according to the confidence.
[0014] In a third aspect, the present disclosure provides a model training apparatus, including:
[0015] A first acquisition module, configured to acquire a sample file, a start text, and an end text, where the sample file includes at least two sample texts and serial numbers of each of the sample texts;
[0016] A first construction module, configured to construct a plurality of text pairs, each text pair being constructed from two of the following texts: the start text, the end text, and the at least two sample texts;
[0017] A first input module, configured to respectively input each of the text pairs into a model, and the model outputs a confidence level that the two texts constructing each of the text pairs have a context relationship;
[0018] An adjustment module, configured to compare the confidence level with the actual relationship between the two texts constructing the text pair, and adjust the parameters of the model according to the comparison result.
[0019] In a fourth aspect, the present disclosure provides a sorting device, including:
[0020] A second acquisition module, configured to acquire a start text, an end text, and at least two texts to be sorted;
[0021] A second construction module, configured to construct a plurality of first text pairs, each first text pair being constructed from two of the following texts: the start text, the end text, and the at least two texts to be sorted;
[0022] A second input module, configured to respectively input the plurality of first text pairs into a pre-trained model, and the pre-trained model outputs a confidence level that the two texts constructing each of the first text pairs have a context relationship;
[0023] A second determination module, configured to determine the order of the at least two texts to be sorted according to the confidence level.
[0024] In a fifth aspect, the present disclosure provides an electronic device, including:
[0025] At least one processor; and
[0026] A memory communicatively connected to the at least one processor; wherein,
[0027] The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the methods provided in the first aspect or the second aspect above.
[0028] In a sixth aspect, the present disclosure provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the methods provided in the first aspect or the second aspect above.
[0029] In a seventh aspect, the present disclosure provides a computer program product including a computer program which, when executed by a processor, implements the method provided in the first aspect or the second aspect above.
[0030] The model training method provided by the present disclosure can reorder text with disordered sequence, and can improve the text processing efficiency and the accuracy of the sorting result.
[0031] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0033] Figure 1 is a schematic flowchart of a model training method according to an embodiment of the present disclosure;
[0034] Figure 2 is a schematic diagram of a sorting solution for disordered PDF scans according to an embodiment of the present disclosure;
[0035] Figure 3 is a schematic flowchart of data preprocessing in a model training method according to an embodiment of the present disclosure;
[0036] Figure 4 is a schematic diagram of an application scenario of a text sorting method according to an embodiment of the present disclosure;
[0037] Figure 5 is a flowchart of the implementation of a text sorting method according to an embodiment of the present disclosure;
[0038] Figure 6 is a schematic diagram of the input and output content of a model in a text sorting method according to an embodiment of the present disclosure;
[0039] Figure 7 is a schematic diagram of the structure of a model training apparatus 700 according to an embodiment of the present disclosure;
[0040] Figure 8 is a schematic diagram of the structure of a model training apparatus 800 according to an embodiment of the present disclosure;
[0041] Figure 9 is a schematic diagram of the structure of a sorting apparatus 900 according to an embodiment of the present disclosure;
[0042] Figure 10 is a schematic diagram of the structure of a sorting apparatus 1000 according to an embodiment of the present disclosure;
[0043] Figure 11 It is a structural block diagram of an electronic device 1100 according to an embodiment of the present disclosure. Detailed implementation manners
[0044] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, descriptions of well-known functions and structures are omitted below for clarity and conciseness.
[0045] In related applications, the text order often gets scrambled. For example, in the scenario of scanning paper materials into a Portable Document Format (PDF) document for storage, sometimes due to too much data and operation errors, the pages of the saved PDF document are in the wrong order, that is, there are multiple scanned pictures in a PDF, but they are not sorted in the reading order. The current common solution is to manually identify the text order and sort it correctly. This method is inefficient and has a low accuracy rate. The present disclosure proposes a solution that uses Artificial Intelligence (AI) technology to reorder scrambled text, which can improve the processing efficiency and the accuracy rate of the sorting result.
[0046] The present disclosure proposes a model training method. Figure 1 It is a schematic flowchart of the model training method according to an embodiment of the present disclosure, including:
[0047] S101: Obtain a sample file, a start text, and an end text. The sample file includes at least two sample texts and the serial number of each sample text;
[0048] S102: Construct multiple text pairs, and each text pair is constructed by two of the following texts: the start text, the end text, and at least two sample texts;
[0049] S103: Input each text pair into the model respectively, and the model outputs the confidence that the two texts constructing each text pair have a context relationship;
[0050] S104: For each text pair, compare the confidence with the actual relationship between the two texts constructing the text pair, and adjust the parameters of the model according to the comparison result.
[0051] In a possible implementation, the above sample file may include at least two pages, and each page is a sample text; the serial number of the sample text can be manually marked, and this serial number represents the actual position of the sample text in the sample file. For example, the sample file includes n pages (i.e., n sample texts, where n is a positive integer); assuming that the serial numbers from 1 to n represent the order of the sample file from front to back, then the serial numbers of each sample text can be manually marked according to the actual position, such as marked as p1, p2,... pn respectively, where p1, p2,... pn are all within the range of [1, n]. The embodiments of the present disclosure can receive the serial numbers of the sample texts manually marked.
[0052] In a possible implementation, the information of the sample text can be carried in the form of a picture. Correspondingly, obtaining the sample file may include:
[0053] Obtaining pictures of each sample text included in the sample file;
[0054] Performing optical character recognition (OCR) on each picture respectively to obtain each sample text, and obtaining the serial number of each sample text.
[0055] Taking a PDF scanned copy as an example, performing PDF scanning on a paper material containing multiple pages, and the obtained PDF file is the sample file; each page in the PDF file can be considered as a picture of the sample text. Performing OCR on each page in the PDF file respectively can obtain each sample text; each sample text may include multiple characters, and the characters included in the sample text represent the content of that page. If a certain page in the PDF scanned copy includes an image and / or a table, the embodiments of the present disclosure can also recognize the characters, symbols, etc. in the image and / or the table by using the OCR technology, so as to obtain the corresponding sample text. By adopting this method, the embodiments of the present disclosure can conveniently obtain the sample text and the serial numbers of each sample text, and can use a large number of PDF files to generate the sample file, improving the richness of the sample file, thereby improving the model training efficiency.
[0056] The model involved in the embodiments of the present disclosure may include a Bidirectional Encoder Representation from Transformers (BERT) model or an Enhanced Representation from kNowledge IntEgration (ERNIE) model. In the following embodiments, the BERT model will be taken as an example for detailed introduction.
[0057] Figure 2Schematic diagram of the out-of-order PDF scan sorting solution according to an embodiment of the present disclosure Figure 2 Involves the following processing procedures:
[0058] First, data preprocessing (such as mapping):
[0059] The multi-page PDF scan is split, and each page undergoes general text OCR processing to obtain multiple texts. These texts are used both in model training and in the subsequent sorting process using the trained model. For ease of distinction, in the embodiments of this application, the text involved in the model training process is called training text, and the text involved in the sorting process using the model is called text to be sorted.
[0060] Second, model training:
[0061] The embodiments of this application propose an automated modeling solution, which involves BERT binary classification model training, parameter tuning, and model release functions. It can be trained using the sample texts obtained from the above-mentioned first part of data preprocessing and the serial numbers of each sample text.
[0062] Third, model inference (Reduce):
[0063] When using the trained model for inference, the results output by the model are aggregated. The data input to this model is the text to be sorted obtained after the data preprocessing process.
[0064] Fourth, PDF sorting process
[0065] Using the results obtained after the above-mentioned third part of model inference, the multi-pages of the PDF scan are sorted.
[0066] The model training method proposed in the embodiments of this application involves the above-mentioned first and second processing procedures. The embodiments of this application also propose a method for sorting using the model, and this sorting method involves the above-mentioned first to fourth processing procedures. The model training method and sorting method proposed in the embodiments of this application are introduced below, and the relevant processing procedures mentioned above are involved in the training method and sorting method.
[0067] Figure 3 Schematic diagram of the data preprocessing process in the model training method according to an embodiment of the present disclosure Figure 3 Taking the data preprocessing for the PDF scan as an example, first, the PDF scan is split, and each page of the scan is an image of a text. Then, each text image undergoes OCR conversion to obtain multiple texts; the preprocessed text is used for model training, so this "text" can be called "sample text".
[0068] In some embodiments, the embodiments of the present application may send a general OCR service request to an OCR service platform, and the OCR service platform performs OCR conversion on the picture to complete the conversion of picture information into text information; or OCR conversion may also be performed locally. The embodiments of the present application do not limit the entity that implements OCR conversion.
[0069] Taking Figure 3 as an example, after splitting a PDF scan, n pictures are obtained, which are respectively the Figure 3 1st page, 2nd page,..., nth page pictures in Figure 3 ; then OCR is performed on each picture respectively, and n corresponding texts are obtained, which are respectively the Figure 3 text 1, text 2,..., text n in Since the pictures in the PDF scan may be out of order, text 1, text 2,..., text n only represent the identifiers of each text, and do not represent the order of the texts in the PDF file. The embodiments of the present application may further receive the serial numbers of each text manually marked. For example, the 1st serial number, the 2nd serial number,..., respectively represent each text from front to back; among them, the text with the 1st serial number corresponds to the 1st page of the PDF scan sorted correctly according to the content, and the text with the 2nd serial number corresponds to the 1st page of the PDF scan sorted correctly according to the content, and so on.
[0070] The embodiments of the present application may also add a start text and an end text (such as the SE sequence). The start text or the end text may be a text including multiple characters. The embodiments of the present application do not limit the character content included in the start text and the end text. For example, the embodiments of the present application may construct some tokens as the start text (such as the Figure 3 "Start (beginning) sequence" in Figure 3 ), and construct some tokens as the end text (such as the Figure 3 "End (ending) sequence" in Figure 3 ). Taking Figure 3 as an example, the total number of texts is n + 2; among them, after splitting and performing OCR on the PDF scan, n texts are obtained, plus the start text and the end text, and the total number of texts for the sequence model is n + 2 in total.
[0071] The model used for sorting in the embodiments of the present application may select language models such as the BERT model and the ERNIE model. Taking the BERT model as an example, the BERT model can understand the relationship between sentences and implement the downstream sentence relationship judgment task. The Next Sentence Prediction in the BERT model is for this purpose. Its principle is to pre-train a binary next sentence prediction task. Specifically, sentences A and B are selected as training samples, and the model needs to correctly predict whether the next sentence of A is B. The input of the BERT model is:
[0072] [CLS]Previous text[SEP]Next text;
[0073] Among them, [CLS] is the token of the first sentence, used to represent the first sentence; [SEP] is the token between two sentences, used to separate two sentences.
[0074] The previous text and the next text can come from different texts, and the purpose is to determine whether there is a context relationship between any two pages, that is, whether any two pages are adjacent texts.
[0075] In some possible implementation manners, a plurality of text pairs are constructed by using a start text, an end text, and at least two sample texts, including:
[0076] Construct X text pairs by using the start text and each of the sample texts respectively, where X is the number of sample texts included in the sample file;
[0077] Construct text pairs by using each of the sample texts and each of the other sample texts or the end text respectively.
[0078] Take Figure 3 as an example, use text 1, text 2,..., text n, Start sequence, and End sequence to construct a plurality of text pairs, including:
[0079] Use the content in the Start sequence as the previous text, and use the content in text 1, text 2,..., text n as the subsequent text respectively to construct n text pairs, and the n text pairs respectively correspond to [Start sequence, text 1], [Start sequence, text 2],..., [Start sequence, text n];
[0080] Use the content in text 1 as the previous text, and use the content in text 2,..., text n, End sequence as the subsequent text respectively to construct n text pairs, and the n text pairs respectively correspond to [text 1, text 2], [text 1, text 3],..., [text 1, text n], [text 1, End sequence];
[0081] In the same way, use the content in text 2,..., text n as the previous text respectively, and use the content in other texts, End sequence as the subsequent text to construct other text pairs.
[0082] Adopt the above method, in Figure 3 the example shown, a total of n×(n + 1) text pairs are constructed.
[0083] The present disclosure can automatically construct text pairs for training a model in the above manner, thereby avoiding manually constructing text pairs, reducing manual workload, and improving the model training efficiency.
[0084] In addition, there may be a length limit for the input data of the BERT model. To meet this length limit, in the model training method proposed in the embodiments of the present application, the method for constructing text pairs may include:
[0085] When the number of characters in the prior text used to construct the text pair is greater than M, the last M characters of the prior text are intercepted, and the intercepted content is used as the prior content for constructing the text pair; when the number of characters in the prior text used to construct the text pair is less than or equal to M, all the content of the prior text is used as the prior content for constructing the text pair; where M is an integer;
[0086] When the number of characters in the subsequent text used to construct the text pair is greater than N, the first N characters of the subsequent text are intercepted, and the intercepted content is used as the subsequent content for constructing the text pair; when the number of characters in the subsequent text used to construct the text pair is less than or equal to N, all the content of the subsequent text is used as the subsequent content for constructing the text pair; where N is an integer;
[0087] The above prior content and subsequent content are combined in sequence to obtain a text pair.
[0088] In some possible implementation manners, the specific values of M and N above can be determined according to the length of the input content of the model.
[0089] Wherein, a text pair is constructed by two texts (which can be the start text, end text or sample text). There is a sequence between the two texts. The text with the prior sequence can be called the prior text for constructing the text pair, and the text with the subsequent sequence can be called the subsequent text for constructing the text pair. Taking the construction of a text pair using [Text 1, Text 3] as an example, there is a sequence between the two texts used to construct the text pair. The previous text can be called the prior text for constructing the text pair. For example, Text 1 is the prior text for constructing the text pair "[Text 1, Text 3]"; the subsequent text can be called the subsequent text for constructing the text pair. For example, Text 3 is the subsequent text for constructing the text pair "[Text 1, Text 3]". Taking the length limit of the text pair as not exceeding 512 characters as an example, the last 255 characters can be intercepted from the prior text (if the number of characters in the prior text is less than 255, all the characters in the prior text are taken), and the first 254 characters can be intercepted from the subsequent text (if the number of characters in the subsequent text is less than 254, all the characters in the subsequent text are taken), and the two parts of content obtained from the prior text and the subsequent text are combined in sequence to obtain the content of the text pair, and the content of this text pair can meet the length requirement of the input content of the model. In some implementation manners, specific Tokens specified in the input content of the BERT model, such as the above "[CLS]" and "[SEP]", are also added when composing the content of the text pair.
[0090] To implement the training of the model, the training method proposed in the embodiments of this application further includes: determining the actual relationship between the two texts for constructing each text pair according to the serial numbers of each sample text.
[0091] For example, the serial numbers of each sample text are the 1st serial number to the Xth serial number, which are used to represent the order of each sample text in the sample file, such as the actual order after sorting according to the file content;
[0092] Determining the actual relationship between the two texts for constructing each text pair (this actual relationship may refer to whether there is a context relationship between the two texts for constructing this text pair) includes at least one of the following:
[0093] For the text pair constructed by the starting text and the sample text with the 1st serial number, determine that the two texts for constructing this text pair have a context relationship;
[0094] For the text pair constructed by the starting text and the sample text without the 1st serial number, determine that the two texts for constructing this text pair do not have a context relationship;
[0095] For the text pair constructed by the sample texts with two adjacent serial numbers before and after respectively, determine that the two texts for constructing this text pair have a context relationship;
[0096] For the text pair constructed by the sample texts without two adjacent serial numbers before and after, determine that the two texts for constructing this text pair do not have a context relationship;
[0097] For the text pair constructed by the sample text with the Xth serial number and the ending text, determine that the two texts for constructing this text pair have a context relationship;
[0098] For the text pair constructed by the sample text without the Xth serial number and the ending text, determine that the two texts for constructing this text pair do not have a context relationship.
[0099] In some possible implementation manners, if it is determined that the two texts for constructing a certain text pair have a context relationship, the label of this text pair can be set to 1; if it is determined that the two texts for constructing a certain text pair do not have a context relationship, the label of this text pair can be set to 0. The foregoing labels are only examples, and the embodiments of the present disclosure do not limit the specific content of the labels.
[0100] For example, if the actual [Text 3] is the first page of a PDF scan and [Text 1] is the last page of the PDF scan, then during annotation:
[0101] The label of the text pair constructed by the text pair "[CLS]Start sequence[SEP]Text 3" can be set to 1, indicating that there is a context relationship in the content of this text pair; the label of the text pair constructed by the Start sequence and other texts is set to 0, indicating that there is no context relationship;
[0102] The label of the text pair constructed by the text pair "[CLS]Text 1[SEP]End sequence" can be set to 1, indicating that there is a context relationship in the content of this text pair; the label of the text pair constructed by other texts and the End sequence is set to 0, indicating that there is no context relationship.
[0103] In the embodiments of the present disclosure, the actual relationship between the two texts that make up the text pair is automatically determined in the above manner, and the label of the text pair is determined according to the actual relationship, which can avoid manually annotating the labels of each text pair, thereby reducing the manual workload and improving the model training efficiency.
[0104] The constructed text pair is used as the input content of the BERT model, and the BERT model predicts whether there is a context relationship in the content of the text pair. For example, the confidence level that there is a context relationship in the content of the text pair is output. When the confidence level is greater than or equal to the first threshold, it can be considered that the BERT model predicts that there is a context relationship in the content of this text pair; when the confidence level is less than or equal to the second threshold, it can be considered that the BERT model predicts that there is no context relationship in the content of this text pair. Among them, the aforementioned first threshold and second threshold can be the same or different. By comparing the confidence level output by the BERT model and the actual relationship between the two texts that make up the text pair (represented by the label of the above text pair), it can be judged whether the prediction result of the BERT model is accurate, so as to adjust the parameters of the BERT model to achieve the training of the BERT model. In some possible implementation manners, when the accuracy rate of the prediction result of the BERT model is greater than the preset threshold, it is considered that the training of the BERT model is completed, and the training of the BERT model ends.
[0105] Due to the existence of the constructed [Start sequence] and [End sequence], the model supports parallel processing during inference: that is, first, the probability that each text page corresponds to other text pages as a prior text page is obtained through parallel inference. After obtaining all the probability results, they are sorted. Instead of calculating the probability that the first text page corresponds to other text pages, finding the text page with the highest probability, and then iterating in sequence for a serial inference and sorting.
[0106] Different from the conventional modeling of upper and lower sentence tasks, in the embodiments of the present disclosure, when training the model, by constructing a starting text (such as a Start sequence) and an ending text (such as an End sequence), the start page and the end page in the text can be uniquely determined according to the probability during inference. For example, the text having a context relationship with the starting text is the start page, and the text having a context relationship with the ending text is the end page; moreover, in the embodiments of the present disclosure, only a small amount of information needs to be annotated to automatically generate sample pairs for inputting into the model to train the model (that is, only the actual serial number of each text in a document needs to be annotated, without annotating whether any text pair has a context relationship), thereby improving the training efficiency of the model and the accuracy of sorting using the model. The present disclosure splits multi-page PDF scans by page dimension, performs OCR conversion on each page to the corresponding text, constructs context samples for different pages, supports parallel training, and greatly enhances the training performance.
[0107] Using the model trained by the above training method, the present disclosure proposes a sorting method. Figure 4 FIG. is a schematic diagram of an application scenario of the sorting method according to the embodiments of the present disclosure. As Figure 4 shown, a scanner scans a paper document with multiple pages to obtain an initial PDF file, and the order of each page in the initial PDF file is chaotic. The sorting device obtains the initial PDF file, sends an OCR conversion request to the OCR platform, and the OCR platform performs OCR conversion on each page in the initial PDF file to obtain the corresponding text for each page. The sorting device constructs text pairs using the corresponding text for each page, the starting text, and the ending text, inputs the constructed text pairs into a pre-trained model, and the model outputs the confidence that the texts constructing each text pair have a context relationship; using this confidence, the sorting device determines the start page, the end page, and the texts having a context relationship in these texts, and finally determines the correct order, rearranges the texts in the correct order, and obtains a sorted PDF file.
[0108] Figure 5 FIG. is a flowchart for implementing the sorting method according to the embodiments of the present disclosure, including:
[0109] S501: Obtain a starting text, an ending text, and at least two texts to be sorted;
[0110] S502: Construct a plurality of first text pairs, and each first text pair is constructed by two of the following texts: the starting text, the ending text, and at least two texts to be sorted;
[0111] S503: Input the plurality of first text pairs into a pre-trained model respectively, and the pre-trained model outputs the confidence that the two texts constructing each first text pair have a context relationship;
[0112] S504: Determine the order of the at least two texts to be sorted according to the confidence level.
[0113] In some possible implementation manners, obtaining at least two texts to be sorted may include:
[0114] Obtain the pictures respectively corresponding to the at least two texts to be sorted;
[0115] Perform optical character recognition on each picture respectively to obtain at least two texts to be sorted.
[0116] Taking a PDF scanned copy as an example, perform PDF scanning on a paper material containing multiple pages to obtain a PDF file, and each page in the PDF file is a picture of the text to be sorted; each sample text may include multiple characters, and the characters included in the sample text represent the content of that page. If a certain page in the PDF scanned copy includes an image and / or a table, the OCR technology adopted in the embodiments of the present disclosure can also recognize the characters, symbols, etc. in the image and / or the table, so as to obtain the corresponding text to be sorted.
[0117] The model adopted in the embodiments of the present disclosure can be pre-trained by using the above model training method, and the specific training method will not be elaborated here.
[0118] In some possible implementation manners, constructing multiple text pairs by using the start text, the end text, and the at least two texts to be sorted includes:
[0119] Construct Y first text pairs by using the start text and each of the texts to be sorted respectively, where Y is the number of texts to be sorted;
[0120] Construct first text pairs by using each of the texts to be sorted and each of the other texts to be sorted or the end text respectively.
[0121] In this embodiment, the manner of constructing text pairs adopts the same manner as that in the above model training method.
[0122] Figure 6 It is a schematic diagram of the input and output content of the model in the sorting method according to an embodiment of the present disclosure. As Figure 6 shown, after splitting and OCR processing the original PDF scanned copy, texts 1, 2,..., n are obtained. Since the pictures in the PDF scanned copy may be out of order, texts 1, 2,..., n only represent the identifiers of each text, and do not represent the order of the texts in the PDF file. Then add the start text (such as Figure 6 the Start sequence in Figure 6 and the end text (such as
[0123] Construct text pairs using the above n + 2 texts. For example,
[0124] Take the content in the Start sequence as the prior text, and take the content in Text 1, Text 2, …, Text n as the posterior text respectively to construct n text pairs, which are [Start sequence, Text 1], [Start sequence, Text 2], …, [Start sequence, Text n];
[0125] Take the content in Text 1 as the prior text, and take the content in Text 2, …, Text n, End sequence as the posterior text respectively to construct n text pairs, which are [Text 1, Text 2], [Text 1, Text 3], …, [Text 1, Text n], [Text 1, End sequence];
[0126] In the same way, take the content in Text 2, …, Text n as the prior text respectively, and take the content in other texts, End sequence as the posterior text to construct other text pairs.
[0127] Using the above method, a total of n×(n + 1) text pairs are constructed.
[0128] When constructing text pairs, it is necessary to consider the length limit of the model input data, and intercept the characters not exceeding the length limit to construct text pairs.
[0129] After that, input the constructed multiple text pairs into a pre-trained model (such as the BERT model) respectively, and the pre-trained model outputs the confidence levels that the two texts constructing each text pair have a context relationship. In Figure 6 In the shown example, the value range of the confidence level is [0, 1]. The higher the confidence level, the greater the possibility that the two texts have a context relationship. For example, Figure 6As shown, the confidence levels of the context relationships between the two texts involved in n text pairs with the Start sequence as the prior text in the model output (including [Start sequence, Text 1], [Start sequence, Text 2], … [Start sequence, Text n]) are 0.04, 0.98, 0.12, …, 0.01 respectively. Similarly, the confidence levels of the context relationships between the two texts involved in n text pairs with Text 1 as the prior text in the model output (including [Text 1, Text 2], [Text 1, Text 3], … [Text 1, End sequence]) are 0.15, 0.04, …, 0.99, 0.02 respectively. Also, the confidence levels of the context relationships between the two texts involved in n text pairs with Text n as the prior text in the model output (including [Text n, Text 1], [Text n, Text 2], … [Text n, End sequence]) are 0.08, 0.03, …, 0.91, 0.01 respectively.
[0130] In some possible implementation manners, according to the above confidence levels, determining the order of the at least two texts to be sorted includes:
[0131] For each text, determining each first text pair with this text as the prior text; obtaining the confidence levels of the context relationships between the two texts that form each of the determined first text pairs, and determining the two texts with the highest confidence level as having a context relationship;
[0132] According to multiple pairs of texts with context relationships, determining the order of the at least two texts to be sorted.
[0133] Taking Figure 6 as an example, for the starting text (i.e., the Start sequence), obtaining the confidence levels corresponding to each text pair with the Start sequence as the prior text, which are 0.04, 0.98, 0.12, …, 0.01 respectively; the text pair with the highest confidence level is [Start sequence, Text 2], and the two texts that form this text pair have a context relationship. Similarly, any two texts with a context relationship can also be found, such as Figure 6 [Text 1, End sequence], [Text n, Text n - 1] in the example.
[0134] According to the above text pairs with context relationships, the order of the texts to be sorted can be determined. Taking Figure 6Taking an example, since it has been determined that [Start sequence, Text 2] has a contextual relationship, it can be determined that Text 2 is the start page; then, find the text pairs with a contextual relationship that have Text 2 as the prior text, and the second page can be determined; and so on, until a text pair with a contextual relationship that has the termination text (i.e., End sequence) as the posterior text is found, and the end page can be determined, thus obtaining the correct order of each page of the complete PDF scan. Sort and restore the PDF scan according to this order. For example, the pictures of the corresponding pages can be obtained in order, and then use the fitz package of Python to organize them into a sorted PDF file.
[0135] In some embodiments, the above confidence levels may exist in the form of an array. For example, among multiple text pairs with the same text as the prior text, the confidence levels of the two texts that make up the text pair having a contextual relationship can form an array. For example,
[0136] For the text pairs constructed from [Start sequence] and all other page texts, the confidence levels of the two texts that make up these text pairs having a contextual relationship form an array as:
[0137] [0.04, 0.98, 0.12,..., 0.01]
[0138] For the text pairs constructed from [Text 1] and all other page texts, the confidence levels of the two texts that make up these text pairs having a contextual relationship form an array as:
[0139] [0.15, 0.04,..., 0.99, 0.02] ...
[0140] For the text pairs constructed from [Text n] and all other page texts, the confidence levels of the two texts that make up these text pairs having a contextual relationship form an array as:
[0141] [0.08, 0.03,..., 0.91, 0.01]
[0142] Correspondingly, when determining text pairs with a contextual relationship, the confidence level with the largest value in each array can be determined respectively, and the text pair corresponding to this confidence level can be determined to be a text pair with a contextual relationship.
[0143] As can be seen from the above process, due to the existence of the [Start sequence] and [End sequence] constructed in the embodiments of the present disclosure, the model supports parallelization during inference: that is, first, parallel inference is performed to obtain the probabilities of each text page corresponding to other text pages, and after all probability results are obtained, sorting is performed. Instead of calculating the probability of the first text page corresponding to other text pages, finding the text page with the highest probability, and then iterating sequentially for serial inference and sorting. This parallel inference method of the present disclosure greatly improves the sorting efficiency.
[0144] The above embodiments of the present disclosure use the BERT model for text sorting. The BERT model is a natural language processing (NLP, Natural Language Process) model and is one of the most breakthrough technologies in the NLP field. The BERT model can be more efficient through pre-training and fine-tuning, and can capture semantic dependencies over longer distances, and capture true bidirectional context information. The embodiments of the present disclosure use the BERT model for text sorting, which can accurately determine the context relationship between texts, thereby improving the text sorting effect.
[0145] In addition, the embodiments of the present disclosure can also use the ERNIE model for text sorting. The ERNIE model can unify the modeling of lexical structures and semantic information in the training data, rather than just focusing on the learning at the character or English word granularity, and can greatly enhance the general semantic representation ability. The embodiments of the present disclosure use the ERNIE model to achieve text sorting, which has a better effect on determining the context relationship between Chinese texts and can improve the text sorting effect, especially the sorting effect of Chinese texts.
[0146] The embodiments of the present disclosure also propose a model training device, Figure 7 which is a schematic structural diagram of a model training device 700 according to the embodiments of the present disclosure, and includes:
[0147] A first acquisition module 710, configured to acquire a sample file, a start text, and an end text, where the sample file includes at least two sample texts and the serial numbers of each sample text;
[0148] A first construction module 720, configured to construct a plurality of text pairs, and each text pair is constructed by two of the following texts: the start text, the end text, and the at least two sample texts;
[0149] A first input module 730, configured to respectively input each of the text pairs into the model, and the model outputs the confidence that the two texts constructing each text pair have a context relationship;
[0150] An adjustment module 740, configured to compare the confidence level with the actual relationship between the two texts that form the text pair, and adjust the parameters of the model according to the comparison result.
[0151] An embodiment of the present disclosure also provides a model training device. Figure 8 As shown in the schematic structural diagram of a model training device 800 according to an embodiment of the present disclosure, it includes:
[0152] A first acquisition module 710, a first construction module 720, a first input module 730, an adjustment module 740, and a first determination module 850. Among them, the first acquisition module 710, the first construction module 720, the first input module 730, and the adjustment module 740 are the same as the corresponding modules above and will not be described in detail here.
[0153] In a possible implementation manner, the above-mentioned first construction module 720 is configured to: respectively construct X text pairs by using the starting text and each of the sample texts, where X is the number of sample texts included in the sample file; respectively construct text pairs by using each of the sample texts and other sample texts or the ending text.
[0154] In a possible implementation manner, the above-mentioned first determination module 850 is configured to determine the actual relationship between the two texts that form each text pair according to the serial numbers of each of the sample texts.
[0155] In a possible implementation manner, the serial numbers of the above-mentioned sample texts are respectively the first serial number to the X serial number, which are used to represent the sequence of each sample text in the sample file;
[0156] The above-mentioned first determination module 850 is configured to:
[0157] For the text pair constructed by the starting text and the sample text with the first serial number, determine that the two texts that form the text pair have a context relationship;
[0158] For the text pair constructed by the starting text and the sample text without the first serial number, determine that the two texts that form the text pair do not have a context relationship;
[0159] For the text pair constructed by sample texts with two adjacent serial numbers before and after, determine that the two texts that form the text pair have a context relationship;
[0160] For the text pair constructed by sample texts without two adjacent serial numbers before and after, determine that the two texts that form the text pair do not have a context relationship;
[0161] For a text pair constructed from the sample text with the Xth serial number and the termination text, the two texts constructing the text pair are determined to have a contextual relationship;
[0162] For a text pair constructed from the sample text without the Xth serial number and the termination text, the two texts constructing the text pair are determined not to have a contextual relationship.
[0163] In a possible implementation, the above-mentioned first construction module 720 is configured to: when the number of characters of the prior text for constructing the text pair is greater than M, intercept the last M characters of the prior text, and use the intercepted content as the prior content for constructing the text pair; when the number of characters of the prior text for constructing the text pair is less than or equal to M, use all the content of the prior text as the prior content for constructing the text pair; M is a positive integer; when the number of characters of the subsequent text for constructing the text pair is greater than N, intercept the first N characters of the subsequent text, and use the intercepted content as the subsequent content for constructing the text pair; when the number of characters of the subsequent text for constructing the text pair is less than or equal to N, use all the content of the subsequent text as the subsequent content for constructing the text pair; N is a positive integer; combine the prior content and the subsequent content in sequence to obtain the text pair.
[0164] In a possible implementation, the above-mentioned first acquisition module 710 is configured to: acquire pictures of each sample text included in the sample file; perform optical character recognition on each picture respectively to obtain each sample text, and acquire the serial number of each sample text.
[0165] In a possible implementation, the above-mentioned model includes a BERT model or an ERNIE model.
[0166] The embodiments of the present disclosure also propose a sorting device, Figure 9 which is a schematic structural diagram of a sorting device 900 according to the embodiments of the present disclosure, including:
[0167] A second acquisition module 910, configured to acquire a starting text, a termination text, and at least two texts to be sorted;
[0168] A second construction module 920, configured to construct a plurality of first text pairs, and each first text pair is constructed from two of the following texts: the starting text, the termination text, and the at least two texts to be sorted;
[0169] A second input module 930, configured to input the plurality of first text pairs into a pre-trained model respectively, and the pre-trained model outputs the confidence that the two texts constructing each of the first text pairs have a contextual relationship;
[0170] A second determination module 940, configured to determine the order of at least two texts to be sorted according to the confidence level;
[0171] Wherein, the above-mentioned pre-trained model can be trained in any manner in the model training method proposed in the present disclosure.
[0172] An embodiment of the present disclosure also provides a sorting device. Figure 10 According to the structural schematic diagram of a sorting device 1000 according to an embodiment of the present disclosure, as Figure 10 shown, in a possible implementation manner, the above-mentioned second determination module 940 includes:
[0173] A text pair determination sub-module 941, configured to, for each text, determine each first text pair with the text as the prior text; obtain the confidence level that the two texts for constructing each of the first text pairs have a context relationship, and determine the two texts with the highest confidence level as having a context relationship;
[0174] An order determination sub-module 942, configured to determine the order of the at least two texts to be sorted according to multiple pairs of texts with a context relationship.
[0175] In a possible implementation manner, the above-mentioned second construction module 920 is configured to: respectively construct Y first text pairs with the starting text and each of the texts to be sorted, where Y is the number of texts to be sorted; respectively construct first text pairs with each of the texts to be sorted and each other text to be sorted or the ending text.
[0176] In a possible implementation manner, the above-mentioned second construction module 920 is configured to:
[0177] When the number of characters of the prior text for constructing the first text pair is greater than M, intercept the last M characters of the prior text, and use the intercepted content as the prior content for constructing the first text pair; when the number of characters of the prior text for constructing the first text pair is less than or equal to M, use all the content of the prior text as the prior content for constructing the first text pair; M is a positive integer;
[0178] When the number of characters of the subsequent text for constructing the first text pair is greater than N, intercept the first N characters of the subsequent text, and use the intercepted content as the subsequent content for constructing the first text pair; when the number of characters of the subsequent text for constructing the first text pair is less than or equal to N, use all the content of the subsequent text as the subsequent content for constructing the first text pair; N is a positive integer;
[0179] Combine the prior content and the subsequent content in sequence to obtain the first text pair.
[0180] In a possible implementation, the second obtaining module 910 is configured to: obtain pictures corresponding to at least two texts to be sorted respectively; perform optical character recognition on each of the pictures to obtain at least two texts to be sorted.
[0181] For the functions of the modules and / or units in the apparatus embodiments of the present disclosure, reference may be made to the relevant descriptions in the above method embodiments of the present disclosure, which will not be elaborated herein.
[0182] In the technical solution of the present disclosure, the acquisition, storage, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0183] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0184] Figure 11 FIG. shows a schematic block diagram of an exemplary electronic device 1100 that can be used to implement the embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0185] As Figure 11 shown, the device 1100 includes a computing unit 1101, which can execute various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1102 or a computer program loaded from a storage unit 1108 into a random access memory (RAM) 1103. In the RAM 1103, various programs and data required for the operation of the device 1100 can also be stored. The computing unit 1101, the ROM 1102, and the RAM 1103 are connected to each other through a bus 1104. An input / output (I / O) interface 1105 is also connected to the bus 1104.
[0186] Multiple components in device 1100 are connected to I / O interface 1105, including: an input unit 1106, such as a keyboard, a mouse, etc.; an output unit 1107, such as various types of displays, speakers, etc.; a storage unit 1108, such as a magnetic disk, an optical disc, etc.; and a communication unit 1109, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1109 allows device 1100 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0187] The computing unit 1101 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1101 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1101 executes the various methods and processes described above, such as a model training method or a sorting method. For example, in some embodiments, the model training method or the sorting method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 1108. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 1100 via the ROM 1102 and / or the communication unit 1109. When the computer program is loaded into the RAM 1103 and executed by the computing unit 1101, one or more steps of the model training method or the sorting method described above can be executed. Alternatively, in other embodiments, the computing unit 1101 can be configured to execute the model training method or the sorting method in any other suitable manner (e.g., by means of firmware).
[0188] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0189] The program code for implementing the methods of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general purpose computer, a special purpose computer, or other programmable data processing device, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on the remote machine or server.
[0190] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0191] In order to provide interaction with a user, the systems and techniques described herein may be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).
[0192] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.
[0193] A computer system can include a client and a server. The client and the server are generally far from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, can also be a server of a distributed system, or a server incorporating a blockchain.
[0194] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this is not limited herein.
[0195] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. A model training method, including: obtaining a sample file, a start text, and an end text, where the sample file includes at least two sample texts and the serial number of each sample text; constructing a plurality of text pairs, each text pair being constructed by two of the following texts: the start text, the end text, and the at least two sample texts; wherein, constructing the plurality of text pairs includes: when the number of characters of the prior text for constructing the text pair is greater than M, intercepting the last M characters of the prior text and using the intercepted content as the prior content for constructing the text pair; when the number of characters of the prior text for constructing the text pair is less than or equal to M, using all the content of the prior text as the prior content for constructing the text pair; M is a positive integer; when the number of characters of the subsequent text for constructing the text pair is greater than N, intercepting the first N characters of the subsequent text and using the intercepted content as the subsequent content for constructing the text pair; when the number of characters of the subsequent text for constructing the text pair is less than or equal to N, using all the content of the subsequent text as the subsequent content for constructing the text pair; N is a positive integer; combining the prior content and the subsequent content in sequence to obtain the text pair; inputting each of the text pairs into the model, and outputting, by the model, the confidence that the two texts constructing each text pair have a context relationship; for each text pair, comparing the confidence with the actual relationship between the two texts constructing the text pair, and adjusting the parameters of the model according to the comparison result.
2. The method according to claim 1, wherein, constructing the plurality of text pairs includes: using the start text to construct X text pairs with each of the sample texts respectively, where X is the number of sample texts included in the sample file; using each of the sample texts to construct text pairs with other sample texts or the end text respectively.
3. The method according to claim 2, further including: determining the actual relationship between the two texts constructing each text pair according to the serial numbers of each sample text.
4. The method according to claim 3, wherein, the serial numbers of each of the sample texts are the 1st serial number to the Xth serial number respectively, which are used to represent the sequence of each sample text in the sample file; determining the actual relationship between the two texts constructing each text pair includes at least one of the following: for the text pair constructed by the start text and the sample text with the 1st serial number, determining that the two texts constructing the text pair have a context relationship; for the text pair constructed by the start text and the sample text without the 1st serial number, determining that the two texts constructing the text pair do not have a context relationship; for the text pair constructed by sample texts with two adjacent serial numbers before and after respectively, determining that the two texts constructing the text pair have a context relationship; for the text pair constructed by sample texts without two adjacent serial numbers before and after, determining that the two texts constructing the text pair do not have a context relationship; For a text pair constructed from a sample text with the X - th serial number and the termination text, the two texts constructing the text pair are determined to have a context relationship. For a text pair constructed from a sample text without the X - th serial number and the termination text, the two texts constructing the text pair are determined not to have a context relationship.
5. According to the method described in any one of claims 1 to 4, wherein, the obtaining of the sample file includes: obtaining pictures of each sample text included in the sample file; performing optical character recognition on each of the pictures respectively to obtain each sample text, and obtaining the serial number of each sample text.
6. A text sorting method, including: obtaining a starting text, a termination text, and at least two texts to be sorted; constructing a plurality of first text pairs, each first text pair being constructed from two of the following texts: the starting text, the termination text, and the at least two texts to be sorted; wherein, the constructing of the plurality of first text pairs includes: when the number of characters of the prior text for constructing the first text pair is greater than M, intercepting the last M characters of the prior text, and using the intercepted content as the prior content for constructing the first text pair; when the number of characters of the prior text for constructing the first text pair is less than or equal to M, using all the content of the prior text as the prior content for constructing the first text pair; M is a positive integer; when the number of characters of the subsequent text for constructing the first text pair is greater than N, intercepting the first N characters of the subsequent text, and using the intercepted content as the subsequent content for constructing the first text pair; when the number of characters of the subsequent text for constructing the first text pair is less than or equal to N, using all the content of the subsequent text as the subsequent content for constructing the first text pair; N is a positive integer; combining the prior content and the subsequent content in sequence to obtain the first text pair; inputting the plurality of first text pairs into a pre - trained model respectively, and outputting by the pre - trained model the confidence that the two texts constructing each first text pair have a context relationship; determining the sequence of the at least two texts to be sorted according to the confidence.
7. According to the method described in claim 6, wherein, the constructing of the plurality of first text pairs includes: using the starting text to construct Y first text pairs with each of the texts to be sorted respectively, Y being the number of texts to be sorted; using each of the texts to be sorted to construct first text pairs with other texts to be sorted or the termination text respectively.
8. According to the method described in claim 7, wherein, the determining of the sequence of the at least two texts to be sorted according to the confidence includes: for each text, determining each first text pair with this text as the prior text; obtaining the confidence that the two texts constructing each determined first text pair have a context relationship, and determining the two texts with the highest confidence as having a context relationship; determining the sequence of the at least two texts to be sorted according to multiple pairs of texts having a context relationship.
9. The method according to claim 8, wherein, obtaining the at least two texts to be sorted includes: obtaining the pictures respectively corresponding to the at least two texts to be sorted; performing optical character recognition on each of the pictures respectively to obtain the at least two texts to be sorted.
10. The method according to any one of claims 6 to 9, wherein, the pre-trained model is trained by using the method according to any one of claims 1 to 5.
11. A model training device, comprising: a first obtaining module, configured to obtain a sample file, a starting text, and an ending text, where the sample file includes at least two sample texts and the serial numbers of each of the sample texts; a first constructing module, configured to construct a plurality of text pairs, and each text pair is constructed by two of the following texts: the starting text, the ending text, and the at least two sample texts; wherein, when the number of characters of the prior text for constructing the text pair is greater than M, intercept the last M characters of the prior text, and use the intercepted content as the prior content for constructing the text pair; when the number of characters of the prior text for constructing the text pair is less than or equal to M, use all the content of the prior text as the prior content for constructing the text pair; M is a positive integer; when the number of characters of the subsequent text for constructing the text pair is greater than N, intercept the first N characters of the subsequent text, and use the intercepted content as the subsequent content for constructing the text pair; when the number of characters of the subsequent text for constructing the text pair is less than or equal to N, use all the content of the subsequent text as the subsequent content for constructing the text pair; N is a positive integer; combine the prior content and the subsequent content in sequence to obtain the text pair; a first input module, configured to input each of the text pairs into the model respectively, and the model outputs the confidence degrees that the two texts for constructing each of the text pairs have a context relationship; an adjustment module, configured to compare the confidence degrees with the actual relationships of the two texts for constructing the text pairs, and adjust the parameters of the model according to the comparison results.
12. The device according to claim 11, wherein, the first constructing module is configured to: construct X text pairs by using the starting text and each of the sample texts respectively, where X is the number of sample texts included in the sample file; construct text pairs by using each of the sample texts and each of the other sample texts or the ending text respectively.
13. The device according to claim 12, further comprising: a first determining module, configured to determine the actual relationships of the two texts for constructing each of the text pairs according to the serial numbers of each of the sample texts.
14. The device according to claim 13, wherein, the serial numbers of each of the sample texts are the first serial number to the X serial number respectively, and are used to represent the sequence of each of the sample texts in the sample file; the first determining module is configured to: for the text pair constructed by the starting text and the sample text having the first serial number, determine that the two texts for constructing the text pair have a context relationship; For a text pair constructed from the starting text and a sample text that does not have the first serial number, the two texts constructing the text pair are determined to have no context relationship; For a text pair constructed from sample texts with two adjacent serial numbers before and after respectively, the two texts constructing the text pair are determined to have a context relationship; For a text pair constructed from a sample text that does not have two adjacent serial numbers before and after, the two texts constructing the text pair are determined to have no context relationship; For a text pair constructed from a sample text with the X serial number and the ending text, the two texts constructing the text pair are determined to have a context relationship; For a text pair constructed from a sample text that does not have the X serial number and the ending text, the two texts constructing the text pair are determined to have no context relationship.
15. The apparatus according to any one of claims 11 to 14, wherein, the first acquisition module is configured to: acquire pictures of each sample text included in the sample file; perform optical character recognition on each of the pictures respectively to obtain each sample text, and acquire the serial number of each sample text.
16. A sorting apparatus, comprising: a second acquisition module, configured to acquire a starting text, an ending text, and at least two texts to be sorted; a second construction module, configured to construct a plurality of first text pairs, each first text pair being constructed from two of the following texts: the starting text, the ending text, and the at least two texts to be sorted; wherein, when the number of characters of the prior text for constructing the first text pair is greater than M, intercept the last M characters of the prior text, and use the intercepted content as the prior content for constructing the first text pair; when the number of characters of the prior text for constructing the first text pair is less than or equal to M, use all the content of the prior text as the prior content for constructing the first text pair; when the number of characters of the subsequent text for constructing the first text pair is greater than N, intercept the first N characters of the subsequent text, and use the intercepted content as the subsequent content for constructing the first text pair; when the number of characters of the subsequent text for constructing the first text pair is less than or equal to N, use all the content of the subsequent text as the subsequent content for constructing the first text pair; N is a positive integer; combine the prior content and the subsequent content in sequence to obtain the first text pair; a second input module, configured to input the plurality of first text pairs into a pre-trained model respectively, and the pre-trained model outputs the confidence that the two texts constructing each first text pair have a context relationship; a second determination module, configured to determine the sequence of the at least two texts to be sorted according to the confidence.
17. The apparatus according to claim 16, wherein, the second construction module is configured to: construct Y first text pairs by using the starting text and each of the texts to be sorted respectively, Y being the number of texts to be sorted; construct first text pairs by using each of the texts to be sorted and each of the other texts to be sorted or the ending text respectively.
18. The apparatus according to claim 17, wherein, according to the confidence level, the second determination module includes: a text pair determination sub-module, configured to determine, for each text, each first text pair with the text as the prior text; obtain the confidence level that the two texts for constructing each of the first text pairs have a context relationship, and determine the two texts with the highest confidence level as having a context relationship; an order determination sub-module, configured to determine the order of the at least two texts to be sorted according to multiple pairs of texts having a context relationship.
19. The apparatus according to claim 18, wherein, the second acquisition module is configured to: acquire the pictures respectively corresponding to the at least two texts to be sorted; perform optical character recognition on each of the pictures to obtain the at least two texts to be sorted.
20. The apparatus according to any one of claims 16 to 19, wherein, the pre-trained model is trained by using the apparatus according to any one of claims 11 to 15.
21. An electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method according to any one of claims 1-10.
22. A non-transitory computer-readable storage medium storing computer instructions, wherein, the computer instructions are used to cause the computer to execute the method according to any one of claims 1-10.
23. A computer program product, comprising a computer program, where the computer program, when executed by a processor, implements the method according to any one of claims 1-10.
Citation Information
Patent Citations
Method and device for sorting litigation documents
CN112785464A
Mapping Documents to Associated Outcome based on Sequential Evolution of Their Contents
US20160155067A1