Text processing method and device and model training method and device

By converting PDF files into text format and splitting characters, deleting lost table text, and using a large language model to repair text content, solving the problems of table loss and formula errors in PDF file parsing, improving the prediction accuracy of the model.

CN120387425APending Publication Date: 2025-07-29INFLY TECH (SHANGHAI) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510500365.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

In the prior art, PDF file parsing methods have problems with inaccurate text extraction, especially table loss and formula recognition errors, which affect the use of subsequent model optimization tasks.

Method used

Convert the pending format file into a text format file, divide it into multiple segmented texts through preset splitters, delete the lost table text, use the large language model to repair text, obtain high-quality target benchmark segmented text, and store it in the target sample collection.

Benefits of technology

Improve the accuracy of PDF file parsing, ensure that the model optimization task uses high-quality text, and improves the prediction accuracy of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120387425A_ABST
    Figure CN120387425A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a text processing method and device and a model training method and device.The text processing method comprises the steps that a to-be-processed format file is converted into a text format file, the text format file is segmented according to a preset segmentation symbol, and a plurality of segmented texts are obtained; determining the number of table characters contained in each segmented text and the number of tables corresponding to each segmented text, and deleting lost table texts in the plurality of segmented texts according to the number of table characters and the number of tables to obtain an initial reference segmented text; determining a repair prompt word corresponding to the initial reference segmented text, and inputting the repair prompt word and the initial reference segmented text into a large language model for text repair processing to obtain a target reference segmented text; and storing the target reference segmented text to a target sample set, the target sample set being used for executing a model optimization task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this specification relate to the technical field of text processing, and particularly to text processing methods and devices, model training methods and devices. Background Art

[0002] With the development of computer and Internet technologies, PDF files are applied in more and more scenarios. As a document format for file exchange and presentation, PDF files have the characteristics of cross-platform, cross-device, and cross-application, and can maintain a consistent presentation effect on different operating systems, different devices, and different software. And the parsing of PDF files is an important means of processing PDF files. In the prior art, the principles for parsing PDF files are mainly divided into rule-based parsing methods and deep learning-based parsing methods. Among them, the rule-based parsing method mainly relies on predefined rules to identify different elements in the document, such as text, tables, and images. And the deep learning-based parsing method is mainly implemented using deep learning models, such as object detection and OCR recognition technologies, to extract text, images, and tables in PDF files. However, both the rule-based parsing method and the deep learning-based parsing method have the problem of inaccurate text extraction, and even problems such as formula recognition errors and table loss may occur, resulting in the parsed file being inconvenient for downstream tasks to use. Therefore, an effective solution is urgently needed to solve the above problems. Summary of the Invention

[0003] In view of this, the embodiments of this specification provide a text processing method. One or more embodiments of this specification also relate to a model training method, a text processing device, a model training device, a computing device, a computer-readable storage medium, and a computer program product to solve the technical defects existing in the prior art.

[0004] According to the first aspect of the embodiments of this specification, a text processing method is provided, including: Converting a file in a to-be-processed format into a text format file, and splitting the text format file according to a preset delimiter to obtain a plurality of split texts; Determining the number of table delimiters included in each split text and the number of tables corresponding to each split text, and deleting the missing table texts in the plurality of split texts according to the number of table delimiters and the number of tables to obtain an initial reference split text; Determining a repair prompt word corresponding to the initial reference split text, and inputting the repair prompt word and the initial reference split text into a large language model for text repair processing to obtain a target reference split text; Store the target reference segmentation text in a target sample set, where the target sample set is used to perform model optimization tasks.

[0005] Optionally, splitting the text format file according to a preset delimiter to obtain multiple split texts, including: Detect the delimiter in the text format file, and determine the target delimiter included in the text format file according to the detection result; Use the target delimiter as the preset delimiter, and perform the step of splitting the text format file according to the preset delimiter to obtain multiple split texts.

[0006] Optionally, determining the number of tab characters included in any one of the multiple split texts includes: Detect whether a preset tab character is included in the first split text; If so, count the tab characters included in the first split text, and determine the number of tab characters included in the first split text according to the statistical result; Among them, determining the number of tables corresponding to any one of the multiple split texts includes: Determine table keywords in the second split text, determine the numerical value corresponding to the table keywords, and use the numerical value corresponding to the table keywords as the number of tables corresponding to the second split text.

[0007] Optionally, deleting missing table text from the multiple split texts according to the number of tab characters and the number of tables to obtain an initial reference segmentation text, including: Determine the i-th split text in the multiple split texts, and detect whether the number of tab characters in the i-th split text is less than the number of tables in the i-th split text, where i starts from 1 and i is a positive integer; If so, use the i-th split text as the missing table text; Increment i in sequence until i increments to n. Delete the missing table text from the multiple split texts, and use the split texts other than the missing table text in the multiple split texts as the initial reference segmentation text, where n is the number of texts in the multiple split texts.

[0008] Optionally, determining the repair prompt word corresponding to the initial reference segmentation text, and inputting the repair prompt word and the initial reference segmentation text into a large language model for text repair processing to obtain a target reference segmentation text, including: Determine the table prompt word corresponding to the initial reference segmentation text, and input the table prompt word and the initial reference segmentation text into a large language model for table repair processing to obtain an intermediate reference segmentation text; Determine the formula prompt word corresponding to the intermediate reference segmentation text, and input the formula prompt word and the intermediate reference segmentation text into a large language model for formula repair processing to obtain the target reference segmentation text.

[0009] Optionally, after the step of inputting the repair prompt word and the initial reference segmentation text into a large language model for text repair processing to obtain the target reference segmentation text, the following is further included: Determine the original segmentation text corresponding to the target reference segmentation text among the multiple segmentation texts, and compare the original segmentation text with the target reference segmentation text; When it is determined according to the comparison result that the target reference segmentation text does not meet the storage condition, optimize the target reference segmentation text by using the original segmentation text; Store the optimized target reference segmentation text into the target sample set.

[0010] Optionally, after the step of storing the target reference segmentation text into the target sample set, the following is further included: Extract the target sample text from the target sample set in response to a model training request; Input the target sample text into the language model to be trained for processing to obtain a prediction text; Optimize the language model to be trained based on the sample label corresponding to the target sample text and the prediction text until a target language model that meets the training stop condition is obtained.

[0011] According to the second aspect of the embodiments of the present specification, a model training method is provided, including: Extract the target sample text from the target sample set in response to a model training request, where the target sample set is determined by the above method; Input the target sample text into the language model to be trained for processing to obtain a prediction text; Optimize the language model to be trained based on the sample label corresponding to the target sample text and the prediction text until a target language model that meets the training stop condition is obtained.

[0012] According to the third aspect of the embodiments of the present specification, a text processing device is provided, including: A conversion module configured to convert a file in a to-be-processed format into a text format file, and segment the text format file according to a preset delimiter to obtain a plurality of segmented texts; A determination module, configured to determine the number of tab characters included in each segmented text and the number of tables corresponding to each segmented text, and delete the missing table text from the multiple segmented texts according to the number of tab characters and the number of tables to obtain initial reference segmented texts; A repair module, configured to determine a repair prompt word corresponding to the initial reference segmented text, and input the repair prompt word and the initial reference segmented text into a large language model for text repair processing to obtain target reference segmented texts; A storage module, configured to store the target reference segmented texts into a target sample set, where the target sample set is used to perform a model optimization task.

[0013] According to a fourth aspect of the embodiments of this specification, there is provided a model training device, including: An extraction module, configured to extract target sample texts from a target sample set in response to a model training request, where the target sample set is determined by the above method; A processing module, configured to input the target sample texts into a language model to be trained for processing to obtain predicted texts; An optimization module, configured to optimize the language model to be trained based on the sample labels corresponding to the target sample texts and the predicted texts until a target language model that meets the training stop condition is obtained.

[0014] According to a fifth aspect of the embodiments of this specification, there is provided a computing device, including: A memory and a processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the above text processing method or model training method are implemented.

[0015] According to a sixth aspect of the embodiments of this specification, there is provided a computer-readable storage medium, which stores computer-executable instructions. When the instructions are executed by a processor, the steps of the above text processing method or model training method are implemented.

[0016] According to a seventh aspect of the embodiments of this specification, there is provided a computer program product, including a computer program or instructions. When the computer program or instructions are executed by a processor, the steps of the above text processing method or model training method are implemented.

[0017] The text processing method provided in this embodiment, in order to improve the parsing accuracy of the file to be processed and avoid losing content and affecting subsequent model optimization tasks, can first convert the file in the to-be-processed format into a text format file. On this basis, the text format file is segmented according to a preset delimiter to obtain multiple segmented texts. Then, the number of table symbols contained in each segmented text and the number of tables corresponding to each segmented text can be determined, and the lost table texts are deleted from the multiple segmented texts according to the number of table symbols and the number of tables, so as to realize the elimination of the segmented texts that cannot be repaired, and only the initial benchmark segmented texts with complete information are retained. On this basis, considering that there may be problems such as formula errors and missing table headers in the initial benchmark segmented texts, and such problems do not affect the essential content of the text, but will affect the accuracy of the text content. Therefore, in order to obtain high-quality text for subsequent use, the repair prompt words corresponding to the initial benchmark segmented texts can be determined, and the repair prompt words and the initial benchmark segmented texts are input into the large language model for text repair processing to obtain the target benchmark segmented texts; realize the repair of the initial benchmark segmented texts by using the large language model. After that, the target benchmark segmented texts can be stored in the target sample set, so that the model optimization task can be executed according to the target sample set subsequently, thereby ensuring that the model optimization task can complete training or evaluation using high-quality text to improve the prediction accuracy of the model. Description of the Drawings

[0018] Figure 1 is a schematic diagram of a text processing method provided by an embodiment of this specification; Figure 2 is a flowchart of a text processing method provided by an embodiment of this specification; Figure 3 is a flowchart of a model training method provided by an embodiment of this specification; Figure 4 is a flowchart of the processing process of a text processing method provided by an embodiment of this specification; Figure 5 is a schematic structural diagram of a text processing device provided by an embodiment of this specification; Figure 6 is a schematic structural diagram of a model training device provided by an embodiment of this specification; Figure 7 is a block diagram of the structure of a computing device provided by an embodiment of this specification. Detailed Description of the Embodiment

[0019] Numerous specific details are set forth in the following description to provide a thorough understanding of this specification. However, this specification can be implemented in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the connotation of this specification. Therefore, this specification is not limited by the specific implementations disclosed below.

[0020] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a", "the", and "said" used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and encompasses any and all possible combinations of one or more of the associated listed items.

[0021] It should be understood that although the terms first, second, etc. may be used in one or more embodiments of this specification to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".

[0022] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data that have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of the relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entrances are provided for the user to select authorization or rejection.

[0023] First, the noun terms involved in one or more embodiments of this specification are explained.

[0024] PDF (Portable Document Format) is a document format for file exchange and presentation. PDF files have the characteristics of cross-platform, cross-device, and cross-application, and can maintain a consistent presentation effect on different operating systems, different devices, and different software.

[0025] In this specification, a text processing method is provided. One or more embodiments of this specification simultaneously relate to a model training method, a text processing device, a model training device, a computing device, a computer-readable storage medium, and a computer program product, which will be described in detail one by one in the following embodiments.

[0026] Referring to Figure 1 the schematic diagram shown, for the text processing method provided in this embodiment, in order to improve the parsing accuracy of the file to be processed and avoid losing content and affecting the subsequent model optimization task, the file in the to-be-processed format can be first converted into a text format file. On this basis, the text format file is segmented according to a preset delimiter to obtain a plurality of segmented texts; then, the number of table symbols included in each segmented text and the number of tables corresponding to each segmented text can be determined, and the lost table texts are deleted from the plurality of segmented texts according to the number of table symbols and the number of tables, so as to eliminate the segmented texts that cannot be repaired, and only retain the initial reference segmented texts with complete information; on this basis, considering that the initial reference segmented texts may have problems such as formula errors and missing headers, and such problems do not affect the essential content of the text, but will affect the accuracy of the text content. Therefore, in order to obtain high-quality text for subsequent use, the repair prompt words corresponding to the initial reference segmented texts can be determined, and the repair prompt words and the initial reference segmented texts are input into a large language model for text repair processing to obtain the target reference segmented texts; the repair of the initial reference segmented texts is completed by using the large language model. After that, the target reference segmented texts can be stored in the target sample set, so that the model optimization task can be executed according to the target sample set subsequently, thereby ensuring that the model optimization task can use high-quality text to complete training or evaluation, so as to improve the prediction accuracy of the model.

[0027] Referring to Figure 2 , Figure 2 shows a flowchart of a text processing method provided according to an embodiment of this specification, which specifically includes the following steps.

[0028] Step S202, convert the file in the to-be-processed format into a text format file, and segment the text format file according to a preset delimiter to obtain a plurality of segmented texts.

[0029] The text processing method provided in this embodiment is applied to the task of parsing PDF files in any scenario, so as to perform subsequent processing according to the parsing results of PDF files. For example, after parsing a PDF file to obtain text, the obtained text is used as a sample in the field to which the content of the corresponding PDF file belongs, for training or evaluating a neural network model in this field. It is also possible to use the text obtained by parsing a PDF file for subsequent text review, such as test paper grading.

[0030] In this embodiment, a PDF file is taken as an example of the test paper in the CFA (Chartered Financial Analyst) qualification exam. For the text processing method, it is used to parse the PDF-format test paper in the CFA qualification exam. After obtaining high-quality parsing results, it can be used as a training set or an evaluation set to train a large model for processing test papers in the CFA qualification exam, thereby improving the professional ability of the large model in the field of CFA exams. For example, using this model for test paper marking, generating test answers, generating answer analysis, etc. For the parsing of PDF files in other scenarios, refer to the same or corresponding descriptions in this embodiment, and this embodiment does not make any limitations here.

[0031] Specifically, the file in the format to be processed specifically refers to a PDF-format file. Correspondingly, the text-format file specifically refers to a file in which the content obtained after converting the file in the format to be processed exists in text format. Correspondingly, the delimiter specifically refers to a symbol for splitting the text content existing in the text-format file, such as "\n", etc., for splitting the text-format file into multiple split texts, so that each split text can be processed subsequently to complete the parsing operation of the file in the format to be processed. Correspondingly, the split text is the text obtained after splitting the text-format file, and each split text is a separate text unit. For example, if the file in the format to be processed is a PDF-format test paper in the CFA exam field, after converting it to text format and splitting it according to the preset delimiter, the multiple split texts obtained are each question in the test paper. At this time, each delimiter is the delimiter for dividing questions.

[0032] Based on this, in order to improve the parsing accuracy of the file to be processed and avoid losing content and affecting the subsequent use of the model optimization task, the file in the to-be-processed format can be first converted into a text format file. On this basis, the text format file is segmented according to a preset delimiter to obtain multiple segmented texts. Then, the number of tab characters included in each segmented text and the number of tables corresponding to each segmented text can be determined, and the lost table texts are deleted from the multiple segmented texts according to the number of tab characters and the number of tables, so as to remove the segmented texts that cannot be repaired and only retain the initial benchmark segmented texts with complete information. On this basis, considering that there may be problems such as formula errors and missing table headers in the initial benchmark segmented texts, and such problems do not affect the substantial content of the text but will affect the accuracy of the text content. Therefore, in order to obtain high-quality text for subsequent use, the repair prompt words corresponding to the initial benchmark segmented texts can be determined, and the repair prompt words and the initial benchmark segmented texts are input into the large language model for text repair processing to obtain the target benchmark segmented texts, thus realizing the repair of the initial benchmark segmented texts by using the large language model. After that, the target benchmark segmented texts can be stored in the target sample set, so that the model optimization task can be executed according to the target sample set subsequently, thereby ensuring that the model optimization task can complete training or evaluation using high-quality text to improve the prediction accuracy of the model.

[0033] Further, when segmenting the text format file according to the preset delimiter, considering that different files contain different delimiters, in order to ensure that the segmented texts obtained after segmentation meet the subsequent usage requirements, the delimiter can be determined by detecting the file. In this embodiment, the specific implementation method is as follows: Perform delimiter detection on the text format file, and determine the target delimiter included in the text format file according to the detection result; use the target delimiter as the preset delimiter, and perform the step of segmenting the text format file according to the preset delimiter to obtain multiple segmented texts.

[0034] Specifically, the delimiter detection specifically refers to the processing operation of detecting whether the text format file contains a delimiter. Correspondingly, the target delimiter specifically refers to the delimiter that the text format file may contain. Based on this, after converting the file in the to-be-processed format into a text format file, in order to be able to segment the text format file and ensure that the obtained segmented texts exist in units of text, the delimiter detection can be first performed on the text format file. When the target delimiter included in the text format file is determined according to the detection result, the target delimiter can be used as the preset delimiter, and the step of segmenting the text format file according to the preset delimiter to obtain multiple segmented texts can be performed. Thus, the accurate segmentation operation of the text format file is realized.

[0035] For example, when parsing the PDF-format test papers in the CFA qualification exam, conventional PDF parsing tools can be used to convert the PDF-format test papers into text format. On this basis, considering that most of the questions in the CFA qualification exam contain multiple tables, if the rule-based parsing method is adopted, it may lead to the loss of tables. If used for the training or evaluation of large models, it will affect the prediction accuracy of the models. Therefore, before that, the questions with lost tables need to be removed. In this process, the text-format test papers need to be segmented first. It can be detected whether the text-format test papers contain the delimiter "\n\n\n\n". If it exists, the text-format test papers can be segmented according to this delimiter, and then the segmented text corresponding to each question can be obtained. In addition, if this delimiter does not exist, the text-format test papers can also be segmented according to the delimiter "\n\n". After obtaining the segmented text corresponding to each question, operations such as detecting the text with lost tables and text repair can be carried out.

[0036] In summary, by detecting the delimiter in the text-format file and segmenting it according to the detected delimiter, it can be ensured that the segmented text obtained after segmentation is segmented in units of the objects to be processed subsequently, which is convenient for the subsequent construction of the text used for model training or evaluation.

[0037] Step S204: Determine the number of table delimiters contained in each segmented text and the number of tables corresponding to each segmented text, and delete the text with lost tables in the multiple segmented texts according to the number of table delimiters and the number of tables to obtain the initial benchmark segmented text.

[0038] Specifically, after obtaining the multiple segmented texts above, considering that the format file to be processed may contain multiple tables, and due to format conversion, the problem of table loss may occur. If the impact of this problem is ignored, it may seriously affect the subsequent sample quality. Therefore, the number of table delimiters contained in each segmented text and the number of tables corresponding to each segmented text can be determined, so as to conveniently determine the text with lost tables with table loss problems in the multiple segmented texts by comparing the number of table delimiters and the number of tables. At this time, the text with lost tables can be deleted, and the remaining segmented texts can be used as the initial benchmark segmented text for subsequent sample construction with the initial benchmark segmented text with complete information.

[0039] Among them, the number of tab characters specifically refers to the total number of tab characters representing the existence in each segmented text, and the number of tables specifically refers to the true value of the number of tables included in each segmented text. By comparing the number of tab characters collected in real time with the true number of tables, it is possible to determine whether there is a problem of missing tables in the segmented text, and then complete the screening of the segmented text corresponding to the perfect information. Correspondingly, the missing table text is the segmented text that has lost tables among multiple segmented texts. The initial reference segmented text specifically refers to the segmented text with perfect information and no missing tables.

[0040] Further, in order to ensure the accuracy of determining the number of tab characters, the determination of the number of tab characters included in any one of the multiple segmented texts can be achieved through the following method: Detect whether the first segmented text contains a preset tab character; if so, count the tab characters included in the first segmented text, and determine the number of tab characters included in the first segmented text according to the statistical result; correspondingly, the determination of the number of tables corresponding to any one of the multiple segmented texts includes: determining table keywords in the second segmented text, determining the value corresponding to the table keywords, and taking the value corresponding to the table keywords as the number of tables corresponding to the second segmented text.

[0041] Specifically, the first segmented text and the second segmented text specifically refer to any one of the multiple segmented texts. Correspondingly, the preset tab character specifically refers to a symbol that can reflect the existence of a table in the text, such as "\t", "|", etc. In specific implementation, it can be set according to actual needs, and this embodiment does not make any limitation here. Correspondingly, the table keyword specifically refers to the keyword "Exhibit" used to record the number of tables after converting the PDF file into a text format file.

[0042] Based on this, for any one of the multiple segmented texts, it is possible to first detect whether the first segmented text contains a preset tab character; if not, it means that the number of tab characters it contains is zero; if so, it means that the first segmented text contains a table, and at this time, the tab characters included in the first segmented text can be counted, and the number of tab characters included in the first segmented text can be determined according to the statistical result.

[0043] For any one of the multiple segmented texts, it is also possible to determine table keywords in the second segmented text and read the value corresponding to the table keywords. This value is used to represent the number of tables corresponding to the segmented text in the actual situation. Therefore, the value corresponding to the table keywords can be taken as the number of tables corresponding to the second segmented text. To realize subsequent comparison with the number of tab characters to determine whether there is a problem of missing tables.

[0044] In summary, by uniformly determining the number of tables included in the current state of the segmented text for table delimiters, and by reading the values of keywords to determine the actual number of tables included in the segmented text, it is possible to determine whether there is a problem of missing tables by comparing the two, so as to ensure that the text with complete information obtained subsequently is used as the initial reference segmented text.

[0045] Furthermore, when determining the initial reference segmented text, in fact, it is to screen the segmented text with complete information, so it can be achieved by means of iterative detection. In this embodiment, the specific implementation method is as follows: Determine the i-th segmented text among the multiple segmented texts, and detect whether the number of table delimiters of the i-th segmented text is less than the number of tables of the i-th segmented text, where i starts from 1 and i is a positive integer; if so, regard the i-th segmented text as the missing table text; i increases sequentially until i increases to n, delete the missing table text among the multiple segmented texts, and regard the segmented texts other than the missing table text among the multiple segmented texts as the initial reference segmented text, where n is the number of texts of the multiple segmented texts.

[0046] Based on this, after determining the number of table delimiters and the number of tables corresponding to each segmented text, at this time, missing table detection can be performed for each segmented text respectively. Specifically, the i-th segmented text can be determined among the multiple segmented texts first, and it is detected whether the number of table delimiters of the i-th segmented text is less than the number of tables of the i-th segmented text, where i starts from 1 and i is a positive integer; if not, it means that the number of table delimiters and the number of tables included in the segmented text are the same, and further indicates that it has not lost tables, so it can be used as the initial reference segmented text.

[0047] If so, it means that there is a situation of missing tables in the i-th segmented text, so the i-th segmented text can be regarded as the missing table text; at this time, i can be incremented sequentially, and the above operations can be repeated. Until i increases to n, at this time, the missing table text can be deleted among the multiple segmented texts, and the segmented texts other than the missing table text among the multiple segmented texts can be used as the initial reference segmented text, which is convenient for subsequent sample construction operations using the initial reference segmented text.

[0048] Continuing with the above example, after obtaining the segmented text corresponding to each test question, for the segmented text corresponding to any test question, it is possible to detect whether it contains table delimiters such as "\t" or "|". If present, the detected table delimiters can be counted to determine the number of table delimiters contained in the segmented text, thereby reflecting the number of tables contained in the current format of the segmented text. At the same time, it is also possible to detect the keyword "Exhibit" contained in the segmented text and use its corresponding value as the actual number of tables contained in the segmented text. Subsequently, the two table numbers can be compared. If the number of table delimiters is less than the number of tables, it indicates that there is a problem of missing tables in the segmented text. If not less, it indicates that there is no problem of missing tables in the segmented text. By analogy, after performing the above detections for the segmented text corresponding to each test question, if it is determined that there are m segmented texts with missing tables, at this time, the m segmented texts with missing tables can be removed, and then the segmented texts corresponding to the remaining test questions without missing tables can be used as the standard segmented texts for subsequent repair processing.

[0049] In summary, by comparing the number of table delimiters and the number of tables in each segmented text, the completeness of the table information in the segmented text can be accurately determined, and then the table texts with complete information can be selected for use, thereby improving the sample quality.

[0050] Step S206: Determine the repair prompt words corresponding to the initial reference segmented text, and input the repair prompt words and the initial reference segmented text into a large language model for text repair processing to obtain the target reference segmented text.

[0051] Specifically, after determining the initial reference segmented text without missing tables, in order to effectively improve the quality of the text, text repair processing can also be combined with a large language model. In order to improve the text repair effect of the model, the repair prompt words corresponding to the initial reference segmented text can be determined first, and the repair direction can be determined through the repair prompt words. On this basis, the repair prompt words and the initial reference segmented text can be input into the large language model for text repair processing, so that a target reference segmented text with higher quality and more accurate information can be obtained according to the repair result.

[0052] Among them, the repair prompt words specifically refer to the prompt words used to clearly repair the initial reference segmented text, including but not limited to header repair prompt words (adding missing headers of the table), formula repair prompt words (adjusting incorrect letters or calculation relationships in the formula), character cell repair prompt words (correcting typos or deleting redundant character cells), etc. Correspondingly, the large language model can be any model with text repair capabilities, and this embodiment does not make any limitations here. Correspondingly, the target reference segmented text specifically refers to a segmented text with higher quality obtained after repairing the initial reference segmented text.

[0053] Further, after obtaining the initially benchmarked segmented text with complete information, in order to improve the text quality and ensure that the content is error-free while maintaining its complete information, a large language model can be used for text repair processing. In this embodiment, the specific implementation method is as follows: Determine the table prompt words corresponding to the initially benchmarked segmented text, input the table prompt words and the initially benchmarked segmented text into the large language model for table repair processing to obtain an intermediate benchmarked segmented text; determine the formula prompt words corresponding to the intermediate benchmarked segmented text, and input the formula prompt words and the intermediate benchmarked segmented text into the large language model for formula repair processing to obtain the target benchmarked segmented text.

[0054] Specifically, the table prompt words specifically refer to the prompt words that enable the model to repair the tables in the initially benchmarked segmented text, such as adding headers for the tables. Correspondingly, the formula prompt words specifically refer to the prompt words that enable the model to repair the formulas in the initially benchmarked segmented text, such as correcting incorrect letters in the formulas.

[0055] Based on this, when using the large language model for the repair processing of the initially benchmarked segmented text, the table prompt words corresponding to the initially benchmarked segmented text can be determined first, and the table prompt words and the initially benchmarked segmented text can be input into the large language model for table repair processing to complete the repair of the tables in the text by using the large language model, such as adding missing headers, and then obtain the intermediate benchmarked segmented text; thereafter, the formula prompt words corresponding to the intermediate benchmarked segmented text can be determined, and the formula prompt words and the intermediate benchmarked segmented text can be input into the large language model for formula repair processing to complete the repair of the formulas in the text by using the large language model, such as unifying letters, and then obtain the target benchmarked segmented text.

[0056] Continuing with the above example, after deleting the segmented text of the missing table, the standard segmented text without the missing table will be obtained at this time. On this basis, considering that the headers of the tables in the test questions may be lost during text format conversion, and the table format often uses tabulation characters such as "\t" or "|", which do not conform to the standard table format. To solve this problem, the standard segmented text corresponding to each test question can be repaired by using the large language model (such as GPT4) combined with the prompt words for repairing tables, and the table format can be unified into Markdown. At this time, the standard segmented text with the table repair completed can be obtained.

[0057] Furthermore, considering that formula format errors may occur during text format conversion in test questions, such as incorrect superscripts and subscripts. And the formulas do not adopt the standard format. To solve this problem, the standard segmented text corresponding to each test question can be formula-fixed through a large language model combined with prompt words for formula repair, and the formula format can be unified into latex. At this time, the standard segmented text with the table repaired and the formula repaired can be obtained.

[0058] In summary, by performing table repair and formula repair on the initial benchmark segmented text, the segmented text can not only be informationally complete but also ensure content accuracy, making it more convenient for subsequent construction of high-quality text for model optimization tasks.

[0059] Step S208, store the target benchmark segmented text into the target sample set, where the target sample set is used to perform model optimization tasks.

[0060] Specifically, after the above-mentioned repair of the initial benchmark segmented text to obtain the target benchmark segmented text, the target benchmark segmented text can be stored into the target sample set, realizing the storage of high-quality and informationally complete samples through the target sample set, so that it is convenient for downstream tasks to use the target sample set to perform model optimization tasks, such as training or evaluating the model. Among them, the target sample set specifically refers to a set composed of a large number of target benchmark segmented texts, and this set can be used to perform model optimization tasks, such as training the model or evaluating the model.

[0061] Further, after obtaining the target benchmark segmented text, in order to ensure that the repaired text is consistent with the original segmented text in content, it can be detected before adding it to the set and optimized according to the detection results. In this embodiment, the specific implementation method is as follows: Determine the original segmented text corresponding to the target benchmark segmented text among the multiple segmented texts, and compare the original segmented text with the target benchmark segmented text; when it is determined according to the comparison result that the target benchmark segmented text does not meet the storage conditions, optimize the target benchmark segmented text using the original segmented text; store the optimized target benchmark segmented text into the target sample set.

[0062] Specifically, the original segmented text specifically refers to the segmented text corresponding to the target benchmark segmented text before repair processing after the text format file is segmented. Correspondingly, the storage condition specifically refers to the detection condition for detecting whether the target benchmark segmented text is the same as the original segmented text. Correspondingly, optimization specifically refers to processing operations such as deleting redundant information and modifying questions for the target benchmark segmented text.

[0063] Based on this, after obtaining the target reference segmentation text using the large language model, in order to avoid redundant information generated by the repair process for this problem, such as the formula not being in a simple form, loss of question information, etc., the original segmentation text corresponding to the target reference segmentation text can be determined among multiple segmentation texts, and the original segmentation text can be compared with the target reference segmentation text; when it is determined that the target reference segmentation text does not meet the storage conditions according to the comparison result, it indicates that there is content that can be optimized in the target reference segmentation text. Therefore, the target reference segmentation text can be optimized using the original segmentation text to complete operations such as deleting redundant information and supplementing questions in the target reference segmentation text. Finally, the optimized target reference segmentation text can be stored in the target sample set.

[0064] In addition, in addition to the above optimizations, optimization operations that meet the scenario requirements can also be performed. For example, the questions in the CFA exam are in English. After obtaining the target reference segmentation text, Chinese characters in the questions can also be deleted, and extra line breaks can be cleared. Specifically in implementation, different adjustment rules can be set for different scenarios, and this embodiment does not make any limitations here.

[0065] In summary, by further optimizing the target reference segmentation text, the results output by the model can be further corrected to supplement or adjust the areas where the model prediction accuracy is insufficient, so that the quality of the text finally stored in the set is higher.

[0066] In addition, after the construction of the target sample set is completed, the model with specified prediction capabilities can be trained using this sample set to meet the business deployment requirements. In this embodiment, the specific implementation method is as follows: Extract the target sample text from the target sample set in response to the model training request; input the target sample text into the language model to be trained for processing to obtain the prediction text; optimize the language model to be trained based on the sample label corresponding to the target sample text and the prediction text until the target language model that meets the training stop conditions is obtained.

[0067] Specifically, the model training request specifically refers to the request triggered when training the language model to be trained. Correspondingly, the language model to be trained specifically refers to the model associated with the target sample field, which is used to process information or generate text in this field. Correspondingly, the prediction text specifically refers to the prediction result obtained by using the language model to be trained to process the target sample text. The training stop conditions specifically refer to the conditions for stopping the training of the language model to be trained, including but not limited to the loss value comparison condition, the validation set verification condition, or the iteration number condition. In practical applications, it can be set according to actual needs, and this embodiment does not make any limitations here.

[0068] Based on this, after obtaining a target sample set with higher quality, the set can be used for model training. Specifically, target sample texts can be extracted from the target sample set in response to a model training request; the target sample texts are input into the language model to be trained for processing to obtain predicted texts; at this time, the language model to be trained can be optimized based on the sample labels corresponding to the target sample texts and the predicted texts, and it is detected whether the optimized model meets the training stop condition. If not, new samples can be selected to continue training until a target language model that meets the training stop condition is obtained and deployed in the business scenario for use.

[0069] Continuing with the above example, after completing the formula repair and table repair processing, high-quality standard segmented texts will be obtained. Thereafter, in combination with the set rules, it can be detected whether there are un-simplified formulas or Chinese characters in the standard segmented texts. Furthermore, the standard segmented texts corresponding to the test questions can be processed into sample texts that meet the model training requirements according to the set rules and stored in the training set. When training a large language model applied to the CFA exam field, samples can be randomly collected from the training set to train the model, so that the trained model can have strong professional knowledge in the CFA exam field. For example, for marking test papers and providing answer analysis, etc.

[0070] For the text processing method provided in this embodiment, in order to improve the parsing accuracy of the file to be processed and avoid losing content and affecting the subsequent model optimization task, the file in the to-be-processed format can be first converted into a text format file. On this basis, the text format file is segmented according to a preset delimiter to obtain multiple segmented texts; then, the number of table delimiters included in each segmented text and the number of tables corresponding to each segmented text can be determined, and the lost table texts are deleted from the multiple segmented texts according to the number of table delimiters and the number of tables, so as to remove the segmented texts that cannot be repaired and only retain the initial benchmark segmented texts with complete information; on this basis, considering that the initial benchmark segmented texts may have problems such as formula errors and missing table headers, and such problems do not affect the substantial content of the text but will affect the accuracy of the text content. Therefore, in order to obtain high-quality texts for subsequent use, the repair prompt words corresponding to the initial benchmark segmented texts can be determined, and the repair prompt words and the initial benchmark segmented texts are input into the large language model for text repair processing to obtain the target benchmark segmented texts; the repair of the initial benchmark segmented texts is completed by using the large language model. Thereafter, the target benchmark segmented texts can be stored in the target sample set, so that the model optimization task can be subsequently executed according to the target sample set, thereby ensuring that the model optimization task can use high-quality texts to complete training or evaluation to improve the prediction accuracy of the model.

[0071] See Figure 3 , Figure 3The flowchart of a model training method provided according to an embodiment of this specification is shown, which specifically includes the following steps.

[0072] Step S302, extract the target sample text from the target sample set in response to a model training request, where the target sample set is determined by the above method; Step S304, input the target sample text into the language model to be trained for processing to obtain a predicted text; Step S306, optimize the language model to be trained based on the sample label corresponding to the target sample text and the predicted text until a target language model that meets the training stop condition is obtained.

[0073] The above is a schematic solution of a model training method of this embodiment. It should be noted that the technical solution of this model training method and the technical solution of the above text processing method belong to the same concept. For the details not described in detail in the technical solution of the model training method, reference can be made to the description of the technical solution of the above text processing method.

[0074] The following combines the attached Figure 4 , taking the application of the text processing method provided in this specification in the PDF file processing scenario as an example, to further illustrate the text processing method. Among them, Figure 4 The processing process flowchart of a text processing method provided according to an embodiment of this specification is shown, which specifically includes the following steps.

[0075] Step S402, convert the file in the format to be processed into a text format file, and perform delimiter detection on the text format file, and determine the target delimiter included in the text format file according to the detection result.

[0076] Step S404, use the target delimiter as the preset delimiter, and split the text format file according to the preset delimiter to obtain a plurality of split texts.

[0077] Step S406, determine the number of table delimiters included in each split text and the number of tables corresponding to each split text, and delete the missing table text in the plurality of split texts according to the number of table delimiters and the number of tables to obtain an initial reference split text.

[0078] Among them, the determination of the number of tab characters included in any one of the multiple segmented texts includes: detecting whether a preset tab character is included in the first segmented text; if so, counting the tab characters included in the first segmented text, and determining the number of tab characters included in the first segmented text according to the statistical result; among them, the determination of the number of tables corresponding to any one of the multiple segmented texts includes: determining table keywords in the second segmented text, determining the numerical value corresponding to the table keywords, and taking the numerical value corresponding to the table keywords as the number of tables corresponding to the second segmented text.

[0079] Step S408: Determine the table prompt word corresponding to the initial reference segmented text, input the table prompt word and the initial reference segmented text into the large language model for table repair processing, and obtain the intermediate reference segmented text.

[0080] Step S410: Determine the formula prompt word corresponding to the intermediate reference segmented text, input the formula prompt word and the intermediate reference segmented text into the large language model for formula repair processing, and obtain the target reference segmented text.

[0081] Step S412: Determine the original segmented text corresponding to the target reference segmented text among the multiple segmented texts, and compare the original segmented text with the target reference segmented text.

[0082] Step S414: When it is determined that the target reference segmented text does not meet the storage condition according to the comparison result, optimize the target reference segmented text by using the original segmented text.

[0083] Step S416: Store the optimized target reference segmented text in the target sample set.

[0084] Step S418: When a training request for the to-be-trained language model is received, extract the target sample text from the target sample set according to the training request.

[0085] Step S420: Input the target sample text into the to-be-trained language model for processing to obtain a prediction text.

[0086] Step S422: Optimize the to-be-trained language model based on the sample label corresponding to the target sample text and the prediction text until a target language model that meets the training stop condition is obtained.

[0087] The text processing method provided in this embodiment can improve the parsing accuracy of the file to be processed and avoid losing content, which may affect the subsequent model optimization tasks. First, the file in the to-be-processed format can be converted into a text format file. On this basis, the text format file can be segmented according to a preset delimiter to obtain multiple segmented texts. Then, the number of table symbols included in each segmented text and the number of tables corresponding to each segmented text can be determined. According to the number of table symbols and the number of tables, the lost table texts can be deleted from the multiple segmented texts, so as to remove the segmented texts that cannot be repaired and only retain the initial benchmark segmented texts with complete information. On this basis, considering that the initial benchmark segmented texts may have problems such as formula errors and missing headers, although these problems do not affect the essential content of the text, they will affect the accuracy of the text content. Therefore, in order to obtain high-quality text for subsequent use, the repair prompt words corresponding to the initial benchmark segmented texts can be determined, and the repair prompt words and the initial benchmark segmented texts can be input into the large language model for text repair processing to obtain the target benchmark segmented texts, thus realizing the repair of the initial benchmark segmented texts by using the large language model. After that, the target benchmark segmented texts can be stored in the target sample set, so that the model optimization task can be executed according to the target sample set subsequently, thereby ensuring that the model optimization task can be completed for training or evaluation using high-quality text to improve the prediction accuracy of the model.

[0088] Corresponding to the above method embodiment, this specification also provides an embodiment of a text processing device. Figure 5 The following shows a schematic structural diagram of a text processing device provided in an embodiment of this specification. As Figure 5 shown, the device includes: A conversion module 502, configured to convert a file in the to-be-processed format into a text format file, and segment the text format file according to a preset delimiter to obtain multiple segmented texts; A determination module 504, configured to determine the number of table symbols included in each segmented text and the number of tables corresponding to each segmented text, and delete the lost table texts from the multiple segmented texts according to the number of table symbols and the number of tables to obtain the initial benchmark segmented texts; A repair module 506, configured to determine the repair prompt words corresponding to the initial benchmark segmented texts, and input the repair prompt words and the initial benchmark segmented texts into the large language model for text repair processing to obtain the target benchmark segmented texts; A storage module 508, configured to store the target benchmark segmented texts in the target sample set, where the target sample set is used to execute the model optimization task.

[0089] In an optional embodiment, the conversion module 502 is further configured to: Perform delimiter detection on the text format file, and determine the target delimiter included in the text format file according to the detection result; use the target delimiter as the preset delimiter, and perform the step of splitting the text format file according to the preset delimiter to obtain multiple split texts.

[0090] In an optional embodiment, the determination of the number of table delimiters included in any one of the multiple split texts includes: Detect whether the first split text contains a preset table delimiter; if so, count the table delimiters included in the first split text, and determine the number of table delimiters included in the first split text according to the statistical result; wherein, the determination of the number of tables corresponding to any one of the multiple split texts includes: determining a table keyword in the second split text, determining the value corresponding to the table keyword, and using the value corresponding to the table keyword as the number of tables corresponding to the second split text.

[0091] In an optional embodiment, the determination module 504 is further configured to: Determine the i-th split text among the multiple split texts, and detect whether the number of table delimiters in the i-th split text is less than the number of tables in the i-th split text, where i starts from 1 and i is a positive integer; if so, use the i-th split text as the missing table text; i increments sequentially until i increments to n, delete the missing table text from the multiple split texts, and use the split texts other than the missing table text in the multiple split texts as the initial reference split texts, where n is the number of texts in the multiple split texts.

[0092] In an optional embodiment, the repair module 506 is further configured to: Determine the table prompt word corresponding to the initial reference split text, input the table prompt word and the initial reference split text into a large language model for table repair processing to obtain an intermediate reference split text; determine the formula prompt word corresponding to the intermediate reference split text, input the formula prompt word and the intermediate reference split text into a large language model for formula repair processing to obtain a target reference split text.

[0093] In an optional embodiment, the device further includes: An optimization module, configured to determine the original split text corresponding to the target reference split text among the multiple split texts, and compare the original split text with the target reference split text; when it is determined according to the comparison result that the target reference split text does not meet the storage condition, optimize the target reference split text by using the original split text; store the optimized target reference split text in the target sample set.

[0094] In an alternative embodiment, the apparatus further includes: A model training module, configured to extract target sample texts from the target sample set in response to a model training request; input the target sample texts into a language model to be trained for processing to obtain predicted texts; and optimize the language model to be trained based on the sample labels corresponding to the target sample texts and the predicted texts until a target language model that meets the training stop condition is obtained.

[0095] For the text processing apparatus provided in this embodiment, in order to improve the parsing accuracy of the file to be processed and avoid losing content and affecting subsequent model optimization tasks, the file in the to-be-processed format can be first converted into a text format file. On this basis, the text format file is segmented according to a preset delimiter to obtain multiple segmented texts. Then, the number of table delimiters included in each segmented text and the number of tables corresponding to each segmented text can be determined, and the segmented texts with lost tables are deleted from the multiple segmented texts to achieve eliminating the segmented texts that cannot be repaired, and only retaining the initial reference segmented texts with complete information. On this basis, considering that the initial reference segmented texts may have problems such as formula errors and missing table headers, and such problems do not affect the substantial content of the text, but will affect the accuracy of the text content. Therefore, in order to obtain high-quality text for subsequent use, the repair prompt words corresponding to the initial reference segmented texts can be determined, and the repair prompt words and the initial reference segmented texts are input into a large language model for text repair processing to obtain target reference segmented texts; realizing the repair of the initial reference segmented texts by using the large language model. After that, the target reference segmented texts can be stored in the target sample set, so that the model optimization task can be executed according to the target sample set subsequently, thereby ensuring that the model optimization task can be completed for training or evaluation using high-quality text to improve the prediction accuracy of the model.

[0096] The above is a schematic solution of a text processing apparatus according to this embodiment. It should be noted that the technical solution of this text processing apparatus and the technical solution of the above text processing method belong to the same concept. For the details not described in the technical solution of the text processing apparatus, reference can be made to the description of the technical solution of the above text processing method.

[0097] Corresponding to the above method embodiment, this specification also provides an embodiment of a model training apparatus. Figure 6 The structure diagram of a model training apparatus provided by an embodiment of this specification is shown. As Figure 6 shown, the apparatus includes: An extraction module 602, configured to extract target sample text from a target sample set in response to a model training request, where the target sample set is determined by the above method; A processing module 604, configured to input the target sample text into a language model to be trained for processing to obtain predicted text; An optimization module 606, configured to optimize the language model to be trained based on the sample label corresponding to the target sample text and the predicted text until a target language model that meets the training stop condition is obtained.

[0098] The above is a schematic solution of a model training device according to this embodiment. It should be noted that the technical solution of this model training device and the technical solution of the above model training method belong to the same concept. For the details not described in detail in the technical solution of the model training device, reference can be made to the description of the technical solution of the above model training method.

[0099] Figure 7 The structural block diagram of a computing device 700 according to an embodiment of this specification is shown. The components of the computing device 700 include, but are not limited to, a memory 710 and a processor 720. The processor 720 is connected to the memory 710 through a bus 730, and a database 750 is used to store data.

[0100] The computing device 700 further includes an access device 740, and the access device 740 enables the computing device 700 to communicate via one or more networks 760. Examples of these networks include a Public Switched Telephone Network (PSTN), a Local Area Network (LAN), a Wide Area Network (WAN), a Personal Area Network (PAN), or a combination of communication networks such as the Internet. The access device 740 may include one or more of any type of wired or wireless network interfaces (for example, a network interface card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC).

[0101] In one embodiment of this specification, the above components of the computing device 700 and Figure 7 other components not shown therein may also be connected to each other, for example, via a bus. It should be understood that Figure 7 the block diagram of the computing device shown is for illustrative purposes only and is not a limitation on the scope of this specification. Those skilled in the art can add or replace other components as needed.

[0102] The computing device 700 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs, Personal Computers). The computing device 700 can also be a mobile or stationary server.

[0103] Among them, the processor 720 is used to execute the following computer-executable instructions, and when the computer-executable instructions are executed by the processor, the steps of the above text processing method or model training method are implemented.

[0104] The above is a schematic solution of a computing device in this embodiment. It should be noted that the technical solution of this computing device and the technical solutions of the above text processing method or model training method belong to the same concept. For the details not described in detail in the technical solution of the computing device, reference can be made to the descriptions of the technical solutions of the above text processing method or model training method.

[0105] One embodiment of this specification also provides a computer-readable storage medium, which stores computer-executable instructions, and when the computer-executable instructions are executed by the processor, the steps of the above text processing method or model training method are implemented.

[0106] The above is a schematic solution of a computer-readable storage medium in this embodiment. It should be noted that the technical solution of this storage medium and the technical solutions of the above text processing method or model training method belong to the same concept. For the details not described in detail in the technical solution of the storage medium, reference can be made to the descriptions of the technical solutions of the above text processing method or model training method.

[0107] One embodiment of this specification also provides a computer program product, including a computer program or instructions, and when the computer program or instructions are executed by the processor, the steps of the above text processing method or model training method are implemented.

[0108] The above is a schematic solution of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product and the technical solutions of the above text processing method or model training method belong to the same concept. For the details not described in the technical solution of the computer program product, reference can be made to the descriptions of the technical solutions of the above text processing method or model training method.

[0109] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the drawings do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0110] The computer instructions include computer program code, which may be in the form of source code, object code, executable files, or some intermediate forms, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0111] It should be noted that for the foregoing method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of this specification are not limited by the described order of actions, because according to the embodiments of this specification, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments of this specification.

[0112] In the above embodiments, the descriptions of each embodiment have their own focuses. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0113] The preferred embodiments of the present specification disclosed above are only used to help explain the present specification. The alternative embodiments do not describe all the details in detail, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and changes can be made according to the content of the embodiments of the present specification. These embodiments are selected and specifically described in the present specification to better explain the principles and practical applications of the embodiments of the present specification, so that those skilled in the art can well understand and utilize the present specification.

Claims

1. A text processing method, characterized in that, Including: Convert the file in the format to be processed into a text format file, and split the text format file according to a preset delimiter to obtain multiple split texts; Determine the number of tab characters included in each split text and the number of tables corresponding to each split text, and delete the missing table texts from the multiple split texts according to the number of tab characters and the number of tables to obtain an initial reference split text; Determine the repair prompt words corresponding to the initial reference split text, and input the repair prompt words and the initial reference split text into a large language model for text repair processing to obtain a target reference split text; Store the target reference split text in a target sample set, where the target sample set is used to perform a model optimization task.

2. The text processing method according to claim 1, characterized in that, The step of splitting the text format file according to a preset delimiter to obtain multiple split texts includes: Detect the delimiter of the text format file, and determine the target delimiter included in the text format file according to the detection result; Use the target delimiter as the preset delimiter, and perform the step of splitting the text format file according to the preset delimiter to obtain multiple split texts.

3. The text processing method according to claim 1, wherein The determination of the number of tab characters included in any one of the multiple split texts includes: Detect whether a preset tab character is included in the first split text; If so, count the tab characters included in the first split text, and determine the number of tab characters included in the first split text according to the statistical result; Among them, the determination of the number of tables corresponding to any one of the multiple split texts includes: Determine the table keyword in the second split text, determine the value corresponding to the table keyword, and use the value corresponding to the table keyword as the number of tables corresponding to the second split text.

4. The text processing method according to claim 1, characterized in that The step of deleting the missing table texts from the multiple split texts according to the number of tab characters and the number of tables to obtain an initial reference split text includes: Determine the i-th split text in the multiple split texts, and detect whether the number of tab characters in the i-th split text is less than the number of tables in the i-th split text, where i starts from 1 and i is a positive integer; If so, use the i-th split text as the missing table text; Increment i in sequence until i increments to n. Delete the missing table texts from the multiple split texts, and use the split texts other than the missing table texts in the multiple split texts as the initial reference split text, where n is the number of texts in the multiple split texts.

5. The text processing method according to claim 1, wherein The step of determining the repair prompt words corresponding to the initial reference split text, and inputting the repair prompt words and the initial reference split text into a large language model for text repair processing to obtain a target reference split text includes: Determine the table prompt words corresponding to the initial reference split text, and input the table prompt words and the initial reference split text into a large language model for table repair processing to obtain an intermediate reference split text; Determine the formula prompt word corresponding to the intermediate reference segmentation text, and input the formula prompt word and the intermediate reference segmentation text into a large language model for formula repair processing to obtain the target reference segmentation text.

6. The text processing method according to claim 1, characterized in that After the step of inputting the repair prompt word and the initial reference segmentation text into the large language model for text repair processing to obtain the target reference segmentation text, the following is further included: Determine the original segmentation text corresponding to the target reference segmentation text in the multiple segmentation texts, and compare the original segmentation text with the target reference segmentation text; When it is determined according to the comparison result that the target reference segmentation text does not meet the storage condition, optimize the target reference segmentation text by using the original segmentation text; Store the optimized target reference segmentation text into the target sample set.

7. The text processing method according to any one of claims 1 to 6, characterized in that After the step of storing the target reference segmentation text into the target sample set, the following is further included: Extract the target sample text from the target sample set in response to a model training request; Input the target sample text into the language model to be trained for processing to obtain a prediction text; Optimize the language model to be trained based on the sample label corresponding to the target sample text and the prediction text until a target language model that meets the training stop condition is obtained.

8. A model training method, characterized in that, Including: Extract the target sample text from the target sample set in response to a model training request, where the target sample set is determined by the method according to any one of claims 1 to 7; Input the target sample text into the language model to be trained for processing to obtain a prediction text; Optimize the language model to be trained based on the sample label corresponding to the target sample text and the prediction text until a target language model that meets the training stop condition is obtained.

9. A text processing device, characterized in that, Including: A conversion module configured to convert a file in a to-be-processed format into a text format file, and segment the text format file according to a preset delimiter to obtain multiple segmentation texts; A determination module configured to determine the number of table delimiters included in each segmentation text and the number of tables corresponding to each segmentation text, and delete the missing table text in the multiple segmentation texts according to the number of table delimiters and the number of tables to obtain the initial reference segmentation text; A repair module configured to determine the repair prompt word corresponding to the initial reference segmentation text, and input the repair prompt word and the initial reference segmentation text into a large language model for text repair processing to obtain the target reference segmentation text; A storage module configured to store the target reference segmentation text into the target sample set, where the target sample set is used to perform a model optimization task.

10. A model training device, characterized in that, Including: An extraction module configured to extract the target sample text from the target sample set in response to a model training request, where the target sample set is determined by the method according to any one of claims 1 to 7; A processing module configured to input the target sample text into the language model to be trained for processing to obtain a prediction text; An optimization module, configured to optimize the to-be-trained language model based on the sample label corresponding to the target sample text and the prediction text until a target language model that meets the training stop condition is obtained.

11. A computing device, characterized in that, Comprising: A memory and a processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the method according to any one of claims 1 to 8 are implemented.

12. A computer-readable storage medium, characterized in that, It stores computer-executable instructions, and when the computer-executable instructions are executed by the processor, the steps of the method according to any one of claims 1 to 8 are implemented.

13. A computer program product, characterized in that, Comprising a computer program or instruction, and when the computer program or instruction is executed by the processor, the steps of the method according to any one of claims 1 to 8 are implemented.