Automatic large language model fine tuning sample generation method based on low-carbon energy text
By converting and cleaning the low-carbon energy text format, generating a training sample set and training the base large language model, the problem of lack of training samples in the low-carbon energy field is solved, and the accuracy and efficiency of the model are improved.
Patent Information
- Application Number
- CN202510778412.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-09-12
AI Technical Summary
The large language model based on the low-carbon energy sector suffers from poor accuracy due to a lack of training samples, and in particular suffers from serious "hallucination" problems when dealing with vertical fields.
By obtaining the low-carbon energy text in portable file format, converting the format into MD format, and using the initial base large language model for cleaning and sample division, a training sample set is generated, and finally the initial base large language model is trained to generate the target base large language model.
It improves the accuracy and data volume of training samples, reduces time and financial costs, improves the training accuracy and efficiency of the base large language model, and reduces the impact of dirty data.
Smart Images

Figure CN120633714A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of large language models, and in particular to a method for generating automatic large language model fine-tuning samples based on low-carbon energy text. Background Art
[0002] With the rapid development of large language models in recent years, they have become ubiquitous across various industries. While the pedestal large language model has excellent general capabilities, it suffers from severe "hallucination" issues in some vertical fields, particularly those with proprietary text and knowledge, as this data has not been used for large model training. In particular, in the low-carbon energy sector, the pedestal large language model's accuracy is poor due to a lack of training samples. Summary of the Invention
[0003] The present application aims to solve one of the technical problems in the related art at least to a certain extent.
[0004] To this end, the first purpose of this application is to propose an automated large language model fine-tuning sample generation method based on low-carbon energy text, so as to improve the accuracy of training sample acquisition and improve the accuracy of base large language model training.
[0005] The second purpose of this application is to propose an automated large language model fine-tuning sample generation device based on low-carbon energy text.
[0006] The third objective of this application is to provide a server.
[0007] The fourth object of this application is to provide a computer-readable storage medium.
[0008] A fifth object of this application is to provide a computer program product.
[0009] To achieve the above objectives, the first embodiment of the present application proposes a method for generating automatic large language model fine-tuning samples based on low-carbon energy text, including:
[0010] Get the low-carbon energy text in Portable Document Format (PDF);
[0011] Performing format conversion on the low-carbon energy text in the portable document format to obtain the low-carbon energy text in the MD format;
[0012] Inputting the MD format text into the initial base large language model for cleaning to obtain the cleaned training text;
[0013] Dividing the cleaned training text into samples to obtain a divided text set, and adding the divided text set to the training sample set;
[0014] The initial base large language model is trained using the training sample set to obtain a target base large language model.
[0015] According to some embodiments, inputting the MD format text into the initial base large language model for cleaning to obtain the cleaned training text includes:
[0016] Recognizing the text in the MD format using an initial base large language model to obtain text in the MD format that does not meet the text requirements;
[0017] The initial base large language model is used to clean the text in the MD format that does not meet the text requirements to obtain cleaned training text.
[0018] According to some embodiments, the method further comprises:
[0019] Recognizing the text in the MD format using an initial base large language model, determining that there is no text in the MD format that does not meet the text requirements, and outputting the text in the MD format;
[0020] The text in the MD format is used as the cleaned training text.
[0021] According to some embodiments, converting the low-carbon energy text in the portable document format to obtain the low-carbon energy text in the MD format includes:
[0022] Obtaining text demand information corresponding to the low-carbon energy text in the portable document format;
[0023] A format converter is used to convert the low-carbon energy text in the portable file format, and when the text requirement information includes text extraction information, optical character recognition (OCR) technology is used to extract the low-carbon energy text in the portable file format to obtain the low-carbon energy text in MD format.
[0024] According to some embodiments, dividing the cleaned training text into samples to obtain a divided text set includes:
[0025] The cleaned training text is recognized by using semantic recognition technology, and the cleaned training text is divided into samples to obtain a divided text set.
[0026] According to some embodiments, adding the divided text set to a training sample set includes:
[0027] According to the segmented text set, a second segmented text is added to the first segmented text, and when it is determined that the total memory size corresponding to the first segmented text and the second segmented text is less than the video memory size, the first segmented text and the second segmented text are used as a training sample, wherein the first segmented text is the text with the highest text order in at least one segmented text, the at least one segmented text is a text in the segmented text set that has not been added to the training sample set, and the second segmented text is the next segmented text adjacent to the first segmented text;
[0028] The first segmented text and the second segmented text are merged, and the merged segmented text is added to the training sample set.
[0029] According to some embodiments, adding the second divided text to the first divided text includes:
[0030] When it is determined that the total memory size corresponding to the first divided text and the second divided text is larger than the video memory size, the first divided text is used as a training sample and added to the training sample set.
[0031] To achieve the above objectives, the second embodiment of the present application proposes an automatic large language model fine-tuning sample generation device based on low-carbon energy text, comprising:
[0032] A text acquisition unit, used for acquiring low-carbon energy text in a portable document format;
[0033] The text acquisition unit is further used to convert the low-carbon energy text in the portable document format to obtain the low-carbon energy text in the MD format;
[0034] The text acquisition unit is further configured to input the MD format text into the initial base large language model for cleaning, and obtain the cleaned training text;
[0035] A set acquisition unit, configured to divide the cleaned training text into samples, obtain a divided text set, and add the divided text set to a training sample set;
[0036] The model acquisition unit is used to train the initial base large language model using the training sample set to obtain the target base large language model.
[0037] To achieve the above-mentioned purpose, a third embodiment of the present application provides a server, comprising: a processor, and a memory communicatively connected to the processor;
[0038] The memory stores computer-executable instructions;
[0039] The processor executes the computer-executable instructions stored in the memory to implement the method as described in the first aspect above.
[0040] To achieve the above-mentioned purpose, the fourth embodiment of the present application proposes a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, they are used to implement the method described in the first aspect above.
[0041] To achieve the above-mentioned purpose, the fifth embodiment of the present application proposes a computer program product, including a computer program, which implements the method described in the first aspect when executed by a processor.
[0042] The present application provides an automated large language model fine-tuning sample generation method based on low-carbon energy text, which comprises the following steps: obtaining low-carbon energy text in a portable file format; performing format conversion on the low-carbon energy text in the portable file format to obtain low-carbon energy text in MD format; inputting the text in MD format into an initial base large language model for cleaning to obtain cleaned training text; performing sample segmentation on the cleaned training text to obtain a segmented text set, and adding the segmented text set to a training sample set; and using the training sample set to train the initial base large language model to obtain a target base large language model. This method solves the problem of poor acquisition accuracy of the base large language model due to a lack of training samples. By performing format conversion and cleaning on the low-carbon energy text, dirty data in the low-carbon energy text can be removed. Using the base large language model for cleaning can reduce restrictions on cleaning scenarios and reduce situations where dirty data cannot be removed, thereby improving the accuracy of training sample acquisition. Without the need for manual text cleaning, the amount of training sample data can be increased while reducing time and financial costs, thereby improving the accuracy and training efficiency of the base large language model training.
[0043] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0045] Figure 1 A flowchart of a method for generating automatic large language model fine-tuning samples based on low-carbon energy text provided in an embodiment of the present application;
[0046] Figure 2 A flowchart of another method for generating automatic large language model fine-tuning samples based on low-carbon energy text provided in an embodiment of the present application;
[0047] Figure 3 A flowchart of a method for generating automatic large language model fine-tuning samples based on low-carbon energy text according to an embodiment of the present application; and
[0048] Figure 4 This is a structural diagram of an automated large language model fine-tuning sample generation device based on low-carbon energy text in an embodiment of the present application. DETAILED DESCRIPTION
[0049] The following describes in detail embodiments of the present application, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.
[0050] The following describes the method and device for generating automatic large language model fine-tuning samples based on low-carbon energy text in an embodiment of the present application with reference to the accompanying drawings.
[0051] According to some embodiments, the present application provides an automated large language model fine-tuning sample generation method based on low-carbon energy text, which can increase the amount of training sample data while reducing time and financial costs, and can improve the accuracy and training efficiency of the base large language model training, such as Figure 1 As shown, the method for generating samples for automatic large language model fine-tuning based on low-carbon energy text includes the following steps:
[0052] Step 101, obtaining a low-carbon energy text in a portable document format;
[0053] According to some embodiments, the execution entity of the technical solution of the present embodiment may be, for example, a server. The server does not specifically refer to a fixed server. For example, when the server identifier changes, the server may also change accordingly. For example, when the server structure changes, the server may also change accordingly. The server may be, for example, a single server or a server cluster, which is not limited in the present embodiment.
[0054] In some embodiments, a portable document format (PDF) is used to indicate the format of the retrieved text. The portable document format may be an electronic file format. The portable document format is a way to share and view documents independently of applications, hardware, and operating systems.
[0055] According to some embodiments, the low-carbon energy text may be, for example, a text associated with the field of low-carbon energy. The low-carbon energy may be, for example, a type of energy, such as an energy source with low or zero emissions. The low-carbon energy text does not specifically refer to a fixed text, but the format of the low-carbon energy text is PDF format. For example, when the text content corresponding to the low-carbon energy text changes, the low-carbon energy text may also change accordingly. For example, when the text memory corresponding to the low-carbon energy text changes, the low-carbon energy text may also change accordingly.
[0056] In some embodiments, a low-carbon energy document in a portable document format may be obtained. The method for obtaining the low-carbon energy document is not limited. For example, the document may be obtained through a display interface or from another server via a network.
[0057] Step 102: Convert the low-carbon energy text in the portable document format to obtain a low-carbon energy text in the MD format.
[0058] In some embodiments, format conversion can be, for example, processing the format of a low-carbon energy document to obtain the required format for model training. This format conversion does not specify a fixed conversion process. For example, if the format of the converted text changes, the format conversion can also change accordingly.
[0059] In some embodiments, the MD format file is a file in Markdown format. Markdown is a lightweight markup language with concise typesetting syntax, which allows more focus on the content itself rather than the typesetting. It uses an easy-to-read and easy-to-write plain text format to write documents, which can be mixed with Hypertext Markup Language (HTML), and can export HTML, PDF, and its own .md format files. Because of its simplicity, efficiency, readability, and ease of writing, Markdown can be used by a large number of cloud-based servers. The pre-training of large model bases all uses the Markdown format.
[0060] According to some embodiments, the low-carbon energy text in MD format can be used to indicate the text after format conversion. The low-carbon energy text in MD format can be used to train the base large language model, and a base large language model that meets the model requirements can be obtained.
[0061] In some embodiments, for example, the low-carbon energy text in the portable document format may be converted to obtain the low-carbon energy text in the MD format.
[0062] Step 103: Input the MD format text into the initial base large language model for cleaning to obtain the cleaned training text;
[0063] In some embodiments, data cleaning can be performed using a text cleaner, for example. In embodiments of the present application, for example, a prompt word project can be used to transform the base language model into a text cleaner to clean the input text. The cleaning process in embodiments of the present application does not specifically refer to a fixed process. For example, if the data to be cleaned changes, the cleaning process can also change accordingly.
[0064] According to some embodiments, the base large language model refers to a model that has been pre-trained on a large amount of data, and the features and knowledge learned can serve as the basis for other tasks. Among them, the initial base large language model can be, for example, a large language model that has not been fine-tuned. The initial base large language model can be, for example, a model that has been trained or a model that has not been trained, and the embodiments of the present application do not limit this. The initial base large language model of the embodiments of the present application does not specifically refer to a fixed model. For example, when the model type corresponding to the initial base large language model changes, the initial base large language model can also change accordingly. For example, when the model parameters corresponding to the initial base large language model change, the initial base large language model can change accordingly. The base large language model can also be called a base large model, and the embodiments of the present application do not limit this.
[0065] In some embodiments, the text in the MD format may be input into an initial base large language model for cleaning to obtain cleaned training text.
[0066] The cleaned training text can be, for example, the text used to train the initial base language model. The cleaned training text does not specifically refer to a fixed text. For example, when the cleaning method changes, the cleaned training text may also change accordingly. For example, when the acquired text changes, the cleaned training text may also change accordingly.
[0067] Step 104: dividing the cleaned training text into samples, obtaining a divided text set, and adding the divided text set to the training sample set;
[0068] In some embodiments, the segmented text set can be, for example, a collection of at least one segmented text. The segmented text set is not specifically a fixed set. For example, if the cleaned training text changes or the segmentation method changes, the segmented text set may also change accordingly.
[0069] In some embodiments, a training sample set refers to a collection of training samples. The training sample set may include, for example, a partitioned text set, or other training samples. This embodiment of the present application does not limit this.
[0070] In some embodiments, the cleaned training text may be divided into samples to obtain a divided text set, and the divided text set may be added to the training sample set.
[0071] Step 105: Use the training sample set to train the initial base large language model to obtain a target base large language model.
[0072] According to some embodiments, the target base large language model can be, for example, a large language model that has been adjusted or trained. The target base large language model does not specifically refer to a fixed model. For example, when the training sample set changes, the target base large language model can also change accordingly. For example, when the training method changes, the target base large language model can also change accordingly.
[0073] In some embodiments, the training sample set may be used to train the initial base large language model to obtain a target base large language model.
[0074] The present application provides an automated large language model fine-tuning sample generation method based on low-carbon energy text, which comprises the following steps: obtaining low-carbon energy text in a portable file format; performing format conversion on the low-carbon energy text in the portable file format to obtain low-carbon energy text in MD format; inputting the text in MD format into an initial base large language model for cleaning to obtain cleaned training text; performing sample segmentation on the cleaned training text to obtain a segmented text set, and adding the segmented text set to a training sample set; and using the training sample set to train the initial base large language model to obtain a target base large language model. This method solves the problem of poor acquisition accuracy of the base large language model due to a lack of training samples. By performing format conversion and cleaning on the low-carbon energy text, dirty data in the low-carbon energy text can be removed. Using the base large language model for cleaning can reduce restrictions on cleaning scenarios and reduce situations where dirty data cannot be removed, thereby improving the accuracy of training sample acquisition. Without the need for manual text cleaning, the amount of training sample data can be increased while reducing time and financial costs, thereby improving the accuracy and training efficiency of the base large language model training. In addition, the low-carbon energy text in the portable file format can be processed to reduce the situation where the text cannot be read directly from the low-carbon energy text in the portable file format, and reduce the situation where dirty data cannot be removed, resulting in poor quality of training samples.
[0075] This embodiment provides another method for generating automatic large language model fine-tuning samples based on low-carbon energy text, such as Figure 2 As shown, the method for generating automatic large language model fine-tuning samples based on low-carbon energy text may include the following steps:
[0076] Step 201, obtaining a low-carbon energy text in a portable document format;
[0077] Among them, the relevant descriptions have been mentioned above and will not be repeated here.
[0078] In some embodiments, Figure 3 This is a flow chart of a method for generating automatic large language model fine-tuning samples based on low-carbon energy text according to an embodiment of the present application. Figure 3 As shown, the PDF text to be trained can be obtained.
[0079] Step 202: Convert the low-carbon energy text in the portable document format to obtain a low-carbon energy text in the MD format.
[0080] Among them, the relevant descriptions have been mentioned above and will not be repeated here.
[0081] According to some embodiments, converting the low-carbon energy text in the portable document format to obtain the low-carbon energy text in the MD format includes:
[0082] Obtaining text demand information corresponding to the low-carbon energy text in the portable document format;
[0083] A format converter is used to convert the low-carbon energy text in the portable document format. If the text requirement information includes text extraction information, optical character recognition (OCR) technology is used to extract text from the low-carbon energy text in the portable document format to obtain the low-carbon energy text in the MD format. Therefore, format conversion can be performed based on the text requirement information, improving the accuracy of text conversion.
[0084] According to some embodiments, Marker can, for example, quickly and accurately convert text in PDF format to text in Markdown format. The Marker can, for example, support multiple types of documents (optimized for books and scientific papers), support unlimited language types, remove headers, footers, and other miscellaneous items, format tables and code blocks, extract and save images to Markdown, convert most formulas to LaTeX, and support GPU, CPU, or MPS operation. Marker is a pipeline of deep learning models: it can extract text; determine whether to perform OCR based on text requirement information; detect page layout and find the reading order; clean and format each text block; combine text blocks and post-process the complete text; it uses models only when necessary, thereby improving speed and accuracy.
[0085] Among them, the embodiments of the present application can, for example, use PDF as a unified format for fine-tuning training samples, which can greatly improve the scope of application of the method for generating fine-tuning samples of an automated large language model based on low-carbon energy text, and there is no need to consider the algorithm compatibility issues caused by different formats of Latex and Word, which can improve the standardization of solution application and reduce the difficulty of use.
[0086] Step 203: Using the initial base large language model to recognize the text in the MD format, and obtaining text in the MD format that does not meet the text requirements;
[0087] Among them, the relevant descriptions have been mentioned above and will not be repeated here.
[0088] Step 204: Using the initial base large language model, clean the text in the MD format that does not meet the text requirements to obtain cleaned training text;
[0089] Among them, the relevant descriptions have been mentioned above and will not be repeated here.
[0090] In some embodiments, text cleaning may include, for example, cleaning dirty data in the text so that the text meets text requirements, where dirty data may include, for example, garbled characters caused by transcoding failure, paper citation format (which is fine in the context of the paper, but does not meet normal language standards), etc.
[0091] In some embodiments, regular expressions are used to clean text, wherein the specific method used for cleaning can be determined based on the cleaning requirements. For example, regular expressions can be used for cleaning when the cleaning complexity is less than the complexity threshold, and a base language model can be used for cleaning when the cleaning complexity is greater than the complexity threshold.
[0092] Among them, the situation where the cleaning complexity is less than the complexity threshold may include, for example, removing footnotes, deleting reference lists, deleting pictures and tables, etc. The situation where the cleaning complexity is greater than the complexity threshold may be, for example, a situation that cannot be solved by regular expressions. For example, there are various formats for document citations, some are "(author, year)", some are directly "(reference number)" or "[reference number]", and some are superscript numbers or superscript "[number]". Therefore, the text corresponds to the problem of citation format, and there may also be some garbled characters or the content of the header and footer may be inserted into a paragraph of text due to conversion and recognition problems. Semantic recognition needs to be introduced to solve this problem. Among them, the base large model to be trained in the embodiment of the present application is a ready-made semantic recognition model. The base large model can be turned into a text cleaner through the prompt project, the original text is input to the base large model, the cleaned text is output, and finally it is used for the training of the base large model itself. Among them, the cleaning process does not include the judgment and processing process of the text by the base large language model, that is, the cleaning process only cleans the text data, and the output is only the cleaned text, which does not include content other than the original text, such as the interpretation of a certain word in the original text.
[0093] Here, prompt can be as follows:
[0094] 1. Next, we will input a paragraph of paper text into the base language model, and we can clean this paragraph of text and output content without problems.
[0095] 2. For example, if a normal word is followed by a number, or a number is contained in small brackets, or the author's name and year are contained in brackets, this is the citation format of the paper, and the citation format can be removed.
[0096] 3. In addition, due to problems with the PDF conversion algorithm, author information (such as name, school, etc.) will be inserted into a normal paragraph. Remove this interfering content.
[0097] 4. Figure 1 and Table 1 also need to be removed.
[0098] 5. In another case, the input may be a reference or pure author information. If this is the case, it will directly return "Invalid information, no need to return".
[0099] 6. Due to problems with the conversion algorithm, there may be some semantically unclear letters or numbers, which may be residual formulas that failed to convert. The rows containing these semantically unclear fields must also be cleared.
[0100] 7. If the content contains formulas, regardless of whether it is written in the markdown format surrounded by $, it should be rewritten as the MD formula format. The letters should also be translated into the latex escape format instead of using UTF direct encoding. For example, rewrite \u2207 to $\\nabula$.
[0101] 8. Pay attention to the difference between in-line formulas and inter-line formulas, and do not use double $ for in-line formulas.
[0102] 9. Finally, if the text contains a table, delete the table and do not output it. If the entire text is a table, output "Invalid information, no need to return".
[0103] 10. Remember, your task is to clean the text data. Do not return the thinking process of the base language model or the judgment of whether there is a problem with the text. Only return the results of text cleaning.
[0104] 11. Even if the text does not contain content that needs to be cleaned, just return the original content directly without adding the judgment of the base language model (such as "no content that needs to be cleaned" and other content added by the base language model). The returned content cannot contain content other than the original text, and using brackets to explain it is not acceptable.
[0105] According to some embodiments, the method further comprises:
[0106] Recognizing the text in the MD format using an initial base large language model, determining that there is no text in the MD format that does not meet the text requirements, and outputting the text in the MD format;
[0107] The text in the MD format is used as the cleaned training text.
[0108] According to some embodiments, text cleaning can also determine whether to clean the text based on the similarity between texts. For example, if the similarity between the acquired text and any text in the training sample set is greater than a similarity threshold, the acquired text can be eliminated. Specifically, for example, an acquired text can be compared with at least one text that already meets the text requirements, and whether to add the text to the training sample set can be determined based on the similarity. Since repeated text entering the model training is equivalent to a sample being trained multiple times, it is easy to cause problems such as overfitting and training failure. Among them, semantic deduplication can use "Simhash". Simhash is a hash algorithm that can calculate document similarity. Through Simhash, a text can be mapped to 64 bits, and then the 64-bit Hamming distance of the two texts can be compared to know the similarity of the articles. If the Hamming distance k of the two texts is less than or equal to 3, it can be considered that the two texts are very similar and can be considered as duplicate texts. In actual use, for example, k=10 can be used.
[0109] Step 205: dividing the cleaned training text into samples, obtaining a divided text set, and adding the divided text set to the training sample set;
[0110] Among them, the relevant descriptions have been mentioned above and will not be repeated here.
[0111] In some embodiments, for example, the maximum number of tokens may be used to divide the text, wherein the specific text division method may be determined according to the user's instructions, for example, and the text may also be identified and determined.
[0112] According to some embodiments, dividing the cleaned training text into samples to obtain a divided text set includes:
[0113] The cleaned training text is recognized by using semantic recognition technology, and the cleaned training text is divided into samples to obtain a divided text set.
[0114] According to some embodiments, adding the divided text set to a training sample set includes:
[0115] According to the segmented text set, a second segmented text is added to the first segmented text, and when it is determined that the total memory size corresponding to the first segmented text and the second segmented text is less than the video memory size, the first segmented text and the second segmented text are used as a training sample, wherein the first segmented text is the text with the highest text order in at least one segmented text, the at least one segmented text is a text in the segmented text set that has not been added to the training sample set, and the second segmented text is the next segmented text adjacent to the first segmented text;
[0116] The first segmented text and the second segmented text are merged, and the merged segmented text is added to the training sample set.
[0117] According to some embodiments, the step of adding the second segmented text to the first segmented text based on the automated large language model fine-tuning sample of the low-carbon energy text includes:
[0118] If the total memory size corresponding to the first and second segmented texts is determined to be greater than the video memory size, the first segmented text is used as a training sample and added to the training sample set. Therefore, each training sample can be determined based on semantic information and video memory size, reducing the likelihood of incomplete semantics and inaccurate fine-tuning of the base language model, thereby improving the accuracy of base language model training.
[0119] According to some embodiments, for example, a natural paragraph represents a complete semantic segment. In papers, formulas, images, tables, and other content are often inserted into the entire paragraph, resulting in the complete semantic segment being divided into multiple segments. Therefore, if there is sufficient video memory, the entire document can be treated as a sample, and the contextual capabilities of the transformer can be used to learn the entire content. Due to server video memory limitations, it is possible to include as much text as possible within a limited maximum sequence length.
[0120] Therefore, we can recursively add the next paragraph of content to the current text. If the maximum sequence length is not exceeded, we continue adding until it is exceeded or the next sentence starts with "#" (which represents a paragraph title in Markdown, indicating that the next paragraph begins with other semantics). Then we return the result of the previous step as a sample. This can preserve the complete semantics as much as possible and utilize the semantic segmentation function of the text itself to achieve a balance between semantic integrity and video memory limitations.
[0121] According to some embodiments, a title recognition operation may be added so that the title appears at the beginning of the sample and is not added after other samples, thereby improving the accuracy of obtaining training samples.
[0122] Step 206: Use the training sample set to train the initial base large language model to obtain a target base large language model.
[0123] Among them, the relevant descriptions have been mentioned above and will not be repeated here.
[0124] In this embodiment, the initial base large language model is used to recognize the text in the MD format to obtain the text in the MD format that does not meet the text requirements; the initial base large language model is used to clean the text in the MD format that does not meet the text requirements to obtain the cleaned training text; the cleaned training text is divided into samples to obtain a divided text set, and the divided text set is added to the training sample set. Therefore, the base large language model can be used to clean the data, improve the accuracy of training sample acquisition, and improve the accuracy of base large language model training.
[0125] In order to implement the above embodiments, the present application also proposes an automated large language model fine-tuning sample generation device based on low-carbon energy text.
[0126] Figure 4 A schematic diagram of the structure of an automated large language model fine-tuning sample generation device based on low-carbon energy text provided in an embodiment of the present application.
[0127] like Figure 4 As shown, the automatic large language model fine-tuning sample generation device based on low-carbon energy text includes:
[0128] The text acquisition unit 401 is used to acquire a low-carbon energy text in a portable document format;
[0129] The text acquisition unit 401 is further configured to convert the low-carbon energy text in the portable document format to obtain the low-carbon energy text in the MD format;
[0130] The text acquisition unit 401 is further configured to input the MD format text into the initial base large language model for cleaning, thereby obtaining the cleaned training text;
[0131] A set acquisition unit 402 is configured to divide the cleaned training text into samples, obtain a divided text set, and add the divided text set to a training sample set;
[0132] The model acquisition unit 403 is configured to train the initial base large language model using the training sample set to acquire a target base large language model.
[0133] Furthermore, in a possible implementation of the embodiment of the present application, the text acquisition unit 401 is configured to input the MD format text into the initial base large language model for cleaning, and to obtain the cleaned training text, specifically for:
[0134] Recognizing the text in the MD format using an initial base large language model to obtain text in the MD format that does not meet the text requirements;
[0135] The initial base large language model is used to clean the text in the MD format that does not meet the text requirements to obtain cleaned training text.
[0136] Furthermore, in a possible implementation of the embodiment of the present application, the text acquisition unit 401 is further configured to:
[0137] Recognizing the text in the MD format using an initial base large language model, determining that there is no text in the MD format that does not meet the text requirements, and outputting the text in the MD format;
[0138] The text in the MD format is used as the cleaned training text.
[0139] Furthermore, in a possible implementation of the embodiment of the present application, the text acquisition unit 401 is configured to convert the low-carbon energy text in the portable document format to obtain the low-carbon energy text in the MD format, specifically to:
[0140] Obtaining text demand information corresponding to the low-carbon energy text in the portable document format;
[0141] A format converter is used to convert the low-carbon energy text in the portable file format, and when the text requirement information includes text extraction information, optical character recognition (OCR) technology is used to extract the low-carbon energy text in the portable file format to obtain the low-carbon energy text in MD format.
[0142] Furthermore, in a possible implementation of the embodiment of the present application, the set acquisition unit 402 is configured to divide the cleaned training text into samples, and when obtaining the divided text set, specifically to:
[0143] The cleaned training text is recognized by using semantic recognition technology, and the cleaned training text is divided into samples to obtain a divided text set.
[0144] Furthermore, in a possible implementation of the embodiment of the present application, the set acquisition unit 402, when used to add the divided text set to the training sample set, is specifically used to:
[0145] According to the segmented text set, a second segmented text is added to the first segmented text, and when it is determined that the total memory size corresponding to the first segmented text and the second segmented text is less than the video memory size, the first segmented text and the second segmented text are used as a training sample, wherein the first segmented text is the text with the highest text order in at least one segmented text, the at least one segmented text is a text in the segmented text set that has not been added to the training sample set, and the second segmented text is the next segmented text adjacent to the first segmented text;
[0146] The first segmented text and the second segmented text are merged, and the merged segmented text is added to the training sample set.
[0147] Furthermore, in a possible implementation of the embodiment of the present application, the set acquisition unit 402 is configured to add the second segmented text to the first segmented text, specifically to:
[0148] When it is determined that the total memory size corresponding to the first divided text and the second divided text is larger than the video memory size, the first divided text is used as a training sample and added to the training sample set.
[0149] It should be noted that the above explanation of the embodiment of the method for generating automatic large language model fine-tuning samples based on low-carbon energy text is also applicable to the device for generating automatic large language model fine-tuning samples based on low-carbon energy text in this embodiment, and will not be repeated here.
[0150] In the embodiment of the present application, a text acquisition unit is used to acquire a low-carbon energy text in a portable file format; the text acquisition unit is also used to convert the low-carbon energy text in the portable file format to acquire a low-carbon energy text in an MD format; the text acquisition unit is also used to input the text in the MD format into an initial base large language model for cleaning to acquire a cleaned training text; a set acquisition unit is used to divide the cleaned training text into samples, acquire a divided text set, and add the divided text set to a training sample set; a model acquisition unit is used to use the training sample set to The initial base large language model is trained to obtain the target base large language model, which solves the problem of poor acquisition accuracy of the base large language model due to the lack of training samples. By converting the format of the low-carbon energy text and cleaning it, the dirty data in the low-carbon energy text can be removed, and the use of the base large language model for cleaning can reduce the restrictions on the cleaning scenarios and reduce the situation where dirty data cannot be removed, thereby improving the accuracy of the training samples obtained, and eliminating the need for manual text cleaning. The amount of training sample data can be increased while reducing time and financial costs, thereby improving the accuracy and training efficiency of the base large language model training.
[0151] In order to implement the above embodiments, the present application also proposes a server, comprising: a processor, and a memory communicatively connected to the processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory to implement the method provided by the above embodiments.
[0152] In order to implement the above embodiments, the present application also proposes a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, they are used to implement the methods provided by the above embodiments.
[0153] In order to implement the above embodiments, the present application also proposes a computer program product, including a computer program, which implements the methods provided by the above embodiments when executed by a processor.
[0154] The collection, storage, use, processing, transmission, provision and disclosure of user personal information involved in this application are in compliance with relevant laws and regulations and do not violate public order and good morals.
[0155] It is important to note that personal information collected from users should be used for legitimate and reasonable purposes and should not be shared or sold beyond these legitimate uses. Furthermore, such collection / sharing should be conducted only after receiving the user's informed consent, including but not limited to notifying the user to read the user agreement / user notice and sign an agreement / authorization that includes the relevant user information before using the feature. Furthermore, any necessary steps must be taken to safeguard and secure access to such personal information and ensure that others with access to personal information comply with its privacy policy and procedures.
[0156] This application contemplates providing implementation options for users to selectively block the use or access of personal information data. Specifically, this application contemplates providing hardware and / or software to prevent or block access to such personal information data. Risks can be minimized by limiting data collection and deleting data once it is no longer needed. Furthermore, where applicable, such personal information can be de-identified to protect user privacy.
[0157] In the descriptions of the foregoing embodiments, the reference terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, unless they are mutually inconsistent.
[0158] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of such features. Throughout the description of this application, "plurality" means at least two, for example, two, three, etc., unless otherwise specifically defined.
[0159] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code comprising one or more executable instructions for implementing the steps of a custom logical function or process, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may be performed out of the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present application belong.
[0160] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection with one or more wires (electronic devices), a portable computer disk cartridge (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and programmable read-only memory (EPROM or flash memory), fiber optic devices, and a portable compact disc read-only memory (CDROM). Furthermore, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or processing it in another suitable manner if necessary, and then storing it in a computer memory.
[0161] It should be understood that various parts of the present application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used to implement: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0162] Those skilled in the art will understand that all or part of the steps in the method of the above embodiment can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.
[0163] In addition, the functional units in the various embodiments of the present application may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into a module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0164] The storage medium mentioned above may be a read-only memory, a magnetic disk, or an optical disk, etc. Although the embodiments of the present application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present application. Persons skilled in the art may make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.
Claims
1. A method for generating automatic large language model fine-tuning samples based on low-carbon energy text, characterized in that: include: Get the low carbon energy text in portable document format; Performing format conversion on the low-carbon energy text in the portable document format to obtain the low-carbon energy text in the MD format; Inputting the MD format text into the initial base large language model for cleaning to obtain the cleaned training text; Dividing the cleaned training text into samples to obtain a divided text set, and adding the divided text set to the training sample set; The initial base large language model is trained using the training sample set to obtain a target base large language model.
2. The method according to claim 1, characterized in that Inputting the MD format text into the initial base large language model for cleaning to obtain the cleaned training text includes: Recognizing the text in the MD format using an initial base large language model to obtain text in the MD format that does not meet the text requirements; The initial base large language model is used to clean the text in the MD format that does not meet the text requirements to obtain cleaned training text.
3. The method according to claim 2, characterized in that The method further comprises: Recognizing the text in the MD format using an initial base large language model, determining that there is no text in the MD format that does not meet the text requirements, and outputting the text in the MD format; The text in the MD format is used as the cleaned training text.
4. The method according to claim 1, wherein The step of converting the low-carbon energy text in the portable document format to obtain the low-carbon energy text in the MD format includes: Obtaining text demand information corresponding to the low-carbon energy text in the portable document format; A format converter is used to convert the low-carbon energy text in the portable file format, and when the text requirement information includes text extraction information, optical character recognition (OCR) technology is used to extract the low-carbon energy text in the portable file format to obtain the low-carbon energy text in MD format.
5. The method according to claim 1, characterized in that The step of dividing the cleaned training text into samples to obtain a divided text set includes: The cleaned training text is recognized by using semantic recognition technology, and the cleaned training text is divided into samples to obtain a divided text set.
6. The method according to claim 5, characterized in that The step of adding the divided text set to the training sample set includes: According to the segmented text set, a second segmented text is added to the first segmented text, and when it is determined that the total memory size corresponding to the first segmented text and the second segmented text is less than the video memory size, the first segmented text and the second segmented text are used as a training sample, wherein the first segmented text is the text with the highest text order in at least one segmented text, the at least one segmented text is a text in the segmented text set that has not been added to the training sample set, and the second segmented text is the next segmented text adjacent to the first segmented text; The first segmented text and the second segmented text are merged, and the merged segmented text is added to the training sample set.
7. The method according to claim 6, characterized in that The adding of the second divided text to the first divided text includes: When it is determined that the total memory size corresponding to the first divided text and the second divided text is larger than the video memory size, the first divided text is used as a training sample and added to the training sample set.
8. An automated large language model fine-tuning sample generation device based on low-carbon energy text, characterized in that: include: A text acquisition unit, used for acquiring low-carbon energy text in a portable document format; The text acquisition unit is further used to convert the low-carbon energy text in the portable document format to obtain the low-carbon energy text in the MD format; The text acquisition unit is further configured to input the MD format text into the initial base large language model for cleaning, and obtain the cleaned training text; A set acquisition unit, configured to divide the cleaned training text into samples, obtain a divided text set, and add the divided text set to a training sample set; The model acquisition unit is used to train the initial base large language model using the training sample set to obtain the target base large language model.
9. A server, characterized in that: include: a processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 7 when executed by a processor.