General text quality evaluation method based on large language model
By designing the prompts for two rounds of dialogue, using a large language model to generate text quality evaluation results with reference answers and without reference answers, the problem of poor quality evaluation results in the existing technology is solved, and a higher quality text quality evaluation is achieved.
Patent Information
- Application Number
- PCT/CN2024/131764
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-30
- Filing Date
- 2024-11-13
- Publication Date
- 2025-06-05
AI Technical Summary
When evaluating the text quality generated by large models, traditional indicators such as BLEU and ROUGE ignore content consistency and relevance, resulting in poor quality of evaluation results, especially when setting without reference text, the gap with manual evaluation is large.
The prompts for designing two rounds of dialogue are used to generate evaluation results with reference answers, and the second round is rewritten based on the first round of results to obtain evaluation results without reference answers, thereby improving the quality of evaluation data.
Through this method, the performance of the general text quality evaluation model is improved, so that it can achieve or exceed the evaluation performance of GPT-4 with the settings with reference answers and without reference answers.
Smart Images

Figure CN2024131764_05062025_PF_FP_ABST
Abstract
Description
A general text quality evaluation method based on large language model Technical Field
[0001] The present invention belongs to the technical field of large models and relates to a general text quality evaluation method, in particular to a general text quality evaluation method based on a large language model. Background Art
[0002] Large language models (LLMs), such as ChatGPT, GPT-4, and GLM, have recently developed rapidly, with their generative performance on various tasks gradually approaching human-level performance. Therefore, accurately evaluating the generative performance of large models has become a research hotspot in natural language processing. High-quality text quality evaluation methods can simultaneously provide evaluation scores and evaluation reasons. These evaluation results can serve as feedback signals to continuously optimize the generative performance of large models. Traditional evaluation metrics such as BLEU (Bilingual Evaluation Understudy) and ROUGE (Recall-Oriented Understudy for Gisting Evaluation) mostly focus on the n-gram overlap between generated and reference texts. This largely ignores content issues in the generated text, such as consistency and relevance.
[0003] Most recent research works are based on the design of evaluation methods based on pre-trained models. One series of works uses the language representation ability of pre-trained language models to calculate the similarity scores between generated text and reference text in the semantic space; another series of works uses the generation ability of pre-trained models to obtain evaluation scores through generation probability or generation results.
[0004] Due to the recent rapid development of large models such as ChatGPT and GPT-4, some researchers have transformed evaluation tasks into instruction-following tasks, then used these large models to directly generate evaluation results. When designing prompting schemes, these methods directly use ChatGPT or GPT-4 to obtain evaluation results based on input information and generated text, and then train the evaluation model based on this data. This approach results in poor evaluation data quality, especially in more difficult settings without reference text. Directly using ChatGPT or GPT-4 to generate evaluation results can lead to a significant gap between the results and manual evaluation.
[0005] Therefore, in order to address the defects in the above-mentioned prior art, it is necessary to develop a new universal text quality evaluation method.
[0006] Summary of the Invention
[0007] In order to overcome the defects of the existing technology, the present invention proposes a general text quality evaluation method based on a large language model, which obtains automatic evaluation results with and without reference answers by designing prompts for two rounds of dialogue, so that the evaluation results without reference answers can be generated by referring to the evaluation results with reference answers in the first round of dialogue, thereby improving the quality of the evaluation data and ultimately improving the performance of the general text quality evaluation model.
[0008] In order to achieve the above object, the present invention provides the following technical solutions:
[0009] A general text quality evaluation method based on a large language model, characterized by comprising the following steps:
[0010] 1) Using a large language model to build a general text quality evaluation model;
[0011] 2) Constructing training data: The input of the training data is prompt words and evaluation input, and the output is evaluation results. The prompt words include instructions, scoring rules and output formats. Constructing training data includes constructing training data with reference answers and constructing training data without reference answers.
[0012] 3) Training the general text quality evaluation model using the training data;
[0013] 4) Use the trained text quality evaluation model to evaluate the quality of general text.
[0014] Preferably, the step 2) of constructing the training data containing the reference answers specifically includes:
[0015] 2.1) Determining a first prompt word for the training data containing the reference answer, that is, determining a first instruction, a first scoring rule, and a first output format for the first prompt word;
[0016] 2.2) Obtaining a first evaluation input of the training data containing the reference answer;
[0017] 2.3) Obtain a first evaluation result of the training data containing the reference answer.
[0018] Preferably, the step 2) of constructing training data without reference answers specifically includes:
[0019] 2.4) Determine a user prompt word, so that the user prompt word is to remove all descriptions directly related to the reference answer from the first prompt word, the first evaluation input, and the first evaluation result;
[0020] 2.5) Using the user prompt word, based on the first prompt word, the first evaluation input and the first evaluation result, a second prompt word, a second evaluation input and a second evaluation result of the training data without a reference answer are generated.
[0021] Preferably, the step 2.2) of obtaining the first evaluation input of the training data containing the reference answer specifically includes:
[0022] 2.2.1) Obtaining enhanced user query instructions: collecting multiple initial user query instructions on the public network platform and processing them to obtain multiple enhanced user query instructions;
[0023] 2.2.2) Collecting and generating responses: Inputting the plurality of enhanced user query instructions into a plurality of Chinese open source large language models and API access models to generate a plurality of generated responses for each of the enhanced user query instructions;
[0024] 2.2.3) Collecting reference answers: Input the multiple enhanced user query instructions into the GPT-4 model respectively, and the GPT-4 model generates initial reference answers. Then, the annotators correct the problems in the initial reference answers to obtain the final reference answers;
[0025] 2.2.4) Based on the enhanced user query instruction, a reply and a reference answer are generated to form a first evaluation input of the training data containing the reference answer.
[0026] Preferably, the evaluation result includes evaluation reasons and evaluation scores.
[0027] Preferably, the first evaluation result of obtaining the training data containing the reference answer in step 2.3) is specifically: inputting the first prompt word and the first evaluation input into the GPT-4 model, and the GPT-4 model scores the quality of the generated response relative to the reference answer based on the first scoring rule and generates corresponding evaluation reasons, thereby obtaining the evaluation reasons and evaluation score of the first evaluation result.
[0028] Preferably, obtaining the enhanced user query instruction in step 2.2.1) specifically includes:
[0029] 2.2.1.1) Collecting multiple initial user query instructions on the public network platform and classifying them;
[0030] 2.2.1.2) Based on the multiple initial user query instructions, use ChatGPT to generate more new user query instructions with similar category distribution and different instruction content to the initial user query instructions;
[0031] 2.2.1.3) Based on the text overlap evaluation index, the new instructions asked by the user are screened;
[0032] 2.2.1.4) Perform difficulty screening on the filtered user inquiries for new instructions, so that the proportion of data on user inquiries for new instructions at different difficulty levels is equal;
[0033] 2.2.1.5) Balance the category distribution of the new user-asked instructions after the difficulty screening, so that the data proportion of the new user-asked instructions in each category is equal, thereby obtaining the multiple enhanced user-asked instructions.
[0034] Preferably, in step 2.2.1.1), the classification is divided into eight categories, namely: logical reasoning, comprehensive question and answer, professional ability, basic ability, mathematical calculation, role-playing, text writing and Chinese comprehension.
[0035] Preferably, in step 4), when using the trained text quality evaluation model to evaluate the quality of general text, two decoding strategies, greedy search and self-consistency, are adopted.
[0036] Preferably, in step 1), when a large language model is used to construct a universal text quality evaluation model, the large language model is a ChatGLM2 model with 6 billion, 12 billion or 66 billion parameters.
[0037] Compared with the prior art, the universal text quality assessment method based on a large language model of the present invention has one or more of the following beneficial technical effects:
[0038] 1. The present invention automatically constructs training data for general text quality evaluation by designing prompts, thereby training a general text quality evaluation model to provide evaluation scores and evaluation reasons for the general text generated by the large model.
[0039] 2. When constructing training data, the present invention designs a first round of prompts to obtain automatic evaluation results containing reference answers, and then designs a second round of prompts to rewrite the above evaluation results to obtain evaluation results without reference answers, and uses them to train a general text quality evaluation model, which supports the evaluation of the generated text quality under two settings: with reference answers and without reference answers.
[0040] 3. Quantitative experiments show that the general text quality evaluation model trained by the present invention can reach a level comparable to that of GPT-4 when provided with reference answers, and can surpass the evaluation performance of GPT-4 on three tasks when provided without reference answers. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] FIG1 is a flow chart of a general text quality evaluation method based on a large language model according to the present invention.
[0042] FIG2 is a schematic diagram of an exemplary process of constructing training data and using the constructed training data to train a universal text quality assessment model according to the present invention.
[0043] FIG3 shows a schematic diagram of correlation coefficients between scores generated by different evaluation models and human scores according to an exemplary embodiment of the present invention.
[0044] FIG4 is a schematic diagram showing the classification evaluation results of exemplary different evaluation models of the present invention and the GPT-4 model. DETAILED DESCRIPTION
[0045] Before describing in detail any embodiment of the present invention, it should be understood that the present invention is not limited in its application to the details of construction and arrangement of components set forth in the following description or illustrated in the following drawings. The present invention is capable of other embodiments and can be practiced or carried out in various ways. In addition, it should be understood that the words and terms used herein are for descriptive purposes and should not be considered restrictive. As used herein, "including" or "having" and variations thereof are intended to cover the items listed below and their equivalents as well as additional items.
[0046] Furthermore, in the disclosure of the present invention, the term "one" should be understood as "at least one" or "one or more", that is, in one embodiment, the number of an element may be one, while in another embodiment, the number of the elements may be multiple, and the term "one" should not be understood as a limitation on the quantity.
[0047] The present invention proposes a universal text quality evaluation method based on a large language model, which automatically constructs training data for quality evaluation by designing prompts, thereby training the evaluation model to give evaluation scores and evaluation reasons for the text generated by the large model.
[0048] FIG1 shows a flow chart of a general text quality evaluation method based on a large language model of the present invention. As shown in FIG1 , the general text quality evaluation method based on a large language model of the present invention comprises the following steps:
[0049] 1. Use a large language model to build a general text quality evaluation model.
[0050] In the present invention, various large Chinese language models can be used as general text quality evaluation models. For example, the ChatGLM2 model with 6 billion, 12 billion, or 66 billion parameters can be used as a general text quality evaluation model.
[0051] 2. Build training data.
[0052] To use the universal text quality evaluation model to evaluate the quality of universal text, the most important thing is to construct appropriate training data and use the constructed training data to train the text quality evaluation model.
[0053] In the present invention, the training data inputs are prompt words and evaluation inputs, and the output is the evaluation results. The prompt words include instructions, scoring rules, and output formats. Furthermore, constructing training data includes constructing training data with reference answers and constructing training data without reference answers. This supports evaluating the quality of generated texts in both settings, with and without reference answers.
[0054] In the present invention, when constructing training data, a dialogue-based training data construction method is adopted. The first round of prompts is designed to obtain training data containing reference answers, and then the second round of prompts is designed to rewrite the above results to obtain training data without reference answers.
[0055] Specifically, first construct training data containing reference answers, which includes the following steps:
[0056] 1. Determine the first prompt word of the training data containing the reference answer, that is, determine the first instruction, first scoring rule and first output format of the first prompt word.
[0057] The first instruction is used to instruct the large model how to generate text. For example, in Figure 2, the first instruction is "Please play the role of an impartial judge and judge the quality of the response generated by an AI assistant..."
[0058] The first scoring rule specifies the criteria for evaluating the quality of the generated responses, which not only specifies the score levels, but also specifies the reasons corresponding to the score levels. In the present invention, the evaluation scores are divided into 10 levels of 1-10, where the level of 1-2 corresponds to very confusing and seriously factually incorrect generated responses, the level of 3-4 corresponds to responses with significant deficiencies, such as grammatical correctness, coherence, etc., and the three intervals of 5-6, 7-8, and 9-10 are all responses of better quality. According to the quality of the generated responses relative to the reference answers, they are divided into three levels: inferior, equal, and better.
[0059] The first output format specifies the format for generating a response. In Figure 2, it reads "After evaluating the reasons for each place, you must output the evaluation score in the following format: "[[Rating]]"".
[0060] 2. Obtain a first evaluation input of the training data containing the reference answer.
[0061] In the present invention, the first evaluation input of the training data containing reference answers includes user instructions, reference answers and generated text, wherein the user instructions are processed enhanced user query instructions.
[0062] Therefore, obtaining the first evaluation input of the training data containing the reference answer specifically includes the following steps:
[0063] (1) Obtaining enhanced user query instructions, that is, collecting multiple initial user query instructions on the public network platform and processing them to obtain multiple enhanced user query instructions.
[0064] In the present invention, obtaining multiple enhanced user query instructions specifically includes the following steps:
[0065] First, multiple initial user query instructions on the public network platform are collected and classified.
[0066] In this paper, 706 initial user queries on a public network platform were collected and classified into categories including logical reasoning, comprehensive question-answering, professional skills, basic skills, mathematical calculations, role-playing, text writing, and Chinese comprehension.
[0067] Secondly, based on the multiple initial user inquiry instructions, ChatGPT is used to generate more new user inquiry instructions with similar category distribution to the initial user inquiry instructions but different instruction content.
[0068] In order to obtain more user inquiry instructions to improve the training effect, in the present invention, based on the collected 706 initial user inquiry instructions, 260,000 new user inquiry instructions with similar category distribution and different instruction content to these initial user inquiry instructions were generated through ChatGPT.
[0069] Secondly, based on the text overlap evaluation index, the new instructions asked by the user are screened.
[0070] In order to solve the problem of duplication among the generated new user-asked instructions, the present invention uses ROUGE-L and Self-BLEU, two automatic evaluation indicators based on text overlap, to screen the new user-asked instructions, and 4,223 instructions were screened.
[0071] Next, the difficulty of the filtered user inquiries for new instructions is screened to ensure that the proportion of data of the user inquiries for new instructions within different difficulty levels is equal.
[0072] Since the difficulty of the new instructions asked by the users varies, in order to make the training data include instructions of various difficulty levels and make the data proportion of instructions of various difficulty levels as uniform as possible, in the present invention, ChatGPT is provided with a classification standard of difficulty levels 1-3, allowing ChatGPT to classify the difficulty of the new instructions asked by the users, and then balance the data proportion of each difficulty level, and finally select 3351 instructions from 4223.
[0073] Finally, the category distribution of the new user-asked instructions after the difficulty screening is balanced, so that the data proportion of the new user-asked instructions in each category is equal, thereby obtaining the multiple enhanced user-asked instructions. As mentioned above, the instructions are divided into 8 categories. In order to improve the training effect, it is preferred that the amount of training data contained in the 8 categories is as balanced as possible. Therefore, in the present invention, given the instruction classification, ChatGPT is allowed to classify the instruction types to which the new user-asked instructions after the difficulty screening belong, and then the number of each type is balanced. Taking into account the cost and data set size factors, 1000 instructions are retained from the 3351 instructions. These 1000 instructions are the enhanced user-asked instructions obtained after processing.
[0074] (2) Collect and generate responses.
[0075] The diversity and coverage of generated responses are also crucial to the quality of evaluation data. Therefore, in the present invention, the multiple enhanced user query instructions are input into multiple Chinese open source large language models and API access models to generate multiple generated responses for each of the enhanced user query instructions.
[0076] Specifically, we selected 10 representative Chinese open-source models and API access models as generative models for generating responses: GPT-4, ChatGPT, two versions of ChatGLM, MOSS, Minimax, Sparkdesk, Chinese-Llama2-7B-Chat, Baichuan2-13B-Chat, and Ernie-Bot. Each model generated responses based on the 1,000 augmented user queries, resulting in approximately 10,000 generated data points.
[0077] (3) Collect reference answers.
[0078] In the present invention, the multiple enhanced user query instructions are respectively input into the GPT-4 model, and the GPT-4 model generates initial reference answers. Then, the annotators correct the problems in the initial reference answers to obtain the final reference answers.
[0079] When the annotators correct the questions in the initial reference answers, they correct the factual errors, logical confusions, detail contradictions and other problems in the initial reference answers they generated, and finally make them high-quality reference answers.
[0080] Because GPT-4 is powerful and comprehensive, the initial reference answers generated by it are relatively accurate, and then they are corrected through manual annotation, making the final reference answers more accurate.
[0081] (4) Based on the enhanced user inquiry instruction, a reply and a reference answer are generated to form a first evaluation input of the training data containing the reference answer.
[0082] With the enhanced user inquiry instruction, generated reply and reference answer, they can constitute the first evaluation input of the training data containing reference answers.
[0083] In the present invention, there are about 10,000 first evaluation inputs in total.
[0084] Thus, through the above steps 1 and 2, the input of training data containing reference answers can be obtained, which is combined into a prompt word template, as shown in the upper left corner of Figure 2.
[0085] 3. Obtain a first evaluation result of the training data containing the reference answer.
[0086] In the present invention, the evaluation result includes an evaluation reason and an evaluation score. Furthermore, obtaining the first evaluation result of the training data containing the reference answer specifically involves inputting both the first prompt word and the first evaluation input (i.e., the input of the training data containing the reference answer) into the GPT-4 model, and having the GPT-4 model score the quality of the generated response relative to the reference answer based on the first scoring rule and generate corresponding evaluation reasons, thereby obtaining the evaluation reason and evaluation score of the first evaluation result.
[0087] Specifically, in FIG2 , the first evaluation result of the training data containing the reference answer is “the AI assistant's answer is compared with the reference answer….score [[5]]”.
[0088] Through steps 1-3 above, we have completed the first round of training data collection with reference answers. The next step is how to collect training data without reference answers.
[0089] In the present invention, the first prompt word, the first evaluation input and the first evaluation result are used as the content of the first round of dialogue. On this basis, the prompt word for the second round of dialogue is designed, that is, the user prompt word, so that GPT-4 can modify the evaluation reasons of the first evaluation result, remove all descriptions directly related to the reference answer, and make the new evaluation reasons conform to the corresponding scoring results.
[0090] Specifically, constructing training data without reference answers includes:
[0091] 4. Determine the user prompt word.
[0092] The user prompt words are the first prompt words, the first evaluation input and the first evaluation result, except for all descriptions directly related to the reference answer. For example, in FIG2 , the user prompt words are as follows:
[0093] Please revise your previous evaluation reasons and scores and comply with the following requirements:
[0094] 1. In your revised justification, you should not mention the reference answer...
[0095] …
[0096] 4. Keep the previous format unchanged…”
[0097] 5. Generate, through the user prompt word, a second prompt word, a second evaluation input, and a second evaluation result of the training data without a reference answer based on the first prompt word, the first evaluation input, and the first evaluation result.
[0098] Based on the user prompt word, GPT-4 can modify the evaluation reasons of the first prompt word, the first evaluation input and the first evaluation result, remove all descriptions directly related to the reference answer, and make the new evaluation reasons consistent with the corresponding scoring results.
[0099] Thus, through the above steps 4 and 5, the second round of collection of training data without reference answers is completed.
[0100] 3. Use the training data to train the general text quality evaluation model.
[0101] With the training data, as shown on the right side of FIG2 , the general text quality assessment model (ie, CritiqueLLM in FIG2 ) can be trained using the training data.
[0102] 4. Use the trained text quality evaluation model to evaluate the quality of general text.
[0103] After the text quality evaluation model is trained, the general text is input into the trained text quality evaluation model, and the output of the text quality evaluation model is the result of the general text quality evaluation.
[0104] Among them, when using the trained text quality evaluation model to evaluate the quality of general text, model reasoning is required. During model reasoning, considering that evaluation generation is a step-by-step reasoning process similar to the chain of thought, the present invention mainly adopts two decoding strategies, greedy search and self-consistency decoding, to further improve the quality of evaluation generation. Greedy search directly selects the word with the highest generation probability as the decoding result at each step in the decoding process, while self-consistency decoding first obtains multiple evaluation results through nuclear sampling, and then calculates the average of the scores in these evaluation results as the final score, and selects the evaluation reason corresponding to the evaluation result whose score is closest to the average as the final evaluation reason.
[0105] In order to demonstrate the performance of the universal text quality evaluation method based on a large language model of the present invention, the inventors conducted quantitative experiments.
[0106] Specifically, the present invention uses ChatGLM2 models with three scales of 6 billion, 12 billion, and 66 billion parameters as general text quality evaluation models, and trains evaluation models in two scenarios with and without reference answers. A test set containing 250 user query instructions, each with answers generated by 8 different large language models, was selected. This test set is designed to test the level of alignment of large language models with human instructions in Chinese scenarios, including eight task types: logical reasoning, comprehensive question and answer, professional ability, basic ability, mathematical calculation, role-playing, text writing, and Chinese comprehension. For each answer, annotators were recruited to manually annotate its quality with a score of 1-5.
[0107] This paper compares the correlation coefficients between scores generated by different evaluation models and human scores, including Pearson (r), Spearman (ρ), and Kendall (τ). The results are shown in Figure 3. As shown in Figure 3, the evaluation model CritiqueLLM-66B, trained on the 66-billion-parameter ChatGLM2, has a correlation coefficient with human scores that is almost the same as that of the current most powerful evaluation model GPT-4 in scenarios with reference answers. In scenarios without reference answers, the correlation coefficient between its scores and human scores also reaches approximately 90% of that of GPT-4, significantly exceeding that of other evaluation models.
[0108] In the categorized evaluation, the results are shown in Figure 4. As shown in Figure 4, CritiqueLLM-66B can achieve evaluation performance that exceeds GPT-4 in the logical reasoning task with reference answers, as well as the comprehensive question answering, text writing, and Chinese comprehension tasks without reference answers.
[0109] Finally, it should be noted that the above embodiments are intended only to illustrate the technical solutions of the present invention and are not intended to limit the scope of protection of the present invention. Those skilled in the art may, based on the principles of the present invention, modify or replace the technical solutions of the present invention with equivalents without departing from the essence and scope of the technical solutions of the present invention.
Claims
1. A general text quality evaluation method based on a large language model, characterized in that: The following steps are involved: 1) Use a large language model to build a general text quality evaluation model; 2) Constructing training data: the input of the training data is prompt words and evaluation input, and the output is evaluation results. The prompt words include instructions, scoring rules and output formats, and constructing training data includes constructing training data with reference answers and constructing training data without reference answers; 3) Training the general text quality evaluation model using the training data; 4) Use the trained text quality evaluation model to evaluate the quality of general text.
2. The general text quality evaluation method based on a large language model according to claim 1 is characterized in that: The step 2) of constructing training data containing reference answers specifically includes: 2.1) Determine the first prompt word of the training data containing the reference answer, that is, determine the first instruction, the first scoring rule and the first output format of the first prompt word; 2.2) Obtaining a first evaluation input of the training data containing the reference answer; 2.3) Obtain a first evaluation result of the training data containing the reference answer.
3. The general text quality evaluation method based on a large language model according to claim 2 is characterized in that: The step 2) of constructing training data without reference answers specifically includes: 2.4) determining a user prompt word, so that the user prompt word is to remove all the descriptions directly related to the reference answer from the first prompt word, the first evaluation input and the first evaluation result; 2.5) Generate, through the user prompt word, a second prompt word, a second evaluation input and a second evaluation result of the training data not containing a reference answer based on the first prompt word, the first evaluation input and the first evaluation result.
4. The general text quality evaluation method based on a large language model according to claim 3 is characterized in that: The step 2.2) of obtaining the first evaluation input of the training data containing the reference answer specifically includes: 2.2.1) Obtaining enhanced user query instructions: collecting multiple initial user query instructions on the public network platform and processing them to obtain multiple enhanced user query instructions; 2.2.2) Collecting and generating responses: inputting the plurality of enhanced user query instructions into a plurality of Chinese open source large language models and API access models to generate a plurality of generated responses for each of the enhanced user query instructions; 2.2.3) Collect reference answers: Input the multiple enhanced user query instructions into the GPT-4 model respectively, and the GPT-4 model generates initial reference answers. Then, the annotators correct the problems in the initial reference answers to obtain the final reference answers. 2.2.4) Based on the enhanced user inquiry instruction, a reply and a reference answer are generated to form a first evaluation input of the training data containing the reference answer.
5. The general text quality evaluation method based on a large language model according to claim 4 is characterized in that: The evaluation result includes evaluation reasons and evaluation scores.
6. The general text quality evaluation method based on a large language model according to claim 5 is characterized in that: The first evaluation result of obtaining the training data containing the reference answer in the step 2.3) is specifically: the first prompt word and the first evaluation input are both input into the GPT-4 model, and the GPT-4 model scores the quality of the generated response relative to the reference answer based on the first scoring rule and generates corresponding evaluation reasons, thereby obtaining the evaluation reasons and evaluation score of the first evaluation result.
7. The general text quality evaluation method based on a large language model according to claim 6 is characterized in that: The step 2.2.1) of obtaining the enhanced user query instruction specifically includes: 2.2.1.1) Collect multiple initial user query instructions on the public network platform and classify them; 2.2.1.2) Based on the multiple initial user inquiry instructions, use ChatGPT to generate more new user inquiry instructions with similar category distribution and different instruction contents to the initial user inquiry instructions; 2.2.1.3) Based on the text overlap evaluation index, the new instructions asked by the user are screened; 2.2.1.4) Perform difficulty screening on the screened users’ inquiries for new instructions, so that the proportion of data of users’ inquiries for new instructions at different difficulty levels is equal; 2.2.1.5) Balance the category distribution of the new user inquiry instructions after the difficulty screening, so that the data proportion of the new user inquiry instructions in each category is equal, thereby obtaining the multiple enhanced user inquiry instructions.
8. The general text quality evaluation method based on a large language model according to claim 7, characterized in that: In step 2.2.1.1), the classification is divided into eight categories, namely: logical reasoning, comprehensive question and answer, professional ability, basic ability, mathematical calculation, role-playing, text writing and Chinese comprehension.
9. The general text quality evaluation method based on a large language model according to any one of claims 1 to 8, characterized in that: In the step 4), when the quality of the general text is evaluated using the trained text quality evaluation model, two decoding strategies, greedy search and self-consistency, are adopted.
10. The general text quality evaluation method based on a large language model according to claim 9, characterized in that: In the step 1), when a large language model is used to construct a universal text quality evaluation model, the large language model is a ChatGLM2 model with 6 billion, 12 billion or 66 billion parameters respectively.
Citation Information
Patent Citations
Text generation method based on deep learning
CN116226378A
Large language model based on incremental learning, training method and text generation method
CN116882369A
General text quality evaluation method based on large language model
CN117634468A
Apparatus and method for fine-tuning artificial intelligence model using question-and-answer security data
KR102574645B1
Cited By
Task processing model training method, role playing model training method and task processing method
CN120705532A
Network protocol intelligent extraction method based on large language model and application
CN120725152A
Device for controlling results of model questions and answers through role definition
CN120804273A
Data processing method and related device
CN120952007A
Adaptive instruction induction method, system and equipment based on large language model and storage medium
CN121072788A