Output consistency evaluation method and device based on large language model
By generating interference terms and target multiple choice questions similar to the original answer, evaluating the output consistency of the large language model, the problem of difficulty in evaluating the output consistency of the large language model in the prior art is solved, and a more accurate and practical evaluation method is achieved.
Patent Information
- Application Number
- CN202411903637.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-23
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art is difficult to effectively evaluate the output consistency of large language models in different scenarios, mainly due to the complexity of the model and the diversity of application scenarios, making it difficult to determine a unified output consistency standard.
By obtaining the answers to the original question and generating similar candidate interference terms, the target single-choice question is constructed, and the output consistency of the large language model is evaluated through multiple rounds of verification and strict judgment criteria.
This method can more comprehensively test the discrimination ability of large language models, simulate real application scenarios, improve the practicality and accuracy of evaluation, and accurately judge the output consistency of the model.
Smart Images

Figure CN120069055A_ABST
Abstract
Description
Background Art
[0002] With the rapid development of artificial intelligence technology, large language models (LLMs) have been widely applied in various fields, such as natural language processing, intelligent customer service, text generation, etc. However, the output consistency problem of large language models has gradually become one of the key factors restricting their further application and development. Output consistency means that under the same or similar input conditions, the large language model should be able to stably output the same or similar results. In practical applications, users expect the large language model to provide reliable and consistent answers to meet the requirements of various tasks. For example, in the field of intelligent customer service, for the same question repeatedly asked by users, the model should give a consistent solution; in information retrieval, for similar queries, the model should return similar results to ensure that users can obtain accurate and stable information.
[0003] Currently, there are many challenges in evaluating the output consistency of large language models. On the one hand, the complexity of large language models makes their output behavior difficult to fully predict. Factors such as the training data, parameter settings, and training algorithms of the model will all affect its output, and even a small change may lead to differences in the output results. On the other hand, the diversity and complexity of real-world application scenarios increase the difficulty of evaluation. Different fields, question types, contexts, etc. will all affect the output of the model, and the input methods and expression ways of users are also different, which makes it difficult to determine a unified output consistency standard.
[0004] How to evaluate the output consistency of large language models in different scenarios is a technical problem that needs to be solved currently. Summary of the Invention
[0005] The present invention provides a method and device for evaluating the output consistency based on a large language model to solve the defects existing in the prior art.
[0006] The present invention provides a method for evaluating the output consistency based on a large language model, including the following steps: Obtain the original question, input the original question into the large language model to obtain the original answer; Generate multiple candidate distractors for the original answer through natural language processing; wherein, the candidate distractor is: answer information similar to the original answer; Generate a target single-choice question corresponding to the original question based on the original answer and a preset number of the candidate distractors; Verify the large language model a preset number of times based on the target single-choice question corresponding to the original question to obtain a verification result, and evaluate the output consistency of the large language model based on the verification result.
[0007] A method for evaluating the output consistency based on a large language model provided by the present invention, generating a plurality of candidate distractors for the original answer through natural language processing, includes: Parsing the original answer through a deep learning algorithm, and generating a plurality of initial distractors for the original answer through natural language processing according to the parsing result; Screening the plurality of initial distractors based on a preset screening condition, and determining the initial distractors that meet the preset screening condition as candidate distractors.
[0008] A method for evaluating the output consistency based on a large language model provided by the present invention, parsing the original answer through a deep learning algorithm, and generating a plurality of initial distractors for the original answer through natural language processing according to the parsing result, includes: Parsing the original answer through a deep learning algorithm to obtain a parsing result; Generating a plurality of initial distractors for the original answer through a replacement algorithm according to the parsing result; Wherein, the replacement algorithm generates a plurality of initial distractors for the original answer through a replacement candidate dictionary, and the replacement candidate dictionary includes: keys and values, the keys represent the words to be replaced, and the values represent the corresponding plurality of replacement words.
[0009] A method for evaluating the output consistency based on a large language model provided by the present invention, generating a target single-choice question corresponding to the original question based on the original answer and a preset number of the candidate distractors, includes: Determining a preset number of candidate distractors among the plurality of candidate distractors, and determining the original answer, the preset number of candidate distractors and a fixed option as the options of the target single-choice question; Determining the scores of each option in the options of the target single-choice question according to an evaluation criterion.
[0010] A method for evaluating the output consistency based on a large language model provided by the present invention, after verifying the large language model a preset number of times based on the target single-choice question corresponding to the original question to obtain a verification result, and evaluating the output consistency of the large language model based on the verification result, the method further includes: Generating explanatory information for each option in the options of the target single-choice question through natural language generation technology, so that a user can understand the evaluation result and the output characteristics of the large language model based on the explanatory information; Wherein, the explanatory information includes: the characteristics of the corresponding option and the difference between the corresponding option and the original answer.
[0011] According to an output consistency evaluation method based on a large language model provided by the present invention, the large language model is verified a preset number of times based on the target single-choice question corresponding to the original question to obtain a verification result, and the output consistency of the large language model is evaluated based on the verification result, including: The large language model is verified a preset number of times based on the target single-choice question corresponding to the original question to obtain a verification result; When the verification result indicates that the probability of the large language model selecting the original answer is greater than the confidence threshold, it is determined that the large language model meets the output consistency.
[0012] According to an output consistency evaluation method based on a large language model provided by the present invention, after the large language model is verified a preset number of times based on the target single-choice question corresponding to the original question to obtain a verification result, the method further includes: When the verification result indicates that the probability of the large language model selecting the original answer is less than or equal to the confidence threshold, it is determined that the large language model does not meet the output consistency; Analyze the large language model through data analysis and mining techniques to determine the type of wrong answer selected and the probability of selecting the wrong answer; Generate adjustment suggestions based on the type of wrong answer selected and the probability of selecting the wrong answer, and adjust the large language model based on the adjustment suggestions.
[0013] The present invention also provides an output consistency evaluation device based on a large language model, including the following modules: An acquisition module, configured to acquire an original question, input the original question into the large language model to obtain an original answer; A first generation module, configured to generate a plurality of candidate distractors for the original answer through natural language processing; wherein, the candidate distractor is: answer information similar to the original answer; A second generation module, configured to generate a target single-choice question corresponding to the original question based on the original answer and a preset number of the candidate distractors; An evaluation module, configured to verify the large language model a preset number of times based on the target single-choice question corresponding to the original question to obtain a verification result, and evaluate the output consistency of the large language model based on the verification result.
[0014] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the computer program, the output consistency evaluation method based on a large language model as described in any one of the above is implemented.
[0015] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the output consistency evaluation method based on a large language model as described in any one of the above.
[0016] The present invention also provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the output consistency evaluation method based on a large language model as described in any one of the above.
[0017] An output consistency evaluation method and device based on a large language model provided by the present invention, by obtaining an original question, inputting the original question into the large language model to obtain an original answer; generating a plurality of candidate interference items for the original answer through natural language processing; wherein, the candidate interference items are: answer information similar to the original answer; generating a target single-choice question corresponding to the original question based on the original answer and a preset number of the candidate interference items; verifying the large language model a preset number of times based on the target single-choice question corresponding to the original question to obtain a verification result, and evaluating the output consistency of the large language model based on the verification result. It can be seen from this that the present invention uses the answer generation step to obtain the original answer of the large language model to public domain questions, providing basic data for subsequent evaluation; constructs interference items similar to the original answer through the interference item generation step, increasing the complexity and challenge of the evaluation, and being able to more comprehensively test the discrimination ability of the large language model; combines the original answer and interference items into the form of a single-choice question through the single-choice question generation step, simulating a real application scenario, and improving the practicality and accuracy of the evaluation; in the output consistency verification step, accurately judges the output consistency of the large language model through multiple rounds of answering and strict judgment criteria. Description of the Drawings
[0018] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0019] Figure 1 It is a schematic flowchart of the output consistency task process of the LLM in the prior art provided by the present invention.
[0020] Figure 2 It is a schematic flowchart of the output consistency evaluation method based on a large language model provided by the present invention.
[0021] Figure 3It is the complete flowchart of the output consistency evaluation method based on the large language model provided by the present invention.
[0022] Figure 4 It is the schematic diagram of the priority of the replacement algorithm provided by the present invention.
[0023] Figure 5 It is the schematic structural diagram of the output consistency evaluation device based on the large language model provided by the present invention.
[0024] Figure 6 It is the schematic structural diagram of the electronic device provided by the present invention. Detailed implementation manners
[0025] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts shall fall within the protection scope of the present invention.
[0026] The following combines Figures 1-6 to describe an output consistency evaluation method and device based on the large language model of the present invention.
[0027] It should be noted that currently, there are many challenges in evaluating the output consistency of large language models. On the one hand, the complexity of large language models makes their output behavior difficult to fully predict. Factors such as the training data, parameter settings, and training algorithms of the model will all affect its output, and even a small change may lead to differences in the output results. On the other hand, the diversity and complexity of real application scenarios increase the difficulty of evaluation. Different fields, problem types, contexts, etc. will all affect the output of the model, and the input methods and expression ways of users are also different, which makes it difficult to determine a unified output consistency standard.
[0028] In addition, the continuous development and update of large language models also pose challenges to output consistency evaluation. As the model is improved and optimized, its performance and output may change, which requires continuous updating and improvement of the evaluation method to adapt to the new model characteristics. At the same time, different application scenarios have different requirements for output consistency. For example, in fields such as medical and finance, the requirements for accuracy and consistency are extremely high, and any small deviation may lead to serious consequences; while in some entertainment or creative generation fields, relatively high flexibility and diversity may be more valued, but a certain degree of consistency cannot be ignored.
[0029] Figure 1 It is the schematic flowchart of the LLM in the prior art in the output consistency task process, asFigure 1 As shown, in previous evaluations of the output consistency of large language models, confidence is the core metric, and self-evaluation was the main method in previous related work. The goal of self-evaluation is to enable the model to express its confidence based on its own knowledge and reasoning ability. The self-evaluation method is relatively simple, that is, to ask whether the answer proposed by the LLM is true or false, and then, extract the confidence score P(True) from the logits of the model.
[0030] Traditional evaluation methods have limitations when faced with the output consistency problem of large language models. Some simple testing methods, such as asking a small number of fixed questions multiple times and comparing the results, can only provide limited information and cannot comprehensively reflect the output consistency of the model in various situations. Moreover, this method is difficult to capture the performance of the model when faced with slightly different inputs or context changes. In addition, existing evaluation metrics often focus on a single dimension such as accuracy and ignore the stability and consistency of the output. For example, the common accuracy metric can only measure the proportion of correct answers of the model, but cannot reflect whether the model remains consistent in multiple answers. There are limitations in the self-evaluation method in existing confidence evaluation techniques. First, the model's own knowledge and reasoning ability are affected by various factors, which may lead to biases in confidence evaluation, such as when faced with ambiguous, vague, or out-of-training scope problems. Second, the method of obtaining the confidence score by only asking whether the answer is true or false is single and rough, without considering the complexity and diversity of the questions and the multi-dimensional characteristics of the answers. Third, the process of extracting the confidence score P(True) from the logits may be problematic, and its calculation method may affect the accuracy and cannot deeply analyze the essence of the model output. Based on this, the present invention provides an output consistency evaluation method based on a large language model to solve at least one of the above problems.
[0031] Figure 2 is a schematic flowchart of the output consistency evaluation method based on a large language model provided by the present invention, as Figure 2 shown, the method includes the following: Step 100, obtain the original question, input the original question into the large language model, and obtain the original answer.
[0032] Figure 3 is a complete flowchart of the output consistency evaluation method based on a large language model provided by the present invention. The following combines Figure 3 to illustrate the output consistency evaluation method based on a large language model provided by the present invention.
[0033] Specifically, in step 100, input a public domain question in the prompt, and the LLM outputs an answer according to the question to obtain the answer of the LLM to this question.
[0034] The selection of open-domain questions covers multiple disciplinary fields and different difficulty levels, including but not limited to science, technology, culture, history, etc., in order to comprehensively examine the breadth and depth of knowledge of the LLM and its ability to handle different types of questions.
[0035] Step 200: Generate multiple candidate distractors for the original answer through natural language processing; wherein, the candidate distractors are: answer information similar to the original answer.
[0036] Specifically, step 200 includes: Step 210: Analyze the original answer through a deep learning algorithm, and generate multiple initial distractors for the original answer through natural language processing according to the analysis result.
[0037] Step 210 specifically includes: Step 211: Analyze the original answer through a deep learning algorithm to obtain an analysis result.
[0038] Step 212: Generate multiple initial distractors for the original answer through a replacement algorithm according to the analysis result; wherein, the replacement algorithm generates multiple initial distractors for the original answer through a replacement candidate dictionary, and the replacement candidate dictionary includes: keys and values, the keys represent the words to be replaced, and the values represent the corresponding multiple replacement words.
[0039] Specifically, in order to generate high-quality distractors, a replacement algorithm is designed, which generates a replacement candidate dictionary, where the keys represent the words to be replaced, and the values are the corresponding potential replacement words. There are three types of replacement algorithms: replacing entities by looking up embeddings, looking up candidate words by looking up embeddings, and looking up candidate words by part of speech. Figure 4 It is the schematic diagram of the priority of the replacement algorithm provided by the present invention, and the specific priority is as Figure 4 shown.
[0040] Step 220: Screen the multiple initial distractors based on a preset screening condition, and determine the initial distractors that meet the preset screening condition as candidate distractors.
[0041] Specifically, the generation of distractors is based on advanced natural language processing technologies, such as semantic analysis and lexical similarity calculation, accurately identifying the key entities and semantic points in the original answer, and making reasonable replacements and adjustments to generate highly similar distractors. The generated distractors are not only similar to the original answer in surface form, but also have a certain degree of confusion in semantic connotation, which can fully test the LLM's ability to distinguish subtle differences.
[0042] Optionally, in addition to entity replacement, a variety of methods such as semantic transformation, syntactic structure adjustment, and lexical replacement are comprehensively used to generate distractors, greatly enriching the diversity and complexity of the distractors. And the generated distractors are strictly screened and evaluated, and machine learning algorithms are used to remove overly similar or unreasonable distractors to ensure the quality and effectiveness of the distractors, thereby improving the accuracy of the evaluation.
[0043] Step 300: Generate a target single-choice question corresponding to the original question based on the original answer and a preset number of the candidate distractors.
[0044] Specifically, step 300 includes: Step 310: Determine a preset number of candidate distractors among the multiple candidate distractors, and determine the original answer, the preset number of candidate distractors, and the fixed option as the options of the target single-choice question.
[0045] Step 320: Determine the score of each option in the options of the target single-choice question according to the evaluation criteria.
[0046] In one embodiment, a certain number of distractors are randomly selected from the candidate distractors as part of the options of the single-choice question, where the options of the single-choice question include the original answer, multiple candidate distractors, and a fixed option such as the "all of the above are incorrect" option.
[0047] The generation of the options of the single-choice question follows strict statistical principles and cognitive psychology laws, ensuring that the original answer and the distractors conform to the probability situation in actual applications in terms of distribution, and the setting of the "all of the above are incorrect" option is carefully considered and has a reasonable appearance frequency to comprehensively increase the difficulty and accuracy of the evaluation. The number and specific content of the options are intelligently and dynamically adjusted according to different evaluation objectives, the types and performance characteristics of the large language models to adapt to diverse evaluation scenarios.
[0048] Optionally, according to the characteristics of the distractors such as semantic similarity and syntactic complexity, as well as the performance of the large language model, an intelligent algorithm is used to reasonably set the score of each option to more accurately evaluate the selection accuracy and confidence of the LLM. At the same time, in order to improve the transparency of the evaluation, an option explanation function is provided. After the evaluation, the characteristics of each option and the differences from the original answer are displayed to the user in detail through natural language generation technology to help the user better understand the evaluation results and the output characteristics of the large language model.
[0049] Step 400: Verify the large language model a preset number of times based on the target single-choice question corresponding to the original question, obtain the verification result, and evaluate the output consistency of the large language model based on the verification result.
[0050] Specifically, step 400 includes: Step 410: Based on the target single-choice question corresponding to the original question, verify the large language model a preset number of times to obtain a verification result.
[0051] Step 420: When the verification result indicates that the probability of the large language model selecting the original answer is greater than the confidence threshold, determine that the large language model meets the output consistency.
[0052] Step 430: When the verification result indicates that the probability of the large language model selecting the original answer is less than or equal to the confidence threshold, determine that the large language model does not meet the output consistency.
[0053] Step 440: Analyze the large language model through data analysis and mining techniques to determine the type of the selected wrong answer and the probability of selecting the wrong answer.
[0054] Step 450: Generate adjustment suggestions based on the type of the selected wrong answer and the probability of selecting the wrong answer, and adjust the large language model based on the adjustment suggestions.
[0055] It should be noted that the maximum number of verifications (i.e., the above preset number) is scientifically set based on a large amount of experimental data and actual application requirements. On the premise of ensuring the accuracy of the evaluation results, the time and resource consumption required for the evaluation are reasonably optimized. During the multi-round answering process, the data is recorded and analyzed in detail. The generated single-choice questions and the selection results of the LLM each time are recorded in detail, and in-depth analysis and statistics are carried out to comprehensively understand the performance trends and stability characteristics of the LLM in different rounds.
[0056] Specifically, a confidence threshold is introduced, and this threshold can be adjusted in real time according to the historical performance of the LLM, the difficulty of the current evaluation question, and other relevant factors. When the probability of the LLM selecting the original answer exceeds the confidence threshold, it is considered to meet the output consistency, which further improves the accuracy and adaptability of the evaluation. For the cases that do not meet the output consistency, data analysis and mining techniques are used for detailed analysis, including the classification of error types, the statistics of error frequencies, etc., to provide targeted and operable suggestions for improving the LLM and promote the continuous improvement of the performance of the large language model.
[0057] The output consistency evaluation method based on large language models provided by the embodiments of the present invention comprehensively and scientifically evaluates the output consistency of large language models, provides strong guarantee for their reliability in practical applications, and helps to promote the wide application and in-depth development of large language models in various fields; through innovative methods of generating distractors and setting single-choice questions, it truly simulates complex application scenarios, improves the accuracy and practicality of the evaluation, and can more effectively discover potential problems in the output consistency of large language models; multiple rounds of output consistency verification and detailed data analysis provide rich information for deeply understanding the performance characteristics and change laws of large language models, and provide valuable reference basis for the optimization and improvement of the models; the evaluation method has high flexibility and scalability, and can adapt to different types and scales of large language models, as well as the changing application requirements and technological development trends.
[0058] The above is the step description of the output consistency evaluation method based on large language models provided by the present invention. From the description of the above steps, it can be seen that according to the output consistency evaluation method based on large language models provided by the present invention, by obtaining the original question, inputting the original question into the large language model, the original answer is obtained; multiple candidate distractors of the original answer are generated through natural language processing; wherein, the candidate distractor is: answer information similar to the original answer; based on the original answer and a preset number of the candidate distractors, a target single-choice question corresponding to the original question is generated; the large language model is verified a preset number of times based on the target single-choice question corresponding to the original question to obtain a verification result, and the output consistency of the large language model is evaluated based on the verification result. It can be seen from this that the present invention uses the answer generation step to obtain the original answer of the large language model to public domain questions, providing basic data for subsequent evaluation; constructs distractors similar to the original answer through the distractor generation step, increasing the complexity and challenge of the evaluation, and can more comprehensively test the discrimination ability of the large language model; combines the original answer and the distractors into the form of a single-choice question through the single-choice question generation step, simulating real application scenarios, and improving the practicality and accuracy of the evaluation; in the output consistency verification step, through multiple rounds of answering and strict judgment criteria, the output consistency of the large language model is accurately judged.
[0059] The output consistency evaluation device based on large language models provided by the present invention will be described below. The output consistency evaluation device based on large language models described below can be mutually corresponding and referred to the output consistency evaluation method described above.
[0060] Figure 5 is the structural schematic diagram of the output consistency evaluation device based on large language models provided by the present invention. As Figure 5 shown, the output consistency evaluation device based on large language models provided by the present invention includes: An acquisition module 501, configured to acquire an original question, input the original question into a large language model, and obtain an original answer; A first generation module 502, configured to generate a plurality of candidate distractors for the original answer through natural language processing; wherein, the candidate distractor is: answer information similar to the original answer; A second generation module 503, configured to generate a target single-choice question corresponding to the original question based on the original answer and a preset number of the candidate distractors; An evaluation module 504, configured to perform a preset number of validations on the large language model based on the target single-choice question corresponding to the original question to obtain a validation result, and evaluate the output consistency of the large language model based on the validation result.
[0061] The output consistency evaluation device based on a large language model provided by the present invention obtains an original question, inputs the original question into a large language model to obtain an original answer; generates a plurality of candidate distractors for the original answer through natural language processing; wherein, the candidate distractor is: answer information similar to the original answer; generates a target single-choice question corresponding to the original question based on the original answer and a preset number of the candidate distractors; performs a preset number of validations on the large language model based on the target single-choice question corresponding to the original question to obtain a validation result, and evaluates the output consistency of the large language model based on the validation result. It can be seen from this that the present invention uses the answer generation step to obtain the original answer of the large language model to public domain questions, providing basic data for subsequent evaluation; constructs distractors similar to the original answer through the distractor generation step, increasing the complexity and challenge of the evaluation, and being able to more comprehensively test the discrimination ability of the large language model; combines the original answer and the distractors into a single-choice question form through the single-choice question generation step, simulating a real application scenario, and improving the practicality and accuracy of the evaluation; in the output consistency verification step, accurately judges the output consistency of the large language model through multiple rounds of answering and strict judgment criteria.
[0062] Based on the above embodiment, in this embodiment, the first generation module 502 is specifically configured to: Parse the original answer through a deep learning algorithm, and generate a plurality of initial distractors for the original answer through natural language processing according to the parsing result; Screen the plurality of initial distractors based on a preset screening condition, and determine the initial distractors that meet the preset screening condition as candidate distractors.
[0063] Based on the above embodiment, in this embodiment, the first generation module 502 is specifically configured to: Parse the original answer through a deep learning algorithm to obtain a parsing result; Generate multiple initial interference items of the original answer through a replacement algorithm according to the parsing result; Among them, the replacement algorithm generates multiple initial interference items of the original answer through a replacement candidate dictionary, and the replacement candidate dictionary includes: keys and values, where the keys represent the words to be replaced, and the values represent the corresponding multiple replacement words.
[0064] Based on the above embodiments, in this embodiment, the second generation module 503 is specifically configured to: Determine a preset number of candidate interference items from the multiple candidate interference items, and determine the original answer, the preset number of candidate interference items, and fixed options as the options of the target single-choice question; Determine the scores of each option in the options of the target single-choice question according to the evaluation criteria.
[0065] Based on the above embodiments, in this embodiment, the device further includes a third generation module, specifically configured to: After verifying the large language model a preset number of times based on the target single-choice question corresponding to the original question to obtain a verification result, and evaluating the output consistency of the large language model based on the verification result, generate explanatory information for each option in the options of the target single-choice question through natural language generation technology, so that users can understand the evaluation result and the output characteristics of the large language model based on the explanatory information; Among them, the explanatory information includes: the characteristics of the corresponding option and the differences between the corresponding option and the original answer.
[0066] Based on the above embodiments, in this embodiment, the evaluation module 504 is specifically configured to: Verify the large language model a preset number of times based on the target single-choice question corresponding to the original question to obtain a verification result; When the verification result indicates that the probability of the large language model selecting the original answer is greater than the confidence threshold, determine that the large language model meets the output consistency.
[0067] Based on the above embodiments, in this embodiment, the device further includes an adjustment module, specifically configured to: After verifying the large language model a preset number of times based on the target single-choice question corresponding to the original question to obtain a verification result, when the verification result indicates that the probability of the large language model selecting the original answer is less than or equal to the confidence threshold, determine that the large language model does not meet the output consistency; Analyze the large language model through data analysis and mining techniques to determine the types of wrong answers selected and the probabilities of selecting wrong answers; Generate adjustment suggestions based on the types of wrong answers selected and the probabilities of selecting wrong answers, and adjust the large language model based on the adjustment suggestions.
[0068] Figure 6 An example of a schematic physical structure diagram of an electronic device is shown as Figure 6 shown. The electronic device can be a robot or other electronic device. The electronic device can include: a processor 610, a communication interface 620, a memory 630, and a communication bus 640. Among them, the processor 610, the communication interface 620, and the memory 630 complete mutual communication through the communication bus 640. The processor 610 can call the logical instructions in the memory 630 to execute the output consistency evaluation method based on the large language model, including: Obtain the original question, input the original question into the large language model, and obtain the original answer; Generate multiple candidate distractors for the original answer through natural language processing; wherein, the candidate distractors are: answer information similar to the original answer; Generate a target single-choice question corresponding to the original question based on the original answer and a preset number of the candidate distractors; Verify the large language model a preset number of times based on the target single-choice question corresponding to the original question, obtain the verification result, and evaluate the output consistency of the large language model based on the verification result.
[0069] In addition, when the logical instructions in the above-mentioned memory 630 can be implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0070] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the output consistency evaluation method based on a large language model provided by each of the above methods, including: Obtain an original question, input the original question into a large language model, and obtain an original answer; Generate a plurality of candidate interference items for the original answer through natural language processing; wherein, the candidate interference items are: answer information similar to the original answer; Generate a target single-choice question corresponding to the original question based on the original answer and a preset number of the candidate interference items; Verify the large language model a preset number of times based on the target single-choice question corresponding to the original question to obtain a verification result, and evaluate the output consistency of the large language model based on the verification result.
[0071] On another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is implemented to execute the output consistency evaluation method based on a large language model provided by each of the above methods, including: Obtain an original question, input the original question into a large language model, and obtain an original answer; Generate a plurality of candidate interference items for the original answer through natural language processing; wherein, the candidate interference items are: answer information similar to the original answer; Generate a target single-choice question corresponding to the original question based on the original answer and a preset number of the candidate interference items; Verify the large language model a preset number of times based on the target single-choice question corresponding to the original question to obtain a verification result, and evaluate the output consistency of the large language model based on the verification result.
[0072] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative labor.
[0073] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0074] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for evaluating output consistency based on a large language model, characterized in that: include: Obtaining an original question, inputting the original question into a large language model, and obtaining an original answer; Generate multiple candidate interference items of the original answer through natural language processing; wherein the candidate interference items are: answer information similar to the original answer; Based on the original answer and a preset number of candidate distractors, generate a target multiple-choice question corresponding to the original question; The large language model is verified a preset number of times based on the target multiple-choice question corresponding to the original question to obtain a verification result, and the output consistency of the large language model is evaluated based on the verification result.
2. The output consistency evaluation method based on a large language model according to claim 1, characterized in that: The multiple candidate interference items for generating the original answer by natural language processing include: Parsing the original answer by using a deep learning algorithm, and generating a plurality of initial interference items of the original answer by using natural language processing according to the parsing result; The multiple initial interference items are screened based on preset screening conditions, and the initial interference items that meet the preset screening conditions are determined as candidate interference items.
3. The output consistency evaluation method based on a large language model according to claim 2 is characterized in that: The method of parsing the original answer by a deep learning algorithm and generating a plurality of initial interference items of the original answer by natural language processing according to the parsing result includes: Parsing the original answer through a deep learning algorithm to obtain a parsing result; Generate multiple initial interference items of the original answer through a replacement algorithm according to the parsing result; The replacement algorithm generates a plurality of initial interference items of the original answer by replacing a candidate dictionary, wherein the candidate replacement dictionary includes: keys and values, wherein the keys represent words to be replaced, and the values represent corresponding plurality of replacement words.
4. The output consistency evaluation method based on a large language model according to claim 1, characterized in that: The step of generating a target multiple-choice question corresponding to the original question based on the original answer and a preset number of candidate distractors includes: Determine a preset number of candidate distractors from the plurality of candidate distractors, and determine the original answer, the preset number of candidate distractors, and fixed options as options for the target single-choice question; The score of each option in the target multiple-choice question is determined according to the evaluation criteria.
5. The output consistency evaluation method based on a large language model according to claim 4 is characterized in that: After verifying the large language model a preset number of times based on the target multiple-choice question corresponding to the original question to obtain a verification result, and evaluating the output consistency of the large language model based on the verification result, the method further includes: Generate explanation information for each option in the target multiple-choice question through natural language generation technology, so that the user can understand the evaluation results and the output characteristics of the large language model based on the explanation information; The explanation information includes: the characteristics of the corresponding option and the difference between the corresponding option and the original answer.
6. The output consistency evaluation method based on a large language model according to claim 1, characterized in that: The verifying the large language model for a preset number of times based on the target multiple-choice question corresponding to the original question to obtain a verification result, and evaluating the output consistency of the large language model based on the verification result, includes: Verifying the large language model a preset number of times based on the target multiple-choice question corresponding to the original question to obtain a verification result; In a case where the verification result indicates that the probability that the large language model selects the original answer is greater than a confidence threshold, it is determined that the large language model meets the output consistency.
7. The output consistency evaluation method based on a large language model according to claim 6, characterized in that: After verifying the large language model a preset number of times based on the target multiple-choice question corresponding to the original question and obtaining the verification result, the method further includes: In a case where the verification result indicates that the probability that the large language model selects the original answer is less than or equal to the confidence threshold, determining that the large language model does not meet output consistency; Analyzing the large language model through data analysis and mining techniques to determine the type of incorrect answer selection and the probability of selecting the incorrect answer; An adjustment suggestion is generated based on the type of the selected incorrect answer and the probability of selecting the incorrect answer, and the large language model is adjusted based on the adjustment suggestion.
8. An output consistency evaluation device based on a large language model, characterized in that: include: An acquisition module is used to acquire the original question, input the original question into the large language model, and obtain the original answer; A first generating module is used to generate a plurality of candidate interference items of the original answer through natural language processing; wherein the candidate interference items are: answer information similar to the original answer; A second generating module, configured to generate a target multiple-choice question corresponding to the original question based on the original answer and a preset number of candidate distractors; An evaluation module is used to verify the large language model a preset number of times based on the target multiple-choice question corresponding to the original question, obtain a verification result, and evaluate the output consistency of the large language model based on the verification result.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the output consistency evaluation method based on the large language model is implemented as described in any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the output consistency evaluation method based on a large language model as described in any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Robot moving operation method and device, electronic equipment, storage medium and computer program product
CN120697017A
Method and device for testing large language model, computer equipment, storage medium and program product
CN120705072A
Method and device for expanding performance evaluation data of large language model
CN120929840A
Dynamic credibility evaluation method and device for large language model
CN121211455A