Large model illusion test and evaluation system and data processing method
By introducing multi-dimensional quantitative indicators and automated evaluation modules into the large-model hallucination test evaluation system, the problem of inaccurate and inefficient test results in the existing technology is solved, and a more efficient and accurate large-model hallucination test evaluation is achieved.
Patent Information
- Application Number
- CN202510457683.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-05-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the prior art, when conducting large-model hallucination tests, it relies on manual annotation and single index testing, and there are subjective deviations, the test results are not accurate enough, the efficiency is not high, and the reusability is poor.
Provides a test and evaluation system for large-scale illusions, including benchmark question answer module, test question adapter module, intelligent question answer system, test answer format module, benchmark answer adapter module, evaluation module and evaluation report module, to automatically test and evaluate through multi-dimensional quantification indicators (relevance, completeness and accuracy).
It improves the accuracy and efficiency of large-scale illusion testing, reduces the cost of manual labeling, realizes multi-dimensional effect evaluation and tuning guidance for intelligent question-and-answer systems, and is suitable for the unified evaluation standards of intelligent question-and-answer systems implemented by different technical components.
Smart Images

Figure CN119990330A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular to a large model hallucination test evaluation system and a data processing method. Background Art
[0002] As the Large Language Model (LLM) technology matures, many users are familiar with the operation of an intelligent question-answering system: user question - large language model - answer. Due to factors such as knowledge defects, outdated knowledge, errors or fabrications used in large model LLM training, abnormal answers appear, which is generally called large model hallucination.
[0003] In order to avoid the illusion of a large model, existing intelligent question answering systems need to use their own knowledge base to update the knowledge of the large model. There are many ways to update the knowledge of a large model, such as fine-tuning the large model, retrieval-augmented generation (RAG) or other methods.
[0004] However, no matter which method is used to implement the intelligent question-answering system, whether problems similar to large model hallucinations will still occur, large model hallucination tests still need to be carried out. The traditional testing process relies on manual labeling and single indicator (such as accuracy) testing, which has subjective bias, and the test results are not accurate enough, the efficiency is low, and the reusability is poor.
[0005] Therefore, the prior art still needs to be improved and developed. Summary of the invention
[0006] The main purpose of the present invention is to provide a test and evaluation system, a data processing method, a terminal and a computer-readable storage medium for large model hallucinations, aiming to solve the problems in the prior art of large model hallucination testing that rely on manual labeling and single indicator testing, have subjective bias, and have inaccurate test results, low efficiency and poor reusability.
[0007] To achieve the above object, the present invention provides a large model hallucination test and evaluation system, the large model hallucination test and evaluation system comprising: Benchmark question answer module, test question adaptation module, intelligent question answering system, test answer formatting module, benchmark answer adaptation module, evaluation module and evaluation report module; The benchmark question answer module, the test question adaptation module, the intelligent question answering system, the test answer formatting module, the evaluation module and the evaluation report module are connected in sequence, and the benchmark question answer module is also connected to the evaluation module through the benchmark answer adaptation module; The benchmark question answer module is used to send an initial benchmark question to the test question adaptation module, and send an initial benchmark answer corresponding to the initial benchmark question to the benchmark answer adaptation module; The test question adaptation module is used to format the initial benchmark question to obtain a test question, and send the test question to the intelligent question answering system; The intelligent question-answering system is used to perform large-model fine-tuning or retrieval enhancement generation on the test question to obtain an initial test answer, and send the initial test answer and the test question to the test answer formatting module; The test answer formatting module is used to format the initial test answer to obtain a test answer, and send the test answer and the test question to the evaluation module; The benchmark answer adaptation module is used to format the initial benchmark answer to obtain a benchmark answer, and send the benchmark answer to the evaluation module; The evaluation module is used to calculate the relevance, completeness and accuracy based on the test question, the test answer and the benchmark answer, and send the relevance, completeness and accuracy to the evaluation report module; The evaluation report module is used to generate a large model hallucination evaluation result of the intelligent question-answering system based on the relevance, the completeness and the accuracy.
[0008] Optionally, in the test and evaluation system for large model hallucination, the evaluation module includes a relevance evaluation unit, an integrity evaluation unit, an accuracy evaluation unit, a similarity calculation unit, a valid segmentation calculation unit and a segmentation unit; The correlation evaluation unit is connected to the similarity calculation unit, the integrity evaluation unit and the accuracy evaluation unit are respectively connected to the effective segmentation calculation unit, and the effective segmentation calculation unit is respectively connected to the similarity calculation unit and the segmentation unit.
[0009] Optionally, in the test evaluation system for large model hallucination, the relevance evaluation unit is used to receive the test question and the test answer, generate multiple new questions through LLM according to the test question, and send the multiple new questions and the test question to the similarity calculation unit; The similarity calculation unit is used to receive the multiple new questions and the test questions, calculate multiple similarities between the test questions and the multiple new questions respectively through cosine similarity, and send the multiple similarities to the relevance evaluation unit; The correlation evaluation unit is further configured to receive a plurality of the similarities and calculate the correlation between the test question and the test answer based on the plurality of the similarities.
[0010] Optionally, in the test and evaluation system for large model hallucination, the integrity evaluation unit is used to receive the reference answer and the test answer, and send the reference answer and the test answer to the effective segmentation calculation unit; The effective segmentation calculation unit is used to receive the reference answer and the test answer, and input the reference answer and the test answer into the segmentation unit; The segmentation unit is used to receive the reference answer and the test answer, use the large model LLM to split the reference answer and the test answer respectively, obtain reference answer segments and test answer segments, and send the reference answer segments and the test answer segments to the effective segmentation calculation unit; The effective segmentation calculation unit is further used to receive the reference answer segmentation and the test answer segmentation, and input each segment in the reference answer segmentation and each segment in the test answer segmentation into the similarity calculation unit one by one; The similarity calculation unit is used to receive each segment in the reference answer segment and each segment in the test answer segment, calculate multiple similarities between each segment in the reference answer segment and each segment in the test answer segment by cosine similarity, and send the multiple similarities to the effective segment calculation unit; The effective segmentation calculation unit is further used to receive a plurality of similarities, take the test answer segment corresponding to the similarity reaching a preset threshold as the effective test answer segment, and send the effective test answer segment and the reference answer segment to the integrity assessment unit; The integrity evaluation unit is further configured to receive the valid test answer segments and the reference answer segments, and calculate the completeness of the test answer relative to the reference answer according to the number of the valid test answer segments and the number of the reference answer segments.
[0011] Optionally, in the test and evaluation system for large model hallucination, the accuracy evaluation unit is used to receive the reference answer and the test answer, and send the reference answer and the test answer to the effective segmentation calculation unit; The effective segmentation calculation unit is used to receive the reference answer and the test answer, and input the reference answer and the test answer into the segmentation unit; The segmentation unit is used to receive the reference answer and the test answer, use the large model LLM to split the reference answer and the test answer respectively, obtain reference answer segments and test answer segments, and send the reference answer segments and the test answer segments to the effective segmentation calculation unit; The effective segmentation calculation unit is further used to receive the reference answer segmentation and the test answer segmentation, and input each segment in the reference answer segmentation and each segment in the test answer segmentation into the similarity calculation unit one by one; The similarity calculation unit is used to receive each segment in the reference answer segment and each segment in the test answer segment, calculate multiple similarities between each segment in the reference answer segment and each segment in the test answer segment by cosine similarity, and send the multiple similarities to the effective segment calculation unit; The effective segmentation calculation unit is further used to receive a plurality of similarities, take the test answer segment corresponding to the similarity reaching a preset threshold as the effective test answer segment, and send the effective test answer segment and the test answer segment to the accuracy evaluation unit; The accuracy evaluation unit is further configured to receive the valid test answer segments and the test answer segments, and calculate the accuracy of the test answer relative to the reference answer based on the number of the valid test answer segments and the number of the test answer segments.
[0012] In addition, to achieve the above-mentioned purpose, the present invention also provides a data processing method of a large model hallucination test and evaluation system, wherein the data processing method comprises: The benchmark question answer module sends an initial benchmark question to the test question adaptation module, and sends an initial benchmark answer corresponding to the initial benchmark question to the benchmark answer adaptation module; The test question adaptation module formats the initial benchmark question to obtain a test question, and sends the test question to the intelligent question answering system; The intelligent question-answering system performs large-model fine-tuning or retrieval enhancement generation on the test question to obtain an initial test answer, and sends the initial test answer and the test question to the test answer formatting module; The test answer formatting module formats the initial test answer to obtain a test answer, and sends the test answer and the test question to the evaluation module; The benchmark answer adaptation module formats the initial benchmark answer to obtain a benchmark answer, and sends the benchmark answer to the evaluation module; The evaluation module calculates the relevance, completeness and accuracy based on the test question, the test answer and the benchmark answer, and sends the relevance, completeness and accuracy to the evaluation report module; The evaluation report module generates a large model hallucination evaluation result of the intelligent question-answering system based on the relevance, the completeness and the accuracy.
[0013] The benchmark question answer module sends an initial benchmark question to the test question adaptation module, and sends an initial benchmark answer corresponding to the initial benchmark question to the benchmark answer adaptation module; The test question adaptation module formats the initial benchmark question to obtain a test question, and sends the test question to the intelligent question answering system; The intelligent question-answering system performs large-model fine-tuning or retrieval enhancement generation on the test question to obtain an initial test answer, and sends the initial test answer and the test question to the test answer formatting module; The test answer formatting module formats the initial test answer to obtain a test answer, and sends the test answer and the test question to the evaluation module; The benchmark answer adaptation module formats the initial benchmark answer to obtain a benchmark answer, and sends the benchmark answer to the evaluation module; The evaluation module calculates the relevance, completeness and accuracy based on the test question, the test answer and the benchmark answer, and sends the relevance, completeness and accuracy to the evaluation report module; The evaluation report module generates a large model hallucination evaluation result of the intelligent question-answering system based on the relevance, the completeness and the accuracy.
[0014] Optionally, the data processing method of the large model hallucination test and evaluation system, wherein the benchmark answer adaptation module formats the initial benchmark answer to obtain the benchmark answer, specifically includes: The benchmark answer adaptation module splits and modifies the initial benchmark answer through AI reasoning to obtain multiple initial benchmark answer segments, wherein each initial benchmark answer segment contains a fact; The benchmark answer adaptation module combines a plurality of the initial benchmark answers in segments to obtain the benchmark answer.
[0015] Optionally, the data processing method of the large model hallucination test evaluation system, wherein the evaluation module calculates the relevance, completeness and accuracy according to the test question, the test answer and the benchmark answer, specifically includes: The evaluation module calculates the correlation between the test question and the test answer based on the test question and the test answer; The evaluation module calculates the completeness and accuracy of the test answer relative to the benchmark answer based on the benchmark answer and the test answer.
[0016] Optionally, the data processing method of the large model hallucination test evaluation system, wherein the evaluation module calculates the correlation between the test question and the test answer based on the test question and the test answer, specifically includes: The evaluation module generates a plurality of new questions according to the test question through LLM, and respectively calculates a plurality of similarities between the test question and the plurality of new questions through cosine similarity; The evaluation module calculates the correlation between the test question and the test answer according to the multiple similarities through an average calculation function.
[0017] Optionally, the data processing method of the large model hallucination test evaluation system, wherein the evaluation module calculates the completeness and accuracy of the test answer relative to the benchmark answer based on the benchmark answer and the test answer, specifically includes: The evaluation module uses the large model LLM to split the benchmark answer and the test answer respectively to obtain benchmark answer segments and test answer segments, and sends the benchmark answer segments and the test answer segments to the effective segmentation calculation unit; The evaluation module calculates multiple similarities between each segment in the reference answer segment and each segment in the test answer segment by cosine similarity, and sends the multiple similarities to the effective segment calculation unit; The evaluation module uses the test answer segment corresponding to the similarity reaching the preset threshold as the valid test answer segment; The evaluation module calculates the completeness of the test answer relative to the benchmark answer based on the number of the valid test answer segments and the number of the benchmark answer segments; The evaluation module calculates the accuracy of the test answer relative to the benchmark answer according to the number of the valid test answer segments and the number of the test answer segments.
[0018] The present invention discloses a test and evaluation system and data processing method for large model hallucinations, the system comprising: a benchmark question answer module, a test question adaptation module, an intelligent question and answer system, a test answer formatting module, a benchmark answer adaptation module, an evaluation module and an evaluation report module. The benchmark question answer module, the test question adaptation module, the intelligent question and answer system, the test answer formatting module, the evaluation module and the evaluation report module are connected in sequence, and the benchmark question answer module is also connected to the evaluation module through the benchmark answer adaptation module. The present invention provides a test and evaluation system for the large model hallucination problem of the intelligent question and answer system, which can realize the effect evaluation and tuning guidance of the intelligent question and answer system based on multi-dimensional quantitative indicators, and improve the accuracy of the large model hallucination test. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 It is the overall architecture diagram of the test and evaluation system of the large model hallucination of the present invention; Figure 2 It is a structural diagram of an evaluation module in the test evaluation system of the large model hallucination of the present invention; Figure 3 It is a flow chart of a preferred embodiment of a data processing method of a testing and evaluation system for large model hallucination of the present invention; Figure 4 It is a structural diagram of a preferred embodiment of the terminal of the present invention. DETAILED DESCRIPTION
[0020] In order to make the purpose, technical solution and advantages of the present invention clearer and more specific, the present invention is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0021] To solve the problems in the prior art, this embodiment provides a large model hallucination test and evaluation system, such as Figure 1 As shown, the test and evaluation system for large model hallucination includes: a benchmark question answer module, a test question adaptation module, an intelligent question and answer system, a test answer formatting module, a benchmark answer adaptation module, an evaluation module and an evaluation report module.
[0022] Among them, the benchmark question answer module, the test question adaptation module, the intelligent question and answer system, the test answer formatting module, the evaluation module and the evaluation report module are connected in sequence, and the benchmark question answer module is also connected to the evaluation module through the benchmark answer adaptation module.
[0023] The benchmark question answer module is used to send an initial benchmark question to the test question adaptation module, and send an initial benchmark answer corresponding to the initial benchmark question to the benchmark answer adaptation module; The test question adaptation module is used to format the initial benchmark question to obtain a test question, and send the test question to the intelligent question answering system; The intelligent question-answering system is used to perform large-model fine-tuning or retrieval enhancement generation on the test question to obtain an initial test answer, and send the initial test answer and the test question to the test answer formatting module; The test answer formatting module is used to format the initial test answer to obtain a test answer, and send the test answer and the test question to the evaluation module; The benchmark answer adaptation module is used to format the initial benchmark answer to obtain a benchmark answer, and send the benchmark answer to the evaluation module; The evaluation module is used to calculate the relevance, completeness and accuracy based on the test question, the test answer and the benchmark answer, and send the relevance, completeness and accuracy to the evaluation report module; The evaluation report module is used to generate a large model hallucination evaluation result of the intelligent question-answering system based on the relevance, the completeness and the accuracy.
[0024] It is understandable that the intelligent question-answering system is one type of question-answering system. It accurately locates the question knowledge required by website users in the form of one question and one answer, and provides personalized information services to website users by interacting with website users. Big model hallucination refers to the phenomenon that the content generated by the model is inconsistent with real-world facts or user input. In order to avoid big model hallucination, it is necessary to use a proprietary knowledge base to update the knowledge of the big model of the intelligent question-answering system. To this end, the present invention provides a test and evaluation system for big model hallucination, which can realize automated testing and evaluation, improve efficiency, reduce manual annotation costs, and is applicable to the unified evaluation standards of intelligent question-answering systems (such as embedding algorithms and retrieval frameworks) implemented by different technical components.
[0025] Furthermore, the present invention describes in detail a data processing unit included in the evaluation module, which is used to perform corresponding data collection and processing operations.
[0026] Specifically, Figure 2 As shown, the evaluation module includes a correlation evaluation unit, an integrity evaluation unit, an accuracy evaluation unit, a similarity calculation unit, a valid segmentation calculation unit and a segmentation unit; The correlation evaluation unit is connected to the similarity calculation unit, the integrity evaluation unit and the accuracy evaluation unit are respectively connected to the effective segmentation calculation unit, and the effective segmentation calculation unit is respectively connected to the similarity calculation unit and the segmentation unit.
[0027] Further, the relevance evaluation unit is used to receive the test question and the test answer, generate multiple new questions through LLM according to the test question, and send the multiple new questions and the test question to the similarity calculation unit; The similarity calculation unit is used to receive the multiple new questions and the test questions, calculate multiple similarities between the test questions and the multiple new questions respectively through cosine similarity, and send the multiple similarities to the relevance evaluation unit; The correlation evaluation unit is further configured to receive a plurality of the similarities and calculate the correlation between the test question and the test answer based on the plurality of the similarities.
[0028] It can be understood that the relevance evaluation unit adopts the AI generation technology based on the large model LLM to evaluate the relevance of the test answer XA and the test question XQ. The basic idea of evaluating the relevance of the test answer XA and the test question XQ is that if the generated answer XA accurately solves the test question XQ, then the large model LLM can generate a question that matches XQ from XA.
[0029] Evaluate relevance, focusing on assessing how relevant the test answer XA is to the test question XQ. XAs that are incomplete or contain redundant information will receive a lower score. Relevance is calculated by comparing XA and XQ, and it ranges from 0 to 1, where higher scores indicate better relevance. XA is considered relevant when it directly and appropriately addresses XQ. Importantly, the assessment of XA relevance does not take into account the true situation, but rather penalizes the situation where XA lacks completeness or contains redundant details.
[0030] In this embodiment, in order to calculate the relevance score, the LLM large model AI is used to generate multiple new questions according to XA. As an example, here it is assumed that there are 5 new questions, named XQ1, XQ2, XQ3, XQ4, and XQ5 respectively, then: Send XQ and XQ1 to the similarity calculation unit to obtain a first similarity C(XQ, XQ1); send XQ and XQ2 to the similarity calculation unit to obtain a second similarity C(XQ, XQ2); send XQ and XQ3 to the similarity calculation unit to obtain a third similarity C(XQ, XQ3); send XQ and XQ4 to the similarity calculation unit to obtain a fourth similarity C(XQ, XQ4); send XQ and XQ5 to the similarity calculation unit to obtain a fifth similarity C(XQ, XQ5).
[0031] The similarity calculation unit is used to receive the message from the correlation evaluation unit as input, and calculate the correlation between the two contained in the input. The similarity calculation is generally constructed based on the cosine similarity method. The output similarity is expressed as C(M,N), where M represents the test question, N represents the new question, and C(M,N) is a value between 0 and 1. The larger the value, the more similar the two are.
[0032] The correlation C between the test question and the test answer is calculated based on the above five similarities: C=AVG(C(XQ,XQ1)+C(XQ,XQ2)+C(XQ,XQ3)+C(XQ,XQ4)+C(XQ,XQ5)).
[0033] Among them, AVG represents the average calculation function, the correlation C ranges from 0 to 1, and the higher the correlation C is, the higher the accuracy of XA is.
[0034] Further, the integrity evaluation unit is used to receive the reference answer and the test answer, and send the reference answer and the test answer to the valid segment calculation unit; The effective segmentation calculation unit is used to receive the reference answer and the test answer, and input the reference answer and the test answer into the segmentation unit; The segmentation unit is used to receive the reference answer and the test answer, use the large model LLM to split the reference answer and the test answer respectively, obtain reference answer segments and test answer segments, and send the reference answer segments and the test answer segments to the effective segmentation calculation unit; The effective segmentation calculation unit is further used to receive the reference answer segmentation and the test answer segmentation, and input each segment in the reference answer segmentation and each segment in the test answer segmentation into the similarity calculation unit one by one; The similarity calculation unit is used to receive each segment in the reference answer segment and each segment in the test answer segment, calculate multiple similarities between each segment in the reference answer segment and each segment in the test answer segment by cosine similarity, and send the multiple similarities to the effective segment calculation unit; The effective segmentation calculation unit is further used to receive a plurality of similarities, take the test answer segment corresponding to the similarity reaching a preset threshold as the effective test answer segment, and send the effective test answer segment and the reference answer segment to the integrity assessment unit; The integrity evaluation unit is further configured to receive the valid test answer segments and the reference answer segments, and calculate the completeness of the test answer relative to the reference answer according to the number of the valid test answer segments and the number of the reference answer segments.
[0035] It can be understood that the completeness output by the completeness evaluation unit is used to evaluate the degree of inclusion of the test answer XA in the reference answer KA. Specifically, the completeness evaluation unit receives the reference answer KA and the test answer XA, and sends XA and KA to the effective segment calculation unit.
[0036] The effective segmentation calculation unit receives the test answer XA and the reference answer KA from the completeness evaluation unit as input. XA is sent to the segmentation unit to obtain all segments of XA (test answer segments), and the number of test answer segments is represented by S(XA). KA is input to the segmentation unit to obtain all segments of KA (reference answer segments), and the number of reference answer segments is represented by S(KA).
[0037] The completeness evaluation unit then performs the following operations on each segment of the test answer segment one by one: each segment in the benchmark answer segment and each segment in the benchmark answer segment are input into the similarity calculation unit one by one to obtain multiple similarities between each segment in the benchmark answer segment and each segment in the test answer segment.
[0038] Sending the plurality of similarities to the effective segmentation calculation unit, the effective segmentation calculation unit will determine the plurality of similarities and take the test answer segment corresponding to the similarity reaching a preset threshold as the effective test answer segment; Among them, the segmentation unit is used to receive messages from the effective segmentation determination unit as input, adopt the large model LLM, and use the reasoning ability of AI to split the input. Each fact description contained in the input is split into a segment, and the output is all the segments, which are sent to the effective segmentation determination unit.
[0039] Finally, the valid segment determination unit outputs the number of valid test answer segments V(XA) and the number of reference answer segments S(KA), and sends them to the completeness evaluation unit. The completeness evaluation unit calculates the completeness I of the test answer relative to the reference answer based on the number of valid test answer segments V(XA) and the number of reference answer segments S(KA): I=V(XA) / S(KA). The range of completeness I is between 0 and 1. The higher the completeness I, the higher the completeness of the test answer XA relative to the reference answer KA.
[0040] Furthermore, the accuracy evaluation unit is used to receive the reference answer and the test answer, and send the reference answer and the test answer to the effective segmentation calculation unit; The effective segmentation calculation unit is used to receive the reference answer and the test answer, and input the reference answer and the test answer into the segmentation unit; The segmentation unit is used to receive the reference answer and the test answer, use the large model LLM to split the reference answer and the test answer respectively, obtain reference answer segments and test answer segments, and send the reference answer segments and the test answer segments to the effective segmentation calculation unit; The effective segmentation calculation unit is further used to receive the reference answer segmentation and the test answer segmentation, and input each segment in the reference answer segmentation and each segment in the test answer segmentation into the similarity calculation unit one by one; The similarity calculation unit is used to receive each segment in the reference answer segment and each segment in the test answer segment, calculate multiple similarities between each segment in the reference answer segment and each segment in the test answer segment by cosine similarity, and send the multiple similarities to the effective segment calculation unit; The effective segmentation calculation unit is further used to receive a plurality of the similarities, take the test answer segment corresponding to the similarity reaching a preset threshold as the effective test answer segment, and send the effective test answer segment and the test answer segment to the accuracy evaluation unit; The accuracy evaluation unit is further configured to receive the valid test answer segments and the test answer segments, and calculate the accuracy of the test answer relative to the reference answer based on the number of the valid test answer segments and the number of the test answer segments.
[0041] In this embodiment, the accuracy evaluation unit is used to evaluate the degree of overlap between the test answer XA and the reference answer KA. The accuracy evaluation unit receives the reference answer KA and the test answer XA, sends XA and KA to the effective segmentation calculation unit, and obtains the number of test answer segments S(XA) and the number of effective test answer segments V(XA). The process here is similar to the process of obtaining effective test answer segments in the above text, and will not be repeated here.
[0042] The accuracy A of the test answer relative to the benchmark answer is calculated based on the number of valid test answer segments and the number of test answer segments: A=V(XA) / S(XA), and the accuracy A ranges from 0 to 1. The higher the accuracy A, the higher the accuracy of XA.
[0043] It can be seen from the above that the present invention provides multi-dimensional quantitative evaluation indicators for the testing of large model hallucinations, realizes automated testing and evaluation, improves efficiency, reduces manual labeling costs, and has strong reusability. It is suitable for unified evaluation standards of intelligent question-answering systems (such as embedding algorithms and retrieval frameworks) implemented by different technical components.
[0044] Based on the large model hallucination test and evaluation system described in the above embodiment, the present invention also provides a data processing method for the large model hallucination test and evaluation system, specifically as follows: Figure 3 As shown in , the data processing method of the large model hallucination test and evaluation system includes the following steps: Step S10: The benchmark question answer module sends an initial benchmark question to the test question adaptation module, and sends an initial benchmark answer corresponding to the initial benchmark question to the benchmark answer adaptation module.
[0045] Specifically, the benchmark question answer module extracts a benchmark question KQ and sends it to the test question adaptation module, and sends the benchmark answer KA corresponding to the benchmark question KQ to the benchmark answer adaptation module.
[0046] It should be noted that before the benchmark question answer module sends the initial benchmark question to the test question adaptation module, the operator needs to set a test value when all modules are powered on. The test value is used to control the number of tests to achieve automated testing and evaluation. During initialization, the number of tests is set to 0, and the number of tests is increased by 1 each time a test is completed.
[0047] Step S20: The test question adaptation module formats the initial benchmark question to obtain a test question, and sends the test question to the intelligent question-answering system.
[0048] The test question adaptation module receives the message from the benchmark question answer module and sends it as a test question XQ to the intelligent question and answer system. Before sending it to the intelligent question and answer system, the test question adaptation module will format it into a message format that can be received by the intelligent question and answer system.
[0049] Step S30: The intelligent question-answering system performs large-model fine-tuning or retrieval enhancement generation on the test question to obtain an initial test answer, and sends the initial test answer and the test question to the test answer formatting module.
[0050] It is understandable that large-scale model fine-tuning is to input more information into the model and optimize the specific functions of the model. By inputting data sets in specific fields, the model can learn knowledge in the field, thereby optimizing the performance of large models in NLP tasks in specific fields, such as sentiment analysis, entity recognition, text classification, dialogue generation, etc. Retrieval-enhanced generation combines language models and information retrieval technology. When the model needs to generate text or answer questions, it will first retrieve relevant information from a large collection of documents, and then use the retrieved information to guide text generation, thereby improving the quality and accuracy of predictions.
[0051] After the intelligent question-answering system performs macro-model fine-tuning or retrieval enhancement on the test question, it generates an initial test answer, and the initial test answer may be abnormal due to the existence of macro-model hallucination.
[0052] Step S40: The test answer formatting module formats the initial test answer to obtain a test answer, and sends the test answer and the test question to the evaluation module.
[0053] In this embodiment, the test answer formatting module receives the initial test answers generated from the intelligent question and answer system, formats the initial test answers, generates test answers in a format that can be processed by the evaluation module, and sends the test answers to the evaluation module for further processing.
[0054] Step S50: The benchmark answer adaptation module formats the initial benchmark answer to obtain a benchmark answer, and sends the benchmark answer to the evaluation module.
[0055] The benchmark answer adaptation module formats the initial benchmark answer to obtain a benchmark answer, specifically including: The benchmark answer adaptation module splits and modifies the initial benchmark answer through AI reasoning to obtain multiple initial benchmark answer segments, wherein each initial benchmark answer segment contains a fact; The benchmark answer adaptation module combines a plurality of the initial benchmark answers in segments to obtain the benchmark answer.
[0056] As an example, suppose there is an intelligent question-answering system whose benchmark question answer module contains the following test questions and their benchmark answers: Benchmark question: Where is the capital of a certain country? Benchmark answer: A certain city is the capital of a certain country. After the test question adaptation module adapts and inputs it into the intelligent question-answering system, the generated test answer is: "A certain city is the capital of a certain country and is located in the north of a certain plain." The test answer formatting module will split it into two facts and format it as: "A certain city is the capital of a certain country. A certain city is located in the north of a certain plain."
[0057] Step S60: The evaluation module calculates the relevance, completeness and accuracy based on the test questions, the test answers and the benchmark answers, and sends the relevance, completeness and accuracy to the evaluation report module.
[0058] The evaluation module calculates the relevance, completeness and accuracy according to the test question, the test answer and the benchmark answer, specifically including: The evaluation module calculates the correlation between the test question and the test answer based on the test question and the test answer; the evaluation module calculates the completeness and accuracy of the test answer relative to the benchmark answer based on the benchmark answer and the test answer.
[0059] Furthermore, the evaluation module calculates the correlation between the test question and the test answer based on the test question and the test answer, specifically including: The evaluation module generates a plurality of new questions according to the test question through LLM, and respectively calculates a plurality of similarities between the test question and the plurality of new questions through cosine similarity; The evaluation module calculates the correlation between the test question and the test answer according to the multiple similarities through an average calculation function.
[0060] Furthermore, the evaluation module calculates the completeness and accuracy of the test answer relative to the benchmark answer based on the benchmark answer and the test answer, specifically including: The evaluation module uses the large model LLM to split the benchmark answer and the test answer respectively to obtain benchmark answer segments and test answer segments, and sends the benchmark answer segments and the test answer segments to the effective segmentation calculation unit; The evaluation module calculates multiple similarities between each segment in the reference answer segment and each segment in the test answer segment by cosine similarity, and sends the multiple similarities to the effective segment calculation unit; The evaluation module uses the test answer segment corresponding to the similarity reaching the preset threshold as the valid test answer segment; The evaluation module calculates the completeness of the test answer relative to the benchmark answer based on the number of the valid test answer segments and the number of the benchmark answer segments; The evaluation module calculates the accuracy of the test answer relative to the benchmark answer according to the number of the valid test answer segments and the number of the test answer segments.
[0061] It can be understood that the evaluation module receives messages from the test answer adaptation module and the benchmark answer adaptation module, and sends them internally and synchronously to the similarity evaluation unit, the accuracy evaluation unit and the integrity evaluation unit. The evaluation unit, the accuracy evaluation unit and the integrity evaluation unit further call the similarity calculation unit, the effective segmentation calculation unit and the segmentation unit to perform data calculation and processing, and finally obtain the correlation between the test question and the test answer, the completeness of the test answer relative to the benchmark answer and the accuracy of the test answer relative to the benchmark answer, and send the correlation, completeness and accuracy to the evaluation report module.
[0062] Step S70: The evaluation report module generates a large model hallucination evaluation result of the intelligent question-answering system according to the relevance, the completeness and the accuracy.
[0063] Specifically, the evaluation report module generates an evaluation record based on the multi-dimensional scoring results (the relevance, the completeness and the accuracy). The evaluation record is a question-answer record in one row. At the same time, the count of the tested quantity is increased by 1, and it is checked whether the tested quantity is equal to or greater than the test value. If so, the entire workflow is terminated, and an evaluation report is generated for all records and submitted to the operator; if not, test questions and benchmark answers are extracted from the benchmark question answer module to perform test evaluation of the large model illusion.
[0064] According to the evaluation report finally generated, the specific values of the intelligent question-answering system in terms of relevance, accuracy and completeness can be displayed, and improvement measures can be provided. For example, in order to increase relevance and accuracy, responses that are not relevant to the question can be reduced.
[0065] Furthermore, if Figure 4 As shown, based on the data processing method of the test and evaluation system of the large model illusion, the present invention also provides a terminal accordingly, and the terminal includes a processor 10, a memory 20 and a display 30. Figure 4 Only some components of the terminal are shown, but it should be understood that it is not required to implement all of the components shown, and more or fewer components may be implemented instead.
[0066] In some embodiments, the memory 20 may be an internal storage unit of the terminal, such as a hard disk or memory of the terminal. In other embodiments, the memory 20 may also be an external storage device of the terminal, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the terminal. Further, the memory 20 may also include both an internal storage unit of the terminal and an external storage device. The memory 20 is used to store application software and various types of data installed on the terminal, such as the program code of the installation terminal. The memory 20 may also be used to temporarily store data that has been output or is to be output. In one embodiment, the memory 20 stores a data processing program 40 of a test and evaluation system for large model illusions, and the data processing program 40 of the test and evaluation system for large model illusions can be executed by the processor 10, thereby realizing the data processing method of the test and evaluation system for large model illusions in the present application.
[0067] In some embodiments, the processor 10 may be a central processing unit (CPU), a microprocessor or other data processing chip, used to run the program code or process data stored in the memory 20, such as executing the data processing method of the test and evaluation system of the large model illusion.
[0068] In some embodiments, the display 30 may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, an OLED (Organic Light-Emitting Diode) touch device, etc. The display 30 is used to display information on the terminal and to display a visual user interface. The components 10-30 of the terminal communicate with each other via a system bus.
[0069] In one embodiment, when the processor 10 executes the data processing program 40 of the large model hallucination test and evaluation system in the memory 20, the following steps are implemented: The benchmark question answer module sends an initial benchmark question to the test question adaptation module, and sends an initial benchmark answer corresponding to the initial benchmark question to the benchmark answer adaptation module; The test question adaptation module formats the initial benchmark question to obtain a test question, and sends the test question to the intelligent question answering system; The intelligent question-answering system performs large-model fine-tuning or retrieval enhancement generation on the test question to obtain an initial test answer, and sends the initial test answer and the test question to the test answer formatting module; The test answer formatting module formats the initial test answer to obtain a test answer, and sends the test answer and the test question to the evaluation module; The benchmark answer adaptation module formats the initial benchmark answer to obtain a benchmark answer, and sends the benchmark answer to the evaluation module; The evaluation module calculates the relevance, completeness and accuracy based on the test question, the test answer and the benchmark answer, and sends the relevance, completeness and accuracy to the evaluation report module; The evaluation report module generates a large model hallucination evaluation result of the intelligent question-answering system based on the relevance, the completeness and the accuracy.
[0070] The benchmark answer adaptation module formats the initial benchmark answer to obtain a benchmark answer, which specifically includes: The benchmark answer adaptation module splits and modifies the initial benchmark answer through AI reasoning to obtain multiple initial benchmark answer segments, wherein each initial benchmark answer segment contains a fact; The benchmark answer adaptation module combines a plurality of the initial benchmark answers in segments to obtain the benchmark answer.
[0071] The evaluation module calculates the relevance, completeness and accuracy based on the test question, the test answer and the benchmark answer, specifically including: The evaluation module calculates the correlation between the test question and the test answer based on the test question and the test answer; The evaluation module calculates the completeness and accuracy of the test answer relative to the benchmark answer based on the benchmark answer and the test answer.
[0072] The evaluation module calculates the correlation between the test question and the test answer based on the test question and the test answer, specifically including: The evaluation module generates a plurality of new questions according to the test question through LLM, and respectively calculates a plurality of similarities between the test question and the plurality of new questions through cosine similarity; The evaluation module calculates the correlation between the test question and the test answer according to the multiple similarities through an average calculation function.
[0073] The evaluation module calculates the completeness and accuracy of the test answer relative to the benchmark answer based on the benchmark answer and the test answer, specifically including: The evaluation module uses the large model LLM to split the benchmark answer and the test answer respectively to obtain benchmark answer segments and test answer segments, and sends the benchmark answer segments and the test answer segments to the effective segmentation calculation unit; The evaluation module calculates multiple similarities between each segment in the reference answer segment and each segment in the test answer segment by cosine similarity, and sends the multiple similarities to the effective segment calculation unit; The evaluation module uses the test answer segment corresponding to the similarity reaching the preset threshold as the valid test answer segment; The evaluation module calculates the completeness of the test answer relative to the benchmark answer based on the number of the valid test answer segments and the number of the benchmark answer segments; The evaluation module calculates the accuracy of the test answer relative to the benchmark answer according to the number of the valid test answer segments and the number of the test answer segments.
[0074] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a data processing program for a test and evaluation system for large model illusion, and when the data processing program for a test and evaluation system for large model illusion is executed by a processor, the steps of the data processing method for a test and evaluation system for large model illusion as described above are implemented.
[0075] In summary, the present invention provides a test evaluation system and data processing method for large model hallucination, the system comprising: a benchmark question answer module, a test question adaptation module, an intelligent question and answer system, a test answer formatting module, a benchmark answer adaptation module, an evaluation module and an evaluation report module. The benchmark question answer module, the test question adaptation module, the intelligent question and answer system, the test answer formatting module, the evaluation module and the evaluation report module are connected in sequence, and the benchmark question answer module is also connected to the evaluation module through the benchmark answer adaptation module. The present invention provides a test evaluation system for the large model hallucination problem of the intelligent question and answer system, which can realize the effect evaluation and tuning guidance of the intelligent question and answer system based on multi-dimensional quantitative indicators, can realize automated testing and evaluation, improve efficiency, reduce manual labeling costs, and is applicable to the unified evaluation standard of the intelligent question and answer system implemented by different technical components, thereby improving the accuracy of large model hallucination testing.
[0076] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or terminal. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the presence of other identical elements in the process, method, article or terminal including the element.
[0077] Of course, those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing related hardware (such as a processor, a controller, etc.) through a computer program, and the program can be stored in a computer-readable storage medium that can be read by a computer, and the program can include the processes of the above-mentioned method embodiments when executed. The computer-readable storage medium can be a memory, a disk, an optical disk, etc.
[0078] It should be understood that the application of the present invention is not limited to the above examples. For ordinary technicians in this field, improvements or changes can be made based on the above description. All these improvements and changes should fall within the scope of protection of the claims attached to the present invention.
Claims
1. A test and evaluation system for large model hallucination, characterized in that: The test and evaluation system for the large model hallucination includes: a benchmark question answer module, a test question adaptation module, an intelligent question and answer system, a test answer formatting module, a benchmark answer adaptation module, an evaluation module and an evaluation report module; The benchmark question answer module, the test question adaptation module, the intelligent question answering system, the test answer formatting module, the evaluation module and the evaluation report module are connected in sequence, and the benchmark question answer module is also connected to the evaluation module through the benchmark answer adaptation module; The benchmark question answer module is used to send an initial benchmark question to the test question adaptation module, and send an initial benchmark answer corresponding to the initial benchmark question to the benchmark answer adaptation module; The test question adaptation module is used to format the initial benchmark question to obtain a test question, and send the test question to the intelligent question answering system; The intelligent question-answering system is used to perform large-model fine-tuning or retrieval enhancement generation on the test question to obtain an initial test answer, and send the initial test answer and the test question to the test answer formatting module; The test answer formatting module is used to format the initial test answer to obtain a test answer, and send the test answer and the test question to the evaluation module; The benchmark answer adaptation module is used to format the initial benchmark answer to obtain a benchmark answer, and send the benchmark answer to the evaluation module; The evaluation module is used to calculate the relevance, completeness and accuracy based on the test question, the test answer and the benchmark answer, and send the relevance, completeness and accuracy to the evaluation report module; The evaluation report module is used to generate a large model hallucination evaluation result of the intelligent question-answering system based on the relevance, the completeness and the accuracy.
2. The large model hallucination test and evaluation system according to claim 1, characterized in that: The evaluation module includes a correlation evaluation unit, an integrity evaluation unit, an accuracy evaluation unit, a similarity calculation unit, a valid segmentation calculation unit and a segmentation unit; The correlation evaluation unit is connected to the similarity calculation unit, the integrity evaluation unit and the accuracy evaluation unit are respectively connected to the effective segmentation calculation unit, and the effective segmentation calculation unit is respectively connected to the similarity calculation unit and the segmentation unit.
3. The large model hallucination test and evaluation system according to claim 2, characterized in that: The relevance evaluation unit is used to receive the test question and the test answer, generate multiple new questions through LLM according to the test question, and send the multiple new questions and the test question to the similarity calculation unit; The similarity calculation unit is used to receive the multiple new questions and the test questions, calculate multiple similarities between the test questions and the multiple new questions respectively through cosine similarity, and send the multiple similarities to the relevance evaluation unit; The correlation evaluation unit is further configured to receive a plurality of the similarities and calculate the correlation between the test question and the test answer based on the plurality of the similarities.
4. The large model hallucination test and evaluation system according to claim 2, characterized in that: The integrity evaluation unit is used to receive the reference answer and the test answer, and send the reference answer and the test answer to the valid segment calculation unit; The effective segmentation calculation unit is used to receive the reference answer and the test answer, and input the reference answer and the test answer into the segmentation unit; The segmentation unit is used to receive the reference answer and the test answer, use the large model LLM to split the reference answer and the test answer respectively, obtain reference answer segments and test answer segments, and send the reference answer segments and the test answer segments to the effective segmentation calculation unit; The effective segmentation calculation unit is further used to receive the reference answer segmentation and the test answer segmentation, and input each segment in the reference answer segmentation and each segment in the test answer segmentation into the similarity calculation unit one by one; The similarity calculation unit is used to receive each segment in the reference answer segment and each segment in the test answer segment, calculate multiple similarities between each segment in the reference answer segment and each segment in the test answer segment by cosine similarity, and send the multiple similarities to the effective segment calculation unit; The effective segmentation calculation unit is further used to receive a plurality of similarities, take the test answer segment corresponding to the similarity reaching a preset threshold as the effective test answer segment, and send the effective test answer segment and the reference answer segment to the integrity assessment unit; The integrity evaluation unit is further configured to receive the valid test answer segments and the reference answer segments, and calculate the completeness of the test answer relative to the reference answer according to the number of the valid test answer segments and the number of the reference answer segments.
5. The large model hallucination test and evaluation system according to claim 2, characterized in that: The accuracy evaluation unit is used to receive the reference answer and the test answer, and send the reference answer and the test answer to the effective segmentation calculation unit; The effective segmentation calculation unit is used to receive the reference answer and the test answer, and input the reference answer and the test answer into the segmentation unit; The segmentation unit is used to receive the reference answer and the test answer, use the large model LLM to split the reference answer and the test answer respectively, obtain reference answer segments and test answer segments, and send the reference answer segments and the test answer segments to the effective segmentation calculation unit; The effective segmentation calculation unit is further used to receive the reference answer segmentation and the test answer segmentation, and input each segment in the reference answer segmentation and each segment in the test answer segmentation into the similarity calculation unit one by one; The similarity calculation unit is used to receive each segment in the reference answer segment and each segment in the test answer segment, calculate multiple similarities between each segment in the reference answer segment and each segment in the test answer segment by cosine similarity, and send the multiple similarities to the effective segment calculation unit; The effective segmentation calculation unit is further used to receive a plurality of the similarities, take the test answer segment corresponding to the similarity reaching a preset threshold as the effective test answer segment, and send the effective test answer segment and the test answer segment to the accuracy evaluation unit; The accuracy evaluation unit is further configured to receive the valid test answer segments and the test answer segments, and calculate the accuracy of the test answer relative to the reference answer based on the number of the valid test answer segments and the number of the test answer segments.
6. A data processing method for a test and evaluation system of a large model illusion based on any one of claims 1 to 5, characterized in that: The data processing method comprises: The benchmark question answer module sends an initial benchmark question to the test question adaptation module, and sends an initial benchmark answer corresponding to the initial benchmark question to the benchmark answer adaptation module; The test question adaptation module formats the initial benchmark question to obtain a test question, and sends the test question to the intelligent question answering system; The intelligent question-answering system performs large-model fine-tuning or retrieval enhancement generation on the test question to obtain an initial test answer, and sends the initial test answer and the test question to the test answer formatting module; The test answer formatting module formats the initial test answer to obtain a test answer, and sends the test answer and the test question to the evaluation module; The benchmark answer adaptation module formats the initial benchmark answer to obtain a benchmark answer, and sends the benchmark answer to the evaluation module; The evaluation module calculates the relevance, completeness and accuracy based on the test question, the test answer and the benchmark answer, and sends the relevance, completeness and accuracy to the evaluation report module; The evaluation report module generates a large model hallucination evaluation result of the intelligent question-answering system based on the relevance, the completeness and the accuracy.
7. The data processing method of the large model hallucination test and evaluation system according to claim 6, characterized in that: The benchmark answer adaptation module formats the initial benchmark answer to obtain a benchmark answer, specifically including: The benchmark answer adaptation module splits and modifies the initial benchmark answer through AI reasoning to obtain multiple initial benchmark answer segments, wherein each initial benchmark answer segment contains a fact; The benchmark answer adaptation module combines a plurality of the initial benchmark answers in segments to obtain the benchmark answer.
8. The data processing method of the large model hallucination test and evaluation system according to claim 6, characterized in that: The evaluation module calculates the relevance, completeness and accuracy according to the test question, the test answer and the benchmark answer, specifically including: The evaluation module calculates the correlation between the test question and the test answer based on the test question and the test answer; The evaluation module calculates the completeness and accuracy of the test answer relative to the benchmark answer based on the benchmark answer and the test answer.
9. The data processing method of the large model hallucination test and evaluation system according to claim 8, characterized in that: The evaluation module calculates the correlation between the test question and the test answer according to the test question and the test answer, specifically including: The evaluation module generates a plurality of new questions according to the test question through LLM, and respectively calculates a plurality of similarities between the test question and the plurality of new questions through cosine similarity; The evaluation module calculates the correlation between the test question and the test answer according to the multiple similarities through an average calculation function.
10. The data processing method of the large model hallucination test and evaluation system according to claim 8, characterized in that: The evaluation module calculates the completeness and accuracy of the test answer relative to the benchmark answer based on the benchmark answer and the test answer, specifically including: The evaluation module uses the large model LLM to split the benchmark answer and the test answer respectively to obtain benchmark answer segments and test answer segments, and sends the benchmark answer segments and the test answer segments to the effective segmentation calculation unit; The evaluation module calculates multiple similarities between each segment in the reference answer segment and each segment in the test answer segment by cosine similarity, and sends the multiple similarities to the effective segment calculation unit; The evaluation module uses the test answer segment corresponding to the similarity reaching the preset threshold as the valid test answer segment; The evaluation module calculates the completeness of the test answer relative to the benchmark answer based on the number of the valid test answer segments and the number of the benchmark answer segments; The evaluation module calculates the accuracy of the test answer relative to the benchmark answer according to the number of the valid test answer segments and the number of the test answer segments.
Citation Information
Patent Citations
Big language model-based illusion detection method and system and storage medium
CN117688164A
Answer quality evaluation method and system based on divide-and-conquer agent
CN118551851A
Assessment method and device of large language model system and related equipment
CN119179631A
Method and system for evaluating quality of RAG knowledge base driven by large language model
CN119226753A
Question and answer system evaluation method and device, electronic equipment and storage medium
CN119440884A