Automatic calculation method and system for technical support question-answering system
The ClozeFact and StepRestore methods evaluate the key terms and operation steps of the technically supported question-and-answer system are solved, and the answer evaluation in existing systems is achieved, which achieves higher accuracy and reliability, and is suitable for large-scale technical support scenarios.
Patent Information
- Application Number
- CN202510616940.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-08-15
AI Technical Summary
The existing technical support question and answer system has shortcomings in evaluating the accuracy of answers, especially the inability to effectively detect key term matching, step sequence and integrity of answers, resulting in the generation of misleading answers.
The large language evaluation model is used to ensure the accuracy of key terms and the correctness of the sequence of steps through key terms cloze-blanks and operational step reconstruction techniques, including the ClozeFact and StepRestore methods, respectively, to evaluate the accuracy of key terms and the sequence and completeness of the operational steps.
It improves the evaluation accuracy and reliability of the technically supported Q&A system, reduces the errors caused by LLM generation illusions, ensures the operability of the answers, and improves the credibility and user experience of the system.
Smart Images

Figure CN120493905A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of large language models, and in particular to an automated calculation method and system for a technical support question-answering system. Background Art
[0002] In the field of IT technical support, technical support question-and-answer (TSA) systems are designed to analyze technical questions raised by users (such as enterprise customers, product end users, or developers) and generate accurate solutions based on relevant documentation. They are widely used in scenarios such as enterprise technical support, online customer service, and troubleshooting. Traditional technical support relies primarily on human experts to analyze issues. Experts analyze issues based on their expertise and experience, and then retrieve information from knowledge bases, product manuals, or troubleshooting guides to provide users with solutions. However, this manually driven model has many limitations, such as slow response times, high labor costs, and difficulty efficiently handling large numbers of user requests. In recent years, with the rapid development of artificial intelligence (AI) technology, automated technical support QA systems based on large language models (LLMs) and retrieval-augmented generation (RAG) have gained widespread application.
[0003] The technical support question-and-answer system based on LLM and RAG mainly generates answers and returns them to users through the following steps: Information retrieval (Retrieval): Extract paragraphs related to the question from technical documents or databases; Answer generation (Generation): Use LLM to combine the retrieved information to generate the final answer, organize the answer content according to the logical structure of technical support, and provide specific solutions to help users solve technical problems.
[0004] Figure 1 This paper demonstrates the typical processing flow of a technical support question-answering system: the system receives a user-entered question, retrieves the K most relevant paragraphs (in this case, K is 5) from reference documents, and generates a step-by-step answer based on the LLM. However, existing technical support question-answering systems still face severe accuracy issues, primarily due to the following reasons: hallucination of large language models: LLMs may generate seemingly reasonable but actually incorrect answers; and imprecise knowledge retrieval: the system fails to retrieve the key information needed to answer the question, resulting in incomplete or misleading answers.
[0005] These issues can cause technical support Q&A systems to generate misleading solutions, preventing users from effectively resolving their issues and potentially leading to more serious system failures or business interruptions. Therefore, a reliable automated evaluation framework is crucial for detecting incorrect answers, providing an effective benchmark for optimizing technical support Q&A systems.
[0006] However, existing technical support question-answering evaluation methods struggle to reliably assess the accuracy of answers generated by technical support question-answering systems. Table 1 lists the types of errors that can occur in technical support question-answering systems. Due to the limitations of existing evaluation methods, some incorrect answers may not be accurately identified, causing the optimization direction of the technical support question-answering system to deviate from expectations. For example, incorrect answers may be misclassified as correct, which in turn affects system training and parameter adjustments, making it difficult for improvement measures to effectively improve answer quality.
[0007] Table 1
[0008]
[0009] Existing Technologies: RefChecker uses a knowledge graph-based triple extraction method to convert the output of a question-answering system into triples. The triples are then compared with reference answers for factual consistency verification, thereby determining the authenticity of the generated content. RAGAS is primarily used to evaluate the answer quality of retrieval-augmented generation (RAG) systems. This framework proposes multiple evaluation metrics, including accuracy, answer relevance, and contextual consistency, to comprehensively measure the performance of RAG-generated answers across different dimensions. RAGQuestEval generates key factual questions based on reference answers and uses LLM to generate answers (i.e., key facts) based on the Q&A system's answers. The answers are then compared to determine if the key facts match the reference answers, thereby assessing the factual consistency of the answers. FActScore extracts atomic facts from text and verifies them item by item to measure the factual accuracy of long text generation tasks, thereby assessing the credibility of content generated by Q&A systems. BERTScore, based on BERT word embeddings, measures the semantic proximity between generated text and reference answers by calculating the cosine similarity of text embedding vectors. It is used for automatic evaluation of text generation and natural language processing tasks. ROUGE is an evaluation method based on n-gram word matching, widely used in automatic summarization and text generation tasks. It measures text similarity by calculating the n-gram overlap between the generated text and the reference answer. BLEU is an automatic evaluation metric used in machine translation and text generation tasks. This method measures text similarity by calculating the n-gram co-occurrence rate between the generated text and the standard answer. Patent CN202311585325.6 describes this method, which constructs an evaluation prompt and takes the question, the LLM-generated answer, the standard answer, or the reference text as input, directly asking the LLM to make an accuracy judgment. The prompt sets the LLM to compare the answers as a scorer and output only "correct" or "incorrect."
[0010] Disadvantages of existing technologies: 1. Methods based on vocabulary matching (ROUGE, BLEU) only calculate text similarity through n-grams and cannot identify synonym substitutions, word order changes or the correctness of key facts. As a result, the question-answering system may mistakenly judge answers with different expressions but the same meaning as wrong, or judge answers containing key factual errors as correct when evaluating answers. 2. Methods based on semantic matching (BERTScore) focus on overall semantic similarity, while ignoring the inspection of specific key terms in technical support questions and answers (such as parameter names, command formats, configuration paths, etc.), which may cause answers with overall semantic similarity but key factual errors to be misjudged, thereby affecting the reliability of the evaluation results. 3. Methods based on RAG evaluation (RAGAS, RAGQuestEval) mainly focus on whether the answer comes from the retrieved document, rather than evaluating whether the answer omits key steps or information. Therefore, it cannot effectively detect the step sequence and completeness of the answer. 4. LLM-based black-box scoring methods (patent CN202311585325.6) rely on the LLM to directly output scores. However, due to hallucinations, the LLM may mistakenly identify some incorrect answers as correct, or judge answers that are correct but expressed differently as incorrect, resulting in unreliable evaluation results. Furthermore, this method is significantly affected by the capabilities of the underlying LLM model used for evaluation.
[0011] Although LLM-based methods for verifying fine-grained facts (RefChecker, RAGQuestEval, FActScore) can detect the factual accuracy of answers, they cannot evaluate whether the execution logic of the operation steps is reasonable. This may cause answers that contain all necessary information but in the wrong order to be mistakenly judged as correct. In technical support scenarios, incorrect execution order will make it impossible to effectively solve user problems. Summary of the Invention
[0012] The present invention aims to solve one of the technical problems in the related art at least to a certain extent.
[0013] To this end, the present invention proposes an automated calculation method for a technical support question-and-answer system. Given a technical support question Q and the corresponding reference answer (i.e., the correct operating steps) GT, as well as an answer A generated by the technical support question-and-answer system to be evaluated, the goal of the evaluation task is to calculate a score S to measure the quality of A.
[0014] Another object of the present invention is to provide an automated computing system for a technical support question-answering system.
[0015] To achieve the above objectives, the present invention provides an automated calculation method for a technical support question-answering system, comprising:
[0016] Build a large language evaluation model and use technical support questions, reference answers, and answers generated by the technical support question-and-answer system as model input data;
[0017] Through the key term cloze test of the large language evaluation model, the accuracy of the key terms in the answer is verified and the set of key terms that failed to match is obtained;
[0018] The large language evaluation model is used to reconstruct the operation steps in the reference answer based on the answer, and calculate whether the answer contains the correctness and completeness of all the operation steps mentioned in the reference answer to obtain the calculation results of the correctness of the step sequence and the calculation results of the step completeness;
[0019] The final calculation score is obtained based on the set of key terms that failed to match, the calculation results of the correctness of the step sequence, and the calculation results of the step completeness.
[0020] The automated calculation method for the technical support question-answering system according to the embodiment of the present invention may also have the following additional technical features:
[0021] In one embodiment of the present invention, a key term cloze test using a large language assessment model is performed to verify the accuracy of the key terms in the answer and obtain a set of key terms that failed to match, including:
[0022] Identify key terms from the reference answers based on predefined rules and replace them with placeholders<BLANK[i]> , to obtain the reference answer GT′ with the key terms hidden;
[0023] The reference answer GT′ without the key terms and the answer A are concatenated and input into the evaluation model LLM Eval , and make LLM Eval Fill in all GT's according to A<BLANK[i]> the contents of the office;
[0024] Determine whether the filled content matches the original key term exactly or fuzzily. If the filled content does not match the original key term at some point, it will be counted as a key term that failed to match.
[0025] In one embodiment of the present invention, the key term set defining the reference answer is K = {k1, k2, ..., k n}, the calculation process is as follows:
[0026] GT′=MaskKeyTerms(GT)
[0027] A′=FillBlanks(GT′,A;LLM Eval )
[0028] ε key=MatchKeyTerms(A′,K)
[0029] Among them, ε key is the set of key terms that failed to match, and A′ is the sequence of steps after filling.
[0030] In one embodiment of the present invention, the large language evaluation model is used to reconstruct the operation steps in the reference answer based on the answer, and the correctness and completeness of all the operation steps mentioned in the reference answer are calculated to obtain the step sequence correctness calculation results and the step completeness calculation results, including:
[0031] Randomly shuffle the operation steps in GT to obtain a random operation step list GT″;
[0032] By evaluating the model LLM Eval Select and rearrange the steps in GT″ according to A, so that GT″ matches the logical order of the answer, and obtain the reconstructed operation steps A rec ;
[0033] Calculate A rec Whether all the operation steps in GT are included to obtain the step completeness calculation results;
[0034] Calculate A rec Whether the relative order of the operation steps in GT is consistent, so as to obtain the calculation result of the correctness of the step sequence.
[0035] In one embodiment of the present invention, the method further includes:
[0036] GT″=ShuffleSteps(GT)
[0037] A rec =ReorderSteps(GT″,A;LLM Eval )
[0038] S comp =CheckCompleteness(A rec ,GT)
[0039] S ord =CheckOrder(A rec ,GT)
[0040] Among them, S comp is the step integrity calculation result, S ord Calculate the results for step sequence correctness.
[0041] In one embodiment of the present invention, the final calculation score is obtained based on the set of key terms that failed to match, the calculation results of the correctness of the step sequence, and the calculation results of the completeness of the step, including:
[0042]
[0043] In one embodiment of the present invention, a technical support question Q and a related reference document set D are preset. The technical support question answering system retrieves the most relevant paragraph D′ and uses LLM to QA Generate answer A:
[0044] D′=Retrieve K (Q,D)
[0045] A=Generate(D′,Q;LLM QA )
[0046] Take the answer A, the reference answer GT and the question Q as input to use the evaluation model LLM Eval Calculate the accuracy score S of A:
[0047] S=Evaluate(A,GT,Q;LLM Eval )
[0048] The optimization goal of the evaluation model is to maximize the ability to distinguish correct and incorrect answers, measured by the AUC metric:
[0049] maxAUC(S,S * )
[0050] Among them, S * For real score.
[0051] To achieve the above objectives, the present invention further provides an automated computing system for a technical support question-answering system, comprising:
[0052] A model building module is used to build a large language assessment model and use technical support questions, reference answers, and answers generated by the technical support question-and-answer system as model input data;
[0053] A key term accuracy calculation module is used to perform matching verification calculations on the accuracy of key terms in the answers using a large language evaluation model to obtain a set of key terms that failed to match;
[0054] The step sequence and completeness calculation module is used to reconstruct the operation steps in the reference answer based on the answer through the large language evaluation model, and calculate whether the answer contains the correctness and completeness of all the operation steps mentioned in the reference answer, so as to obtain the calculation results of the correctness of the step sequence and the calculation results of the step completeness;
[0055] The final score calculation module is used to obtain the final calculation score based on the key term set that failed to match, the step sequence correctness calculation results, and the step completeness calculation results.
[0056] The automated calculation method and system for a technical support question-and-answer system, as described in embodiments of the present invention, can effectively improve the accuracy and reliability of technical support question-and-answer system assessments. This improves the overall reliability of the system, and the use of automated scoring makes the assessment results of different question-and-answer systems more comparable, facilitating system optimization and improvement while reducing manual assessment costs. This improves the credibility of the system and enhances user experience, making it suitable for large-scale technical support scenarios.
[0057] Additional aspects and advantages of the present invention will be set forth in part in the description which follows and, in part, will be obvious from the description which follows, or may be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0059] Figure 1 This is a typical workflow diagram of the existing technical support question-answering system based on LLM and RAG;
[0060] Figure 2 is a flow chart of an automated calculation method for a technical support question-answering system according to an embodiment of the present invention;
[0061] Figure 3 is an architectural diagram of an automated computing method for a technical support question-answering system according to an embodiment of the present invention;
[0062] Figure 4 4 is a structural diagram of an automated computing system for a technical support question-answering system according to an embodiment of the present invention. DETAILED DESCRIPTION
[0063] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments of the present invention can be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0064] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0065] The following describes an automated calculation method and system for a technical support question-answering system according to an embodiment of the present invention with reference to the accompanying drawings.
[0066] Figure 2 is a flow chart of an automated calculation method for a technical support question-answering system according to an embodiment of the present invention. Figure 2 As shown, the method includes but is not limited to the following steps:
[0067] S1, builds a large language evaluation model and uses technical support questions, reference answers, and answers generated by the technical support question-answering system as model input data;
[0068] S2, through the key term cloze test of the large language evaluation model, to verify the accuracy of the key terms in the answer and obtain the set of key terms that failed to match;
[0069] S3, using the large language evaluation model to reconstruct the operation steps in the reference answer based on the answer, and calculate whether the answer contains the correctness and completeness of all the operation steps mentioned in the reference answer, so as to obtain the calculation results of the correctness of the step sequence and the calculation results of the step completeness;
[0070] S4, obtains the final calculation score based on the set of key terms that failed to match, the calculation results of the correctness of the step sequence, and the calculation results of the step completeness.
[0071] The overall structure of the present invention is as follows Figure 3 As shown in the figure, the present invention formally defines the evaluation problem of the technical support question answering system: given a technical support question Q and a set of related reference documents D, the technical support question answering system retrieves the most relevant K paragraphs D′ and evaluates the results based on the LLM. QA Generate answer A:
[0072] D′=Retrieve K (Q,D)
[0073] A=Generate(D′,Q;LLM QA )
[0074] The goal of the problem is to design an evaluation model Evaluate, which takes the answer A, the reference answer GT (containing all necessary solution steps) and the question Q as input, and uses LLM Eval Calculate the accuracy score S of A:
[0075] S=Evaluate(A,GT,Q;LLM Eval )
[0076] The optimization goal of the evaluation model is to maximize the ability to distinguish correct and incorrect answers, measured by the AUC (Area Under Curve) metric:
[0077] maxAUC(S,S * )
[0078] Among them, S * For real score.
[0079] In one embodiment of the present invention, a key term accuracy assessment (ClozeFact) is performed. ClozeFact assesses the accuracy of key terms (such as commands, parameters, file paths, configuration options, etc.) in the answer by having the LLM perform a key term cloze test. The detailed implementation steps are as follows:
[0080] S21, Extract key terms (MaskKeyTerms): Identify key terms from the reference answer GT based on predefined rules (whether they match the domain term list) and replace them with placeholders<BLANK[i]> , and obtain the reference answer GT′ with the key terms hidden.
[0081] S22, FillBlanks: The processed reference answer GT' and answer A are concatenated and input into the evaluation model LLM Eval , and requires LLM Eval Fill in all GT's according to A<BLANK[i]> The content of the place.
[0082] S23, Match Verification (MatchKeyTerms): Determine whether the filled content matches the original key term exactly or fuzzily (the tolerance threshold can be customized as needed). If the filled content does not match the original key term at any point, it will be counted as a key term that failed to match.
[0083] Formally, the key term set of the reference answer is defined as K = {k1, k2, ..., k n}, the calculation and evaluation process of this stage can be expressed as follows:
[0084] GT′=MaskKeyTerms(GT)
[0085] A′=FillBlanks(GT′,A;LLM Eval )
[0086] ε key =MatchKeyTerms(A′,K)
[0087] Among them, ε key is the set of key terms that failed to match, and A′ is the sequence of steps after filling.
[0088] The prompt used by FillBlanks(·) is as follows:
[0089] Given the provided text,replace each placeholder BLANK with the corresponding key term based on the given response.If the requiredinformation is not explicitly mentioned,return "Unanswerable".Ensure that the filled terms exactly match those in the reference.
[0090] In one embodiment of the present invention, the operation step sequence and integrity assessment (StepRestore) is performed. StepRestore allows the LLM to reconstruct the operation steps based on the answer, evaluate whether the answer contains all the operation steps mentioned in the GT, and ensure that the step sequence is correct. The detailed implementation steps are as follows:
[0091] S31, shuffle the operation steps in the reference answer (ShuffleSteps): randomly shuffle the operation steps in GT to obtain a random operation step list GT″.
[0092] S32, Step ReorderSteps: Using the Evaluation Model LLM Eval , requiring them to select and rearrange the steps in GT″ according to A to match the logical order of the answer, and obtain the reconstructed operation steps A rec .
[0093] S33, Check Completeness: Calculate and verify A rec Whether all operation steps in GT are included to obtain the step completeness calculation result. If yes, return 1, otherwise return 0.
[0094] S34, Check Order: Calculate and verify A rec Whether the relative order of the operation steps in GT is consistent with that in GT, so as to obtain the correctness calculation result of the step sequence. If yes, it returns 1, otherwise it returns 0.
[0095] Formally, the computational evaluation process at this stage can be expressed as follows:
[0096] GT″=ShuffleSteps(GT)
[0097] A rec =ReorderSteps(GT″,A;LLM Eval )
[0098] S comp =CheckCompleteness(A rec ,GT)
[0099] S ord =CheckOrder(A rec ,GT)
[0100] Among them, S comp is the step integrity calculation result, S ord Calculate the results for step sequence correctness.
[0101] The prompt used by ReorderSteps(·) is as follows:
[0102] Based on the provided text, identify and arrange the mentioned steps in the correct logical execution order. Only include steps explicitly stated in the text, and ignore any steps not mentioned, as they are misleading options. Only use the steps listed in the given options.
[0103] Furthermore, in the actual evaluation process, since the above two evaluation stages have no interdependence, they will be triggered in parallel to reduce the total calculation time. After both evaluation stages are completed, the output evaluation score S can be calculated. The default calculation method is as follows:
[0104]
[0105] The [cond] notation indicates a Boolean test of the condition cond, returning 1 if cond holds, and 0 otherwise.
[0106] The default product calculation method is used because the technical support Q&A scenario is a serious one. Any errors in key facts in the answers will make it difficult for users to effectively solve the problem. Therefore, errors in any evaluation stage will result in the final output evaluation score being 0. This invention also supports custom output evaluation score calculation methods to accommodate other scenarios with less stringent requirements.
[0107] The automated calculation method for a technical support question-and-answer system, based on the innovative ClozeFact and StepRestore technologies, effectively improves the accuracy and reliability of technical support question-and-answer system evaluation. ClozeFact ensures correct key term matching, reducing errors caused by LLM generation illusions; StepRestore verifies the order and completeness of operation steps, preventing misleading answers due to incorrect or omitted steps. Experimental results demonstrate that this method is more accurate and reliable than all investigated methods, achieving a 7.6% improvement in Area Under Correspondence (AUC) over the currently best-performing method, RefChecker. This method better meets the evaluation requirements of technical support question-and-answer systems. Compared to existing evaluation methods that primarily measure text similarity, this method focuses on the execution logic of technical support answers, accurately determining whether answers contain all necessary steps and arrange them in the correct order, thereby ensuring the operability of the answers. This method can improve the overall reliability of technical support question-and-answer systems. Its automated scoring approach makes evaluation results from different question-and-answer systems more comparable, facilitating system optimization and improvement. It also reduces manual evaluation costs, enhances the credibility of technical support question-and-answer systems, and improves user experience. It is suitable for large-scale technical support scenarios. The present invention has broad applicability, applicable to both traditional knowledge-based question-answering systems and LLM-based retrieval-augmented generation (RAG) systems. It can provide stable and reliable evaluation scores for question-answering systems based on different principles, and offer accurate feedback for system optimization. The present invention offers the advantages of high efficiency and low cost, achieving high computational efficiency while ensuring high accuracy. Compared to existing evaluation methods, it has lower computational overhead and is suitable for large-scale enterprise-level question-answering evaluations, thereby ensuring the optimization efficiency of technical support question-answering systems.
[0108] In order to implement the above embodiment, Figure 4 As shown, this embodiment also provides an automated computing system 10 for a technical support question-answering system, including:
[0109] A model building module 100 is used to build a large language evaluation model and use technical support questions, reference answers, and answers generated by the technical support question-answering system as model input data;
[0110] A key term accuracy calculation module 200 is configured to perform a matching verification calculation on the accuracy of key terms in the answer by using a key term cloze test of a large language evaluation model to obtain a set of key terms that failed to match;
[0111] The step sequence and completeness calculation module 300 is used to reconstruct the operation steps in the reference answer based on the answer using the large language evaluation model, and calculate whether the answer contains the correctness and completeness of all the operation steps mentioned in the reference answer, so as to obtain the step sequence correctness calculation results and the step completeness calculation results;
[0112] The final score calculation module 400 is used to obtain a final calculation score based on the set of key terms that failed to match, the calculation results of the correctness of the step sequence, and the calculation results of the completeness of the step.
[0113] Furthermore, the key term accuracy calculation module 200 is further configured to:
[0114] Identify key terms from the reference answers based on predefined rules and replace them with placeholders<BLANK[i]> , to obtain the reference answer GT′ with the key terms hidden;
[0115] The reference answer GT′ without the key terms and the answer A are concatenated and input into the evaluation model LLM Eval , and make LLM Eval Fill in all GT's according to A<BLANK[i]> the contents of the office;
[0116] Determine whether the filled content matches the original key term exactly or fuzzily. If the filled content does not match the original key term at some point, it will be counted as a key term that failed to match.
[0117] Furthermore, the step sequence and integrity calculation module 300 is further configured to:
[0118] Randomly shuffle the operation steps in GT to obtain a random operation step list GT″;
[0119] By evaluating the model LLM Eval Select and rearrange the steps in GT″ according to A, so that GT″ matches the logical order of the answer, and obtain the reconstructed operation steps A rec ;
[0120] Calculate A rec Whether all the operation steps in GT are included to obtain the step completeness calculation results;
[0121] Calculate A rec Whether the relative order of the operation steps in GT is consistent, so as to obtain the calculation result of the correctness of the step sequence.
[0122] According to the automated computing system for a technical support question-and-answer system according to an embodiment of the present invention, the present invention focuses on the execution logic of the technical support answer, can accurately determine whether the answer contains all the necessary steps, and arranges them in the correct order, thereby ensuring the operability of the answer. The present invention can improve the overall reliability of the technical support question-and-answer system, and adopts an automated scoring method to make the evaluation results of different question-and-answer systems more comparable, facilitate system optimization and improvement, and reduce the cost of manual evaluation, improve the credibility and user experience of the technical support question-and-answer system, and is suitable for large-scale technical support scenarios. The present invention has a wide range of applicability and can be applied to traditional question-and-answer systems based on knowledge bases, as well as to retrieval enhancement generation (RAG) systems based on LLMs. It can provide stable and reliable evaluation scores for question-and-answer systems based on different principles, and provide accurate feedback for system optimization. The present invention has the advantages of high efficiency and low cost, and has high computational efficiency while ensuring high accuracy. Compared with existing evaluation methods, it has lower computational overhead and is suitable for large-scale enterprise-level question-and-answer evaluation, thereby ensuring the optimization efficiency of the technical support question-and-answer system.
[0123] In the description of this specification, the reference terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and features of different embodiments or examples without contradiction.
[0124] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one such feature. In the description of the present invention, "plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.
Claims
1. An automated calculation method for a technical support question-answering system, characterized in that: include: Build a large language evaluation model and use technical support questions, reference answers, and answers generated by the technical support question-and-answer system as model input data; Through the key term cloze test of the large language evaluation model, the accuracy of the key terms in the answer is verified and the set of key terms that failed to match is obtained; The large language evaluation model is used to reconstruct the operation steps in the reference answer based on the answer, and calculate whether the answer contains the correctness and completeness of all the operation steps mentioned in the reference answer to obtain the calculation results of the correctness of the step sequence and the calculation results of the step completeness; The final calculation score is obtained based on the set of key terms that failed to match, the calculation results of the correctness of the step sequence, and the calculation results of the step completeness.
2. The method according to claim 1, characterized in that Through the key term cloze test of the large language evaluation model, the accuracy of the key terms in the answer is verified and matched, and the set of key terms that failed to match is obtained, including: Identify key terms from the reference answers based on predefined rules and replace them with placeholders<BLANK[i]> , to obtain the reference answer GT′ with the key terms hidden; The reference answer GT′ without the key terms and the answer A are concatenated and input into the evaluation model LLM Eval , and make LLM Eval Fill in all GT's according to A<BLANK[i]> the contents of the office; Determine whether the filled content matches the original key term exactly or fuzzily. If the filled content does not match the original key term at some point, it will be counted as a key term that failed to match.
3. The method according to claim 2, characterized in that Define the key term set of the reference answer as K = {k1, k2, ..., k n }, the calculation process is as follows: GT′=MaskKeyTerms(GT) A′=FillBlanks(GT′,A;LLM Eval ) ε key =MatchKeyTerms(A′,K) Among them, ε key is the set of key terms that failed to match, and A′ is the sequence of steps after filling.
4. The method according to claim 3, characterized in that The large language evaluation model is used to reconstruct the operation steps in the reference answer based on the answer, and calculate whether the answer contains the correctness and completeness of all the operation steps mentioned in the reference answer, so as to obtain the calculation results of the correctness of the step sequence and the completeness of the step, including: Randomly shuffle the operation steps in GT to obtain a random operation step list GT″; By evaluating the model LLM Eval Select and rearrange the steps in GT″ according to A, so that GT″ matches the logical order of the answer, and obtain the reconstructed operation steps A rec ; Calculate A rec Whether all the operation steps in GT are included to obtain the step completeness calculation results; Calculate A rec Whether the relative order of the operation steps in GT is consistent, so as to obtain the calculation result of the correctness of the step sequence.
5. The method according to claim 4, characterized in that The method further comprises: GT″=ShuffleSteps(GT) A rec =ReorderSteps(GT″,A;LLM Eval ) S comp =CheckCompleteness(A rec ,GT) S ord =CheckOrder(A rec ,GT) Among them, S comp is the step integrity calculation result, S ord Calculate the results for step sequence correctness.
6. The method according to claim 5, characterized in that The final calculation score is obtained based on the set of key terms that failed to match, the calculation results of the correctness of the step sequence, and the calculation results of the step completeness, including:
7. The method according to claim 6, characterized in that Preset technical support question Q and related reference document set D, the technical support question answering system retrieves the most relevant paragraph D′ and uses LLM QA Generate answer A: D′=Retrieve K (Q,D) A=Generate(D′,Q;LLM QA ) Take the answer A, the reference answer GT and the question Q as input to use the evaluation model LLM Eval Calculate the accuracy score S of A: S=Evaluate(A,GT,Q;LLM Eval ) The optimization goal of the evaluation model is to maximize the ability to distinguish correct and incorrect answers, measured by the AUC metric: max AUC(S,S * ) Among them, S * For real score.
8. An automated computing system for a technical support question-answering system, characterized in that: include: A model building module is used to build a large language assessment model and use technical support questions, reference answers, and answers generated by the technical support question-and-answer system as model input data; A key term accuracy calculation module is used to perform matching verification calculations on the accuracy of key terms in the answers using a large language evaluation model to obtain a set of key terms that failed to match; The step sequence and completeness calculation module is used to reconstruct the operation steps in the reference answer based on the answer through the large language evaluation model, and calculate whether the answer contains the correctness and completeness of all the operation steps mentioned in the reference answer, so as to obtain the calculation results of the correctness of the step sequence and the calculation results of the step completeness; The final score calculation module is used to obtain the final calculation score based on the key term set that failed to match, the step sequence correctness calculation results, and the step completeness calculation results.
9. The system according to claim 8, characterized in that Key term accuracy calculation module, also used for: Identify key terms from the reference answers based on predefined rules and replace them with placeholders<BLANK[i]> , to obtain the reference answer GT′ with the key terms hidden; The reference answer GT′ without the key terms and the answer A are concatenated and input into the evaluation model LLM Eval , and make LLM Eval Fill in all GT's according to A<BLANK[i]> the contents of the office; Determine whether the filled content matches the original key term exactly or fuzzily. If the filled content does not match the original key term at some point, it will be counted as a key term that failed to match.
10. The system according to claim 9, characterized in that The step sequence and integrity calculation module is also used to: Randomly shuffle the operation steps in GT to obtain a random operation step list GT″; By evaluating the model LLM Evai Select and rearrange the steps in GT″ according to A, so that GT″ matches the logical order of the answer, and obtain the reconstructed operation steps A rec ; Calculate A rec Whether all the operation steps in GT are included to obtain the step completeness calculation results; Calculate A rec Whether the relative order of the operation steps in GT is consistent, so as to obtain the calculation result of the correctness of the step sequence.
Citation Information
Patent Citations
Question answering system evaluation method, device, computing device and storage medium
CN117290694B