Interactive artificial intelligence question answering evaluation method and system for education field
Through the interactive artificial intelligence Q&A evaluation method for the field of education, using Q&A big model for simulated Q&A and multi-dimensional evaluation, the problem of low efficiency and inaccurate results in the existing technology Q&A system evaluation is solved, and efficient and accurate Q&A evaluation is achieved, reducing costs and improving the objectivity of the evaluation.
Patent Information
- Application Number
- CN202510275568.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-27
AI Technical Summary
Existing interactive Q&A systems have shortcomings in assessing the quality of responses, handling complex questions, and maintaining response consistency and relevance, and relying on manual assessments has high cost, inefficiency, and subjectivity problems.
An interactive artificial intelligence Q&A evaluation method for the education field is proposed. By obtaining real Q&A data, using Q&A big model to simulate Q&A, and using model answers and teacher answers as benchmarks, multi-dimensional evaluation of grammatical readability, content quality and interactive ability is carried out, and overall scores and visual reports are finally generated.
It improves evaluation efficiency and accuracy, reduces manual intervention and costs, provides more objective and impartial evaluation results, and supports accurate feedback in educational applications to ensure the educational value of the model output.
Smart Images

Figure CN120219121A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of intelligent teaching, and particularly relates to an interactive artificial intelligence question answering and evaluation method and system for the education field. Background Art
[0002] In the education field, interactive question answering systems (also known as intelligent question answering systems or virtual teaching assistants) have been widely used. These systems mainly rely on natural language processing (NLP) and machine learning technologies. Especially in recent years, the application of large language models (such as the GPT series of OpenAI, BERT of Google, and various language models specifically designed for the education field) in intelligent question answering has achieved remarkable progress.
[0003] However, existing interactive question answering systems face the following technical challenges: (1) Lack of standardization in answer quality evaluation: Most systems lack consistent criteria to measure the quality of question answering and cannot accurately evaluate the performance of the model. Current evaluation methods usually rely on manual annotation or review, which is inefficient and difficult to scale. (2) Limited ability to handle complex questions: Large language models can generate fluent answers, but when dealing with complex questions or multi-round reasoning, they may output inaccurate or incomplete answers, and there is a lack of effective evaluation methods to detect these problems. (3) Poor response consistency and relevance: The evaluation methods in existing technologies usually cannot fully measure the model's ability to understand context, resulting in inconsistent performance of the model in multi-round conversations and possibly providing irrelevant or incorrect answers.
[0004] Currently, the evaluation of interactive question answering systems mainly still relies on manual work. Manual evaluation methods usually rely on human reviewers to manually check the system's answers, and the evaluation criteria include accuracy, relevance, language fluency, etc. This method can fully simulate the evaluation scenarios in actual applications. Especially when dealing with complex interactive questions and answers, it can comprehensively consider user needs and context. However, manual evaluation has problems such as high cost, low efficiency, and strong subjectivity, and it cannot meet the needs of real-time evaluation in large-scale applications.
[0005] To make up for the deficiencies of manual evaluation, automated evaluation methods have emerged. These methods usually evaluate by comparing the similarity between the model output and the reference answer, using metrics such as BLEU and ROUGE. These methods are suitable for standardized question answering tasks, but when dealing with multi-round interactions or more complex question answering scenarios, they cannot comprehensively capture the actual performance of the model, often ignoring the model's understanding of user intentions and context adaptability, resulting in evaluation results deviating from the actual effects.
[0006] The current automatic evaluation methods are still in the development stage. Although they perform well in some standardized tasks, they still cannot replace the comprehensiveness and accuracy of manual evaluation in complex interactive Q&A scenarios. Therefore, how to improve the evaluation efficiency and objectivity while ensuring the evaluation quality has become an important research direction. Summary of the Invention
[0007] The present invention aims to solve the deficiencies of the prior art and provides the following solutions:
[0008] An interactive artificial intelligence Q&A evaluation method for the education field, comprising the following steps:
[0009] Obtain the real Q&A data and process the real Q&A data to obtain test data, where the test data includes: background questions, student questions, and teacher answers;
[0010] According to the background questions and the student questions, use the Q&A large model to perform simulated Q&A to obtain the model answers of the Q&A large model;
[0011] Input the model answers and the test data into the evaluation model, and use the teacher answers as the benchmark to obtain evaluation results in different dimensions;
[0012] Integrate and weight the evaluation results in each dimension to obtain an overall score, and perform visualization processing to obtain a final evaluation report.
[0013] Preferably, the processing method includes:
[0014] Transcribe the real Q&A data in audio form into text form data;
[0015] Use prompt to extract information from the text form data to obtain preliminary extracted data;
[0016] Manually modify and screen the preliminary extracted data to obtain the test data.
[0017] Preferably, the method for obtaining the evaluation results in different dimensions includes: grammar readability dimension evaluation, content quality dimension evaluation, and interaction ability dimension evaluation;
[0018] The method for grammar readability dimension evaluation includes: calculating the matching degree between the model answer and the actual sentence to obtain perplexity; calculating the average number of words in each clause of the model answer and the proportion of adverbs and conjunctions in each sentence, and then calculating the readability index; calculating the grammar readability dimension evaluation result based on the perplexity and the readability index;
[0019] The method for evaluating the content quality dimension includes: inputting the model answer and the test data into the pre-trained evaluation model for evaluation to obtain content accuracy and content relevance; calculating the average value of the content accuracy and the content relevance to obtain the evaluation result of the content quality dimension;
[0020] The method for evaluating the interaction ability dimension includes: inputting the model answer and the test data into the pre-trained evaluation model for evaluation to obtain interaction quality evaluation, interaction rate evaluation, and emotional support evaluation, and integrating them to obtain the evaluation result of the interaction ability dimension.
[0021] Preferably, the method for calculating the evaluation result of the grammar readability dimension includes:
[0022]
[0023] where S represents the evaluation result of the grammar readability dimension, P represents the perplexity, w P represents the weighted exponent of the perplexity, R represents the readability index, w R represents the weighted exponent of the readability index.
[0024] The present invention also provides an interactive artificial intelligence question answering evaluation system for the education field. The system applies the method described in any one of the above, and includes: a data acquisition module, a simulation answer module, an evaluation module, and an integration module;
[0025] The data acquisition module is used to acquire the real question answering data and process the real question answering data to obtain test data, and the test data includes: background questions, student questions, and teacher answers;
[0026] The simulation answer module uses the question answering large model to perform simulation question answering according to the background question and the student question to obtain the model answer of the question answering large model;
[0027] The evaluation module is used to input the model answer and the test data into the evaluation model, and take the teacher answer as a benchmark to obtain evaluation results in different dimensions;
[0028] The integration module is used to integrate and weight the evaluation results in each dimension to obtain an overall score, and perform visualization processing to obtain a final evaluation report.
[0029] Preferably, the data acquisition module includes: a conversion unit, an extraction unit, and a screening unit;
[0030] The conversion unit is used to transcribe the real question answering data in audio form into text form data;
[0031] The extraction unit uses prompts to extract information from the text-form data, obtaining the preliminarily extracted data;
[0032] The screening unit is used to manually modify and screen the preliminarily extracted data to obtain the test data.
[0033] Preferably, the evaluation module includes: a grammar readability dimension evaluation unit, a content quality dimension evaluation unit, and an interaction ability dimension evaluation unit;
[0034] The workflow of the grammar readability dimension evaluation unit includes: calculating the matching degree between the model answer and the actual sentence to obtain the perplexity; calculating the average number of words in each clause of the model answer and the proportion of adverbs and conjunctions in each sentence, and then calculating the readability index; calculating the grammar readability dimension evaluation result based on the perplexity and the readability index;
[0035] The workflow of the content quality dimension evaluation unit includes: inputting the model answer and the test data into the pre-trained evaluation model for evaluation to obtain content accuracy and content relevance; calculating the average value of the content accuracy and the content relevance to obtain the content quality dimension evaluation result;
[0036] The workflow of the interaction ability dimension evaluation unit includes: inputting the model answer and the test data into the pre-trained evaluation model for evaluation to obtain interaction quality evaluation, interaction rate evaluation, and emotional support evaluation, and integrating to obtain the interaction ability dimension evaluation result.
[0037] Preferably, the method for calculating the grammar readability dimension evaluation result includes:
[0038]
[0039] where S represents the grammar readability dimension evaluation result, P represents the perplexity, w P represents the weighted index of the perplexity, R represents the readability index, w R represents the weighted index of the readability index.
[0040] Compared with the prior art, the beneficial effects of the present invention are:
[0041] (1) Improved evaluation efficiency: Compared with traditional manual evaluation, the automated evaluation method of the present invention can greatly reduce manual intervention and shorten the evaluation time; according to experimental observations, when the present invention completes the detection of a sample (excluding the waiting time for the model answer), the average time consumption is about 50 seconds, while experts need at least 120 seconds to complete the evaluation task of the same dimension scale; compared with manual evaluation, the evaluation efficiency of the present invention is increased by at least one time, and no manual intervention is required at all;
[0042] (2) Evaluation accuracy and quality improvement: By adopting specific evaluation metrics and standardized evaluation processes, the present invention can ensure a high degree of consistency and accuracy of evaluation results. Especially with the use of specific prompt designs, it can guide large language models to focus on key factors during the evaluation process, making the scoring process more interpretable and consistent. The present invention not only effectively handles complex multi-round conversations and context information, making the evaluation results more comprehensive, but also reduces human bias compared to manual evaluation, ensuring more objective and fair scoring. Compared with existing automatic evaluation methods, the present invention can provide more accurate evaluations, especially in educational applications, and can more accurately reflect the actual performance of the model.
[0043] (3) Cost savings and reduction of human intervention: Manual evaluation usually requires the participation of multiple reviewers, and due to the large workload, the cost of hiring experts is relatively high. In contrast, using large language models for automated evaluation, although it also requires certain computing resources, the overall cost is much lower than hiring experts for manual evaluation. Through the automated evaluation method of the present invention, the evaluation cost can be significantly reduced while maintaining the objectivity and accuracy of the evaluation results.
[0044] (4) Support for accurate feedback in educational applications: By combining specific knowledge accuracy, logic, and practicality indicators in the educational field, the present invention provides an evaluation tool that is more suitable for educational scenarios. The present invention can effectively evaluate the rationality of model outputs in educational applications, avoid deviations or misguidance in educational content in the answers generated by the model, and ensure the educational value of the generated results.
[0045] In summary, the present invention not only significantly improves the efficiency and accuracy of the evaluation of the interactive Q&A system, but also has significant technical advantages in terms of cost savings, reduction of human intervention, and improvement of the actual effect in educational applications, meeting the urgent needs of the current education industry for intelligent evaluation tools. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] To more clearly illustrate the technical solutions of the present invention, the following briefly introduces the drawings required for the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0047] Figure 1 It is a schematic flowchart of the method according to the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0048] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0049] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.
[0050] Embodiment 1
[0051] In this embodiment, as Figure 1 shown, an interactive artificial intelligence question-answering evaluation method for the education field includes the following steps:
[0052] S1. Obtain the real question-answering data and process the real question-answering data to obtain the test data, where the test data includes: background questions, student questions, and teacher answers.
[0053] The selection and design of test data are the key links for the successful implementation of the evaluation method of the present invention. Reasonable test data can ensure the effectiveness and representativeness of the evaluation results, directly affecting the accuracy and generalization ability of the evaluation model. Therefore, when screening teacher-student Q&A conversations in real educational scenarios, the selection and design of test data must follow the following basic principles: (1) Clear direction of solution: Each test data must contain a clear and solvable knowledge point or problem. The data should cover teacher-student conversations with clear answers to ensure that the model can provide effective and educationally compliant feedback during the test. (2) Comprehensiveness and balance: The test data should cover multiple subjects, different knowledge areas, and various difficulty levels to ensure the good performance of the evaluation model in multi-field and multi-dimensional applications. During implementation, the test data will cover all basic disciplines at all stages of compulsory education to ensure that there are corresponding samples for each subject and grade. For any grade and subject, the number of samples should not be less than N (N = 10). At the same time, to avoid data bias, strict balance control will be carried out on the test data to ensure that the proportion of each subject in the data does not exceed 25%. (3) Design targeting the pain points of the model: In the design of test data, special attention should be paid to the weak links commonly existing in existing large language models. For example, the data should include problems involving mathematical reasoning, spatial imagination, and practical application of knowledge to test whether the model has the ability to effectively solve these complex problems. In specific implementation, such "difficult" test data should account for about 30% of the overall data. (4) Detection of emotional support ability: Some data should also include dialogue scenarios showing students' low or depressed moods to evaluate the model's performance in emotional support.
[0054] The processing methods include: transcribing the real Q&A data in audio form into text form data; using prompts to extract information from the text form data to obtain the initially extracted data; and manually modifying and screening the initially extracted data to obtain the test data.
[0055] In this embodiment, the data processing method refers to how to convert the teacher-student Q&A dialogue in the real education scenario into the format required for evaluation. Generally, teacher-student dialogue data exists in the form of text or audio. If it is audio, it is first necessary to transcribe the audio into text. Subsequently, using a specific prompt, information extraction is performed on the transcribed text with the help of a large language model. Given that there may be uncontrollable factors in the output of the large language model, the extracted text needs to be manually modified and screened. During the modification process, the following principles need to be followed: (1) Student questions: The extracted student questions should be the key questions raised by the students in the original material, ensuring the representativeness and accuracy of the questions. (2) Teacher answers (benchmark answers): The extracted teacher answers should accurately reflect the original intention of the teacher, and if the answers contain key points of step-by-step solutions, this information must be completely retained in the modified text. Ensure the accuracy of this part of the information during manual screening. (3) Background questions: The extracted background questions should be accurate and have sufficient breadth to summarize the core problems and related knowledge points faced by the current students.
[0056] S2. According to the background questions and student questions, use the Q&A large model to conduct simulated Q&A to obtain the model answers of the Q&A large model.
[0057] In this embodiment, according to the background questions and student conversations, simulate multiple rounds of Q&A and obtain the responses of the Q&A large model in each round, that is, the model answers. If the large model to be evaluated is not specifically for the field of educational Q&A, a system prompt: "You are a teacher good at answering students' questions" is uniformly added to the input to guide the model to generate answers that meet the requirements of the educational scenario.
[0058] S3. Input the model answers and test data into the evaluation model, and use the teacher answers as the benchmark to obtain evaluation results in different dimensions.
[0059] The methods for obtaining evaluation results in different dimensions include: evaluation in the dimension of grammar readability, evaluation in the dimension of content quality, and evaluation in the dimension of interaction ability.
[0060] The methods for evaluating the grammar readability dimension include: calculating the matching degree between the model answer and the actual sentence to obtain the perplexity; calculating the average number of words in each clause of the model answer and the proportion of adverbs and conjunctions in each sentence, and then calculating the readability index; calculating the evaluation result of the grammar readability dimension based on the perplexity and the readability index.
[0061] In this embodiment, when evaluating the grammar readability dimension, the present invention combines two metrics, namely Perplexity and Readability Index, to comprehensively reflect the grammar quality of the text. Perplexity is used to measure the matching degree between the language model generating the text and the actual sentence, reflecting the naturalness of the sentence structure. Specifically, the lower the perplexity, the closer the sentence is to the word order that might be generated in the large language model, indicating that the grammar structure is more natural and fluent. The perplexity of each sentence is calculated using the GPT-2 model, and the output value of the perplexity of GPT-2 represents the degree of understanding of the sentence by the model. According to the perplexity analysis of the content generated by the large model, the reasonable range of perplexity is determined to be between 8 and 15. The lower the perplexity, the more compliant the sentence grammar is with the norms of natural language. The Readability Index measures the readability and clarity of a sentence through two metrics, specifically the average number of words per clause and the ratio of adverbs to conjunctions in the sentence: (1) readability1: the average number of words per clause. This metric reflects the suitability of the sentence length, and longer sentences may affect readability; (2) readability2: the ratio of adverbs and conjunctions in each sentence. Adverbs and conjunctions help enhance the fluency of the sentence, but too many may make the sentence become complex and difficult to understand. Based on the above metrics, the Readability Index is calculated as follows:
[0062] Readability Index = readability1 + readability2 * 10
[0063] According to the readability analysis of primary and secondary school textbooks, the reasonable range of the Readability Index is determined to be between 40 and 75.
[0064] Combining the calculation results of perplexity and the Readability Index, the comprehensive score for the grammar dimension is finally obtained, and this comprehensive score should be within the range of 0 to 100. According to the normal range, the perplexity (P) is usually between 8 and 15, and the readability metric (R) is between 40 and 75. Therefore, a formula is needed to linearly compress the normal ranges of these two metrics to between 60 and 100. To this end, the method for calculating the evaluation result of the grammar readability dimension includes:
[0065]
[0066] where S represents the evaluation result of the grammar readability dimension, P represents perplexity, w P represents the weighted index of perplexity, R represents the Readability Index, w R represents the weighted index of the Readability Index, w P + w R = 1, and here w P = w R = 0.5.
[0067] The methods for evaluating the content quality dimension include: inputting the model answer and test data into a pre-trained evaluation model for evaluation to obtain content accuracy and content relevance; calculating the average value of content accuracy and content relevance to obtain the evaluation result of the content quality dimension.
[0068] In the calculation of the content quality dimension of this embodiment, the main evaluation indicators include accuracy and relevance, and these two indicators will be evaluated by calling a large language model.
[0069] In the evaluation of the content quality dimension, accuracy is a key indicator, mainly evaluated by a large language model. The core of accuracy evaluation is to judge the objective accuracy of the answer generated by the model in terms of knowledge content and whether it completely contains the key points of the reference prompt in the teacher's answer. Specifically, in the data processing stage, the student question, background question, and reference answer are extracted, and the answer of the question-answering large model is obtained using the student question and background question. Subsequently, the answer to be evaluated by the model, the reference answer, the student question, and the background question are input into the evaluation large model together to let it judge whether the reply of the question-answering large model meets the reference requirements. The requirements for meeting the reference include: (1) the answer must meet the necessary requirements prompted in the reference, such as providing the definition of a certain concept or performing a necessary calculation step; (2) the final result must be correct, such as selecting the correct answer in a true or false question, or giving the correct numerical value in a question related to calculation. To ensure the objectivity and accuracy of the evaluation, before the judgment, the evaluation large model will receive some scoring examples to help it understand the scoring criteria. The evaluation large model will score the input answer, with the scoring range from 1 to 5 points, and provide detailed judgment reasons during the scoring process to ensure that the evaluation process is transparent and interpretable.
[0070] Relevance evaluation is carried out by evaluating whether the answer is closely related to the content of the question, including entity consistency and context coherence. The evaluation model mainly determines by referring to two aspects: (1) entity consistency, which means whether the answer and the question are centered around the same knowledge point or entity to ensure that the model does not deviate from the topic; (2) context coherence, which means whether the content of the answer is logically consistent with the question to ensure that there is no digression or misunderstanding. In the evaluation of this dimension, only the relevance between the answer and the question is concerned, and accuracy is not involved. Therefore, during the evaluation process, only the background question, the student question, and the model answer need to be input into the evaluation model. Similar to the accuracy evaluation, to ensure the consistency and objectivity of the evaluation, the evaluation model will first receive some scoring examples to help it understand the scoring criteria. Finally, the evaluation large model will score the answer, with the scoring range from 1 to 5 points, and provide detailed judgment basis when scoring to ensure that the evaluation process is transparent and interpretable.
[0071] Calculate the average of content accuracy and content relevance: In this evaluation method, one student question and one model answer are regarded as one round of conversation. As mentioned before, each evaluation of accuracy and relevance is carried out for a single round of conversation, that is, whenever the model answers a question raised by the student, the answer will be scored for accuracy and relevance. In practical applications, a sample usually contains multiple rounds of conversations. Therefore, the average of the accuracy and relevance scores of all single-round conversations will be used as the comprehensive score of the sample. And the average of the accuracy and relevance scores of all samples represents the overall score of the model in terms of accuracy and relevance. In addition, the average score of the i-th round of conversation for each sample can be further calculated to obtain the performance of the model in the i-th round of conversation. By observing whether there is a downward trend in the accuracy and relevance scores of the model in multiple rounds of conversations, the stability of the model can be judged, and thus its performance in dealing with continuous interactions can be understood. When presenting, the scores of accuracy and relevance under the content dimension will be presented separately, so they need to be percentile-normalized separately. Since the values output by the evaluation model are in the range of 1-5, multiplying them by 20 will obtain the corresponding percentile score, that is:
[0072] Percentile score = Average score output by the evaluation model * 20
[0073] The methods for evaluating the interaction ability dimension include: inputting the model answer and test data into a pre-trained evaluation model for evaluation to obtain the evaluation of interaction quality, interaction rate, and emotional support, and integrating them to obtain the evaluation result of the interaction ability dimension.
[0074] In this embodiment, the interaction ability dimension mainly evaluates the interaction effect of the model in the educational scenario, covering three core indicators: interaction quality, interaction rate, and emotional support.
[0075] The interactive quality assessment focuses on whether the model can effectively guide students to gradually delve deeper into knowledge points during the conversation and perform appropriate knowledge transfer. Specifically, during the assessment process, it is necessary to analyze the knowledge guidance performance of the model in each round of conversation, especially whether the model can help students think and master new knowledge through questioning, guiding statements, or prompts. In this evaluation, the evaluation model will score the complete conversation of each sample according to the preset criteria, with the scoring range from 1 to 5 points, and give the reasons for the evaluation. The evaluation model will provide corresponding scoring examples for reference to ensure the consistency and transparency of the scoring process. During the evaluation process, in order to facilitate the analysis of the context by the evaluation model and conduct a comprehensive evaluation, it is necessary to input the conversation content between the student and the Q&A model of the entire sample. The evaluation model will score the interactive quality according to the following criteria: (1) Knowledge guidance: Whether the model can clearly put forward guiding questions related to the problem to help students clarify concepts or reason out answers; (2) Progressive advancement: Whether the model can gradually guide students to deeply understand knowledge points in multiple rounds of conversation, avoiding simple repetition; (3) Knowledge transfer: Whether the model can effectively transfer the learned knowledge to new problem situations to help students establish cross-domain connections.
[0076] The interaction rate assessment mainly measures the frequency of the model's guiding behavior in the conversation, especially whether the model enhances the interactivity of the conversation by asking questions, encouraging students to think, or suggesting examples. The scoring of the interaction rate will be based on the following criteria: (1) Guiding questions: For example, "Can you explain it again?" "How do you understand this concept?" etc.; (2) Prompting thinking or giving examples: For example, "Think about this problem", "Give an example to illustrate this concept". In each round of conversation, the frequency and effectiveness of guiding questions will directly affect the scoring of the interaction rate. A high interaction rate indicates that the model inspires students to think more in the conversation, promotes students' active participation, and thus improves the teaching effect. Specifically, the evaluation model will calculate the total number of sentences in all model answers, then calculate the number of sentences containing guiding behaviors, and then obtain the proportion of guiding sentences in the total number of sentences, that is, the interaction rate:
[0077]
[0078] The emotional support evaluation focuses on whether the model can effectively provide emotional support in the conversation, especially in helping students maintain a positive learning attitude. The evaluation model will score the emotional support performance of each round of conversation according to the preset criteria, with a score range of 1 to 5, and provide corresponding scoring cases for reference. To facilitate the comprehensive analysis of the evaluation model, the conversation content between students and the Q&A model of the entire sample needs to be input during the evaluation, so that the model can accurately evaluate based on the context of the conversation. The evaluation model will evaluate the scoring emotional support dimension from the following aspects: (1) Positivity: Whether the model uses encouraging language to motivate students, such as "very good", "well done", etc., to help students build confidence; (2) Warm and friendly tone: Whether the model expresses care through a warm and friendly tone, avoiding cold or negative wording; (3) Emotional care: Whether the model pays attention to the emotional state of students and provides psychological encouragement at appropriate times, such as "Don't worry, you're doing great!", "It's normal to encounter difficulties, keep going!" etc.
[0079] As can be seen from the above, for each round of conversation, the large model will give a score of 1 to 5 in terms of interaction quality and emotional support dimension. In order to convert the score into a range of 0 to 100, the percentage score is also used for score conversion:
[0080] Percentage score = average score output by the evaluation model * 20
[0081] For the interaction rate, a standard of 10% is set as 100 points. The interaction rate is converted through the following formula. Assuming the interaction rate is x%, then the score Y of the interaction rate is:
[0082] Y = min(100, 100 * (x * 10) 1 / γ )
[0083] Among them, γ represents the parameter that controls the smoothness of the curve rise. The larger it is, the smoother it is. Here, γ = 2 is taken.
[0084] S4. Integrate and weight the evaluation results of each dimension to obtain an overall score, and perform visualization processing to obtain a final evaluation report.
[0085] In this embodiment, in the last step of the evaluation process, it is first necessary to integrate the scoring results of each dimension and sub-dimension to generate a comprehensive score. These scores not only reflect the performance of the model in each independent dimension, but also obtain the overall score through weighted calculation to comprehensively evaluate the Q&A quality of the model. Subsequently, the scoring results are presented to the user through visualization techniques to help the user intuitively understand the performance of each dimension. For example, the score distribution of several dimensions can be displayed through a radar chart to intuitively reflect the advantages and disadvantages of the model in different fields. Based on these evaluation results, the system automatically generates a feedback report. The report content includes detailed scores, analysis of evaluation dimensions, and scores of each indicator.
[0086] Embodiment 2
[0087] In this embodiment, an interactive artificial intelligence Q&A evaluation system for the education field includes: a data acquisition module, a simulation answering module, an evaluation module, and an integration module.
[0088] The data acquisition module is used to obtain real Q&A data and process the real Q&A data to obtain test data. The test data includes: background questions, student questions, and teacher answers. The data acquisition module includes: a conversion unit, an extraction unit, and a screening unit. The conversion unit is used to transcribe the real Q&A data in audio form into text form data; the extraction unit uses prompts to extract information from the text form data to obtain initially extracted data; the screening unit is used to manually modify and screen the initially extracted data to obtain test data.
[0089] The simulation answering module uses the Q&A large model to perform simulated Q&A based on the background questions and student questions to obtain the model answers of the Q&A large model.
[0090] The evaluation module is used to input the model answers and test data into the evaluation model, and take the teacher answers as the benchmark to obtain evaluation results in different dimensions. The evaluation module includes: a grammar readability dimension evaluation unit, a content quality dimension evaluation unit, and an interaction ability dimension evaluation unit.
[0091] The work process of the grammar readability dimension evaluation unit includes: calculating the matching degree between the model answer and the actual sentence to obtain the perplexity; calculating the average number of words in each clause of the model answer and the proportion of adverbs and conjunctions in each sentence, and then calculating the readability index; calculating the grammar readability dimension evaluation result based on the perplexity and the readability index; the method for calculating the grammar readability dimension evaluation result includes:
[0092]
[0093] Among them, S represents the grammar readability dimension evaluation result, P represents the perplexity, w PThe weighted index representing perplexity, R represents the readability index, w R The weighted index representing the readability index.
[0094] The workflow of the content quality dimension evaluation unit includes: inputting the model answer and test data into a pre-trained evaluation model for evaluation to obtain content accuracy and content relevance; calculating the average value of content accuracy and content relevance to obtain the content quality dimension evaluation result;
[0095] The workflow of the interaction ability dimension evaluation unit includes: inputting the model answer and test data into a pre-trained evaluation model for evaluation to obtain interaction quality evaluation, interaction rate evaluation, and emotional support evaluation, and integrating them to obtain the interaction ability dimension evaluation result.
[0096] The integration module is used to integrate and weight the evaluation results of each dimension to obtain an overall score, and perform visualization processing to obtain a final evaluation report.
[0097] The embodiments described above are only descriptions of the preferred embodiments of the present invention, and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.
Claims
1. An interactive artificial intelligence question-answering evaluation method for the education field, characterized in that: The following steps are involved: Acquire real question-answering data, and process the real question-answering data to obtain test data, wherein the test data includes: background questions, student questions, and teacher answers; According to the background questions and the student questions, the question-answering model is used to simulate question-answering to obtain a model answer of the question-answering model; Inputting the model answers and the test data into the evaluation model, and taking the teacher answers as a benchmark, obtaining evaluation results of different dimensions; The evaluation results of each dimension are integrated and weighted to obtain an overall score, which is then visualized to obtain a final evaluation report.
2. According to claim 1, an interactive artificial intelligence question-answering evaluation method for the education field is characterized in that: The processing method includes: Transcribing the question-answering real data in audio form into text form data; Use prompt to extract information from the textual data to obtain preliminary extracted data; The initially extracted data are manually modified and screened to obtain the test data.
3. According to claim 1, an interactive artificial intelligence question-answering evaluation method for the education field is characterized in that: The method for obtaining evaluation results of different dimensions includes: grammatical readability dimension evaluation, content quality dimension evaluation and interactive ability dimension evaluation; The method for evaluating the grammatical readability dimension includes: calculating the degree of matching between the model answer and the actual sentence to obtain the perplexity; calculating the average number of words in each sentence in the model answer and the proportion of adverbs and conjunctions in each sentence to calculate the readability index; calculating the grammatical readability dimension evaluation result based on the perplexity and the readability index; The method for evaluating the content quality dimension includes: inputting the model answer and the test data into the pre-trained evaluation model for evaluation to obtain content accuracy and content relevance; calculating the average value of the content accuracy and the content relevance to obtain a content quality dimension evaluation result; The method for evaluating the interactive ability dimension includes: inputting the model answer and the test data into the pre-trained evaluation model for evaluation, obtaining an interactive quality evaluation, an interactive rate evaluation and an emotional support evaluation, and integrating them to obtain an interactive ability dimension evaluation result.
4. According to claim 3, an interactive artificial intelligence question-answering evaluation method for the education field is characterized in that: The method for calculating the grammatical readability dimension evaluation result includes: Among them, S represents the evaluation result of the grammatical readability dimension, P represents the perplexity, and w P represents the weighted index of perplexity, R represents the readability index, and w R A weighted index representing the readability index.
5. An interactive artificial intelligence question-answering and evaluation system for the education field, the system applying the method described in any one of claims 1 to 4, characterized in that: include: Data acquisition module, simulated answer module, evaluation module and integration module; The data acquisition module is used to acquire real question-answering data and process the real question-answering data to obtain test data, wherein the test data includes: background questions, student questions and teacher answers; The simulated answering module uses the big question-answering model to simulate question-answering according to the background questions and the student questions, and obtains the model answer of the big question-answering model; The evaluation module is used to input the model answer and the test data into the evaluation model, and obtain evaluation results of different dimensions based on the teacher's answer; The integration module is used to integrate and weight the evaluation results of each dimension to obtain an overall score, and perform visualization to obtain a final evaluation report.
6. The interactive artificial intelligence question-answering and evaluation system for the education field according to claim 5, characterized in that: The data acquisition module includes: a conversion unit, an extraction unit and a screening unit; The conversion unit is used to transcribe the question-answering real data in audio form into text form data; The extraction unit uses prompt to extract information from the text form data to obtain preliminary extracted data; The screening unit is used to manually modify and screen the initially extracted data to obtain the test data.
7. The interactive artificial intelligence question answering and evaluation system for the education field according to claim 5, characterized in that: The evaluation module includes: a grammatical readability dimension evaluation unit, a content quality dimension evaluation unit and an interactive capability dimension evaluation unit; The workflow of the grammatical readability dimension evaluation unit includes: calculating the matching degree between the model answer and the actual sentence to obtain the perplexity; calculating the average number of words in each sentence in the model answer and the proportion of adverbs and conjunctions in each sentence to calculate the readability index; calculating the grammatical readability dimension evaluation result based on the perplexity and the readability index; The workflow of the content quality dimension evaluation unit includes: inputting the model answer and the test data into the pre-trained evaluation model for evaluation to obtain content accuracy and content relevance; calculating the average value of the content accuracy and the content relevance to obtain a content quality dimension evaluation result; The workflow of the interactive ability dimension evaluation unit includes: inputting the model answers and the test data into the pre-trained evaluation model for evaluation, obtaining an interactive quality evaluation, an interactive rate evaluation and an emotional support evaluation, and integrating them to obtain an interactive ability dimension evaluation result.
8. The interactive artificial intelligence question answering and evaluation system for the education field according to claim 7, characterized in that: The method for calculating the grammatical readability dimension evaluation result includes: Among them, S represents the evaluation result of the grammatical readability dimension, P represents the perplexity, and w P represents the weighted index of perplexity, R represents the readability index, and w R A weighted index representing the readability index.
Citation Information
Cited By
Monitoring system and method based on enterprise WeChat group answering service
CN120612211A