Evaluation method and device of education large model, terminal equipment and storage medium

By constructing the mapping relationship between the predicted Q&A pair of classroom transcripts of the education model and the standard Q&A pair, determining the evaluation strategy and calculating the evaluation index value, the problem that the existing technology cannot accurately evaluate the performance of the education model is solved, and efficient and accurate model evaluation is achieved.

CN120030116APending Publication Date: 2025-05-23NANJING HONGHE ARTIFICIAL INTELLIGENCE TECHNOLOGY RESEARCH INSTITUTE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411999929.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The prior art evaluation methods for educational big models focus on general competency testing, lack consideration of vertical field tasks, and cannot accurately evaluate the performance of educational big models used for classroom question-and-answer analysis.

Method used

By obtaining the predicted Q&A pair of classroom transcripts from the educational model, a mapping relationship is constructed based on the similarity between the predicted Q&A pair and the standard Q&A pair, the evaluation strategy is determined, and the evaluation index value is calculated to accurately evaluate the performance of the educational model.

Benefits of technology

It realizes an accurate evaluation of the performance of educational large models, improves the degree of fit between the evaluation results and actual applications, and can flexibly adapt to different evaluation needs and conducts fine-grained model evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030116A_ABST
    Figure CN120030116A_ABST
Patent Text Reader

Abstract

The invention is suitable for the technical field of artificial intelligence, and provides an evaluation method and device of an education large model, terminal equipment and a storage medium, and the method comprises the steps: obtaining a prediction question and answer pair obtained by predicting a classroom live record text through the education large model; according to the similarity between the predicted question and answer pair and a standard question and answer pair corresponding to the classroom live record text, constructing a mapping relation between the predicted question and answer pair and the standard question and answer pair, and determining an evaluation strategy based on the mapping relation; calculating an evaluation index value based on the evaluation strategy; and evaluating the large education model according to the evaluation index value. According to the method, the performance of the education large model for classroom question and answer pair analysis can be accurately and effectively evaluated, and the integrating degree of an evaluation result and actual application is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to an evaluation method, apparatus, terminal device and storage medium for a large educational model. Background Art

[0002] With the rapid development of artificial intelligence technology, especially the increasing maturity of deep learning models, educational big models refer to large language models (LLM) in the vertical field of education. The application of educational big models in the field of education is gradually becoming a hot spot in research and practice.

[0003] The educational big model used to analyze classroom questions and answers has broad application prospects in the field of education. However, the existing technical evaluation schemes for educational big models focus on general ability tests and lack consideration of vertical field tasks. The existing model evaluation methods have obvious limitations in the evaluation of educational big models used for classroom question and answer analysis and cannot accurately evaluate the performance of the model.

[0004] Therefore, how to accurately and effectively evaluate the performance of the educational big model used for classroom question-answer pair analysis and improve the fit between the evaluation results and practical applications is a problem that needs to be considered at present. Summary of the invention

[0005] The embodiments of the present application provide a method, apparatus, terminal device and storage medium for evaluating an educational big model, which can accurately and effectively evaluate the performance of the educational big model used for classroom question-answer pair analysis and improve the compatibility of the evaluation results with actual applications.

[0006] In a first aspect, the present application embodiment provides an evaluation method for a large education model, including:

[0007] Obtain predicted question-answer pairs obtained by predicting classroom transcripts using a large education model;

[0008] According to the similarity between the predicted question-answer pair and the standard question-answer pair corresponding to the classroom transcript text, a mapping relationship between the predicted question-answer pair and the standard question-answer pair is constructed, and an evaluation strategy is determined based on the mapping relationship;

[0009] Based on the evaluation strategy, calculating the evaluation index value;

[0010] The educational model is evaluated according to the evaluation index value.

[0011] In a possible implementation of the first aspect, the predicted question-answer pair includes a predicted question and a predicted answer, and the standard question-answer pair includes a standard question and a standard answer; and constructing a mapping relationship between the predicted question-answer pair and the standard question-answer pair according to a similarity between the predicted question-answer pair and the standard question-answer pair corresponding to the classroom transcript text, and determining an evaluation strategy based on the mapping relationship, includes:

[0012] Constructing a first mapping relationship according to a first similarity between the predicted question in the predicted question-answer pair and the standard question in the standard answer, and determining a first evaluation strategy based on the first mapping relationship;

[0013] constructing a second mapping relationship according to a second similarity between the predicted answer in the predicted question-answer pair and the standard answer in the standard answer, and determining a second evaluation strategy based on the second mapping relationship;

[0014] A third mapping relationship is constructed according to the first similarity and the second similarity, and a third evaluation strategy is determined based on the third mapping relationship.

[0015] In a possible implementation manner of the first aspect, there is more than one evaluation strategy, and the calculating the evaluation index value based on the evaluation strategy includes:

[0016] Determining a first evaluation index value based on the evaluation strategy, the number of the predicted question-answer pairs, and the number of the standard question-answer pairs;

[0017] A second evaluation index value is determined based on the evaluation strategy, the number of words in the predicted question-answer pair, and the number of words in the standard question-answer pair.

[0018] In a possible implementation of the first aspect, the first evaluation indicator value includes precision, recall, and area under the curve; and determining the first evaluation indicator value based on the evaluation strategy, the number of predicted question-answer pairs, and the number of standard question-answer pairs includes:

[0019] When the evaluation strategy is the first evaluation strategy, a first precision is determined according to the ratio of the number of correctly predicted questions to the total number of predicted questions, a first recall is determined according to the ratio of the number of correctly predicted questions to the total number of standard questions in the standard question-answer pairs, and a first area under the curve is determined according to the first precision and the first recall; or

[0020] When the evaluation strategy is the second evaluation strategy, a second precision is determined according to the ratio of the number of correctly predicted answers to the total number of predicted answers; a second recall is determined according to the ratio of the number of correctly predicted answers to the total number of standard answers in the standard question and answer pairs, and a second area under the curve is determined according to the second precision and the second recall; or,

[0021] When the evaluation strategy is the third evaluation strategy, the third precision is determined according to the ratio of the number of correctly predicted question and answer pairs in the predicted question and answer pairs to the total number of predicted question and answer pairs; the third recall is determined according to the ratio of the number of correctly predicted question and answer pairs to the total number of standard question and answer pairs; and the area under the third curve is determined according to the third precision and the third recall.

[0022] In a possible implementation manner of the first aspect, the second evaluation indicator value includes fidelity; and determining the second evaluation indicator value based on the evaluation strategy, the number of words of the predicted question-answer pair, and the number of words of the standard question-answer pair includes:

[0023] When the evaluation strategy is the first evaluation strategy, the question fidelity is determined according to the first question ratio, the second question ratio and the total number of predicted questions in the predicted question-answer pair, wherein the first question ratio is the ratio of the number of common question words in each group of predicted questions and their corresponding standard questions to the total number of predicted question words in the predicted question-answer pair under the first evaluation strategy, and the second question ratio is the ratio of the number of common question words to the total number of standard question words in the standard question-answer pair; or,

[0024] When the evaluation strategy is the second evaluation strategy, the answer fidelity is determined according to the first answer ratio, the second answer ratio and the total number of predicted answers in the predicted question and answer pair, wherein the first answer ratio is the ratio of the number of common words of answers in each group of predicted answers and their corresponding standard answers to the total number of predicted answer words in the predicted question and answer pair under the second evaluation strategy, and the second answer ratio is the ratio of the number of common words of answers to the total number of standard answer words in the standard question and answer pair; or,

[0025] When the evaluation strategy is the third evaluation strategy, the question fidelity is determined according to the third question ratio, the fourth question ratio and the total number of predicted answer pairs under the third evaluation strategy, wherein the third question ratio is the ratio of the number of common question words in each group of predicted questions and their corresponding standard questions to the total number of predicted question words in the predicted question-answer pair under the third evaluation strategy, and the fourth question ratio is the ratio of the number of common question words under the third evaluation strategy to the total number of standard question words in the standard question-answer pair; the answer fidelity is determined according to the third answer ratio, the fourth answer ratio and the total number of predicted answer pairs in the predicted question-answer pair under the third evaluation strategy, wherein the third answer ratio is the ratio of the number of common answer words in each group of predicted answers and their corresponding standard answers to the total number of predicted answer words in the predicted question-answer pair under the third evaluation strategy, and the fourth answer ratio is the ratio of the number of common answer words under the third evaluation strategy to the total number of standard answer words in the standard question-answer pair.

[0026] In a possible implementation of the first aspect, obtaining predicted question-answer pairs obtained by predicting classroom transcript text using the education macro model includes:

[0027] When receiving the evaluation instruction, searching for the classroom transcript text and the corresponding predicted question-answer pair corresponding to the educational model based on the data index;

[0028] If the data index does not exist, then obtain the classroom transcript text;

[0029] The classroom transcript text and prompts are input into the educational model, and the prompts are used to guide the educational model to predict the predicted question and answer pairs in the classroom transcript text.

[0030] In a possible implementation of the first aspect, obtaining the classroom transcript text includes:

[0031] Obtaining an initial classroom record, and transcribing the initial classroom record to obtain a transcribed text;

[0032] The transcribed text is preprocessed to obtain a classroom transcript text.

[0033] In a second aspect, the embodiment of the present application provides an evaluation device for a large educational model, including:

[0034] A question-answer pair acquisition unit, used to acquire predicted question-answer pairs obtained by predicting classroom transcript texts using the educational big model;

[0035] A mapping strategy determination unit, configured to construct a mapping relationship between the predicted question-answer pair and the standard question-answer pair according to the similarity between the predicted question-answer pair and the standard question-answer pair corresponding to the classroom transcript text, and determine an evaluation strategy based on the mapping relationship;

[0036] An indicator value calculation unit, used to calculate the evaluation indicator value based on the evaluation strategy;

[0037] The model evaluation unit is used to evaluate the educational model according to the evaluation index value.

[0038] In a third aspect, an embodiment of the present application provides a terminal device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the evaluation method for the educational large model as described in the first aspect above is implemented.

[0039] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the evaluation method of the educational big model as described in the first aspect above is implemented.

[0040] In a fifth aspect, an embodiment of the present application provides a computer program product, which, when executed on a terminal device, enables the terminal device to execute the method for evaluating the educational model as described in the first aspect above.

[0041] In the embodiment of the present application, the terminal device obtains the predicted question-answer pair obtained by predicting the classroom transcript text using the educational big model, and then constructs a mapping relationship between the predicted question-answer pair and the standard question-answer pair according to the similarity between the predicted question-answer pair and the standard question-answer pair corresponding to the classroom transcript text, and determines the evaluation strategy based on the mapping relationship, and then calculates the evaluation index value based on the evaluation strategy, and finally evaluates the educational big model according to the evaluation index value. The present application scheme can determine the evaluation strategy based on the constructed mapping relationship so as to flexibly adapt to different evaluation needs. The use of multi-dimensional evaluation indicators is helpful for fine-grained model evaluation, and can accurately and effectively evaluate the performance of the educational big model used for classroom question-answer pair analysis, and comprehensively reflect the application effect of the model in the educational scenario, thereby improving the fit between the evaluation results and actual applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0043] Figure 1 It is a flow chart for implementing the evaluation method of the educational model provided in the embodiment of the present application;

[0044] Figure 2 It is a specific implementation flow chart of step S101 in the evaluation method of the educational model provided in the embodiment of the present application;

[0045] Figure 3 It is a specific implementation flow chart of step S102 in the evaluation method of the educational model provided in the embodiment of the present application;

[0046] Figure 4 It is a specific implementation flow chart of step S103 in the evaluation method of the educational model provided in the embodiment of the present application;

[0047] Figure 5 It is a structural block diagram of the evaluation device of the educational model provided in the embodiment of the present application;

[0048] Figure 6 It is a schematic diagram of a terminal device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0049] In the following description, specific details such as specific system structures, technologies, etc. are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application may also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to prevent unnecessary details from obstructing the description of the present application.

[0050] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, wholes, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or combinations thereof.

[0051] It should also be understood that the term “and / or” used in the specification and appended claims refers to any and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0052] As used in the specification and appended claims of this application, the term "if" can be interpreted as "when" or "uponce" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "uponce it is determined" or "in response to determining" or "uponce [described condition or event] is detected" or "in response to detecting [described condition or event]", depending on the context.

[0053] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.

[0054] References to "one embodiment" or "some embodiments" etc. described in the specification of this application mean that one or more embodiments of the present application include specific features, structures or characteristics described in conjunction with the embodiment. Therefore, the statements "in one embodiment", "in some embodiments", "in some other embodiments", "in some other embodiments", etc. that appear in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "including", "comprising", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized in other ways.

[0055] The existing evaluation methods have great limitations for evaluating the educational big model of this application, and cannot effectively evaluate the ability of the educational big model of this application to analyze classroom questions and answers.

[0056] In view of this, an embodiment of the present application provides an evaluation method for a large educational model. The method is applied to a terminal device. The terminal device determines an evaluation strategy based on the constructed mapping relationship so as to flexibly adapt to different evaluation needs. The use of multi-dimensional evaluation indicators is helpful for fine-grained model evaluation. It can accurately and effectively evaluate the performance of the large educational model used for classroom question-and-answer analysis, and comprehensively reflect the application effect of the model in educational scenarios, thereby improving the fit between the evaluation results and actual applications.

[0057] The evaluation method of the educational model provided in the embodiment of the present application is applicable to terminal devices of various types of data that need to perform the evaluation of the educational model, and the terminal devices may specifically include mobile phones, tablet computers, wearable devices, notebook computers, ultra-mobile personal computers (UMPC), desktop computers, interactive large screens, and servers, etc. The embodiment of the present application does not impose any restrictions on the specific types of terminal devices.

[0058] Figure 1 The implementation process of the evaluation method of the educational model provided in the embodiment of the present application is shown, and the method flow includes steps S101 to S104. The specific implementation principle of each step is as follows:

[0059] Step S101: Obtain predicted question-answer pairs obtained by predicting classroom transcript text using the educational big model.

[0060] The classroom transcript text is a text obtained by transcribing the audio and video of the classroom transcript. The classroom transcript text records the questions and answers that actually occurred in the classroom. The educational big model is a large language model used for classroom question and answer pair analysis in the vertical field of education. The educational big model uses deep learning and natural language processing technology to efficiently understand and parse the question and answer pairs in the classroom, assist teaching staff to better grasp the students' learning status, optimize teaching strategies, improve teaching quality, analyze students' thinking patterns and learning habits, and provide strong support for personalized teaching. In the embodiment of the present application, the classroom transcript text is analyzed using the educational big model to obtain at least one question and at least one answer predicted based on the classroom transcript text. According to the correspondence between the questions and answers generated by the educational big model, the questions and answers are combined into at least one predicted question and answer pair. The predicted question and answer pairs predicted by the educational big model can be one or more, and any predicted question and answer pair includes a predicted question and at least one predicted answer.

[0061] In one possible implementation, the predicted answer corresponding to the predicted question is selected based on the principle of the highest semantic relevance. The educational model performs semantic analysis on the predicted question and each predicted answer, vectorizes the text, converts the sentences of the predicted question and each predicted answer into vector representation, measures the semantic relevance by calculating the cosine similarity between vectors, and selects the predicted answer with the highest semantic similarity to the predicted question to form a predicted question-answer pair.

[0062] In one possible implementation, the predicted answer corresponding to the predicted question is selected based on the keyword matching degree. The core keywords in the predicted question are extracted, and the keywords in each predicted answer are also extracted. The predicted answer containing the keywords with the highest matching degree, the largest number and the most logical relationship with the question keywords is selected to form a question-answer pair.

[0063] In one possible implementation, the context of the classroom transcript text in which the predicted question is located is analyzed, and based on the context, the answer that can most logically follow the predicted question smoothly is selected as the predicted answer corresponding to the predicted question.

[0064] As a possible implementation of this application, Figure 2 A specific implementation process of step S101 in the evaluation method of the educational model provided in the embodiment of the present application is shown, and is described in detail as follows:

[0065] A1: When an evaluation instruction is received, the classroom transcript text corresponding to the educational model and its corresponding predicted question-answer pairs are searched based on the data index.

[0066] The data index is used to indicate the local storage path of the classroom transcript text and its corresponding predicted question and answer pairs in the terminal device, or the specific link address in the cloud disk. The data index can also indicate the storage location on a specific network server. The data index can comprehensively and flexibly help the terminal device locate and obtain the classroom transcript text and its corresponding predicted question and answer pairs, meeting the data management and call requirements in different storage environments and usage scenarios.

[0067] A2: If the data index does not exist, obtain the classroom transcript.

[0068] In a possible implementation, obtaining classroom transcripts includes: obtaining an initial classroom transcript, transcribing the initial classroom transcript to obtain a transcript; and preprocessing the transcript to obtain a classroom transcript. Preprocessing is used to optimize the transcript and remove low-quality text to improve the efficiency of the education model in processing and predicting classroom transcripts. Preprocessing includes denoising, data cleaning, etc.

[0069] In one possible implementation, the quality of the text in the transcribed text can be determined based on the completeness of the text sentences. The higher the sentence completeness, the higher the text quality. Conversely, the lower the sentence completeness, the lower the text quality. According to the specific application scenario, the sentence completeness threshold is automatically set or manually set by the education big model, and the sentences in the transcribed text with sentence completeness lower than the preset completeness threshold are removed to obtain the classroom transcript text.

[0070] A3: Input the classroom transcript text and prompts into the education model, and use the prompts to guide the education model to predict the question-answer pairs in the classroom transcript text.

[0071] The predicted question-answer pairs predicted by the educational model include predicted questions and predicted answers. In the embodiment of the present application, the powerful text comprehension ability of the educational model is used to find the content of the classroom question-answer pairs from the classroom transcript text. The question-answer pairs can be teacher-student question-answer pairs, that is, the teacher asks questions and the students answer.

[0072] In one possible implementation, the predicted question-answer pair output by the education model must meet the preset requirements. If the predicted question-answer pair output by the education model can be correctly parsed as JSON data, it is determined that the output predicted by the model meets the preset requirements, otherwise it is determined that it does not meet the preset requirements.

[0073] In this embodiment, the terminal device stores the classroom transcript text and the predicted question-answer pairs in the classroom transcript text predicted by the education model to a designated storage location, and establishes a data index according to the storage path of the classroom transcript text and the predicted question-answer pairs. In a possible implementation, the designated storage location also includes quality information of the classroom transcript text and quality information of the predicted question-answer pairs.

[0074] Step S102: construct a mapping relationship between the predicted question-answer pair and the standard question-answer pair corresponding to the classroom transcript text according to the similarity between the predicted question-answer pair and the standard question-answer pair, and determine an evaluation strategy based on the mapping relationship.

[0075] Predicted question-answer pairs include predicted questions and predicted answers, and standard question-answer pairs include standard questions and standard answers. Standard question-answer pairs are all question-answer pairs in the test data set that correspond to the classroom transcript text and have been manually or authoritatively recognized as accurate and standardized. Standard question-answer benchmarks are the correct combination of all questions and answers that should appear in the classroom transcript scenario under ideal circumstances. According to the similarity between the predicted question-answer pairs and the standard question-answer pairs, a mapping relationship between the predicted question-answer pairs and the standard question-answer pairs is constructed, which can match the output of the educational big model with the standard question-answer pairs one by one, providing the possibility for subsequent accurate comparison and analysis. The evaluation strategy is used to provide clear methods and rules for the evaluation of the educational big model. Different evaluation strategies are determined according to the different mapping relationships and characteristics between the predicted question-answer pairs and the standard question-answer pairs, which can conduct targeted evaluations on different aspects of the educational big model.

[0076] In the embodiment of the present application, by calculating the similarity between the predicted question-answer pair and the standard question-answer pair corresponding to the classroom transcript text, a mapping relationship between the predicted question-answer pair and the standard question-answer pair is constructed based on the similarity. The similarity refers to the degree of similarity between the predicted question-answer pair and the standard question-answer pair.

[0077] The similarity between the predicted question-answer pair and the standard question-answer pair includes structural similarity and semantic similarity. Structural similarity focuses on the structural similarity of the question-answer pair, such as the length of the sentence, sentence pattern, grammatical structure, etc. Semantic similarity refers to the similarity between two or more texts at the semantic level. It measures the closeness of the meaning expressed by the text, not just the similarity of vocabulary or grammatical structure. Due to the diversity of language expression, the same sentence can express the same meaning through different vocabulary and word order. In the embodiment of the present application, by calculating the semantic similarity of the predicted question-answer pair and the standard question-answer pair corresponding to the classroom transcript text, a mapping relationship between the predicted question-answer pair and the standard question-answer pair with the highest semantic similarity is constructed, and based on the constructed mapping relationship, an evaluation strategy is determined. Semantic similarity can break through the limitations of the surface form of the text, deeply understand the actual meaning relationship between words, sentences or texts, and avoid misjudging as different content only because of different expressions, thereby more accurately judging whether the predicted question-answer pair and the standard question-answer pair are semantically consistent or similar.

[0078] As a possible implementation of this application, Figure 3 A specific implementation process of step S102 in the evaluation method of the educational model provided in the embodiment of the present application is shown, and is described in detail as follows:

[0079] B1: constructing a first mapping relationship according to a first similarity between the predicted question in the predicted question-answer pair and the standard question in the standard answer, and determining a first evaluation strategy based on the first mapping relationship.

[0080] The first similarity is the semantic similarity between the predicted question in the predicted question-answer pair and the standard question in the standard answer. The first mapping relationship maps the correspondence between the question in the predicted question-answer pair and the question in the standard answer pair. The construction of the first mapping relationship only considers the first similarity between the predicted question and the standard question. The first evaluation strategy is an evaluation strategy for the prediction performance of the educational big model question based on the first mapping relationship. Specifically, it focuses on the evaluation of the questions output by the educational big model, and the first evaluation strategy helps to accurately judge the ability of the educational big model in the question extraction link. Specifically, the similarity between a predicted question in the predicted question-answer pair and all the standard questions in the standard question-answer pair is calculated, and the standard question with the highest similarity to the predicted question among all the standard questions is determined as the standard question corresponding to the predicted question, and the correspondence between the predicted question and its corresponding standard question is constructed; repeat this step until the determination of the correspondence between all the predicted questions in the predicted question-answer pair and their corresponding standard questions is completed, thereby obtaining the first mapping relationship between the predicted question-answer pair and the standard question-answer pair, and then based on the first mapping relationship, the first evaluation strategy is determined. In the first evaluation strategy, only questions are included in the evaluation. As long as the first similarity between the predicted question in the predicted question-answer pair output by the model and the corresponding standard question reaches the first preset threshold, the predicted question found by the model from the classroom transcript text is considered correct; otherwise, the predicted question output by the model is judged to be wrong. Among them, the first preset threshold is a pre-set measurement standard used to determine whether the similarity between the predicted question in the predicted question-answer pair and the standard question in the standard question-answer pair reaches an acceptable limit value. According to the specific application scenario, the first preset threshold is automatically set or manually set by the education big model.

[0081] When the first similarity between the predicted question in the predicted question-answer pair and its corresponding standard question does not reach the first preset threshold, it means that the predicted question is wrong, and the predicted question is recorded and the error type is marked. For example, if the semantic difference between the predicted question and all standard questions is too large, it can be marked as a "semantic understanding deviation" error; if only some key words are missing or the words are used improperly, resulting in low similarity, it is marked as an "improper use of vocabulary" error. The recording of incorrect predicted questions and the marking of error types will help to optimize the ability of the education model to extract questions in the future.

[0082] B2: Constructing a second mapping relationship according to a second similarity between the predicted answer in the predicted question-answer pair and the standard answer in the standard answer, and determining a second evaluation strategy based on the second mapping relationship.

[0083] The second similarity is the semantic similarity between the predicted answer in the predicted question and answer pair and the standard answer in the standard answer. The second mapping relationship maps the correspondence between the answer in the predicted question and answer pair and the answer in the standard answer pair. The construction of the second mapping relationship only considers the second similarity between the predicted answer and the standard answer. The second evaluation strategy is an evaluation strategy for the prediction performance of the education model answer based on the second mapping relationship. Specifically, it focuses on the evaluation of the answer output by the education model, and the second evaluation strategy helps to accurately judge the ability of the education model in the answer extraction link. Specifically, the similarity between a predicted answer in the predicted question and answer pair and all the standard answers in the standard question and answer pair is calculated, and the standard answer with the highest similarity to the predicted answer among all the standard answers is determined as the standard answer corresponding to the predicted answer, and the correspondence between the predicted answer and its corresponding standard answer is constructed; repeat this step until the determination of the correspondence between all the predicted answers in the predicted question and answer pair and their corresponding standard answers is completed, thereby obtaining the second mapping relationship between the predicted question and answer pair and the standard question and answer pair, and then based on the second mapping relationship, the second evaluation strategy is determined. In the second evaluation strategy, only the responses are included in the evaluation. As long as the second similarity between the predicted responses in the predicted question-answer pair output by the model and the corresponding standard responses reaches the second preset threshold, the predicted responses found by the model from the classroom transcript text are considered correct; otherwise, the predicted responses output by the model are judged to be wrong. Among them, the second preset threshold is a pre-set measurement standard used to determine whether the similarity between the predicted responses in the predicted question-answer pair and the standard responses in the standard question-answer pair reaches an acceptable limit. According to the specific application scenario, the second preset threshold is automatically set or manually set by the educational model.

[0084] When the second similarity between the predicted answer in the predicted question and answer pair and its corresponding standard answer does not reach the second preset threshold, it means that the predicted answer is wrong. The predicted answer is recorded and the error type is marked. The recording of the wrong predicted answer and the marking of the error type will help to optimize the subsequent ability of the education model to extract answers.

[0085] B3: Constructing a third mapping relationship according to a third similarity between the predicted answer in the predicted question-answer pair and the standard answer in the standard answer, and determining a third evaluation strategy based on the third mapping relationship.

[0086] The third similarity is the semantic similarity between the predicted question and answer pair and the standard question and answer pair in the predicted question and answer pair. The average of the first similarity and the second similarity is taken as the third similarity. The third mapping relationship maps the correspondence between the question and answer pairs in the predicted question and answer pair and the question and answer pairs in the standard question and answer pair. The construction of the third mapping relationship not only considers the first similarity between the predicted question in the predicted question and answer pair and the standard question in the standard question and answer pair, but also considers the second similarity between the predicted answer in the predicted question and answer pair and the standard answer in the standard question and answer pair. The third evaluation strategy is an evaluation strategy for the prediction performance of the question and answer pairs of the educational large model based on the third mapping relationship. Specifically, it focuses on the evaluation of the question and answer pairs output by the educational large model. The third evaluation strategy helps to accurately judge the ability of the educational large model in the question and answer pair extraction link. Specifically, the similarity between a predicted question in a predicted question-answer pair and all standard questions in a standard question-answer pair is calculated, and the similarity between a predicted answer in a predicted question-answer pair and all standard answers in a standard question-answer pair is calculated. The average of the two similarities is taken, and the standard question-answer pair with the highest average similarity among all standard question-answer pairs is determined as the standard question-answer pair corresponding to the predicted question-answer pair, and the corresponding relationship between the predicted question-answer pair and its corresponding standard question-answer pair is constructed; this step is repeated until the determination of the corresponding relationship between all predicted question-answer pairs and their corresponding standard question-answer pairs is completed, thereby obtaining a third mapping relationship between the predicted question-answer pair and the standard question-answer pair, and then based on the third mapping relationship, a third evaluation strategy is determined. In the third evaluation strategy, both questions and answers are included in the evaluation, and only when the average similarity between the predicted question and its corresponding standard question in the predicted question-answer pair output by the model reaches a third preset threshold, the predicted question-answer pair found by the model from the classroom transcript text is considered to be correct; otherwise, the predicted question-answer pair output by the model is judged to be wrong. The third preset threshold is a pre-set measurement standard used to determine whether the similarity between the predicted question-answer pair and the standard question-answer pair reaches an acceptable threshold. According to the specific application scenario, the third preset threshold is automatically set by the education model or manually set.

[0087] When the third similarity between the predicted question and answer pair and its corresponding standard question and answer pair does not reach the third preset threshold, it means that the predicted question and answer pair is wrong. The predicted question and answer pair is recorded and the error type is marked. The recording of the incorrect predicted question and answer pairs and the marking of the error type will help to subsequently optimize the ability of the educational large model to extract question and answer pairs.

[0088] Questions and answers in classroom teaching include two parts: questions and answers, for example, teachers ask questions and students answer. In this embodiment, the first evaluation strategy corresponds to questions, the second evaluation strategy corresponds to answers, and the third evaluation strategy corresponds to question-answer pairs including questions and answers.

[0089] The embodiment of the present application accurately constructs a mapping relationship between the predicted question-answer pair and the standard question-answer pair according to the similarity between the predicted question-answer pair and the standard question-answer pair corresponding to the classroom transcript text, and then effectively determines the evaluation strategy based on the mapping relationship, which is conducive to improving the accuracy of the model performance evaluation. At the same time, according to the comparison result of the semantic similarity between the predicted question-answer pair and the standard question-answer pair corresponding to the classroom transcript text and the preset threshold, the correctness of the model output can be effectively measured.

[0090] Step S103: Calculate the evaluation index value based on the evaluation strategy.

[0091] In the embodiment of the present application, there can be more than one mapping relationship, and there can be more than one evaluation strategy determined based on the mapping relationship. For example, the first evaluation strategy is determined based on the mapping relationship between the predicted question and the standard question, the second evaluation strategy is determined based on the mapping relationship between the predicted answer and the standard answer, and the third evaluation strategy is determined based on the mapping relationship between the predicted question and answer pair and the standard question and answer pair in the third mapping relationship. The evaluation index value is a quantitative value used to measure and evaluate the performance of the large education model. Based on the distribution and change trend of the evaluation index value, it can help discover problems and deficiencies in the model.

[0092] In order to effectively evaluate the performance of the educational model, the embodiment of the present application evaluates the model performance by calculating the evaluation index values ​​of multiple dimensions. The evaluation index values ​​of multiple dimensions include a first evaluation index value and a second evaluation index value. The first evaluation index value and the second evaluation index value are used to evaluate the ability of the educational model to extract (predict) question-answer pairs from the classroom transcript text.

[0093] As a possible implementation of this application, Figure 4 A specific implementation process of step S103 of the evaluation method of the educational model provided in the embodiment of the present application is shown, and is described in detail as follows:

[0094] C1: Determine a first evaluation index value based on the evaluation strategy, the number of predicted question-answer pairs, and the number of standard question-answer pairs.

[0095] The first evaluation index value includes precision, recall and area under the curve. Precision refers to the proportion of true positive examples among all samples predicted as positive examples by the model. Recall refers to the proportion of all actual positive samples correctly predicted as positive examples by the model. The area under the curve usually refers to the area under the receiver operating characteristic curve. In the embodiment of the present application, the characteristic curve is a precision-recall curve (Precision-Recall Curve, PR curve). The PR curve is a curve drawn with recall as the horizontal axis and precision as the vertical axis. In one possible implementation, when the evaluation strategy is the first evaluation strategy, based on the evaluation strategy, the number of predicted question-answer pairs and the number of standard question-answer pairs, the first evaluation index value is determined, including: determining the first precision according to the ratio of the number of correctly predicted questions to the total number of predicted questions, determining the first recall according to the ratio of the number of correctly predicted questions to the total number of standard questions in the standard question-answer pairs, and determining the first area under the curve according to the first precision and the first recall. The number of correctly predicted questions refers to the number of predicted questions whose similarity with the corresponding standard questions reaches the first preset threshold. The first precision rate refers to the proportion of correctly predicted questions by the model to all predicted questions. The first recall rate refers to the proportion of correctly predicted questions by the model to all standard questions that should actually be correctly predicted.

[0096] When the evaluation strategy is the first evaluation strategy, the first precision is used to measure whether the predicted questions output by the model have actually occurred in the classroom; the first recall is used to measure whether there are any omissions in the predicted questions output by the model.

[0097] In order to more accurately measure the performance of the model, when executing the first evaluation strategy, a set of thresholds is calculated (for example, 100 threshold points are taken at equal intervals between the value range of 0 and 1), and then the threshold strategy experiment is performed at each threshold point. After completing this set of experiments, a set of data will be obtained, and each data point contains a first precision and a first recall. With the first precision as the vertical axis and the first recall as the horizontal axis, all points are drawn on the coordinate axis, and they are connected to obtain a PR curve, and the area under the first curve is calculated; the larger the area under the first curve, the better the performance of the model output prediction question.

[0098] In the embodiment of the present application, the matching accuracy and recall rate of predicted questions and standard questions are focused on, and the first precision rate is used to accurately measure the extent to which the questions generated by the large education model are consistent with the ideal standard questions, so as to avoid the model from generating a large number of invalid questions that deviate from the topic and do not meet the teaching points, and the first recall rate is used to evaluate the coverage of all correctly predicted questions by the model. The area under the first curve obtained by combining the first precision rate and the first recall rate comprehensively considers the overall performance of the model in question prediction under different threshold settings.

[0099] In a possible implementation, when the evaluation strategy is the second evaluation strategy, the first evaluation index value is determined based on the evaluation strategy, the number of predicted question-answer pairs, and the number of standard question-answer pairs, including: determining the second precision rate according to the ratio of the number of correctly predicted answers to the total number of predicted answers; determining the second recall rate according to the ratio of the number of correctly predicted answers to the total number of standard answers in the standard question-answer pairs, and determining the second area under the curve according to the second precision rate and the second recall rate. The number of correctly predicted answers refers to the number of predicted answers whose similarity with the corresponding standard answers reaches a second preset threshold. The second precision rate refers to the proportion of correctly predicted answers of the model to all predicted answers. The second recall rate refers to the proportion of correctly predicted answers of the model to all standard answers that should actually be correctly predicted.

[0100] When the evaluation strategy is the second evaluation strategy, the second precision is used to measure whether the predicted responses output by the model have actually occurred in the classroom; the second recall is used to measure whether there are omissions in the predicted responses output by the model.

[0101] In order to measure the performance of the model more accurately, when executing the second evaluation strategy, a set of thresholds is calculated (for example, 100 threshold points are taken at equal intervals between the value range of 0 and 1), and then the threshold strategy experiment is performed at each threshold point. After completing this set of experiments, a set of data will be obtained, and each data point contains a second precision and a second recall. With the second precision as the vertical axis and the second recall as the horizontal axis, all points are drawn on the coordinate axis, and they are connected to obtain a PR curve, and the area under the second curve is calculated; the larger the area under the second curve, the better the performance of the model output predictive response.

[0102] In the embodiment of the present application, the matching accuracy and recall rate of the predicted answer and the standard answer are considered, and the accuracy of the model extracting answers is effectively evaluated based on the second precision rate reflecting the model's ability to give correct and effective answers to given questions, and the model's coverage of all correct predicted answers is evaluated based on the second recall rate. The area under the second curve obtained by combining the second precision rate and the second recall rate comprehensively considers the overall performance of the model in predicting answers under different threshold settings.

[0103] In a possible implementation, when the evaluation strategy is the third evaluation strategy, the first evaluation index value is determined based on the evaluation strategy, the number of predicted question-answer pairs, and the number of standard question-answer pairs, including: determining a third precision rate according to the ratio of the number of correctly predicted question-answer pairs in the predicted question-answer pairs to the total number of predicted question-answer pairs; determining a third recall rate according to the ratio of the number of correctly predicted question-answer pairs to the total number of standard question-answer pairs, and determining the area under the third curve according to the third precision rate and the third recall rate. The number of correctly predicted question-answer pairs refers to the number of predicted question-answer pairs whose similarity with the corresponding standard question-answer pairs reaches a third preset threshold. The third precision rate refers to the proportion of the model's correctly predicted question-answer pairs to all predicted question-answer pairs. The third recall rate refers to the proportion of the model's correctly predicted question-answer pairs to all standard question-answer pairs that should actually be correctly predicted.

[0104] When the evaluation strategy is the third evaluation strategy, the third precision is used to measure whether the predicted question-answer pairs output by the model have actually occurred in the classroom; the third recall is used to measure whether there are any omissions in the predicted question-answer pairs output by the model.

[0105] In order to measure the performance of the model more accurately, when executing the third evaluation strategy, a set of thresholds is calculated (for example, 100 threshold points are taken at equal intervals between the value range of 0 and 1), and then the threshold strategy experiment is performed at each threshold point. After completing this set of experiments, a set of data will be obtained, and each data point contains a third precision and a third recall. With the third precision as the vertical axis and the third recall as the horizontal axis, all points are drawn on the coordinate axis, and they are connected to obtain a PR curve, and the area under the curve is calculated; the larger the area under the curve, the better the model performance.

[0106] In the embodiment of the present application, the accuracy of the model question-answer matching is examined from an overall perspective, the synergy of questions and answers is comprehensively considered, and the third precision is used to evaluate whether the model can coherently and accurately complete the extraction of questions and answers from the classroom transcript text to avoid the situation where questions and answers are disconnected. The model's ability to restore the entire standard question-answering situation is evaluated based on the third recall rate. The area under the third curve obtained by combining the third precision and the third recall rate comprehensively considers the overall performance of the model in predicting question-answer pairs under different threshold settings.

[0107] In the embodiment of the present application, both precision and recall are evaluated on the sentence as a whole, which is a macro-level (sentence level) test that measures whether the output of the model is what actually happened in the classroom and whether there are omissions in the output, and determines the optimal threshold corresponding to the model based on the area under the curve.

[0108] C2: Determine a second evaluation index value based on the evaluation strategy, the number of words in the predicted question-answer pair, and the number of words in the standard question-answer pair.

[0109] The second evaluation indicator value includes fidelity. Fidelity is used to measure the ability of the model output to be faithful to the information and semantics contained in the standard question-answer pair. By testing the fidelity of the model, it is measured whether the question-answer pairs extracted by the model are faithful to the classroom transcript.

[0110] In the embodiment of the present application, we go deep into the vocabulary level, segment the sentences into individual words, and then compare whether the vocabulary in the predicted question and answer pairs generated by the model is the same or similar to the vocabulary used in the classroom transcript text. The large education model is evaluated from a micro level (lexical level), which can further improve the effectiveness of the model performance evaluation.

[0111] In a possible implementation, when the evaluation strategy is the first evaluation strategy, a second evaluation index value is determined based on the evaluation strategy, the number of words in the predicted question-answer pair, and the number of words in the standard question-answer pair, including: determining the question fidelity according to the first question ratio, the second question ratio, and the total number of predicted questions in the predicted question-answer pair, wherein the first question ratio is the ratio of the number of common words in each group of predicted questions and their corresponding standard questions to the total number of predicted questions in the predicted question-answer pair under the first evaluation strategy, and the second question ratio is the ratio of the number of common words in the questions to the total number of words in the standard questions in the standard question-answer pair.

[0112] In a possible implementation, when the evaluation strategy is the first evaluation strategy, the question fidelity Question_Fidelity is calculated according to the following formulas (1) and (2):

[0113]

[0114] Among them, N q is the total number of predicted questions under the first evaluation strategy, {question_fidelity} i represents the fidelity of the i-th predicted question, represents the ratio of the number of common words in the question to the total number of words in the predicted question under the first evaluation strategy, It represents the ratio of the number of common words in the question to the total number of words in the standard question under the first evaluation strategy.

[0115] The embodiment of the present application uses standard questions as a reference to calculate the question fidelity of the model under the first evaluation strategy, and effectively measures the ability of the model output to be faithful to the information and semantics contained in the standard questions based on the question fidelity.

[0116] In a possible implementation, when the evaluation strategy is the second evaluation strategy, a second evaluation index value is determined based on the evaluation strategy, the number of words in the predicted question-answer pair, and the number of words in the standard question-answer pair, including: determining the answer fidelity according to the first answer ratio, the second answer ratio, and the total number of predicted answers in the predicted question-answer pair, wherein the first answer ratio is the ratio of the number of common words in the answers in each group of predicted answers and their corresponding standard answers to the total number of predicted words in the predicted question-answer pair under the second evaluation strategy, and the second answer ratio is the ratio of the number of common words in the answers to the total number of words in the standard answers in the standard question-answer pair;

[0117] In a possible implementation, when the evaluation strategy is the second evaluation strategy, the response fidelity Response_Fidelity is calculated according to the following formulas (3) and (4):

[0118]

[0119] Among them, N r is the total number of predicted responses under the second evaluation strategy, {response_fidelity} i represents the fidelity of the ith predicted response, represents the ratio of the number of common words in the response to the total number of words in the predicted response under the second evaluation strategy, It represents the ratio of the number of common words in the response to the total number of words in the standard response under the second evaluation strategy.

[0120] The embodiment of the present application uses the standard response as a reference to calculate the response fidelity of the model under the second evaluation strategy, and effectively measures the ability of the model output to be faithful to the information and semantics contained in the standard response based on the response fidelity.

[0121] In one possible implementation, when the evaluation strategy is the third evaluation strategy, the second evaluation index value is determined based on the evaluation strategy, the number of words in the predicted question-answer pair, and the number of words in the standard question-answer pair, including: determining the question fidelity according to the third question ratio, the fourth question ratio, and the total number of predicted answer pairs under the third evaluation strategy, wherein the third question ratio is the ratio of the number of common words in each group of predicted questions and their corresponding standard questions to the total number of predicted question words in the predicted question-answer pair under the third evaluation strategy, and the fourth question ratio is the third evaluation strategy. The fidelity of the answer is determined according to the third answer ratio, the fourth answer ratio and the total number of predicted answer pairs in the predicted question and answer pairs under the third evaluation strategy, wherein the third answer ratio is the ratio of the number of common words of the answers in each group of predicted answers and their corresponding standard answers under the third evaluation strategy to the total number of predicted answer words in the predicted question and answer pairs, and the fourth answer ratio is the ratio of the number of common words of the answers under the third evaluation strategy to the total number of standard answer words in the standard question and answer pairs.

[0122] In a possible implementation, when the evaluation strategy is the third evaluation strategy, the question fidelity QR_Q_Fidelity under the third evaluation strategy is calculated according to the following formulas (5) and (6):

[0123]

[0124] Wherein, N is the total number of predicted response pairs under the third evaluation strategy, {qr_q_fidelity} i represents the question fidelity in the i-th predicted question-answer pair under the third evaluation strategy, represents the ratio of the number of common words in the question to the total number of words in the predicted question under the third evaluation strategy, It represents the ratio of the number of common words in the question to the total number of words in the standard question under the third evaluation strategy.

[0125] The response fidelity QR_R_Fidelity under the third evaluation strategy is calculated according to the following formulas (7) and (8):

[0126]

[0127] Wherein, N is the total number of predicted response pairs under the third evaluation strategy, {qr_r_fidelity} i represents the answer fidelity in the i-th predicted question-answer pair under the third evaluation strategy, represents the ratio of the number of common words in the response to the total number of words in the predicted response under the third evaluation strategy, It represents the ratio of the number of common words in the response to the total number of words in the standard response under the third evaluation strategy.

[0128] The embodiment of the present application uses the standard question-answer pair as a reference to calculate the question fidelity and answer fidelity of the model under the third evaluation strategy, effectively measuring the ability of the model output to be faithful to the information and semantics contained in the standard question-answer pair.

[0129] Step S104: Evaluate the educational model according to the evaluation index value.

[0130] In this embodiment, the evaluation index value may be visualized, for example, the evaluation index value may be converted into a bar chart for visual display, which may make the evaluation result more intuitive.

[0131] As can be seen from the above, in the embodiment of the present application, the terminal device obtains the predicted question and answer pairs obtained by predicting the classroom transcript text using the educational big model, and then constructs a mapping relationship between the predicted question and answer pairs and the standard question and answer pairs according to the similarity between the predicted question and answer pairs and the standard question and answer pairs corresponding to the classroom transcript text, and determines the evaluation strategy based on the mapping relationship, and then calculates the evaluation index value based on the evaluation strategy, and finally evaluates the educational big model according to the evaluation index value. The present application scheme can determine the evaluation strategy based on the constructed mapping relationship, so as to flexibly adapt to different evaluation needs, and the use of multi-dimensional evaluation indicators is helpful for fine-grained model evaluation, which can accurately and effectively evaluate the performance of the educational big model used for classroom question and answer pair analysis, and comprehensively reflect the application effect of the model in the educational scenario, thereby improving the fit between the evaluation results and actual applications.

[0132] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0133] Corresponding to the evaluation method of the educational model described in the above embodiment, Figure 5 A structural block diagram of an evaluation device for an educational model provided in an embodiment of the present application is shown. For ease of explanation, only the parts related to the embodiment of the present application are shown.

[0134] Reference Figure 5 The evaluation device of the educational model includes: a question-answer pair acquisition unit 51, a mapping strategy determination unit 52, an indicator value calculation unit 53, and a model evaluation unit 54, wherein:

[0135] A question-answer pair acquisition unit 51 is used to acquire predicted question-answer pairs obtained by predicting the classroom transcript text using the educational model;

[0136] A mapping strategy determination unit 52, configured to construct a mapping relationship between the predicted question-answer pair and the standard question-answer pair according to the similarity between the predicted question-answer pair and the standard question-answer pair corresponding to the classroom transcript text, and determine an evaluation strategy based on the mapping relationship;

[0137] An indicator value calculation unit 53, used to calculate the evaluation indicator value based on the evaluation strategy;

[0138] The model evaluation unit 54 is used to evaluate the educational model according to the evaluation index value.

[0139] As a possible implementation manner of the present application, the predicted question-answer pair includes a predicted question and a predicted answer, the standard question-answer pair includes a standard question and a standard answer, and the mapping strategy determination unit 52 includes:

[0140] A first construction and determination module, configured to construct a first mapping relationship according to a first similarity between the predicted question in the predicted question-answer pair and the standard question in the standard answer, and determine a first evaluation strategy based on the first mapping relationship;

[0141] a second construction and determination module, configured to construct a second mapping relationship according to a second similarity between the predicted answer in the predicted question-answer pair and the standard answer in the standard answer, and determine a second evaluation strategy based on the second mapping relationship;

[0142] The third construction and determination module is used to construct a third mapping relationship according to the first similarity and the second similarity, and determine a third evaluation strategy based on the third mapping relationship.

[0143] As a possible implementation of the present application, the index value calculation unit 53 includes:

[0144] A first calculation module, configured to determine a first evaluation index value based on the evaluation strategy, the number of predicted question-answer pairs, and the number of standard question-answer pairs;

[0145] The second calculation module is used to determine a second evaluation index value based on the evaluation strategy, the number of words in the predicted question-answer pair, and the number of words in the standard question-answer pair.

[0146] As a possible implementation of the present application, the first evaluation index value includes precision, recall, and area under the curve; the first calculation module is specifically used for:

[0147] When the evaluation strategy is the first evaluation strategy, a first precision is determined according to the ratio of the number of correctly predicted questions to the total number of predicted questions, a first recall is determined according to the ratio of the number of correctly predicted questions to the total number of standard questions in the standard question-answer pairs, and a first area under the curve is determined according to the first precision and the first recall;

[0148] When the evaluation strategy is the second evaluation strategy, a second precision is determined according to the ratio of the number of correctly predicted answers to the total number of predicted answers; a second recall is determined according to the ratio of the number of correctly predicted answers to the total number of standard answers in the standard question and answer pairs, and a second area under the curve is determined according to the second precision and the second recall;

[0149] When the evaluation strategy is the third evaluation strategy, the third precision is determined according to the ratio of the number of correctly predicted question and answer pairs in the predicted question and answer pairs to the total number of predicted question and answer pairs; the third recall is determined according to the ratio of the number of correctly predicted question and answer pairs to the total number of standard question and answer pairs; and the area under the third curve is determined according to the third precision and the third recall.

[0150] As a possible implementation manner of the present application, the second evaluation index value includes fidelity; and the second calculation module is specifically used for:

[0151] When the evaluation strategy is the first evaluation strategy, the question fidelity is determined according to the first question ratio, the second question ratio and the total number of predicted questions in the predicted question-answer pair, wherein the first question ratio is the ratio of the number of common question words in each group of predicted questions and their corresponding standard questions to the total number of predicted question words in the predicted question-answer pair under the first evaluation strategy, and the second question ratio is the ratio of the number of common question words to the total number of standard question words in the standard question-answer pair;

[0152] When the evaluation strategy is the second evaluation strategy, the answer fidelity is determined according to the first answer ratio, the second answer ratio and the total number of predicted answers in the predicted question and answer pair, wherein the first answer ratio is the ratio of the number of common words of answers in each group of predicted answers and their corresponding standard answers to the total number of predicted answer words in the predicted question and answer pair under the second evaluation strategy, and the second answer ratio is the ratio of the number of common words of answers to the total number of standard answer words in the standard question and answer pair;

[0153] When the evaluation strategy is the third evaluation strategy, the question fidelity is determined according to the third question ratio, the fourth question ratio and the total number of predicted answer pairs under the third evaluation strategy, wherein the third question ratio is the ratio of the number of common question words in each group of predicted questions and their corresponding standard questions to the total number of predicted question words in the predicted question-answer pair under the third evaluation strategy, and the fourth question ratio is the ratio of the number of common question words in the third evaluation strategy to the total number of standard question words in the standard question-answer pair; the answer fidelity is determined according to the third answer ratio, the fourth answer ratio and the total number of predicted answer pairs in the predicted question-answer pair under the third evaluation strategy, wherein the third answer ratio is the ratio of the number of common answer words in each group of predicted answers and their corresponding standard answers to the total number of predicted answer words in the predicted question-answer pair under the third evaluation strategy, and the fourth answer ratio is the ratio of the number of common answer words in the third evaluation strategy to the total number of standard answer words in the standard question-answer pair.

[0154] As a possible implementation of the present application, the question-answer pair acquisition unit 51 includes:

[0155] A search module, for searching, based on a data index, the classroom transcript text corresponding to the educational model and its corresponding predicted question-answer pair when receiving an evaluation instruction;

[0156] A text acquisition module, used for acquiring the classroom transcript text if the data index does not exist;

[0157] The question-answer pair acquisition module is used to input the classroom transcript text and prompts into the education model, and use the prompts to guide the education model to predict the predicted question-answer pairs in the classroom transcript text.

[0158] As a possible implementation of the present application, the text acquisition module is specifically used for:

[0159] Obtaining an initial classroom record, and transcribing the initial classroom record to obtain a transcribed text;

[0160] The transcribed text is preprocessed to obtain a classroom transcript text.

[0161] As can be seen from the above, in the embodiment of the present application, the terminal device obtains the predicted question and answer pairs obtained by predicting the classroom transcript text using the educational big model, and then constructs a mapping relationship between the predicted question and answer pairs and the standard question and answer pairs according to the similarity between the predicted question and answer pairs and the standard question and answer pairs corresponding to the classroom transcript text, and determines the evaluation strategy based on the mapping relationship, and then calculates the evaluation index value based on the evaluation strategy, and finally evaluates the educational big model according to the evaluation index value. The present application scheme can determine the evaluation strategy based on the constructed mapping relationship, so as to flexibly adapt to different evaluation needs, and the use of multi-dimensional evaluation indicators is helpful for fine-grained model evaluation, which can accurately and effectively evaluate the performance of the educational big model used for classroom question and answer pair analysis, and comprehensively reflect the application effect of the model in the educational scenario, thereby improving the fit between the evaluation results and actual applications.

[0162] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / units are based on the same concept as the method embodiment of the present application. Their specific functions and technical effects can be found in the method embodiment part and will not be repeated here.

[0163] The present application also provides a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, Figures 1 to 4 The steps of an evaluation method for any large educational model are represented.

[0164] The embodiment of the present application also provides a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, Figures 1 to 4 The steps of an evaluation method for any large educational model are represented.

[0165] The embodiment of the present application also provides a computer program product, when the computer program product is run on a terminal device, the terminal device executes the following Figures 1 to 4 The steps of an evaluation method for any large educational model are represented.

[0166] Figure 6 Schematic diagram of a terminal device provided by an embodiment of the present application. Figure 6 As shown, the terminal device 6 of this embodiment includes: a processor 60, a memory 61, and a computer program 62 stored in the memory 61 and executable on the processor 60. When the processor 60 executes the computer program 62, the steps in the above-mentioned evaluation method embodiments of each educational model are implemented, for example Figure 1Alternatively, when the processor 60 executes the computer program 62, the functions of each module / unit in the above-mentioned device embodiments are realized, for example Figure 5 The functions of units 51 to 54 are shown.

[0167] Exemplarily, the computer program 62 may be divided into one or more modules / units, which are stored in the memory 61 and executed by the processor 60 to complete the present application. The one or more modules / units may be a series of computer-readable instruction segments capable of completing specific functions, which are used to describe the execution process of the computer program 62 in the terminal device 6.

[0168] The terminal device 6 may include, but is not limited to, a processor 60 and a memory 61. Those skilled in the art will appreciate that Figure 6 It is only an example of the terminal device 6 and does not constitute a limitation of the terminal device 6. It may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, the terminal device 6 may also include input and output devices, network access devices, buses, etc.

[0169] The processor 60 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.

[0170] The memory 61 may be an internal storage unit of the terminal device 6, such as a hard disk or memory of the terminal device 6. The memory 61 may also be an external storage device of the terminal device 6, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the terminal device 6. Further, the memory 61 may also include both an internal storage unit and an external storage device of the terminal device 6. The memory 61 is used to store the computer program and other programs and data required by the terminal device. The memory 61 may also be used to temporarily store data that has been output or is to be output.

[0171] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / units are based on the same concept as the method embodiment of the present application. Their specific functions and technical effects can be found in the method embodiment part and will not be repeated here.

[0172] The technicians in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In practical applications, the above-mentioned function allocation can be completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated in a processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here.

[0173] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, which can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may at least include: any entity or device that can carry the computer program code to the device / terminal device, a recording medium, a computer memory, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), an electric carrier signal, a telecommunication signal, and a software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disk. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electric carrier signals and telecommunication signals.

[0174] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0175] The embodiments described above are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, a person skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. A method for evaluating a large educational model, characterized in that: include: Obtain predicted question-answer pairs obtained by predicting classroom transcripts using a large education model; According to the similarity between the predicted question-answer pair and the standard question-answer pair corresponding to the classroom transcript text, a mapping relationship between the predicted question-answer pair and the standard question-answer pair is constructed, and an evaluation strategy is determined based on the mapping relationship; Based on the evaluation strategy, calculating the evaluation index value; The educational model is evaluated according to the evaluation index value.

2. The method according to claim 1, characterized in that The predicted question-answer pair includes a predicted question and a predicted answer, and the standard question-answer pair includes a standard question and a standard answer; constructing a mapping relationship between the predicted question-answer pair and the standard question-answer pair according to the similarity between the predicted question-answer pair and the standard question-answer pair corresponding to the classroom transcript text, and determining an evaluation strategy based on the mapping relationship, including: constructing a first mapping relationship according to a first similarity between the predicted question in the predicted question-answer pair and the standard question in the standard answer, and determining a first evaluation strategy based on the first mapping relationship; constructing a second mapping relationship according to a second similarity between the predicted answer in the predicted question-answer pair and the standard answer in the standard answer, and determining a second evaluation strategy based on the second mapping relationship; A third mapping relationship is constructed according to the first similarity and the second similarity, and a third evaluation strategy is determined based on the third mapping relationship.

3. The method according to claim 2, characterized in that The calculating of the evaluation index value based on the evaluation strategy includes: Determining a first evaluation index value based on the evaluation strategy, the number of the predicted question-answer pairs, and the number of the standard question-answer pairs; A second evaluation index value is determined based on the evaluation strategy, the number of words in the predicted question-answer pair, and the number of words in the standard question-answer pair.

4. The method according to claim 3, characterized in that The first evaluation index value includes precision, recall, and area under the curve; the first evaluation index value is determined based on the evaluation strategy, the number of predicted question-answer pairs, and the number of standard question-answer pairs, including: When the evaluation strategy is the first evaluation strategy, a first precision is determined according to the ratio of the number of correctly predicted questions to the total number of predicted questions, a first recall is determined according to the ratio of the number of correctly predicted questions to the total number of standard questions in the standard question-answer pairs, and a first area under the curve is determined according to the first precision and the first recall; or When the evaluation strategy is the second evaluation strategy, a second precision is determined according to the ratio of the number of correctly predicted answers to the total number of predicted answers; a second recall is determined according to the ratio of the number of correctly predicted answers to the total number of standard answers in the standard question and answer pairs, and a second area under the curve is determined according to the second precision and the second recall; or, When the evaluation strategy is the third evaluation strategy, the third precision is determined according to the ratio of the number of correctly predicted question and answer pairs in the predicted question and answer pairs to the total number of predicted question and answer pairs; the third recall is determined according to the ratio of the number of correctly predicted question and answer pairs to the total number of standard question and answer pairs; and the area under the third curve is determined according to the third precision and the third recall.

5. The method according to claim 3, characterized in that: The second evaluation index value includes fidelity; the determining the second evaluation index value based on the evaluation strategy, the number of words of the predicted question-answer pair, and the number of words of the standard question-answer pair includes: When the evaluation strategy is the first evaluation strategy, the question fidelity is determined according to the first question ratio, the second question ratio and the total number of predicted questions in the predicted question-answer pair, wherein the first question ratio is the ratio of the number of common question words in each group of predicted questions and their corresponding standard questions to the total number of predicted question words in the predicted question-answer pair under the first evaluation strategy, and the second question ratio is the ratio of the number of common question words to the total number of standard question words in the standard question-answer pair; or, When the evaluation strategy is the second evaluation strategy, the answer fidelity is determined according to the first answer ratio, the second answer ratio and the total number of predicted answers in the predicted question and answer pair, wherein the first answer ratio is the ratio of the number of common words of answers in each group of predicted answers and their corresponding standard answers to the total number of predicted answer words in the predicted question and answer pair under the second evaluation strategy, and the second answer ratio is the ratio of the number of common words of answers to the total number of standard answer words in the standard question and answer pair; or, When the evaluation strategy is the third evaluation strategy, the question fidelity is determined according to the third question ratio, the fourth question ratio and the total number of predicted answer pairs under the third evaluation strategy, wherein the third question ratio is the ratio of the number of common question words in each group of predicted questions and their corresponding standard questions to the total number of predicted question words in the predicted question-answer pair under the third evaluation strategy, and the fourth question ratio is the ratio of the number of common question words in the third evaluation strategy to the total number of standard question words in the standard question-answer pair; the answer fidelity is determined according to the third answer ratio, the fourth answer ratio and the total number of predicted answer pairs in the predicted question-answer pair under the third evaluation strategy, wherein the third answer ratio is the ratio of the number of common answer words in each group of predicted answers and their corresponding standard answers to the total number of predicted answer words in the predicted question-answer pair under the third evaluation strategy, and the fourth answer ratio is the ratio of the number of common answer words in the third evaluation strategy to the total number of standard answer words in the standard question-answer pair.

6. The method according to any one of claims 1 to 5, characterized in that: The step of obtaining predicted question-answer pairs obtained by predicting the classroom transcript text using the educational big model includes: When receiving the evaluation instruction, searching for the classroom transcript text corresponding to the educational model and its corresponding predicted question-answer pair based on the data index; If the data index does not exist, then obtain the classroom transcript text; The classroom transcript text and prompts are input into the educational model, and the prompts are used to guide the educational model to predict the predicted question and answer pairs in the classroom transcript text.

7. The method according to claim 6, characterized in that The obtaining of classroom transcripts includes: Obtaining an initial classroom record, and transcribing the initial classroom record to obtain a transcribed text; The transcribed text is preprocessed to obtain a classroom transcript text.

8. An evaluation device for a large educational model, characterized in that: include: A question-answer pair acquisition unit, used to acquire predicted question-answer pairs obtained by predicting classroom transcript texts using the educational big model; A mapping strategy determination unit, configured to construct a mapping relationship between the predicted question-answer pair and the standard question-answer pair according to the similarity between the predicted question-answer pair and the standard question-answer pair corresponding to the classroom transcript text, and determine an evaluation strategy based on the mapping relationship; An indicator value calculation unit, used to calculate the evaluation indicator value based on the evaluation strategy; The model evaluation unit is used to evaluate the educational model according to the evaluation index value.

9. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, it implements the evaluation method of the educational model as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, it implements the evaluation method of the educational model as described in any one of claims 1 to 7.