A method for enhancing teaching evaluation credibility in an educational large model
By constructing historical performance sequences and dynamic smoothing coefficients in a large-scale education model, the problem of large fluctuations in scoring results is solved, the stability and interpretability of scoring results are achieved, and the reliability of education assessment is improved.
Patent Information
- Application Number
- CN202510909861.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-02
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2045-07-02
AI Technical Summary
In the large-scale education model, the technology of automatically scoring the answer content based on the large language model does not fully consider the changes in the confidence of historical scores or model scores in the time series, resulting in large fluctuations in the scoring results between different time periods or batches, affecting the stability and reliability of education assessment, and the interpretability of the scoring basis is poor.
By acquiring scoring prompts, constructing a teaching knowledge base, pre-training a large language model, generating initial subjective questions, and constructing a historical score sequence under the same scoring criteria, calculating the dynamic smoothing coefficient and the current matching degree numerical edge weight, and obtaining the confidence value of the scoring evaluation information, we can ensure the output of high-quality feedback scoring information.
This approach achieves teaching evaluation with minimal fluctuations in scoring results and good interpretability of scoring criteria, ensuring that scoring results across different batches and times are controllable and providing interpretable scoring criteria.
Smart Images

Figure CN120672216B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of intelligent teaching, and in particular to a method for enhancing the credibility of teaching evaluation in an education large model. BACKGROUND
[0002] At present, in the education large model, the technology of automatically scoring the answer content based on a large language model (such as GPT, BERT, etc.) mainly relies on the semantic judgment of the large language model on a single answer, without fully considering the changes of historical scores or model score confidence in the time sequence, resulting in large fluctuations in the scoring results between different time periods or batches, affecting the stability and reliability of the education evaluation. The processing method of "inputting the answer content of the student into the large language model, then performing semantic understanding on the answer content, and finally generating the scoring result" has poor score basis explanation, and teachers and students cannot effectively understand the scoring result. SUMMARY
[0003] The purpose of the present application is to provide a method for enhancing the credibility of teaching evaluation in an education large model, achieving the effect of small fluctuations in the scoring result and good score basis explanation.
[0004] According to one aspect of the present application, a method for enhancing the credibility of teaching evaluation in an education large model is provided, the method comprising:
[0005] obtaining a scoring prompt, and according to the scoring prompt, obtaining first scoring evaluation information of the answer content;
[0006] obtaining an interactive object corresponding to the first scoring evaluation information, and under the same scoring standard, constructing the scores of different users into a historical score sequence;
[0007] obtaining a dynamic smoothing coefficient according to the historical score sequence;
[0008] obtaining a current matching degree numerical edge weight of the first scoring evaluation information according to the historical score sequence and the dynamic smoothing coefficient;
[0009] obtaining a confidence value of the first scoring evaluation information according to the current matching degree numerical edge weight;
[0010] obtaining a feedback quality of the first scoring evaluation information according to the confidence value;
[0011] in the case that the feedback quality of the first scoring evaluation information is high, determining that the first scoring evaluation information meets the standard, and sending the first scoring evaluation information to a first terminal.
[0012] Optionally, before the step of obtaining a scoring prompt and obtaining first scoring evaluation information of the answer content according to the scoring prompt, the method further comprises:
[0013] obtaining teaching text data of an online teaching system;
[0014] preprocessing the obtained teaching text data to construct a teaching knowledge base, wherein the preprocessing includes text cleaning, text deduplication, and text segmentation;
[0015] pre-training a large language model based on the teaching knowledge base to obtain an initial state LLM large model;
[0016] obtaining a question setting instruction;
[0017] generating an initial subjective question from the initial state LLM large model according to the question setting instruction.
[0018] Optionally, before the obtaining a scoring prompt and the obtaining first scoring evaluation information of the answer content according to the scoring prompt, the method further comprises:
[0019] obtaining an initial subjective question and a scoring scale corresponding to the answer content, wherein the answer content is the answer content of a plurality of users to the initial subjective question;
[0020] The obtaining a scoring prompt and the obtaining first scoring evaluation information of the answer content according to the scoring prompt, comprising:
[0021] obtaining the scoring prompt and the scoring scale;
[0022] inputting the scoring prompt, the scoring scale, and the answer content into the initial state LLM large model to obtain the first scoring evaluation information of the answer content.
[0023] Optionally, the method further comprises:
[0024] obtaining a high-quality data set in the case that the feedback quality of the first scoring evaluation information is low;
[0025] fine-tuning the initial state LLM large model based on the high-quality data set to obtain a fine-tuned LLM large model.
[0026] Optionally, the fine-tuning the initial state LLM large model based on the high-quality data set to obtain a fine-tuned LLM large model, comprising:
[0027] filtering out a text passage containing a scoring standard and an answer quality causal relationship description from the teaching knowledge base;
[0028] constructing a causal relationship data set based on the scoring standard and the answer quality causal relationship description in the text passage;
[0029] Inference and analysis are performed on the input subjective questions by using the causal relationship dataset, and a question and answer pair is constructed;
[0030] Based on the question and answer pair, the initial state LLM large model is fine-tuned to obtain the fine-tuned LLM large model.
[0031] Optionally, the method further comprises:
[0032] Based on the fine-tuned LLM large model, second scoring evaluation information of the answer content is obtained;
[0033] The second scoring evaluation information is sent to the first terminal;
[0034] Alternatively,
[0035] Based on the fine-tuned LLM large model, third scoring evaluation information of the answer content is obtained;
[0036] In the case where the feedback quality of the third scoring evaluation information is determined to be high, the third scoring evaluation information is sent to the first terminal.
[0037] Optionally, the dynamic smoothing coefficient is obtained according to the historical score sequence, comprising:
[0038] The standard deviation of the historical score sequence is calculated;
[0039] The dynamic smoothing coefficient is calculated according to a dynamic coefficient algorithm, wherein the dynamic coefficient algorithm is obtained by the following method:
[0040]
[0041] Wherein, The dynamic smoothing coefficient is represented by; The minimum smoothing factor is represented by; The maximum smoothing factor is represented by; The standard deviation of the historical score sequence is represented by; base represents the upper limit of the score threshold.
[0042] Optionally, the current matching degree numerical edge weight of the first scoring evaluation information is obtained according to the historical score sequence and the dynamic smoothing coefficient, comprising:
[0043] Residual initialization weight is calculated according to a residual weight algorithm;
[0044] The current scoring contribution value is calculated according to a scoring contribution algorithm;
[0045] The current matching degree numerical edge weight is calculated according to a matching strength algorithm, wherein the matching strength algorithm is obtained by the following method:
[0046]
[0047] wherein, represents the current matching degree numerical edge weight; represents the main right control coefficient; represents the current score contribution value; represents the residual initialization weight; represents the initial matching degree numerical edge weight.
[0048] Optionally, the method further comprises: acquiring a confidence value of the first scoring evaluation information according to the current matching degree numerical edge weight; and acquiring a feedback quality of the first scoring evaluation information according to the confidence value, which comprises:
[0049] The confidence value of the first scoring evaluation information is calculated according to a confidence algorithm, and the confidence algorithm is obtained by using the following method:
[0050]
[0051] wherein, Z represents the confidence value; represents the first result of the first scoring evaluation information; represents the expected score;
[0052] determining whether the confidence value belongs to a confidence interval;
[0053] in the case that the confidence value belongs to the confidence interval, determining that the feedback quality of the first scoring evaluation information is high.
[0054] Optionally, after the pre-training of the large language model based on the teaching knowledge base to obtain the initial state LLM large model, the method further comprises:
[0055] constructing a multi-dimensional feature vector set based on the teaching knowledge base, wherein the multi-dimensional feature vector set comprises a dimension coverage vector, a standard description vector and a domain adaptation vector;
[0056] setting an output structure rationality scoring standard according to the dimension coverage vector, setting an operability scoring standard according to the standard description vector, and setting a domain adaptation scoring standard according to the domain adaptation vector;
[0057] calculating a question score value R according to a question scoring algorithm, wherein the question scoring algorithm is obtained by using the following method:
[0058] R= *structure rationality score+ *operability score+ *domain adaptation score;
[0059] wherein, + + =1, and , , Adjust dynamically according to the application scenario. , , All represent rating coefficients.
[0060] Optionally, the method further includes:
[0061] The initial subjective questions, the answer content, and the first score evaluation information with high feedback quality are preprocessed to construct a second teaching knowledge base, wherein the preprocessing includes text cleaning, text deduplication, and text segmentation.
[0062] Based on the second teaching knowledge base, the initial state LLM large model is pre-trained to obtain an advanced version of the LLM large model.
[0063] According to another aspect of the present invention, a system for enhancing the credibility of teaching evaluation in a large-scale education model is also provided, the system comprising an acquisition unit, an assessment unit, and a quality feedback unit;
[0064] The acquisition unit is used to acquire a scoring prompt and, based on the scoring prompt, acquire the first scoring evaluation information of the answer content.
[0065] The evaluation unit is used to: acquire the interactive object corresponding to the first rating evaluation information; construct a historical score sequence of different users' scores under the same rating standard; acquire a dynamic smoothing coefficient based on the historical score sequence; acquire the current matching degree value edge weight of the first rating evaluation information based on the historical score sequence and the dynamic smoothing coefficient; acquire the confidence value of the first rating evaluation information based on the current matching degree value edge weight; and acquire the feedback quality of the first rating evaluation information based on the confidence value.
[0066] The quality feedback unit is used to determine that the first rating evaluation information meets the standard when the feedback quality of the first rating evaluation information is high, and to send the first rating evaluation information to the first terminal.
[0067] Optionally, the system further includes a pre-training unit and a question generation unit;
[0068] The pre-training unit is used to acquire teaching text data from the online teaching system; preprocess the acquired teaching text data to construct a teaching knowledge base, wherein the preprocessing includes text cleaning, text deduplication, and text segmentation; and pre-train a large language model based on the teaching knowledge base to obtain an initial-state LLM large model.
[0069] The question-generating unit is used to obtain question-generating instructions; and to generate initial subjective questions from the initial state LLM large model according to the question-generating instructions.
[0070] According to another aspect of the present invention, an electronic device is also provided, the electronic device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the method described above for enhancing the credibility of teaching evaluation in a large-scale educational model.
[0071] This invention provides a method for enhancing the credibility of teaching evaluation in a large-scale education model. The method involves obtaining scoring prompts, acquiring first-level scoring information based on these prompts, calculating the current matching degree edge weight of the first-level scoring information, calculating the confidence value of the first-level scoring information based on the current matching degree edge weight, and using the confidence value as the criterion to output only the first-level scoring information with high feedback quality. In this technical solution, based on the first-level scoring information, the scores of different users under the same scoring standard are constructed into a historical score sequence. A dynamic smoothing coefficient is calculated based on the historical score sequence. The current matching degree edge weight of the first-level scoring information is obtained based on the historical score sequence and the dynamic smoothing coefficient. The current matching degree edge weight consists of the current scoring contribution value (new information) and the residual initialization weight (old information). The feedback quality of the first-level scoring information is quantitatively evaluated through the real-time change of the dynamic ratio of new and old information. By transforming abstract scoring information into data expression, the true feedback quality is reflected with objective data; through the dynamic adjustment of the ratio of new and old information, the fluctuations in scoring results across different batches and times are ensured to be controllable, forming a highly interpretable scoring basis. Attached Figure Description
[0072] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Furthermore, the same reference numerals denote the same parts throughout the drawings.
[0073] Figure 1 This is a flowchart illustrating a method for enhancing the credibility of teaching evaluation in a large-scale education model provided by the present invention.
[0074] Figure 2 This is a schematic diagram illustrating the association and mapping relationship of the teaching text data provided by the present invention.
[0075] Figure 3 This is a schematic diagram of the scoring table structure for constructing answer content provided by the present invention.
[0076] Figure 4 This is another flowchart illustrating a method for enhancing the credibility of teaching evaluation in a large-scale education model provided by the present invention.
[0077] Figure 5 This is a schematic diagram illustrating the display of subjective questions on a terminal according to the present invention.
[0078] Figure 6 This is a schematic diagram of the composition structure of a system for enhancing the credibility of teaching evaluation in a large-scale education model, as provided by the present invention.
[0079] Figure 7 This is a schematic diagram of the hardware architecture of an electronic device provided by the present invention. Detailed Implementation
[0080] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Many specific details are set forth in the following description to provide a thorough understanding of the present invention. However, the present invention can be practiced in many other ways different from those described herein, and those skilled in the art can make similar modifications without departing from the spirit of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0081] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the description of the invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0082] Firstly, embodiments of the present invention provide a method for enhancing the credibility of teaching evaluation in a large-scale education model, applied to a client (e.g., a first terminal). The client is the executing entity of this method, and the client's processing flow specifically includes the following steps:
[0083] Step S01: Obtain the initial subjective questions.
[0084] In this embodiment, the initial subjective questions are sent to a client, such as a first terminal. The first terminal can be a computer device used by a teacher or a computer device used by a student. Specifically, the initial subjective questions are generated from a pre-trained initial-state LLM large model. Subjective questions are a type of question that requires candidates to answer through independent thinking, organizing language, explaining viewpoints, or creative expression. Their core characteristic is that there is no single correct answer, and there is usually no fixed standard answer; they place greater emphasis on assessing the candidate's comprehensive qualities, including their thought process, analytical ability, knowledge application ability, and language expression ability.
[0085] The full English name of LLM is Large Language Model. For example, a large language model can be any of the following: chatGPT, Wenxin Yiyan, Tongyi Qianwen, or Deepseek.
[0086] Step S02: Obtain the rating scale corresponding to the answer content, wherein the initial subjective questions are the subjective questions corresponding to the answer content.
[0087] In this embodiment, the rating scale, also known as a rating rubric, is a structured rating tool used for the systematic evaluation of subjective questions. By clearly defining rating criteria and rating levels (including scores), the rating scale provides an objective and operational framework for the rating process, reducing subjectivity and improving consistency and fairness. In this embodiment, the initial-state LLM large model can automatically generate a rating scale based on the rating logic in its pre-trained teaching knowledge base. For example, the rating scale can be in JSON / XML format.
[0088] It should be noted that there is a one-to-one correspondence between the rating scales and the subjective questions; that is, each subjective question corresponds to a specific rating scale. For example, subjective question X corresponds to rating scale x; subjective question Y corresponds to rating scale y, and rating scales x and y are not the same.
[0089] Step S03: Obtain the answer content for the initial subjective questions.
[0090] In this embodiment, the initial subjective questions are sent to the computer device used by the student. After viewing the initial subjective questions on the computer device, the student inputs the answers (i.e., the answer content) into the computer device.
[0091] like Figure 1 As shown: After step 03, the client's processing flow further includes the following steps:
[0092] Step S101: Obtain scoring prompts and, based on the scoring prompts, obtain the first scoring evaluation information for the answer content.
[0093] In this embodiment, the step includes: obtaining a scoring prompt and a scoring scale; inputting the scoring prompt, scoring scale, and answer content into the initial-state LLM large-scale model to obtain the first scoring evaluation information of the answer content. The scoring prompt executes a scoring trigger operation based on a prompt word. The scoring prompt represents a suggestive input used to guide scoring behavior or generate a score. For example, the scoring prompt can be a prompt entered by the teacher into the computer device, such as "Generate scoring evaluation information based on the student's input answer content." Optionally, the scoring prompt can be generated by the initial-state LLM large-scale model.
[0094] "Prompt" refers to a "hint word" or "input instruction." In the field of artificial intelligence, especially in Natural Language Processing (NLP) and chatbots, a prompt is the text content entered by the user to guide the AI in generating an answer or performing a task. For example, when user A asks a question to the system, user A's question is the prompt, and the system will generate an answer based on this prompt.
[0095] Upon receiving the scoring prompt, the initial-state LLM model will generate initial scoring information based on the student's input. Specifically, the initial scoring information includes an initial grade and initial feedback; that is, the initial grade is the score given by the initial-state LLM model.
[0096] Step S102: Obtain the interaction object corresponding to the first rating evaluation information, and construct a historical score sequence for different users under the same rating standard.
[0097] In this embodiment, the interactive objects refer to the initial subjective question (object 1), the answer to the initial subjective question (object 2), the student's identity (which can be recorded by student ID, object 3), and the scoring criteria (object 4). The scoring criteria represent the specific angles of scoring, such as... Figure 3 As shown, the scoring criteria include content completeness, language expression, structural organization, and originality. By acquiring the interaction data, the student's identity can be determined, and a correspondence can be established between the student's identity and their answers. This allows for accurate recording of scoring information to the corresponding student after subsequent evaluation.
[0098] Based on the timestamp or rating record identifier in the first rating evaluation information, the current rating time is accurately obtained. Using this time as a dividing point, rating data from the same rating standard for a preset number of times (e.g., 5 times) prior to this time is extracted. This rating data comes from the scores of different users and is defined as historical rating data. For example, if the interactive object includes a "language expression" rating standard, then the rating data from 5 different users belonging to the "language expression" rating standard prior to this time is extracted. Based on the historical rating data, a historical score sequence is constructed. Different users can be multiple students, distinguished by student ID or by the unique identifier (ID) of the terminal device they use; they can also be multiple faculty members, multiple employees, etc.
[0099] Step S103: Obtain the dynamic smoothing coefficient based on the historical score sequence.
[0100] In this embodiment, based on the fluctuation characteristics and trend analysis of the historical performance sequence, the adaptive module calculates a dynamic smoothing coefficient, which is used to quantify the stability of the historical performance.
[0101] Step S104: Obtain the current matching degree value edge weight of the first score evaluation information based on the historical score sequence and dynamic smoothing coefficient.
[0102] Step S105: Obtain the confidence value of the first rating evaluation information based on the current matching degree value edge weight.
[0103] Step S106: Obtain the feedback quality of the first rating evaluation information based on the confidence level.
[0104] In this embodiment, based on the statistical characteristics (standard deviation) and dynamic smoothing coefficient of the historical score sequence, the adaptive module calculates the current matching degree numerical weight, which includes the current score contribution value (new information) and the residual initialization weight (old information). The current score contribution value reflects the real-time contribution of the current score, while the residual initialization weight reflects the cumulative impact of historical scores. The proportions of the current score contribution value and the residual initialization weight in the current matching degree numerical weight change dynamically. Through the real-time change of the dynamic ratio of new and old information, the feedback quality of the first score evaluation information is quantitatively assessed. By transforming abstract score evaluation information into data expression, the true feedback quality is reflected with objective data; through the dynamic adjustment of the new and old information ratio, the fluctuations in the score evaluation information of different batches and different times under the same scoring standard are ensured to be controllable, forming a highly interpretable scoring basis.
[0105] like Figure 3As shown, the scoring criteria include four dimensions: content completeness, language expression, structural organization, and originality. In this embodiment, all the above operation steps are performed based on the same scoring criteria. The first scoring evaluation information refers to the score and feedback on the answer content under a specific scoring criterion (e.g., language expression). This technical solution needs to generate corresponding scoring evaluation information separately for each of the four scoring criteria, and finally sum up the various scoring evaluation information to obtain the comprehensive score.
[0106] Step S107: If the feedback quality of the first rating evaluation information is high, determine that the first rating evaluation information meets the standard, and send the first rating evaluation information to the first terminal.
[0107] In this embodiment, if the feedback quality of the first scoring evaluation information is high (e.g., small fluctuation range, high consistency with historical trends), the first scoring evaluation information is sent to the first terminal, and the valid scoring evaluation information is output. For example, the first scoring evaluation information is displayed at a corresponding position on the first terminal where the answer content is displayed. The first scoring evaluation information includes a first score and first evaluation feedback. The corresponding position on the first terminal where the answer content is displayed can be the next line after the answer content, the top of the answer content, etc.
[0108] This technical solution transforms abstract scoring and evaluation information into data expression through quantitative analysis of historical performance sequences, using objective data to reflect the true quality of feedback. By dynamically adjusting the ratio of new and old information, it ensures that the fluctuations in scoring results across different batches and times are controllable, forming a highly interpretable scoring basis.
[0109] In one embodiment of the present invention, before obtaining the scoring prompt and obtaining the first scoring evaluation information of the answer content based on the scoring prompt, the above method further includes the following steps:
[0110] Step S201: Obtain the teaching text data from the online teaching system.
[0111] In this embodiment, the data to be preprocessed can be at least one of text data, image data, video data, audio data, formula data, and graphic data.
[0112] The introduction focuses on preprocessing based on teaching text data. Specifically, it involves acquiring teaching text data from an online teaching system. This teaching text data includes definitions of concepts such as courses, majors, disciplines, course objectives, syllabi, class hours, training objectives, and graduation requirements, as well as teaching content and the correlation and mapping relationships between various teaching text data. These correlation and mapping relationships include conceptual subordination relationships, content inclusion relationships, and logical derivation relationships.
[0113] like Figure 2As shown, the online teaching system records the definitions of concepts such as course, major, discipline, course objectives, syllabus, class hours, training objectives, and graduation requirements, as well as the recorded teaching content (corresponding to...). Figure 2 The document contains the content of the document and records the relationships between the aforementioned items. These relationships include conceptual subordination, content inclusion, and logical derivation.
[0114] For example, courses are determined based on majors, so the relationship between courses and majors is a conceptual subordination relationship; majors are determined based on disciplines, so the relationship between majors and disciplines is a conceptual subordination relationship.
[0115] For example, since the syllabus covers the course content, the relationship between the syllabus and the content is one of content inclusion.
[0116] For example, graduation requirements are defined based on training objectives, so the relationship between training objectives and graduation requirements is a logical derivation relationship.
[0117] Optionally, the initial LLM model includes a vector database unit, a knowledge base unit, and a crawling unit. The knowledge base unit stores the teaching text data for the online teaching system. The crawling unit retrieves descriptions of training objectives, graduation requirements, syllabi, class hours, subjective questions, and answer content, employing techniques such as web scraping. The vector database unit converts the teaching text data into high-dimensional vectors (Embeddings) using an encoder. Each dimension of the vector represents semantic features (such as keywords, sentiment, and contextual relationships). Based on metrics such as cosine similarity and Euclidean distance between vectors, it quickly finds historical data or knowledge base content that most closely matches the current input semantics.
[0118] Step S202: Preprocess the acquired teaching text data to construct a teaching knowledge base, wherein the preprocessing includes text cleaning, text deduplication, and text segmentation.
[0119] In this embodiment, preprocessing refers to the entire process of systematically processing the raw text data before large-scale model training. Its goal is to transform unstructured text into standardized input that the model can understand. Text cleaning, text deduplication, and text segmentation are all core operations in preprocessing. These three processes respectively address data quality, redundancy, and structural standardization issues, collectively forming the basic framework of preprocessing. Text cleaning, text deduplication, and text segmentation all employ processing methods known in the art. The raw text data can be text from images, text / table extraction from images, speech-to-text conversion, etc.
[0120] Step S203: Based on the teaching knowledge base, pre-train the large language model to obtain the initial state LLM large model.
[0121] Specifically, large-scale model pre-training refers to the process of initially training a deep learning model on large-scale general data, aiming to enable the model to automatically learn common patterns, semantic representations, and structural regularities in the data. The core goal of pre-training is to endow the model with "prior knowledge," giving it the basic ability to handle diverse tasks, similar to how humans accumulate common sense through extensive reading. The essence of pre-training is to use self-supervised learning to allow the large model to mine intrinsic relationships from massive amounts of unlabeled data and form transferable feature representations. In this embodiment, after pre-training, the initial-state LLM large model possesses general knowledge capabilities in the field of education and teaching.
[0122] Step S204: Obtain the question-generating instruction.
[0123] In this embodiment, as Figure 5 As shown, the question-setting requirements input by the user to the client ( Figure 5 The first item in the document serves as the instruction for setting questions.
[0124] Step S205: Generate initial subjective questions from the initial state LLM large model according to the question generation instruction.
[0125] In this embodiment, initial subjective questions are generated from the initial state LLM large model according to the question-generating requirements of the question-generating instruction.
[0126] In one embodiment of the present invention, before obtaining the rating scale, the above method further includes the following steps:
[0127] Based on the content and answer requirements of the initial subjective questions, a scoring table structure for the answer content is constructed. The scoring table structure includes scoring criteria and scoring levels. The scoring criteria are set according to the course objectives, and the scoring levels are set according to the graduation requirements.
[0128] In this embodiment, the scoring sheet structure is a tool for assessing whether the answers meet the requirements of the initial subjective questions and whether the expected training objectives have been achieved. For example... Figure 3 As shown, the scoring criteria indicate the specific angles from which the answer is scored. For example, the scoring criteria include content completeness, language expression, structural organization, and originality. The scoring level indicates the level of performance of the answer on a certain scoring criterion. For example, the scoring levels include excellent, good, average, and poor, with each level specifying the upper limit of the score.
[0129] It should be noted that the scores listed in the rating scale represent the upper limit of that scale; for example, the upper limit for Excellent is 4 points, with a scoring range of 3 < Excellent ≤ 4; the upper limit for Good is 3 points, with a scoring range of 2 < Good ≤ 3; the upper limit for Average is 2 points, with a scoring range of 1 < Average ≤ 2; and the upper limit for Poor is 1 point, with a scoring range of 0 < Poor ≤ 1. Scores can be displayed to two decimal places, for example, 0.78 or 3.56.
[0130] In another embodiment of this application, a scoring table structure for the initial subjective questions can be constructed based on the specific content and requirements of the question-setting instructions. This scoring table structure includes scoring criteria and scoring levels. In this case, the scoring table structure serves as a tool to assess whether the initial subjective questions align with the course objectives and the teaching content.
[0131] In one embodiment of the present invention, the above method further includes the following steps:
[0132] Obtain evaluation responses;
[0133] Based on the evaluation failure information in the evaluation response, it is determined that the feedback quality of the first rating evaluation information is low;
[0134] When the feedback quality of the initial rating assessment is low, a high-quality dataset is obtained. This high-quality dataset consists of historically generated, high-quality data, or a dataset specifically constructed based on the requirements of the rating scale.
[0135] Based on the high-quality dataset, the initial state LLM large model is fine-tuned to obtain the fine-tuned LLM large model.
[0136] In this embodiment, the evaluation response indicates that the feedback quality of the first rating evaluation information is low. When the evaluation response is received, the client automatically marks it as low quality. Based on the evaluation response, the type of quality problem is identified. The types of quality problems include missing scoring criteria for the initial subjective question, ambiguous scoring levels for the initial subjective question, missing scoring criteria for the answer content, and ambiguous scoring levels for the answer content.
[0137] For example, when the initial subjective question has an ambiguous rating level corresponding to the language expression dimension, the rating scale module (Rubric) generates a fine-tuning instruction to "optimize the rating level description of this dimension." The rating scale module transmits the fine-tuning instruction and a high-quality dataset to the fine-tuning module (LORA). Through low-rank matrix update technology, without changing the core parameters of the initial LLM large model, the weights of parameters related to language expression scoring are specifically optimized to achieve accurate calibration of the scoring logic, ensuring that the rating judgment of this dimension is clearer and more objective in subsequent scoring. The low-rank matrix update technology adopts a processing method known in the art.
[0138] For example, when the quality issue type is that the scoring criteria for the content completeness dimension of the answer content is missing, the scoring scale module (Rubric) generates a fine-tuning instruction to "supplement the scoring criterion description for this dimension." The scoring scale module transmits the fine-tuning instruction and the high-quality dataset to the fine-tuning module (LORA). Through low-rank matrix update technology, without changing the core parameters of the initial state LLM model, the parameter weights related to content completeness are specifically supplemented to complete the scoring criteria and ensure that subsequent scoring supports the level determination for this dimension. In one embodiment provided by the present invention, the fine-tuning of the initial state LLM model based on the high-quality dataset to obtain the fine-tuned LLM model includes the following steps:
[0139] Text paragraphs containing descriptions of the causal relationship between scoring criteria and answer quality are selected from the teaching knowledge base. The teaching context information is extracted, the causal relationship tree is manually identified, and a causal relationship dataset S is constructed.
[0140] In the causal relationship dataset S, each data point S = {CTX, SD, QA, CRS}, where...
[0141] CTX={Location,Time}: CTX represents the context attribute, Location represents the location of the test (e.g., Beijing, Shanghai, etc.), and Time represents the time of the test (e.g., the year of the exam, the semester, etc.).
[0142] SD represents a subject;
[0143] QA refers to characteristics related to the quality of answers, including answer format, content completeness, and logical coherence.
[0144] CRS stands for Causal Relationship Dataset, which includes causal variables and outcome variables.
[0145] Provide contextual information for subjective questions based on contextual attributes, subject matter, and features related to answer quality; use causal relationship datasets to reason and analyze subjective questions and construct question-answer pairs.
[0146] Based on question-answer pairs, the initial-state LLM large model is fine-tuned to obtain the fine-tuned LLM large model.
[0147] In one embodiment of the present invention, the above method further includes the following steps:
[0148] The second scoring and evaluation information of the answer content is obtained based on the fine-tuned LLM large model;
[0149] Send the second rating and evaluation information to the first terminal;
[0150] or,
[0151] The third-party evaluation information of the answer content is obtained based on the fine-tuned LLM large model;
[0152] If the feedback quality of the third rating evaluation information is determined to be high, the third rating evaluation information will be sent to the first terminal.
[0153] In this embodiment, the fine-tuned LLM model re-evaluates the answer content to obtain second evaluation information. This second evaluation information is assumed to be of high feedback quality and is sent to the first terminal. Alternatively, in another embodiment, the fine-tuned LLM model re-evaluates the answer content to obtain third evaluation information. This third evaluation information is input into an adaptive module, which determines the feedback quality of the third evaluation information. If the third evaluation information is determined to be of high feedback quality, it is sent to the first terminal.
[0154] In one embodiment of the present invention, obtaining the dynamic smoothing coefficient based on the historical performance sequence includes the following steps:
[0155] Establish a mapping relationship between the answers and the rating scale;
[0156] Calculate the standard deviation of the historical performance series using the standard deviation algorithm;
[0157] The dynamic smoothing coefficients are calculated using a dynamic coefficient algorithm, which is obtained as follows:
[0158]
[0159] in, Indicates the dynamic smoothing coefficient; Represents the minimum smoothing factor; Indicates the maximum smoothing factor; The standard deviation of the historical score series is represented by t, where each historical score contains t terms, and t is a natural number (including 0); base represents the upper limit threshold for the score.
[0160] It should be noted that "base" represents the upper limit threshold for scoring, which refers to the "normalized upper limit threshold" of the standard deviation of the historical score series by the LLM large model. When When the score is ≥ base, the system determines that the score of the large LLM model has already fluctuated significantly. Reaching its maximum value. When If the value exceeds the base, the normalized upper limit is 1.
[0161] In this embodiment, as shown in Tables 1 and 2, the base initialization value is determined to be 0.3.
[0162]
[0163]
[0164] In this embodiment, the standard deviation can be the sample standard deviation. The standard deviation algorithm is obtained using the following method:
[0165]
[0166] Let n represent the historical scores, and n represent the number of historical scores. This represents the sample mean.
[0167] For example, the values of the historical performance sequence are 0.62, 0.80, 0.45, 0.92, and 0.68. = (0.62 + 0.80 + 0.45 + 0.92 + 0.68) ÷ 5 = 0.694; =0.005476+0.011236+0.059536+0.051076+0.000196≈0.1275;
[0168] That is, the standard deviation is 0.18.
[0169] For example, The value is 0.18, and the base value is 0.3. The value is 0.18. If the value of t is 0.27 and the value of t is 4, then... And so on. ; ; ; .
[0170] In this step, based on the changes in the historical score sequence, Adaptive change, i.e. The value of is dynamic and not unique; it can quickly capture changes in teaching and adjust automatically without manual intervention. It dynamically adjusts based on factors such as LLM model scoring confidence, content innovation intensity, and fluctuating student behavior. It allows for more flexible control over the proportions of new and old information (in training sample data). Therefore, based on dynamic expansion points, it is suitable for adaptive update scenarios in instructional graph rubric generation and scoring.
[0171] In one embodiment of the present invention, obtaining the current matching degree value edge weight of the first rating evaluation information based on the historical score sequence and the dynamic smoothing coefficient includes the following steps:
[0172] The residual initial weights are calculated according to the residual weight algorithm;
[0173] The current score contribution value is calculated according to the score contribution algorithm;
[0174] The edge weights of the current matching degree are calculated according to the matching strength algorithm, wherein the matching strength algorithm is obtained by the following method:
[0175]
[0176] in, Indicates the current matching degree value of the edge weight; Indicates the sovereign control coefficient; This represents the current score contribution value; This represents the residual initialization weight; The initial matching degree is represented by the edge weight.
[0177] The residual weight algorithm is obtained using the following method:
[0178]
[0179] The score contribution algorithm is obtained using the following method:
[0180]
[0181] This represents the score given to the student's assignment (i.e., the answer content) by the Rubric dimension (a certain scoring criterion) at the tkth time.
[0182] This indicates the half-life step of historical contributions, controlling the range of the historical scoring window. It is obtained using the following method:
[0183]
[0184] It should be noted that: when If the calculated result is not an integer, it is rounded down. For example, The calculation yields 2.3, which is then rounded down to the nearest integer, resulting in a value of 2. The calculated value is 4.9, which is then rounded down to 4. Rounding down ensures that the historical window does not exceed the number of steps required for half a life, simplifying the calculation and avoiding excessive accumulation of old scores. Furthermore, t is calculated starting from 0.
[0185] The fully expanded formula for the matching strength algorithm is:
[0186]
[0187] This value stems from the empirical principle of balancing "Rubric dynamic grading" and "original instructional structure design" in actual online learning systems (LMS). It was set after multidimensional analysis, including the following dimensions:
[0188] Reference Dimension 1: From the perspective of actual teaching curriculum design models (such as the Association of American Colleges and Universities' Valid Assessment of Learning in Undergraduate Education Rubric scheme; Bloom's Taxonomy of Objectives Rubric scheme), automatic or dynamic scoring structures should not completely cover the teacher's original structure, and should utilize a certain degree of structural memory. Based on rubric modeling practice, it is recommended that: 70% be driven by scoring results (emphasizing dynamic adaptation to student performance); 30% retain the original design structure (reflecting the generality and stability of the course). Therefore, Setting the value to 0.7 by default is more reasonable.
[0189] Reference Dimension 2: 30% of the teaching data requires innovative content in curriculum development, which can lead to significant noise. However, the teaching vertical domain represents 70% of the long-term stable preference. A stratified range of 0.2 < α ≤ 0.35 (indicating 20%–35% is historical information memory) is used to accommodate long-term preferences and smooth short-term fluctuations. Approximately 70%–80% of the data reflects current new information observations. Mapping this to the self-developed formula, λ = 0.7 aligns with the empirical principle of "current dominance, historical stability" for smooth integration.
[0190] Reference dimension 3: When the general model is fine-tuned and used for automatic Rubric scoring, the model itself has a certain risk of score drift or student "sudden performance" phenomenon.
[0191] Training performance after using synchronized numerical values:
[0192] λ=1.0: Completely controlled by scoring, high risk;
[0193] λ=0.5: Slow response to changes in rating;
[0194] λ=0.7: A practical balance is achieved between "responsiveness" and "structural stability"; the structure is relatively stable, allowing for dynamic adjustment of the scoring direction.
[0195] The conclusion is that a numerical strategy of λ is adopted when refining the scene:
[0196] λ=0.7 is a balanced default value in teaching evaluation that considers both "responsiveness" and "structural memory." Simultaneously, the system dynamically adjusts the value of λ based on the teacher's selected preference settings and by monitoring the confidence level of the large model's scoring. This achieves the goal of automatic rolling regression with a flexible and dynamic adjustment strategy.
[0197] also, Other values can also be set according to other application scenarios and application requirements. Example
[0198] Taking the assessment of the feedback quality of the first rating evaluation information based on a single rating criterion dimension (e.g., content completeness) as an example, this embodiment will be further elaborated. In this embodiment, as shown in Table 3, the edge weights of the current matching degree are calculated for each parameter value. Here, t represents the discrete time step. For example, the most recent score before the time indicated by the timestamp in the first rating evaluation information can be used as a historical score. Alternatively, the five most recent scores before the time indicated by the timestamp in the first rating evaluation information can be used as historical scores to construct a historical score sequence.
[0199]
[0200] For example, when t=4 (t starts counting from 0), under the same scoring criteria, The scores are shown in Table 4.
[0201]
[0202] The value is the largest, and its value is 3 when it is taken down from this value. Taking a value of 2 downwards, k = 0, 1, 2 in the following summation formula:
[0203] ;
[0204] Therefore, we can calculate: .
[0205] In one embodiment of the present invention, obtaining the confidence value of the first rating evaluation information based on the edge weight of the current matching degree value, and obtaining the feedback quality of the first rating evaluation information based on the confidence value, includes the following steps:
[0206] The confidence value of the first rating evaluation information is calculated according to a confidence algorithm, which is obtained using the following method:
[0207]
[0208] Where Z represents the confidence level; This indicates the first score of the first rating evaluation information; This indicates the expected score. The first score in the first rating evaluation information refers to the score given under a specific rating criterion.
[0209] Determine whether the confidence value falls within the confidence interval;
[0210] If the confidence value falls within the confidence interval, the feedback quality of the first rating evaluation information is determined to be high.
[0211] It should be noted that the confidence level is calculated for a specific scoring criterion; high feedback quality is a quality confirmation for that specific scoring criterion. In this application, the comprehensive scoring evaluation information can be sent to the first terminal only when the feedback quality of the scoring evaluation information corresponding to each of the four scoring criteria is high; alternatively, the scoring evaluation information can be sent to the first terminal only when the feedback quality of the scoring evaluation information corresponding to a specific scoring criterion is high.
[0212] The expected score is calculated according to a scoring algorithm, which is obtained using the following method:
[0213] ; This indicates the maximum score in a certain scoring criterion.
[0214] For example, It is 30 points. If it is 0.325, then =0.325×30=9.75; If the score is 9.9 (the actual score), then Z = (9.9 - 9.75) ÷ 0.18 = 0.83. The confidence interval is the range from -1.645 (inclusive) to 1.645 (inclusive). To determine if a confidence value falls within the confidence interval, |Z| ≤ 1.645. Since |0.83| < 1.645, the confidence value falls within the confidence interval, indicating that the feedback quality of the first rating evaluation information is high, meaning the rating accuracy is above 90%. In this step, the confidence interval can be set; for example, it can be the range from -1.512 (inclusive) to 1.512 (inclusive), etc.
[0215] In existing technologies, the noise level in scoring data is approximately 40%. By utilizing the steps described in this application, the noise level of the generated scoring information is controlled to below 10%, achieving a scoring accuracy of over 90%.
[0216] Example 2
[0217] In the embodiments of this application, the scoring and evaluation information for each topic (answer content or subjective question) corresponds to multiple different scoring standard dimensions. The following is an example illustrating how "the scoring and evaluation information for the answer content is judged from four scoring standard dimensions (e.g., content completeness, language expression, structural organization, and originality) to determine the feedback quality of the comprehensive scoring and evaluation information."
[0218] Step S31: Calculate and determine the feedback quality of the rating information corresponding to each rating criterion dimension in turn. When the feedback quality of the rating information for all four rating criterion dimensions is low, the overall rating information is determined to be non-compliant, and the non-compliant information is transmitted to the rating scale. Return to fine-tuning, that is, fine-tune the initial state LLM model. The fine-tuning process is consistent with the description in the text and will not be repeated here.
[0219] Step S32: When the rating evaluation information of at least one rating criterion dimension is of high feedback quality, calculate the confidence value Z' of the comprehensive rating evaluation information according to the comprehensive algorithm. The confidence value Z' of the comprehensive rating evaluation information is obtained by the following method:
[0220] ;
[0221] Where Z' represents the confidence value of the comprehensive rating evaluation information; a represents the overall coefficient, 0 < a < 1; M h This represents the sum of confidence scores for each rating criterion dimension indicating high feedback quality; M l This represents the sum of confidence scores for each rating criterion dimension with low feedback quality.
[0222] Step S33: Compare Z' with the comprehensive threshold. If Z' is greater than the comprehensive threshold, the overall rating feedback quality is considered high, and all rating feedback information from the four rating criteria dimensions is output. If Z' is less than the comprehensive threshold, the overall rating feedback quality is considered low, and all rating feedback information from the four rating criteria dimensions is not output, returning to fine-tuning. The comprehensive threshold can be the average confidence value Z of historically obtained high-quality criteria, or it can be any other set value.
[0223] In the embodiments of this application, 'a' can be dynamically adjusted based on the student's ability requirements according to the answer content. When more emphasis is placed on the student's high-confidence items, the confidence level of the scoring information generated by a single item or several items under the corresponding scoring standard dimension is high (as long as the confidence level of the scoring information generated under the corresponding scoring standard dimension is high), which can compensate for the impact of low confidence levels of the scoring information generated under other scoring standard dimensions on the feedback quality. The value range of 'a' is set to 0.7 < a < 1. This increases the proportion of scoring standard dimensions with high feedback quality, aligning with the overall high feedback quality of the scoring information. Here, high confidence means that the confidence value Z corresponding to the scoring information is within the confidence interval; low confidence means that the confidence value Z corresponding to the scoring information is not within the confidence interval.
[0224] When the question content places greater emphasis on students' outstanding abilities in high-confidence areas, only when the confidence level of the rating information generated by a single or several items under the corresponding rating criterion dimension is high (and only when the confidence level of the rating information generated under the corresponding rating criterion dimension reaches a certain proportion (e.g., the confidence level of all three rating criterion dimensions is high)) can the impact of low confidence level rating information under other rating criterion dimensions on the feedback quality be compensated. Setting the value range of 'a' to 0 < a < 0.3 reduces the proportion of rating criterion dimensions with high feedback quality. Only when the ability in high-confidence items is outstanding can the impact of low confidence level abilities on the feedback quality be compensated. This can be used to select answer content that has high requirements for a certain ability.
[0225] When the answer content places greater emphasis on the student's overall assessment criteria, for example, the value range of 'a' can be set to 0.3 < a < 0.7. Only when the score with a low confidence level in the assessment criterion dimension corresponding to the answer content is close to the required confidence level can the impact of low confidence levels in other assessment criterion dimensions on the student's score be compensated. This reflects the reconciliation effect between assessment criterion dimensions with high and low feedback quality. Specifically, for a score with a low confidence level in the assessment criterion dimension corresponding to the answer content to be close to the required confidence level, the difference between the low confidence level and the required confidence level is less than a threshold difference, which can be set according to needs.
[0226] In one embodiment of this application, 'a' can be set according to the ability requirements of the answer content, or it can be based on the analysis of historical data, which is preferentially divided into three categories. Features of each category of answer content are extracted based on a large model. When new answer content is obtained, the features of the new answer content are compared with the features of each category of answer content to automatically obtain the corresponding value of 'a'. This embodiment of the application does not impose any limitations.
[0227] Thus, in the embodiments of this application, based on the characteristics of each type of answer content, the answer content corresponding to the comprehensive scoring and evaluation information that meets the feedback quality can be determined. Students who meet the requirements can be selected from the answer content for further training, rather than simply cultivating students' general abilities for different subjective questions. This is more conducive to the development and cultivation of students' special talents.
[0228] In one embodiment of the present invention, after pre-training the large language model based on the teaching knowledge base to obtain the initial state LLM large model, the method provided by the present invention performs the following steps:
[0229] A multidimensional feature vector set is constructed based on the teaching knowledge base, and the multidimensional feature vector set includes dimension coverage vector, standard description vector, and domain adaptation vector.
[0230] Specifically, constructing a multi-dimensional feature vector set based on a teaching knowledge base involves systematically extracting and quantitatively modeling features from core dimensions of the knowledge system within the teaching knowledge base, such as its structure, descriptive standardization, and domain adaptability. Dimensional coverage vectors extract structural features from the knowledge modules in the teaching knowledge base, including hierarchical relationships, logical connections, and coverage, forming vectors representing the dimensional completeness of the knowledge system. Standard description vectors extract features from the standardization, process-orientation, and practical guidance of knowledge descriptions, forming vectors that measure the quality of knowledge representation. Domain adaptability vectors extract adaptability features of the knowledge content based on the needs, scenario characteristics, and audience characteristics of the target domain, forming domain-customized vectors.
[0231] Furthermore, a reasonableness scoring standard for the output structure is set based on the dimensional coverage vector; an operability scoring standard is set based on the standard description vector; and a domain adaptability scoring standard is set based on the domain adaptability vector.
[0232] The question score R is calculated according to the question scoring algorithm, wherein the question scoring algorithm is obtained using the following method:
[0233] R= *Structural rationality score+ *Feasibility score+ *Domain suitability score;
[0234] in + + =1, and , , Dynamically adjust according to application scenarios, among which... , , All represent rating coefficients.
[0235] Based on multidimensional feature vectors, the output scores for structural rationality, operability, and domain adaptability are determined.
[0236] According to the formula R= *Structural rationality score+ *Feasibility score+ *Domain suitability scoring calculates the question's score; where... + + =1, and , , The system dynamically adjusts based on the application scenario. The technologies used here—Transformer architecture, attention mechanism, and positional encoding—ensure the ability to understand the Rubric language logic.
[0237] It should be noted that, in this application, it is also possible to construct, for example, a topic containing, such as Figure 3 The scoring scale, which includes the scoring criteria and scoring dimensions shown, is used to score the initial subjective questions generated by the initial-state LLM large model based on the question scoring scale.
[0238] In one embodiment of the present invention, the above method further includes the following steps:
[0239] The initial subjective questions, the answer content, and the first score evaluation information with high feedback quality are preprocessed to construct a second teaching knowledge base, wherein the preprocessing includes text cleaning, text deduplication, and text segmentation.
[0240] Based on the second teaching knowledge base, the initial state LLM large model is pre-trained to obtain an advanced version of the LLM large model.
[0241] In this step, based on the second teaching knowledge base, the initial-state LLM model is pre-trained to obtain an advanced version of the LLM model. The advanced version of the LLM model will enhance the ability to reason about knowledge points, optimize the learning of the "question-evaluation" logic, and correct the randomness deviations in question design and difficulty control of the initial-state LLM model.
[0242] like Figure 4As shown, the technical solution of this invention constructs an intelligent scoring system based on a pre-trained initial-state LLM large model, and achieves automation and continuous optimization of question generation and scoring through the following steps:
[0243] Step A: The system first receives the question-generating instructions from the teacher, including parameters such as the scope of knowledge points, difficulty level, and question type requirements. The pre-trained initial-state LLM model generates initial subjective questions and a corresponding scoring scale based on the teacher's input instructions. The scoring scale is used to score the answers to the initial subjective questions. The scoring scale employs a multi-dimensional design, including dimensions such as language expression, structural organization, and content completeness, and defines grading standards and anchor examples for each dimension, forming a structured scoring framework.
[0244] Step B: Validate the rating scale. Rating scale validation involves multi-dimensional testing to confirm whether the rating scale effectively achieves its scoring objectives and avoids subjective bias or rule loopholes. For example, the testing can include: discrimination testing (using a test set to verify whether the rating scale can effectively distinguish between responses of different quality); consistency assessment (using indicators such as the Kappa coefficient to test the consistency of multiple users' use of the rating scale); and robustness testing (implanting pre-set biased answers (such as well-structured but empty text) to test the robustness of the rating scale). If the testing fails, the system will mark the suspicious dimensions and trigger a second calibration by the teacher until the rating scale passes validation.
[0245] Step C: Generate scoring prompts based on the teacher's input instructions, and the initial-state LLM model performs the scoring. For example, after the teacher inputs instructions containing scoring information, the system converts them into a scoring prompt template understandable by the LLM model. The template includes information such as the answer content, scoring dimensions, assessment focus, and scoring criteria. The initial-state LLM model scores the answer content based on this prompt, outputting structured results, including scores for each dimension, textual feedback, and improvement suggestions.
[0246] Step D: Input the rating and evaluation information generated by the initial state LLM large model into the adaptive module, which calculates the confidence value and determines the feedback quality of the rating and evaluation information based on the confidence value.
[0247] Step E: If the feedback quality is high, the scoring and evaluation information is input into the teaching knowledge base. This means the scoring and evaluation information is used as part of the teaching text data and pre-trained on the initial-state LLM model. In this step, the scoring and evaluation information that passes quality verification is stored in the teaching knowledge base, which is categorized and indexed by dimensions such as subject and knowledge point. For example, the system employs an incremental pre-training strategy, periodically extracting data from the knowledge base to continuously train the initial-state LLM model, improving its performance in specific domain scoring tasks.
[0248] Step F: When feedback quality is low, determine the corresponding scoring criteria for the evaluation information. Send the low feedback quality information under this scoring criterion dimension to the scoring scale. The scoring scale generates fine-tuning instructions, which are then sent to the fine-tuning module. The fine-tuning module fine-tunes the initial LLM model to obtain the fine-tuned LLM model. Using the fine-tuned LLM model, regenerate the evaluation information according to the aforementioned scoring criteria dimensions.
[0249] Secondly, embodiments of the present invention provide a system for enhancing the credibility of teaching evaluation in a large-scale education model, such as... Figure 6 As shown, the system includes an acquisition unit 601, an evaluation unit 602, and a quality feedback unit 603;
[0250] The acquisition unit 601 is used to acquire a scoring prompt and, based on the scoring prompt, acquire the first scoring evaluation information of the answer content.
[0251] The evaluation unit 602 is used to: acquire the interactive object corresponding to the first rating evaluation information; construct a historical score sequence for different users' scores under the same rating standard; acquire a dynamic smoothing coefficient based on the historical score sequence; acquire the current matching degree value edge weight of the first rating evaluation information based on the historical score sequence and the dynamic smoothing coefficient; acquire the confidence value of the first rating evaluation information based on the current matching degree value edge weight; and acquire the feedback quality of the first rating evaluation information based on the confidence value.
[0252] The quality feedback unit 603 is used to determine that the first rating evaluation information meets the standard when the feedback quality of the first rating evaluation information is high, and to send the first rating evaluation information to the first terminal.
[0253] In one embodiment of the present invention, the system further includes a pre-training unit and a question generation unit;
[0254] The pre-training unit is used to acquire teaching text data from the online teaching system; preprocess the acquired teaching text data to construct a teaching knowledge base, wherein the preprocessing includes text cleaning, text deduplication, and text segmentation; and pre-train a large language model based on the teaching knowledge base to obtain an initial-state LLM large model.
[0255] The question-generating unit is used to obtain question-generating instructions; and to generate initial subjective questions from the initial state LLM large model according to the question-generating instructions.
[0256] Thirdly, embodiments of the present invention provide an electronic device, such as... Figure 7As shown, the electronic device 70 includes: a memory 701, a processor 702, and a computer program stored in the memory 701 and executable on the processor 702, wherein the processor 702 executes the computer program to implement the steps of the above method.
[0257] The electronic device 70 can be a computing device such as a tablet computer, desktop computer, or cloud server. The electronic device 70 may include, but is not limited to, a processor 702 and a memory 701. Those skilled in the art will understand that... Figure 7 This is merely an example of electronic device 70 and does not constitute a limitation on electronic device 70. It may include more or fewer components than shown, or combine certain components, or different components. For example, it may also include input / output devices, network access devices, etc.
[0258] The processor 702 may be a Central Processing Unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor. The processor 702 may also be a GPU / DCU computing power processor.
[0259] In some embodiments, the memory 701 may be an internal storage unit of the electronic device 70, such as a hard disk or memory of the electronic device 70. In other embodiments, the memory 701 may be an external storage device of the electronic device 70, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the electronic device 700. Furthermore, the memory 701 may include both internal and external storage units of the electronic device 70. The memory 701 is used to store the operating system, applications, bootloader, data, and other programs, such as the program code of the computer program. The memory 701 can also be used to temporarily store data that has been output or will be output.
[0260] In the several embodiments provided in this application, it will be understood that each block in the flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the figures. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved.
[0261] If the aforementioned functions are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0262] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of this application. It should be understood that the above descriptions are merely specific embodiments of this application and are not intended to limit the scope of protection of this application. In particular, it should be noted that any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application for those skilled in the art.
Claims
1. A method for enhancing the credibility of teaching evaluation in a large-scale education model, characterized in that, The method includes: Obtain a scoring prompt, and based on the scoring prompt, obtain the first scoring evaluation information of the answer content; Obtain the interaction object corresponding to the first rating evaluation information, and construct a historical score sequence for different users under the same rating standard; Based on the historical performance sequence, obtain the dynamic smoothing coefficient; Based on the historical score sequence and the dynamic smoothing coefficient, obtain the current matching degree value edge weight of the first score evaluation information; Based on the current matching degree value edge weight, obtain the confidence value of the first rating evaluation information; Based on the confidence level, the feedback quality of the first rating evaluation information is obtained; If the feedback quality of the first rating evaluation information is high, it is determined that the first rating evaluation information meets the standard, and the first rating evaluation information is sent to the first terminal. The step of obtaining the dynamic smoothing coefficient based on the historical performance sequence includes: Calculate the standard deviation of the historical performance series; The dynamic smoothing coefficient is calculated according to the dynamic coefficient algorithm, wherein the dynamic coefficient algorithm is obtained by the following method: ; in, Indicates the dynamic smoothing coefficient; Represents the minimum smoothing factor; Indicates the maximum smoothing factor; The standard deviation of the historical score series is represented by "base"; the upper limit threshold for scoring is represented by "base". The step of obtaining the current matching degree value edge weight of the first rating evaluation information based on the historical score sequence and the dynamic smoothing coefficient includes: The residual initial weights are calculated according to the residual weight algorithm; The current score contribution value is calculated according to the score contribution algorithm; The edge weights of the current matching degree are calculated according to the matching strength algorithm, wherein the matching strength algorithm is obtained by the following method: ; in, Indicates the current matching degree value of the edge weight; Indicates the sovereign control coefficient; This represents the current score contribution value; This represents the residual initialization weight; The initial matching degree is represented by the edge weight.
2. The method according to claim 1, characterized in that, Before obtaining the scoring prompt and, based on the scoring prompt, obtaining the first scoring evaluation information of the answer content, the method further includes: Obtain teaching text data from online teaching systems; The acquired teaching text data is preprocessed to construct a teaching knowledge base, wherein the preprocessing includes text cleaning, text deduplication, and text segmentation. Based on the teaching knowledge base, pre-training is performed on the large language model to obtain the initial state LLM large model; Obtain the question-setting instruction; According to the question-generating instruction, initial subjective questions are generated from the initial state LLM large model.
3. The method according to claim 1, characterized in that, Before obtaining the scoring prompt and, based on the scoring prompt, obtaining the first scoring evaluation information of the answer content, the method further includes: Obtain initial subjective questions and a rating scale corresponding to the answers; wherein, the answers are the responses given by multiple users to the initial subjective questions; The step of obtaining a scoring prompt, and obtaining first scoring evaluation information of the answer content based on the scoring prompt, includes: Obtain the rating prompts and the rating scale; Input the scoring prompts, the scoring scale, and the answer content into the initial state LLM large model to obtain the first scoring evaluation information of the answer content.
4. The method according to claim 2, characterized in that, The method further includes: When the feedback quality of the first rating and evaluation information is low, obtain a high-quality dataset; Based on the high-quality dataset, the initial state LLM large model is fine-tuned to obtain the fine-tuned LLM large model.
5. The method according to claim 4, characterized in that, The process of fine-tuning the initial-state LLM model based on the high-quality dataset to obtain the fine-tuned LLM model includes: Text paragraphs containing descriptions of the causal relationship between scoring criteria and answer quality were selected from the teaching knowledge base. A causal relationship dataset is constructed based on the scoring criteria and causal relationship descriptions of answer quality in the text paragraphs. The causal relationship dataset is used to reason and analyze the input subjective questions to construct question-answer pairs; Based on the question-answer pair, the initial state LLM large model is fine-tuned to obtain the fine-tuned LLM large model.
6. The method according to claim 5, characterized in that, The method further includes: The second scoring and evaluation information of the answer content is obtained based on the finely tuned LLM model. Send the second rating and evaluation information to the first terminal; or, The third scoring and evaluation information of the answer content is obtained based on the fine-tuned LLM large model; If the feedback quality of the third rating evaluation information is determined to be high, the third rating evaluation information is sent to the first terminal.
7. The method according to claim 1, characterized in that, The confidence value of the first rating evaluation information is obtained based on the edge weight of the current matching degree value. Based on the confidence level, the feedback quality of the first rating evaluation information is obtained, including: The confidence value of the first rating evaluation information is calculated according to a confidence algorithm, which is obtained using the following method: ; Where Z represents the confidence level; This indicates the first score of the first rating evaluation information; Indicates the expected score; Determine whether the confidence value falls within the confidence interval; If the confidence value falls within the confidence interval, the feedback quality of the first rating evaluation information is determined to be high.
8. The method according to claim 2, characterized in that, After pre-training the large language model based on the teaching knowledge base to obtain the initial state LLM large model, the method further includes: A multidimensional feature vector set is constructed based on the teaching knowledge base, and the multidimensional feature vector set includes dimension coverage vector, standard description vector, and domain adaptation vector. The output structure rationality scoring criteria are set according to the dimensional coverage vector; the operability scoring criteria are set according to the standard description vector; and the domain adaptability scoring criteria are set according to the domain adaptability vector. The question score R is calculated according to the question scoring algorithm, wherein the question scoring algorithm is obtained using the following method: R= Structural rationality score + Operability score + Domain suitability score; in + + =1, and , , Dynamically adjust according to application scenarios, among which... , , All represent rating coefficients.
9. The method according to claim 1, characterized in that, The method further includes: The initial subjective questions, the answer content, and the first score evaluation information with high feedback quality are preprocessed to construct a second teaching knowledge base, wherein the preprocessing includes text cleaning, text deduplication, and text segmentation. Based on the second teaching knowledge base, the initial state LLM large model is pre-trained to obtain an advanced version of the LLM large model.
10. A system for enhancing the credibility of teaching evaluation in a large-scale education model, characterized in that, The system includes an acquisition unit, an evaluation unit, and a quality feedback unit; The acquisition unit is used to acquire a scoring prompt and, based on the scoring prompt, acquire the first scoring evaluation information of the answer content. The evaluation unit is used to: acquire the interactive object corresponding to the first rating evaluation information; construct a historical score sequence of different users' scores under the same rating standard; acquire a dynamic smoothing coefficient based on the historical score sequence; acquire the current matching degree value edge weight of the first rating evaluation information based on the historical score sequence and the dynamic smoothing coefficient; acquire the confidence value of the first rating evaluation information based on the current matching degree value edge weight; and acquire the feedback quality of the first rating evaluation information based on the confidence value. The quality feedback unit is used to determine that the first rating evaluation information meets the standard when the feedback quality of the first rating evaluation information is high, and to send the first rating evaluation information to the first terminal. The step of obtaining the dynamic smoothing coefficient based on the historical performance sequence includes: Calculate the standard deviation of the historical performance series; The dynamic smoothing coefficient is calculated according to the dynamic coefficient algorithm, wherein the dynamic coefficient algorithm is obtained by the following method: ; in, Indicates the dynamic smoothing coefficient; Represents the minimum smoothing factor; Indicates the maximum smoothing factor; The standard deviation of the historical score series is represented by "base"; the upper limit threshold for scoring is represented by "base". The step of obtaining the current matching degree value edge weight of the first rating evaluation information based on the historical score sequence and the dynamic smoothing coefficient includes: The residual initial weights are calculated according to the residual weight algorithm; The current score contribution value is calculated according to the score contribution algorithm; The edge weights of the current matching degree are calculated according to the matching strength algorithm, wherein the matching strength algorithm is obtained by the following method: ; in, Indicates the current matching degree value of the edge weight; Indicates the sovereign control coefficient; This represents the current score contribution value; This represents the residual initialization weight; The initial matching degree is represented by the edge weight.
11. The system according to claim 10, characterized in that, The system also includes a pre-training unit and a question generation unit; The pre-training unit is used to acquire teaching text data from the online teaching system; preprocess the acquired teaching text data to construct a teaching knowledge base, wherein the preprocessing includes text cleaning, text deduplication, and text segmentation; and pre-train a large language model based on the teaching knowledge base to obtain an initial-state LLM large model. The question-generating unit is used to obtain question-generating instructions; and to generate initial subjective questions from the initial state LLM large model according to the question-generating instructions.
12. An electronic device, the electronic device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor, when executing the computer program, implements the steps of the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Schoolwork grading method, system, schoolwork management system
CN107146176A
Multi-dimensional interpretable subjective question scoring method based on large model
CN120068840A