Open test question evaluation method based on domain large model and related device
Through the open test question evaluation method based on the domain-based large model, the problem of simple feature extraction and lack of semantic understanding in the existing technology is solved, and the deep semantic understanding and accurate automatic scoring of answers to open test questions is achieved.
Patent Information
- Application Number
- CN202510583018.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-05-07
AI Technical Summary
The existing open test evaluation technology has the feature extraction that is too simple and lacks semantic understanding, and cannot accurately analyze the coverage of the viewpoint and capture the candidates' innovative views, and is prone to overfitting problems.
The open test question evaluation method based on the domain-based large model is adopted, and domain knowledge is injected through the incremental training data set and instruction fine-tuning, domain model is constructed, and the prompt template of the viewpoint mining task is fine-tuned, and candidates are extracted and classified, and finally, they are automatically scored based on the standard answers viewpoint set.
A deep understanding of complex semantics, logical relationships and context dependencies in long texts is achieved, overfitting is avoided, viewpoint coverage is accurately analyzed, candidates' innovative views are captured, and automatic scoring is provided at the semantic level.
Smart Images

Figure CN120106176A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of educational examination data mining, and in particular to an open test question evaluation method based on a domain-based large model and related devices. Background Art
[0002] The evaluation of open-ended test answers aims to give a reasonable evaluation by analyzing and assessing students' answers to open-ended questions. In recent years, with the continuous advancement of educational informatization, the way of answering and evaluating test questions has gradually shifted towards automation and intelligence. In the evaluation of open-ended test questions, the traditional manual marking method is inefficient and highly subjective, and is easily affected by the emotional bias and fatigue of the examiner. It is also easy to ignore the innovative answers of some students and even cause misjudgment. The existing open-ended test evaluation technology based on capturing simple text features for pattern matching has problems such as overly simple feature extraction, lack of semantic understanding, and neglect of innovative answers. Therefore, a new open-ended test evaluation technology is urgently needed. In response to the above problems, the present invention proposes an open-ended test answer evaluation method based on a domain-based large model.
[0003] The prior art CN108764074B proposes a subjective question intelligent marking method and system for automatic marking and grading, which mainly includes: first, obtaining the answer sheet image and preprocessing it, dividing the answer area of objective questions and subjective questions and identifying them; then reviewing them separately according to the question type. For objective questions, compare whether the answer content and the standard answer match; for subjective questions, extract text features to train the convolutional neural network so that it can output the scoring range of the subjective question answers; finally, for papers with abnormal scores during the marking process, manual intervention is performed.
[0004] The above-mentioned subjective question intelligent marking method based on deep learning relies on convolutional neural networks to process simple text structural features, ignores the semantic understanding of the text, has a weak understanding of the complex semantics, logical relationships and context dependencies in long texts, and inevitably has the risk of overfitting. It is unable to accurately analyze the coverage of opinions and capture the innovative opinions of candidates. Summary of the invention
[0005] The purpose of the present invention is to provide an open-ended test question evaluation method and related devices based on a domain-based large model, so as to solve the problems of weak understanding of complex semantics, logical relationships and contextual dependencies in long texts, the inevitable risk of overfitting, and the inability to accurately analyze the coverage of viewpoints and capture the innovative viewpoints of examinees.
[0006] To achieve the above object, the present invention adopts the following technical solutions: In a first aspect, the present invention provides an open-ended test question evaluation method based on a domain-based large model, comprising: Obtain incremental training data sets and instruction fine-tuning sets through the subject corpus data, and obtain the candidate's answer set; Based on the incremental training data sets and instruction fine-tuning sets, inject domain knowledge into the preset base model and align the intent of the instructions to obtain the domain model; Based on the candidate answer set, a prompt template for the candidate opinion mining task is constructed, and the obtained domain model is fine-tuned based on the prompt template of the candidate opinion mining task to obtain the fine-tuned domain model, and the fine-tuned domain model is tuned; Extract candidates' answering opinions from the optimized domain model and classify them into a standard answer opinion set and an innovative opinion set. Supplement the standard answer opinion set based on the innovative opinion set to obtain the final standard answer opinion set. Candidates' answers are scored based on the final set of standard answer viewpoints.
[0007] Optionally, the step of obtaining an incremental training data set and an instruction fine-tuning set through the subject corpus data, and obtaining a test taker's answer set, includes: Textbooks, books, papers, and question banks of the disciplines to which the open test questions belong are collected as corpus, low-quality data in the corpus is removed through a rule-based data filtering method, and document deduplication is performed to obtain an incremental training data set; an instruction fine-tuning set is generated through question bank question and answer data and a large model dialogue based on subject documents, and the instruction fine-tuning set includes a single-round dialogue set and a multi-round dialogue set; the candidate answer set is obtained by image segmentation and optical character recognition.
[0008] Optionally, injecting domain knowledge into the preset base large model and aligning instruction intent based on the incremental training data set and the instruction fine-tuning set to obtain a domain model includes: The base large model is incrementally pre-trained using the incremental pre-training set to inject domain knowledge, and then the instruction fine-tuning dataset is used to fine-tune the incremental pre-trained model to align the intention of human instructions, thereby obtaining a discipline-oriented domain large model.
[0009] Optionally, constructing a prompt template for the examinee's viewpoint mining task based on the examinee's answer set, and fine-tuning the obtained domain model based on the prompt template for the examinee's viewpoint mining task to obtain the fine-tuned domain model includes: The prompt template format is defined based on the candidate's answer set, the prompt data is annotated according to the prompt template, and the low-rank adaptive efficient parameter fine-tuning domain model in the parameter efficient fine-tuning package is used to train the model on the samples to adapt to the opinion mining task. The trained domain model is used to output the opinion information and its explanation of the candidate's answer in the format defined by the template.
[0010] Optionally, the fine-tuning operation on the domain model includes: Based on the fine-tuned domain model, access the external data storage system to perform retrieval optimization on the model output.
[0011] Optionally, extracting examinees' answering opinions from the optimized domain model, classifying them into a standard answer opinion set and an innovative opinion set, and supplementing the standard answer opinion set based on the innovative opinion set to obtain a final standard answer opinion set, including: First, for each answer in the candidate's answer set, a prompt is filled in according to the opinion mining prompt template to obtain a prompt, and then the prompt is input into the optimized domain big model to extract the candidate's answer opinions to obtain the initial opinion set corresponding to all candidate's answers, and then all opinions in the initial opinion set are de-duplicated, merged and inferior to obtain the final opinion set. With reference to the preset recommended standard opinion set, opinions that do not appear in the recommended standard opinion set are constructed into an unverified innovative opinion set. Finally, the unverified innovative opinion set is reviewed and evaluated as correct innovative opinions. The correct innovative opinions are expanded to the standard answer opinion set as part of the subsequent marking criteria. The expanded opinion set is used as the final standard answer opinion set.
[0012] Optionally, the step of assigning scores to the examinee's answers based on the final set of standard answer viewpoints includes: Based on the final set of standard answer opinions, after opinion mining is performed on the candidates' answers to obtain the candidates' opinion information, the opinion words are compared with the final set of standard answer opinions. Points are awarded if the comparison is successful, and no points are awarded if the comparison is unsuccessful.
[0013] In a second aspect, the present invention provides an open test question evaluation system based on a domain-based large model, comprising: The domain model building module is used to obtain incremental training data sets and instruction fine-tuning sets through the subject corpus data, and obtain the candidate answer set through text recognition; based on the incremental training data sets and instruction fine-tuning sets, domain knowledge is injected and the intention of human instructions is aligned to obtain the domain model; The domain model fine-tuning module is used to construct a prompt template for the candidate's opinion mining task based on the candidate's answer set, perform prompt fine-tuning on the obtained domain model, obtain the fine-tuned domain model, and optimize the fine-tuned domain model; The module for obtaining the standard answer viewpoint set is used to extract the examinees' answer viewpoints from the optimized domain model, classify them into the standard answer viewpoint set and the innovative viewpoint set, and supplement the standard answer viewpoint set based on the innovative viewpoint set to obtain the final standard answer viewpoint set; The scoring output module is used to assign scores to candidates' answers based on the final set of standard answer viewpoints.
[0014] In a third aspect, the present invention provides a computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the open-ended test question evaluation method based on the domain-based large model when executing the computer program.
[0015] In the first aspect, the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the open test question evaluation method based on the domain-based big model is implemented.
[0016] Compared with the prior art, the present invention has the following technical effects: The present invention realizes the evaluation of answers to open questions by building a domain model, fine-tuning a few sample prompts, mining answer viewpoints, expert feedback on innovative viewpoints, and automatic scoring. First, based on the open source basic large model, large-scale incremental pre-training is carried out using subject resources and other field knowledge, a multi-task template set is designed for instruction fine-tuning, and a large domain model is built; secondly, a viewpoint mining and answer summary prompt learning template is designed to fine-tune the domain model with a few samples; then, the viewpoint information of the answer is analyzed and summarized, and the new viewpoints emerging therein are submitted to experts for review; finally, based on the viewpoint set composed of new and old viewpoints, automatic scoring is performed with reference to the matching degree between the answer viewpoint and the viewpoint set. The present invention fully mines the semantic information of the answers to open questions and simplifies it into the form of a viewpoint set, which is convenient for extracting and reviewing innovative viewpoints, and also provides more fine-grained automatic scoring at the semantic level, with the advantages of sufficient semantic understanding and accurate automatic scoring, which makes the present invention have obvious advantages over other methods based on simple text feature pattern recognition.
[0017] The present invention learns the semantic features of a subject to construct a domain model, which can be reused in all subjects related to the subject after being constructed once, and is not limited to a specific question.
[0018] With the powerful semantic capability of the large domain model, the present invention only needs to construct a small number of samples for prompt fine-tuning to meet the model construction conditions. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 It is a framework diagram of the open-ended test question answering evaluation method based on the domain-based large model of the present invention.
[0020] Figure 2 It is a flow chart of data collection and processing.
[0021] Figure 3 It is a flowchart for constructing a large domain model.
[0022] Figure 4 This is a flowchart for tuning large domain models.
[0023] Figure 5 It is a flow chart of opinion mining for answering open questions.
[0024] Figure 6 It is the automatic scoring flow chart.
[0025] Figure 7 It is a flow chart of the present invention. DETAILED DESCRIPTION
[0026] The following is a detailed description of the embodiments of the present invention in conjunction with the accompanying drawings and embodiments. It should be noted that the embodiments described herein are only used to explain the present invention and are not intended to limit the present invention. In addition, the technical features involved in the embodiments of the present invention can be combined with each other in the absence of conflict.
[0027] Glossary: OpenCV: Open Source Computer Vision Library (Open Source Computer Vision Library); OCR: Optical Character Recognition; Bert: Pre-trained language model (Bidirectional Encoder Representations from Transformers); HTML: HyperText Markup Language; SimHash: Locality-Sensitive Hashing; Block: code block; word2vec: word embedding technology; PEFT: Parameter-Efficient Fine-Tuning; Lora: A parameter efficient fine-tuning technology (Low-Rank Adaptation); NLL: Negative Log-Likelihood Loss.
[0028] Example 1, please refer to Figure 7 The present invention provides an open test question evaluation method based on a domain-based large model, comprising: The incremental training data set and instruction fine-tuning set are obtained through the subject corpus data, and the candidate answer set is obtained through text recognition. Based on the incremental training data set and instruction fine-tuning set, domain knowledge is injected and the intention of human instructions is aligned to obtain the domain model. Based on the candidate's answer set, a prompt template for the candidate's opinion mining task is constructed, the obtained domain model is fine-tuned, the fine-tuned domain model is obtained, and the fine-tuned domain model is optimized; Extract candidates' answering opinions from the optimized domain model and classify them into a standard answer opinion set and an innovative opinion set. Supplement the standard answer opinion set based on the innovative opinion set to obtain the final standard answer opinion set. Candidates' answers are scored based on the final set of standard answer viewpoints.
[0029] The present invention fully mines the semantic information of answers to open-ended questions and simplifies it into the form of a set of opinions, which not only facilitates the extraction and review of innovative opinions, but also provides more fine-grained automatic scoring at the semantic level. It has the advantages of sufficient semantic understanding and accurate automatic scoring, which makes the present invention have obvious advantages over other methods based on simple text feature pattern recognition.
[0030] Example 2, please refer to Figure 1 The present invention provides an open-ended test question evaluation method based on a domain-based large model, specifically comprising: 1. Data collection and processing process: The specific process of data collection and processing is as follows: (1) To obtain and process the examinee’s answer information, we first use OpenCV’s image segmentation technology to extract the image area containing text information. Then, we use the OCR handwriting recognition model to perform text recognition on the segmented answer image to obtain preliminary recognition results. Finally, we use the Bert-based text error correction model to correct the OCR recognition results and obtain accurate examinee answer information.
[0031] (2) In order to obtain and process corpus data in the subject area, we systematically collect multi-source data such as subject open source question banks, educational resource libraries, e-book libraries, and academic resource libraries through web crawler technology. First, we conduct preliminary data screening based on the preset subject keyword tag system; second, we effectively identify and remove non-text information such as HTML tags and hyperlinks through the characteristics of key noise data; then, we comprehensively use quantitative indicators such as perplexity, punctuation distribution characteristics, and sentence length to evaluate and filter the quality of text data; finally, we use the SimHash algorithm to calculate the semantic similarity between documents to achieve accurate deduplication of high-similarity document pairs.
[0032] (3) For document type data expected in the subject area, such as e-books, academic resources, and educational resource libraries, use them as incremental pre-training datasets; (4) For the question bank data and educational resource library in the corpus, the questions in the question bank data are used as instructions and the answers are used as responses to construct a single-round dialogue dataset. Then, the large model instruction template is used to generate multi-round dialogues based on the subject documents in the educational resource library. The large model instruction template is constructed as follows: “Based on the following document content, please generate a multi-turn dialogue. Role 1 and Role 2 can be students and instructors, readers and experts, etc. The goal is to help Role 1 understand the key concepts or ideas in the document and answer related questions.
[0033] Document Content: {Briefly summarize the document's topic, key messages, or main points}; Dialogue starts: Role 1 (student): "{generate a relevant question or doubt about the content of the document}".
[0034] The above steps are as follows Figure 2 As shown, the incremental training data set is obtained , instruction fine-tuning dataset ,in represents a single-round dialogue set, Represents a multi-round dialogue set and a candidate answer data set .
[0035] in, Data on different document types expected in the subject area; Questions and answers for a single round of conversation; Questions and answers for multi-round conversations.
[0036] 2. The process of building a large domain model: The construction of large domain models includes two processes: incremental pre-training and multi-task instruction fine-tuning.
[0037] The incremental pre-training process is as follows: (1) All the corpora of the incremental training dataset D are concatenated into a continuous text data stream. According to the preset block size, the text is divided into multiple fixed-length text blocks. The data in each block will be used for autoregressive training.
[0038] (2) We selected ChatGLM-6B, an open source conversational language model from the Knowledge Engineering Laboratory of Tsinghua University, as the base model and performed unsupervised autoregressive training on each text block, with the goal of predicting the character or word at the next position. The model generates the next character or word by learning the information of the current context, and calculates the loss function based on the difference between the generated word and the real word at the next position. Through multiple rounds of training, the loss function is continuously reduced to update the model parameter weights. The loss function is defined as follows:
[0039] Where NLL is the negative log-likelihood loss, Smoothed is the smoothed loss part. M is the loss normalization factor, N is the input text sequence length, and L is the total number of predicted labels. is the loss smoothing coefficient, is the one-hot encoded true label, For the model The predicted probability of a label, then the NLL loss and Smoothed loss are defined as follows:
[0040] The multi-task instruction fine-tuning process is as follows: (1) For a single-turn dialogue in the instruction fine-tuning set I, the starting dialogue is taken as input x, and the model directly predicts the response, which is defined as , t represents the tth token of the reply. The difference between the reply predicted by the model and the target reply is measured by the cross entropy loss function, which is defined as follows:
[0041] in Indicates that the current model is predicting the sequence number of the token in the target response. Represents the target response The total number, P represents the model's response to a given input x and the generated response Under the condition of , the probability of predicting the tth token is maximized by continuously reducing the loss function to maximize the probability of the model generating the target response.
[0042] (2) For multi-round dialogues in the instruction fine-tuning set I, the model needs to generate responses based on the historical dialogue information h and the current instruction x. Therefore, the loss function needs to consider the historical dialogue information. Compared with the loss function in the case of a single-round dialogue, the multi-round dialogue instruction fine-tuning loss function is defined as follows:
[0043] in Indicates that the current model is predicting the sequence number of the token in the target response. Represents the target response The total number, P represents the model's response to the given historical dialogue h, current instruction x, and generated partial response Under the condition of , the probability of predicting the tth token is maximized by continuously reducing the loss function to maximize the probability of the model generating the target response.
[0044] The flow chart of the construction process of the large model in the above fields is as follows Figure 3 shown.
[0045] 3. Few-sample prompt fine-tuning process: The prompt template for defining the opinion mining task is as follows: "You are a professional scoring assistant, responsible for extracting opinion information from the candidates' open-ended answers. Please read the following answers carefully and output the candidates' opinions and explanations in the format of "Opinion 1: Explanation 1; Viewpoint 2: Explanation 2; Viewpoint 3: Explanation 3".
[0046] Answer content: {Candidate's answer text} Output format: View 1: Explanation 1; View 2: Explanation 2; View 3: Explanation 3. ” By manually annotating a small amount (hundreds) of prompt data according to the prompt template, using the lora efficient parameter fine-tuning configuration in the PEFT package, and training on a small number of samples to adapt to the opinion mining task. The trained model can output the opinion information and explanation of the examinee's answer in the format defined by the template.
[0047] 4. Domain large model tuning process: The tuning of large domain models mainly assists in the generation of domain models by retrieving additional knowledge from external knowledge bases.
[0048] First, an external knowledge base is constructed and vectorized. Since there are problems of knowledge iteration and generation hallucination after domain model training, the knowledge base obtained in step 1 is updated and the external knowledge base is vectorized. Then, the knowledge base The text is divided into blocks , each text block is vectorized and stored in the vector database through the semantic representation model Bert to obtain the external knowledge vector library .
[0049] Then, according to the user's instructions Vector Library After the user enters the command, the user command is vectorized to obtain , calculate its similarity with each text vector in the vector library , extract the one with the highest similarity score Segment text vector, restore it to text data .
[0050] Finally, the domain model is combined with the search data for generation optimization. After the retrieved text data is combined with the user's instructions, it is provided to the domain model for subsequent generation.
[0051]
[0052] The flowchart of the large model tuning process in the above fields is as follows Figure 4 shown.
[0053] 5. Answer point mining process: Each answer a in the candidate's answer set A is filled into the opinion mining prompt template and then input into the model. Get the complete set of opinions corresponding to all n candidates' answers . Then use word2vec to calculate The feature vectors of all opinion words in the dataset are obtained, and the similarity between the two feature vectors is calculated using the cosine similarity calculation formula. For two opinion words with high similarity, the one with the larger total number of occurrences is retained to achieve de-duplication and merging. Then, the wrong words are removed by comparing with the standard dictionary to obtain the final initial opinion set. .
[0054] Answers to standard opinions pre-defined by subject experts as a standard opinion set , The rest are not in The opinions that appear in the article are regarded as a collection of unverified and possibly reasonable innovative opinions. .
[0055] against Views in , query contains viewpoint of candidates answered , and submit it to subject experts for review. If the experts agree that the view is the correct answer, then add the view to The final expanded viewpoint set The final standard answer viewpoint .
[0056] The flowchart of the process of mining opinions for the above open-ended questions is as follows: Figure 5 shown.
[0057] 6. Automatic scoring process: Standard opinion sets by subject matter experts The answer text a of the candidate to be evaluated is mined, and the obtained opinions are compared with the standard opinion set. For comparison, if the opinion word is directly contained in In the case of , the score of the opinion word is directly obtained. Otherwise, word2vec is used to represent the opinion word and then the score is obtained. The cosine similarity score is calculated for the words in . If the score is higher than the threshold p, it is judged as not scored, otherwise it is not scored.
[0058] The above automatic scoring process is as follows Figure 6 shown.
[0059] In yet another embodiment of the present invention, an open-ended test question evaluation system based on a domain-based large model is provided, which can be used to implement the above-mentioned open-ended test question evaluation method based on a domain-based large model. Specifically, the system includes: The domain model building module is used to obtain incremental training data sets and instruction fine-tuning sets through the subject corpus data, and obtain the candidate answer set through text recognition; based on the incremental training data sets and instruction fine-tuning sets, domain knowledge is injected and the intention of human instructions is aligned to obtain the domain model; The domain model fine-tuning module is used to construct a prompt template for the candidate's opinion mining task based on the candidate's answer set, perform prompt fine-tuning on the obtained domain model, obtain the fine-tuned domain model, and optimize the fine-tuned domain model; The module for obtaining the standard answer viewpoint set is used to extract the examinees' answer viewpoints from the optimized domain model, classify them into the standard answer viewpoint set and the innovative viewpoint set, and supplement the standard answer viewpoint set based on the innovative viewpoint set to obtain the final standard answer viewpoint set; The scoring output module is used to assign scores to candidates' answers based on the final set of standard answer viewpoints.
[0060] The division of modules in the embodiments of the present invention is schematic and is only a logical function division. There may be other division methods in actual implementation. In addition, each functional module in each embodiment of the present invention may be integrated into one processor, or may exist physically separately, or two or more modules may be integrated into one module. The above-mentioned integrated modules may be implemented in the form of hardware or in the form of software functional modules.
[0061] In another embodiment of the present invention, a computer device is provided, the computer device including a processor and a memory, the memory is used to store a computer program, the computer program includes program instructions, and the processor is used to execute the program instructions stored in the computer storage medium. The processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc., which are the computing core and control core of the terminal, which are suitable for implementing one or more instructions, and are specifically suitable for loading and executing one or more instructions in the computer storage medium to implement the corresponding method flow or corresponding function; the processor described in the embodiment of the present invention can be used for the operation of the open test question evaluation method based on the domain-based large model.
[0062] In another embodiment of the present invention, the present invention also provides a storage medium, specifically a computer-readable storage medium (Memory), which is a memory device in a computer device for storing programs and data. It is understandable that the computer-readable storage medium here can include both built-in storage media in a computer device and, of course, extended storage media supported by the computer device. The computer-readable storage medium provides a storage space, which stores the operating system of the terminal. In addition, one or more instructions suitable for being loaded and executed by a processor are also stored in the storage space, and these instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory. The processor can load and execute one or more instructions stored in the computer-readable storage medium to implement the corresponding steps of the open test question evaluation method based on a domain-based large model in the above embodiment.
[0063] It will be appreciated by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0064] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0065] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0066] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0067] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the relevant field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.
Claims
1. An open-ended test question evaluation method based on a domain-based large model, characterized in that: include: Obtain incremental training data sets and instruction fine-tuning sets through the subject corpus data, and obtain the candidate answer sets; Based on the incremental training data set and instruction fine-tuning set, domain knowledge is injected into the preset base model and instruction intent is aligned to obtain a domain model. Based on the candidate answer set, a prompt template for the candidate opinion mining task is constructed, and the obtained domain model is fine-tuned based on the prompt template of the candidate opinion mining task to obtain the fine-tuned domain model, and the fine-tuned domain model is tuned; Extract candidates' answering opinions from the optimized domain model and classify them into a standard answer opinion set and an innovative opinion set. Supplement the standard answer opinion set based on the innovative opinion set to obtain the final standard answer opinion set. Candidates' answers are scored based on the final set of standard answer viewpoints.
2. The open-ended test question evaluation method based on a domain-based large model according to claim 1 is characterized in that: The method of obtaining an incremental training data set and an instruction fine-tuning set through the subject corpus data, and obtaining a candidate's answer set, includes: Textbooks, books, papers, and question banks of the disciplines to which the open test questions belong are collected as corpus, low-quality data in the corpus is removed through a rule-based data filtering method, and document deduplication is performed to obtain an incremental training data set; an instruction fine-tuning set is generated through question bank question and answer data and a large model dialogue based on subject documents, and the instruction fine-tuning set includes a single-round dialogue set and a multi-round dialogue set; the candidate answer set is obtained by image segmentation and optical character recognition.
3. The open-ended test question evaluation method based on a domain-based large model according to claim 1 is characterized in that: Based on the incremental training data set and the instruction fine-tuning set, domain knowledge is injected into the preset base large model and the instruction intent is aligned to obtain a domain model, including: The base large model is incrementally pre-trained using the incremental pre-training set to inject domain knowledge, and then the instruction fine-tuning dataset is used to fine-tune the incremental pre-trained model to align the intention of human instructions, thereby obtaining a discipline-oriented domain large model.
4. The open-ended test question evaluation method based on a domain-based large model according to claim 1 is characterized in that: The method of constructing a prompt template for the candidate's viewpoint mining task based on the candidate's answer set, and fine-tuning the obtained domain model based on the prompt template for the candidate's viewpoint mining task to obtain the fine-tuned domain model includes: The prompt template format is defined based on the candidate's answer set, the prompt data is annotated according to the prompt template, and the low-rank adaptive efficient parameter fine-tuning domain model in the parameter efficient fine-tuning package is used to train the model on the samples to adapt to the opinion mining task. The trained domain model is used to output the opinion information and its explanation of the candidate's answer in the format defined by the template.
5. The open-ended test question evaluation method based on a domain-based large model according to claim 4 is characterized in that: The tuning operation of the fine-tuned domain model includes: Based on the fine-tuned domain model, access the external data storage system to perform retrieval optimization on the model output.
6. The open-ended test question evaluation method based on a domain-based large model according to claim 1 is characterized in that: The candidate's answering opinions are extracted from the optimized domain model and classified into a standard answer opinion set and an innovative opinion set, and the standard answer opinion set is supplemented based on the innovative opinion set to obtain the final standard answer opinion set, including: First, for each answer in the candidate's answer set, a prompt is filled in according to the opinion mining prompt template to obtain a prompt, and then the prompt is input into the optimized domain big model to extract the candidate's answer opinions to obtain the initial opinion set corresponding to all candidate's answers, and then all opinions in the initial opinion set are de-duplicated, merged and inferior to obtain the final opinion set. With reference to the preset recommended standard opinion set, opinions that do not appear in the recommended standard opinion set are constructed into an unverified innovative opinion set. Finally, the unverified innovative opinion set is reviewed and evaluated as correct innovative opinions. The correct innovative opinions are expanded to the standard answer opinion set as part of the subsequent marking criteria. The expanded opinion set is used as the final standard answer opinion set.
7. The open-ended test question evaluation method based on a domain-based large model according to claim 6 is characterized in that: The scoring of the examinee's answers based on the final set of standard answer viewpoints includes: Based on the final set of standard answer opinions, after opinion mining is performed on the candidates' answers to obtain the candidates' opinion information, the opinion words are compared with the final set of standard answer opinions. Points are awarded if the comparison is successful, and no points are awarded if the comparison is unsuccessful.
8. The open test question evaluation system based on the domain-based large model is characterized by: include: The domain model building module is used to obtain incremental training data sets and instruction fine-tuning sets through the subject corpus data, and obtain the candidate answer set through text recognition; based on the incremental training data sets and instruction fine-tuning sets, domain knowledge is injected and the intention of human instructions is aligned to obtain the domain model; The domain model fine-tuning module is used to construct a prompt template for the candidate's opinion mining task based on the candidate's answer set, perform prompt fine-tuning on the obtained domain model, obtain the fine-tuned domain model, and optimize the fine-tuned domain model; The module for obtaining the standard answer viewpoint set is used to extract the examinees' answer viewpoints from the optimized domain model, classify them into the standard answer viewpoint set and the innovative viewpoint set, and supplement the standard answer viewpoint set based on the innovative viewpoint set to obtain the final standard answer viewpoint set; The scoring output module is used to assign scores to candidates' answers based on the final set of standard answer viewpoints.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, it implements the open test question evaluation method based on the domain-based large model as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, an open-ended test question evaluation method based on a domain-based large model as described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
A Deep Learning-Based Intelligent Grading Method, System, and Storage Medium for Subjective Questions
CN108764074B
Paper marking method and system, electronic equipment and storage medium
CN117252739A
Domain question-answering system, domain question-answering construction method, electronic equipment and storage medium
CN117909466A
Method and system for evaluating effect of large language model in power field
CN118093371A
KR20220120253A