A method for marking subjective test papers based on large models

Through a large model based on Transformer architecture, the automatic marking of subjective test papers is solved in the existing technology, the problems of low marking efficiency, insufficient scoring fairness, limited semantic analysis ability and slow feedback speed are achieved, and the scoring effect of efficient, fair, in-depth semantic analysis and timely feedback are achieved.

CN118861522BActive Publication Date: 2025-06-03网才科技(广州)集团股份有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410908286.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-08
Publication Date
2025-06-03
Estimated Expiration
2044-07-08

AI Technical Summary

Technical Problem

The inefficient marking of existing subjective test papers, insufficient fairness and reliability of the scoring results, limited semantic analysis capabilities and slow feedback speed affect the effectiveness of teacher training.

Method used

A large model based on Transformer architecture is used to automatically mark subjective test papers. Through data cleaning, format uniformity, anonymization processing and the integration of multi-rate scorer weights, a high-precision marking model is built, and the two-way semantic matching calculation method is used for scoring.

Benefits of technology

It significantly improves the efficiency of marking papers and the fairness of scores, can conduct in-depth semantic analysis, accurately evaluate the content quality and logic of answers, and promptly feedback on the scoring results, improving the effectiveness of teacher training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118861522B_ABST
    Figure CN118861522B_ABST
Patent Text Reader

Abstract

The present invention provides a method for marking subjective test papers based on large models, belonging to the field of teacher examination evaluation. First, collect and preprocess subjective question answer data from the self-owned examination database to construct a high-quality data set; secondly, design a detailed scoring standard, manually annotate dimensions such as the accuracy, logic, and language expression of the answer content to generate weighted average score labels; then, select a marking and scoring large model based on the Transformer architecture, and utilize its powerful semantic representation ability to construct a marking model including modules such as multi-head attention; finally, design a two-way semantic matching calculation method to measure the similarity between the answer and the reference answer as the basis for scoring. Through the above steps, the model can achieve efficient, accurate, and objective automatic marking of subjective questions, improving the intelligent level of teacher evaluation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of teacher examination assessment, and particularly relates to a subjective question paper marking method based on a large model. Background Art

[0002] Currently, the marking of subjective question papers mainly relies on manual operation, which has problems such as low efficiency, high cost, and strong subjectivity in the marking criteria. With the rapid development of artificial intelligence technology, the market's demand for an efficient and accurate automated marking system is becoming increasingly urgent. Such a system can not only significantly improve the marking efficiency, reduce the labor cost, but also reduce the subjective deviation of scoring through a unified standard, realizing the intelligence and fairness of the subjective question marking work, thereby further promoting the modernization of teacher training.

[0003] However, the existing technologies have the following defects:

[0004] 1. Low marking efficiency: The current marking work requires a large amount of manpower input and is difficult to cope with the increasing marking demand, resulting in low efficiency and inability to quickly process a large number of papers.

[0005] 2. Subjectivity of manual scoring: Manual scoring is easily affected by subjective factors, and the scoring criteria are inconsistent, resulting in insufficient fairness and reliability of the scoring results.

[0006] 3. Limited semantic analysis ability: Traditional marking methods cannot conduct comprehensive and detailed semantic analysis on answers, so it is difficult to accurately evaluate the content quality and logic of answers.

[0007] 4. Slow feedback speed: The scoring results cannot be feedback in time, which affects the training effect and delays the timeliness and effectiveness of feedback in the teacher training process.

[0008] Therefore, this proposal presents a subjective question paper marking method based on a large model. This system realizes efficient, objective, and detailed marking effects, not only significantly improving the marking efficiency and fairness of scoring, but also being able to conduct in-depth semantic analysis, accurately evaluating the content quality and logic of answers, and feedback the scoring results in time, thereby improving the training effect of teachers. Summary of the Invention

[0009] In view of the above problems, the first aspect of the present invention proposes a subjective question paper marking method based on a large model, including the following processes:

[0010] Step 1, Data Set Collection and Production: By collecting data from the self-owned examination database, collect real, reliable, and diverse subjective question answer data, covering three types: short answer questions, essay questions, and case analysis questions. After data cleaning, format unification, anonymization processing, expert scoring, and multi-label fusion of the collected raw data, a high-quality subjective question answer data set is finally obtained;

[0011] Step 2, Establishment of a Subjective Question Paper Marking Model for the Marking and Scoring Large Model Based on the Transformer Architecture: Select the marking and scoring large model based on the Transformer architecture as the basic model. Utilizing its powerful semantic representation ability, transfer learning efficiency, and interpretability, an input layer, word embedding layer, position encoding layer, multi-head attention mechanism layer, feed-forward network layer, normalization layer, and prediction output layer are designed, and a bidirectional semantic matching calculation method is adopted to construct a high-precision subjective question marking model;

[0012] Step 3, Model Training of the Subjective Question Paper Marking Model for the Marking and Scoring Large Model Based on the Transformer Architecture: Use the constructed data set to fine-tune and train the marking and scoring large model based on the Transformer architecture. Adopt the cross-entropy loss function, combined with the Adam optimization algorithm and exponential decay learning rate scheduling, to improve the model convergence speed and generalization ability, and ensure the stability of the training process and the accuracy of scoring;

[0013] Step 4, System Deployment and Optimization: Deploy the trained marking and scoring model of the marking and scoring large model based on the Transformer architecture to the cloud server, provide an online scoring interface. At the same time, conduct statistical analysis on the scoring results, discover problems and optimize the model, and continuously improve the accuracy and robustness of the system to meet the requirements of actual application scenarios.

[0014] Preferably, the specific process of the said Step 1 includes the following:

[0015] S1, Raw Data Collection: Collect data from the self-owned examination database, collect real, reliable, and diverse subjective question answer data. Specifically, the collected data contains subjective questions and answer data of three main types:

[0016] 1) Short answer questions: Questions that require a simple answer in one or two sentences, such as "What is a genetic algorithm?";

[0017] 2) Essay questions: Questions that require the respondent to combine the knowledge learned and conduct multi-sentence or multi-paragraph discussions, such as "Please discuss the principle of using the genetic algorithm to solve the TSP problem";

[0018] 3) Case analysis questions: Present an actual case situation and require the respondent to comprehensively analyze and propose solutions, such as "Company XXX is facing xx dilemmas. How would you solve them?";

[0019] The initial data is obtained through the above steps , including the questions and their corresponding answers;

[0020] S2. Data preprocessing: The raw data collected may contain invalid, duplicate, or inconsistently formatted answers. To ensure data quality and model training effectiveness, the data needs to be cleaned, formatted uniformly, and anonymized. These steps help improve data validity and privacy protection:

[0021] 1) Data cleaning: Remove invalid answers (such as empty answers, "none", "?", and other meaningless answers) and duplicate answers. This operation can reduce noisy data and improve the quality of the dataset, enabling the training model to more accurately learn effective features;

[0022] 2) Format unification: Standardize the format, unify the answer format into a standard text format, remove garbled characters and special symbols, and convert all normalized answer texts into plain text in UTF-8 encoding. Each answer is stored as a text file;

[0023] 3) Anonymization: To ensure data privacy and security, remove any information that can identify personal identity from each answer file, such as names, student IDs, and barcodes. Each answer file will only retain: question content, answer content, and score information;

[0024] The preprocessed answer data is obtained through the above steps ;

[0025] S3. Data marking: To train a model that can accurately score subjective question answers, detailed label design and annotation of the answer data are required. By scoring the answer quality from three dimensions: content accuracy, logic, and language expression, the evaluation of answer quality can be refined, thereby improving the model's judgment ability. Specifically, score levels can be set for each dimension:

[0026] 1) Content accuracy (0 - 5 points): Score according to the accuracy of the answer content. 5 points means the content is completely correct, and 0 points means it has nothing to do with the question requirements;

[0027] 2) Logic (0 - 3 points): Judge whether the logic and organizational structure of the answer are reasonable. 3 points means logical and smooth writing, and 0 points means logical chaos and self - contradiction;

[0028] 3) Language expression (0 - 2 points): Evaluate whether the language is used accurately, concisely, and coherently. 2 points indicate fluent and relevant language, while 0 points indicate ungrammatical or irrelevant language;

[0029] According to the three set dimensions, invite experts in the field to independently score each pre - processed answer file. Each answer requires L (L >= 3) raters to score from the three dimensions respectively and give a comprehensive score. The score given by the th rater is denoted as ;

[0030] S4, multi - label fusion: For the same answer, there are certain differences in the scores given by different raters, which are caused by differences in subjective awareness. Therefore, a weighted average method is used to fuse the L scores, which can balance the differences between raters and obtain a more fair and objective scoring result:

[0031] ;

[0032] Among them, is the score after weighted average, is the weight corresponding to the th rater. This weight can be assigned according to factors such as the rater's experience, qualifications, and the consistency of historical scores to ensure that the fused score is more objective and fair. After the above steps, the score label is obtained;

[0033] After these steps, a large number of subjective question answer files are obtained. Each file contains the question content, reference answer, candidate answer data and the score data after weighted average , which will be used as the input of the training dataset for the large model and is denoted as .

[0034] Preferably, the establishment of the model in step 2 specifically includes the following process:

[0035] S1, large model selection: In the subjective question marking system, selecting a suitable large model is crucial for achieving high - precision automatic scoring. Weigh the advantages and disadvantages of different large models in terms of performance, interpretability, and efficiency. A reasonable model selection will lay a foundation for subsequent model architecture design and optimization, ensuring the efficient operation of the overall system;

[0036] As a bidirectional Transformer encoder model, the large-scale model for marking and scoring based on the Transformer architecture can make full use of context semantic information to deeply represent the answer text. This semantic representation of the context helps the large-scale model for marking and scoring based on the Transformer architecture accurately capture the semantic matching relationship between the answer and the reference answer, thus providing a reliable semantic understanding basis for subjective question scoring. In addition, the pre-training of the large-scale model for marking and scoring based on the Transformer architecture on a large-scale general corpus endows it with powerful transfer learning capabilities. When fine-tuning on a specific subjective question answer dataset, the large-scale model for marking and scoring based on the Transformer architecture can quickly converge and achieve excellent scoring performance, effectively saving time and computing resources. More importantly, the self-attention mechanism inside the large-scale model for marking and scoring based on the Transformer architecture gives the model good interpretability, and the attention weights can be analyzed to diagnose scoring errors and optimize the model, thereby continuously improving the accuracy of the subjective question marking system;

[0037] Generally speaking, the large-scale model for marking and scoring based on the Transformer architecture, with its powerful semantic representation ability, transfer learning efficiency and interpretability, makes it an ideal choice for building a high-precision subjective question marking system;

[0038] S2, Input layer: The large-scale model for marking and scoring based on the Transformer architecture allows the maximum length of the input sequence to be 512. To adapt to the input format of the large-scale model for marking and scoring based on the Transformer architecture, the obtained dataset The answer text data corresponding to the fixed questions in is sequentially segmented into token subsequences of length 512 For each subsequence Add the [CLS] mark at the beginning to identify the start of the sequence, and add the [SEP] mark at the end to identify the end of the sequence, thus obtaining the input representation:

[0039] x' i = [CLS]⊕ x i ⊕[SEP] ;

[0040] Among them, is the input token subsequence, is the concatenation operation;

[0041] S3, Word Embedding Layer: The large-scale marking and scoring model based on the Transformer architecture uses WordPiece embeddings to encode the input token subsequences. WordPiece is a subword embedding method based on corpus statistical learning. Compared with traditional Word2Vec, it can fully represent the character-level features within words. The word embedding layer maps each token subsequence into a dense vector space to obtain the word embedding matrix :

[0042] ;

[0043] Among them, is the WordPiece word embedding function;

[0044] S4, Position Encoding Layer: To enable the model to learn sequence information, the large-scale marking and scoring model based on the Transformer architecture adds position encoding after word embedding to encode the absolute position of each token subsequence in the sequence. This is because the Transformer in the large-scale marking and scoring model based on the Transformer architecture does not have a recursive or convolutional structure and cannot directly obtain sequence information. Therefore, it is necessary to explicitly add position information and use sine / cosine functions to encode positions:

[0045] ;

[0046] Among them, is the absolute position embedding matrix, is the position index, is the encoding dimension, is the sequence dimension;

[0047] Embed the obtained position encoding matrix into the word embedding matrix to obtain the final embedding representation Z:

[0048] ;

[0049] S5, Multi-Head Attention Mechanism Layer: The core of the large-scale marking and scoring model based on the Transformer architecture is the multi-head self-attention mechanism, which is used to capture the long-range dependence relationships between words in the sequence. First, self-attention calculation is performed:

[0050] ;

[0051] Among them, are the query sequence, key sequence, and value sequence respectively, is the activation function, is the dimension of the attention head, the The calculation results of h attention heads are expressed as:

[0052] ;

[0053] Among them, are the weight matrices of the linear transformations of query, key, and value respectively. Connecting h attention heads gives the multi-head attention mechanism:

[0054] ;

[0055] Among them, is the weight matrix of the linear transformation output by the multi-head attention mechanism;

[0056] S6, Feed-Forward Network Layer: In the encoder of the large-scale marking and scoring model based on the Transformer architecture, the feed-forward network layer is used to perform deeper feature extraction and transformation on the output of the multi-head attention, enabling the model to better capture complex patterns and semantic information in the input sequence. The feed-forward network includes two linear layers and a ReLU activation function:

[0057] ;

[0058] Among them, is the weight matrix, is the bias vector;

[0059] S7, Normalization Layer: Normalize the output obtained from the feed-forward network layer to avoid the problem of gradient vanishing or explosion and improve the training stability of the model;

[0060] S8, Encoder Stacking: Perform residual connection between the normalized output and the multi-head attention input. This helps the model capture local and global features, accelerate the training process, and by stacking M encoder blocks (including multi-head attention, feed-forward network, and normalization layer), the large-scale marking and scoring model based on the Transformer architecture can capture lexical and syntactic features from the bottom layer to extract complex semantic and context features at the high layer. After passing through M layers of encoders, the large-scale marking and scoring model based on the Transformer architecture obtains the final context representation ;

[0061] S9, Prediction Output Layer: For the final context representation obtained by the large-scale marking and scoring model based on the Transformer architecture, add a softmax layer to output the final prediction output :

[0062] ;

[0063] Among them, and The weight matrix and bias vector of the softmax layer respectively, and the final output is the total predicted score of the large-scale marking and scoring model based on the Transformer architecture for each answer, which includes three dimensions: accuracy score, logical score, and language expression score;

[0064] S10, Semantic matching calculation: To calculate the scoring score, it is necessary to measure the semantic similarity between the answer text and the reference answer. Therefore, the output of the large-scale marking and scoring model based on the Transformer architecture is used as the semantic representation, and a bidirectional matching method is designed. First, concatenate the answer and the reference answer into a sequence:

[0065] Input sequence = [CLS] + answer + [SEP] + reference answer + [SEP];

[0066] Input this sequence into the large-scale marking and scoring model based on the Transformer architecture, and extract the output vector marked with [CLS], denoted as as the semantic representation of the entire sequence. At the same time, for each token in the answer and the reference answer, extract its attention weight with as the importance of this token to the sequence semantics:

[0067] ;

[0068] ;

[0069] Among them, and are respectively the importance of the th token in the answer and the reference answer to the sequence semantics, and are respectively the th token in the answer and the reference answer, then the importance of this token to the final semantic representation is:

[0070] ;

[0071] ;

[0072] Sum the importance of all tokens in the answer and the reference answer respectively to obtain:

[0073] ;

[0074] ;

[0075] Among them, and are the semantic representation vectors of the answer and the reference answer;

[0076] Calculate their cosine similarity as the semantic similarity score:

[0077] ;

[0078] Among them, is the cosine similarity function. The closer this score is to 1, the more similar the semantics of the answer and the reference answer are.

[0079] Preferably, the training of the model in step 3 specifically includes the following process:

[0080] Based on the establishment of the subjective question paper marking model of the marking and scoring large model based on the Transformer architecture in the previous text, input the labeled data set into the model, and perform end-to-end training on the model in an autoregressive manner;

[0081] S1, Training data preparation: Divide the collected and labeled data into a training set and a validation set to prepare for the training and evaluation of the model. Divide it in a ratio of 8:2, use 80% of the data as the training set, and the remaining 20% as the validation set;

[0082] S2, Data preprocessing: Import the training data set in batches, and perform data cleaning, format unification, and anonymization processing on the input data of each batch to improve the generalization ability of the model;

[0083] S3, Fine-tuning the large model: Design an appropriate loss function, and compare the score mapped by the model with the weighted average score manually labeled, so as to construct the loss function :

[0084] ;

[0085] Among them, is the training batch size, and are the scores predicted by the model and the weighted average scores labeled in the th training batch;

[0086] S4, Loss calculation and gradient update;

[0087] ;

[0088] Among them, represents the parameter to be updated, including all weight matrices and bias vectors in the marking and scoring large model based on the Transformer architecture, is the learning rate, which controls the step size of each parameter update, Represents the loss function For the parameters The gradient, which represents the rate of change of the loss function with respect to the parameters. By multiplying by the learning rate , the parameters can be updated according to the direction and magnitude of the gradient , so that the loss function gradually decreases and the model can better adapt to the training data. The updated parameters Will be used in the training process of the next iteration;

[0089] S5, Layer Adaptive Learning Rate: Assign different learning rates to different layers, with the learning rate of the lower layers being less than that of the higher layers, so as to protect the basic features obtained by the lower layers from being destroyed;

[0090] S6, Training Termination Condition: According to the loss function value of the validation set, when the loss cannot be reduced for multiple consecutive epochs, the training process is terminated;

[0091] S7, Model Saving: Save the final model parameters at the end of training;

[0092] After the above model training process, a subjective question paper marking and scoring model with excellent performance is obtained, laying a foundation for the actual automatic marking system.

[0093] Preferably, the normalization layer described in step 2 is characterized in that the specific calculation formula is as follows:

[0094] ;

[0095] Wherein, And Are respectively the mean and standard deviation calculated on a single sample of the output FNN of the feedforward network, And Are learnable scaling and translation parameters.

[0096] Preferably, the loss calculation and gradient update described in step 3 are characterized in that the specific process includes the following:

[0097] Adopt the Adam optimization algorithm to perform gradient update on the parameters of the marking and scoring large model based on the Transformer architecture according to the loss function :

[0098] ;

[0099] Wherein, Represents the parameter to be updated, including all weight matrices and bias vectors in the marking and scoring large model based on the Transformer architecture, Is the learning rate, which controls the step size of each parameter update, Represents the loss function The gradient of the parameter , which represents the rate of change of the loss function with respect to the parameter, can be used to update the parameter by multiplying it with the learning rate . Based on the direction and magnitude of the gradient, the parameter can be updated so that the loss function gradually decreases, enabling the model to better adapt to the training data. The updated parameter will be used in the training process of the next iteration.

[0100] Preferably, the technical implementation of step 4 specifically includes the following process:

[0101] S1, Data collection: Obtain the subjective question answers of candidates from the in-house examination database and store them in a unified database for subsequent processing;

[0102] S2, Data preprocessing: Clean, unify the format, and anonymize the collected raw answer data, remove invalid, duplicate, or inconsistently formatted answers, convert the answer text to plain text in standard UTF-8 encoding, and remove any information that may identify personal identity;

[0103] S3, Data input: Tokenize and segment the preprocessed answer data according to the input requirements of the marking and scoring large model based on the Transformer architecture, and input the processed data into the marking and scoring large model based on the Transformer architecture in batches;

[0104] S4, Model processing: The marking and scoring large model based on the Transformer architecture matches the input answer data with the reference answers, captures context semantic features based on the self-attention mechanism, and performs deep representation learning on the answers using the encoder layer and the feed-forward network;

[0105] S5, Semantic matching and score calculation: Adopt a bidirectional semantic matching strategy, extract the importance of each Token in the answer and the reference answer to the entire sequence, sum them up with weights to obtain the semantic representation vector, and then calculate the semantic matching score through cosine similarity as the scoring result;

[0106] S6, Model output: The marking and scoring large model based on the Transformer architecture outputs the calculated scoring score through the softmax layer;

[0107] S7, Feedback generation and result storage: Analyze and refine the scoring result output by the model, generate a detailed feedback including three dimensions of content accuracy, logical coherence, and language expression, and store the scoring result and the feedback in the database together;

[0108] S8, Display and Query: Provide an API interface and a web-based front-end page, allowing users to query, display, and compare scoring results and give feedback;

[0109] S9, Continuous Optimization: Based on user feedback and scoring results, continuously optimize and iterate the large-scale model for marking and scoring subjective questions and the scoring system based on the Transformer architecture, continuously improving the accuracy and consistency of scoring and the overall performance of the system.

[0110] The second aspect of the present invention also provides a device for marking subjective question papers based on a large-scale model. The device includes at least one processor and at least one memory, and the processor is coupled to the memory; the memory stores a computer executable program of the large-scale model for marking subjective question papers constructed by the construction method as described in the first aspect; when the processor executes the computer executable program stored in the memory, the processor executes a method for marking subjective question papers based on a large-scale model.

[0111] Compared with the prior art, the innovation points and beneficial effects of the present invention are:

[0112] 1. The innovation points of the present invention are:

[0113] 1) Design a refined three-dimensional scoring standard (content accuracy 0-5 points, logic 0-3 points, language expression 0-2 points), and use the multi-rater weight fusion method for annotation. Each answer needs to be independently scored by more than 3 raters from three dimensions, and then the comprehensive score is obtained through weighted average fusion. The weights can be assigned according to factors such as the experience and qualifications of the raters, effectively balancing subjective differences and improving the objectivity and fine-grainedness of the labels;

[0114] 2) Select the large-scale model for marking and scoring based on the Transformer architecture as the core of the subjective question scoring model, which has three advantages: (1) The bidirectional Transformer encoder structure can capture context semantic information, laying a foundation for semantic matching; (2) Through large-scale corpus pre-training, it has strong transfer learning ability, can converge quickly and achieve good performance; (3) The internal attention mechanism gives good interpretability and can be diagnosed and optimized;

[0115] 3) Propose a semantic similarity calculation method based on bidirectional matching. Concatenate the answer and the reference answer and input them into the large-scale model for marking and scoring based on the Transformer architecture. Use the output vector of the [CLS] token as the semantic representation of the entire sequence, calculate the attention weight of each token to the semantic representation as its importance degree, and the sum can measure the semantic similarity between the answer and the reference answer, effectively improving the scoring accuracy.

[0116] 2. Based on the above innovation points, the following technical effects are achieved:

[0117] 1) The automation and high efficiency of subjective question scoring are realized, and the scoring efficiency is greatly improved. Subjective question answer data is collected and preprocessed from the self-owned examination database to construct a high-quality dataset, enabling the trained large-scale marking and scoring model based on the Transformer architecture to score a large number of subjective question answers efficiently and accurately, solving the pain point of low efficiency in manual marking and being applicable to various scenarios such as teacher skills examinations;

[0118] 2) The scoring results are more fair and objective, reducing the influence of human subjective differences. By refining the scoring criteria in three dimensions and using the method of multi-rater weight fusion for annotation, the subjective differences among raters are effectively balanced, making the final scoring labels more objective and fair, reflecting the true quality level of the answers;

[0119] 3) It has strong semantic understanding ability, improving the accuracy of scoring. The large-scale marking and scoring model based on the Transformer architecture, as the core of the scoring model, its bidirectional Transformer structure endows the model with the ability to capture context semantic information. Combined with the proposed bidirectional semantic matching calculation method, it can effectively measure the semantic association between the answer and the reference answer, making the scoring more accurate and reasonable;

[0120] 4) The model has good interpretability, which helps to continuously optimize the system performance. Based on the self-explanatory nature of the internal attention mechanism of the large-scale marking and scoring model based on the Transformer architecture, it can analyze the attention weights, diagnose the reasons for the differences between the scoring results and the actual situation, and optimize the model accordingly to continuously improve the level and performance of the overall marking system. BRIEF DESCRIPTION OF THE DRAWINGS

[0121] In order to more clearly illustrate the technical solutions of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following description is only one embodiment of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0122] Figure 1 A logic flowchart of a method for marking subjective question papers based on a large model.

[0123] Figure 2 The overall network structure diagram of the method of the present invention.

[0124] Figure 3 The overall block diagram of the method of the present invention.

[0125] Figure 4This is a schematic diagram of the subjective question paper marking device based on the large model in Embodiment 2 of the present invention. Detailed implementation manners

[0126] The present invention will be further described below in conjunction with specific embodiments.

[0127] Embodiment 1:

[0128] The present invention designs a subjective question paper marking method based on a large model. The overall logic of the present invention is as Figure 1 shown: First, collect and preprocess the subjective question answer data from the self-owned examination database to construct a high-quality data set; secondly, design a detailed scoring standard, manually annotate dimensions such as the accuracy, logic, and language expression of the answer content to generate a weighted average score label; then, select the marking and scoring large model Bidirectional Encoder Representations from Transformers (BERT) based on the Transformer architecture, and utilize its powerful semantic representation ability to construct a marking model including modules such as multi-head attention; finally, design a bidirectional semantic matching calculation method to measure the similarity between the answer and the reference answer as the basis for scoring.

[0129] 1. Obtain the subjective question answer data set, which specifically includes the following steps:

[0130] S1, Original data collection: Collecting data from the self-owned examination database can ensure the diversity and comprehensiveness of the data. Specifically, the collected data includes three main types of subjective questions and answer data:

[0131] 1) Short answer questions: Questions that require a simple answer in one or two sentences, such as "What is a genetic algorithm?";

[0132] 2) Essay questions: Questions that require the answerer to combine the knowledge learned and conduct multi-sentence or multi-paragraph discussions, such as "Please discuss the principle of using the genetic algorithm to solve the Traveling Salesman Problem (TSP)";

[0133] 3) Case analysis questions: Given an actual case situation, the answerer is required to comprehensively analyze and propose solutions, such as "Company XXX is facing xx dilemmas. How will you solve them?";

[0134] After the above steps, the initial data , including the questions and their corresponding answers, is obtained;

[0135] S2, Data preprocessing: The collected original data may contain invalid, duplicate, or inconsistently formatted answers. To ensure data quality and model training effects, the data needs to be cleaned, format unified, and anonymized. These steps help improve the effectiveness of the data and privacy protection:

[0136] 1) Data cleaning: Remove invalid answers (such as empty answers, meaningless answers like "none", "?", etc.) and duplicate answers. This operation can reduce noisy data and improve the quality of the dataset, enabling the training model to more accurately learn effective features;

[0137] 2) Format unification: Standardize the format, unify the answer format into a standard text format, remove garbled characters and special symbols, and convert all the standardized answer texts into plain text in UTF-8 encoding. Each answer is stored as a text file;

[0138] 3) Anonymization: To ensure data privacy and security, remove any personally identifiable information from each answer file, such as names, student IDs, and barcodes. Each answer file will only retain: question content, answer content, and score information;

[0139] The preprocessed answer data is obtained after the above steps ;

[0140] S3. Data marking: To train a model that can accurately score subjective question answers, detailed label design and annotation are required for the answer data. By scoring on three dimensions: content accuracy, logic, and language expression, the evaluation of answer quality can be refined, thereby enhancing the model's judgment ability. Specifically, score levels can be set for each dimension:

[0141] 1) Content accuracy (0 - 5 points): Score according to the accuracy of the answer content. 5 points means the content is completely correct, and 0 points means it has nothing to do with the question requirements;

[0142] 2) Logic (0 - 3 points): Judge whether the logic and organizational structure of the answer are reasonable. 3 points means rigorous logic and smooth writing, and 0 points means chaotic logic and self - contradiction;

[0143] 3) Language expression (0 - 2 points): Judge whether the language use is accurate, concise, and coherent. 2 points means smooth and relevant language, and 0 points means ungrammatical or irrelevant to the question;

[0144] According to the three set dimensions, invite experts and senior teachers in this field to independently score each preprocessed answer file. Each answer requires L (L >= 3) raters to score from the three dimensions respectively and give a comprehensive score. The score given by the th rater is denoted as ;

[0145] S4, Multi-label Fusion: For the same answer, there are certain discrepancies in the scores given by different graders, which are caused by differences in subjective awareness. Therefore, the L scores are fused using the weighted average method to balance the differences between graders and obtain a more fair and objective scoring result:

[0146] ;

[0147] Among them, is the score after weighted average, is the weight corresponding to the th grader. This weight can be assigned according to factors such as the grader's experience, qualifications, and the consistency of historical scores to ensure that the fused score is more objective and fair. After the above steps, the score label is obtained;

[0148] After these steps, a large number of subjective question answer files are obtained, each containing the question content, reference answer, candidate answer data and the score data after weighted average , which will be used as the input of the dataset for training the large model and is denoted as .

[0149] 2. Establishment of the grading model, which specifically includes the following processes:

[0150] S1, Selection of the large model: In the subjective question paper grading system, selecting a suitable large model is crucial for achieving high-precision automatic grading. Weighing the advantages and disadvantages of different large models in terms of performance, interpretability, and efficiency, a reasonable model selection will lay the foundation for subsequent model architecture design and optimization, ensuring the efficient operation of the overall system;

[0151] As a bidirectional Transformer encoder model, the Transformer-based scoring model can make full use of contextual semantic information to deeply represent the answer text. This contextual semantic representation helps the Transformer-based scoring model to accurately capture the semantic matching relationship between the answer and the reference answer, thereby providing a reliable semantic understanding basis for scoring subjective questions. In addition, the pre-training of the Transformer-based scoring model on a large-scale general corpus gives it a strong transfer learning ability. When fine-tuning a specific subjective question answer data set, the Transformer-based scoring model can quickly converge and achieve excellent scoring performance, effectively saving time and computing resources. More importantly, the self-attention mechanism inside the Transformer-based scoring model gives the model good interpretability, and can analyze the attention weight to diagnose scoring errors and optimize the model, thereby continuously improving the accuracy of the subjective question scoring system.

[0152] In general, the Transformer-based grading model is an ideal choice for building a high-precision subjective test grading system due to its powerful semantic representation ability, transfer learning efficiency, and interpretability.

[0153] S2, input layer: The Transformer-based scoring model allows a maximum input sequence length of 512. In order to adapt to the input format of the Transformer-based scoring model, the obtained dataset The answer text data corresponding to the fixed questions Divide into A token subsequence of length 512 , for each subsequence Add the [CLS] tag at the beginning to mark the start of the sequence and the [SEP] tag at the end to mark the end of the sequence, thus obtaining the input representation:

[0154] x' i = [CLS]⊕ x i ⊕[SEP] ;

[0155] in, is the input token subsequence, For splicing operation;

[0156] S3, Word Embedding Layer: The large-scale model for marking and scoring based on the Transformer architecture uses WordPiece embeddings to encode the input token subsequences. WordPiece is a subword embedding method based on corpus statistical learning. Compared with traditional Word2Vec, it can fully represent the character-level features within words. The word embedding layer maps each token subsequence into a dense vector space to obtain a word embedding matrix :

[0157] ;

[0158] Among them, is the WordPiece word embedding function;

[0159] S4, Position Encoding Layer: To enable the model to learn sequence information, the large-scale model for marking and scoring based on the Transformer architecture adds position encoding after word embedding to encode the absolute position of each token subsequence in the sequence. This is because the Transformer in the large-scale model for marking and scoring based on the Transformer architecture does not have a recursive or convolutional structure and cannot directly obtain sequence information. Therefore, it is necessary to explicitly add position information and use sine / cosine functions to encode positions:

[0160] ;

[0161] Among them, is the absolute position embedding matrix, is the position subscript, is the encoding dimension, is the sequence dimension;

[0162] Embed the obtained position encoding matrix into the word embedding matrix to obtain the final embedding representation Z:

[0163] ;

[0164] S5, Multi-Head Attention Mechanism Layer: The core of the large-scale model for marking and scoring based on the Transformer architecture is the multi-head self-attention mechanism, which is used to capture the long-range dependencies between words in the sequence. First, self-attention calculation is performed:

[0165] ;

[0166] Among them, are the query Query sequence, key Key sequence, and value Value sequence respectively, is the activation function, is the dimension of the attention head. The calculation result of the th attention head is expressed as:

[0167] ;

[0168] Among them, are the linear transformation weight matrices for query, key, and value respectively, and connecting h attention heads gives the multi-head attention mechanism:

[0169] ;

[0170] Among them, is the linear transformation weight matrix for the output of the multi-head attention mechanism;

[0171] S6, Feed-Forward Network Layer: In the encoder of the large-scale marking and scoring model based on the Transformer architecture, the feed-forward network layer is used to perform deeper feature extraction and transformation on the output of the multi-head attention, enabling the model to better capture complex patterns and semantic information in the input sequence. The feed-forward network includes two linear layers and a ReLU activation function:

[0172] ;

[0173] Among them, is the weight matrix, is the bias vector;

[0174] S7, Normalization Layer: Normalize the output obtained from the feed-forward network layer to avoid the problem of gradient vanishing or explosion and improve the training stability of the model:

[0175] ;

[0176] Among them, and are the mean and standard deviation calculated for the output FNN of the feed-forward network on a single sample respectively, and are the learnable scaling and translation parameters;

[0177] S8, Encoder Stacking: Make a residual connection between the normalized output and the multi-head attention input, which helps the model capture local and global features, accelerate the training process, and by stacking M encoder blocks (including multi-head attention, feed-forward network, and normalization layer), the large-scale marking and scoring model based on the Transformer architecture can capture lexical and syntactic features from the bottom layer to extract complex semantic and context features at the high layer. After passing through M layers of encoders, the large-scale marking and scoring model based on the Transformer architecture obtains the final context representation ;

[0178] S9, Prediction Output Layer: For the obtained large-scale marking and scoring model based on the Transformer architecture, obtain the final context representation , add a softmax layer to output the final prediction output :

[0179] ;

[0180] Among them, and are the weight matrix and bias vector of the softmax layer respectively, and the final output is the total predicted score of the large-scale marking and scoring model based on the Transformer architecture for each answer, which includes three dimensions: accuracy score, logicality score, and language expression score;

[0181] S10, Semantic Matching Calculation: To calculate the scoring score, it is necessary to measure the semantic similarity between the answer text and the reference answer. Therefore, the output of the large-scale marking and scoring model based on the Transformer architecture is used as the semantic representation, and a bidirectional matching method is designed. First, concatenate the answer and the reference answer into a sequence:

[0182] Input sequence = [CLS] + answer + [SEP] + reference answer + [SEP];

[0183] Input this sequence into the large-scale marking and scoring model based on the Transformer architecture, and extract the output vector marked with [CLS], denoted as , as the semantic representation of the entire sequence. At the same time, for each token in the answer and the reference answer, extract its attention weight with as the importance of this token to the sequence semantics:

[0184] ;

[0185] ;

[0186] Among them, and are the importance of the th token in the answer and the reference answer to the sequence semantics respectively, and are the th token in the answer and the reference answer respectively, then the importance of this token to the final semantic representation is:

[0187] ;

[0188] ;

[0189] Sum the importance of all tokens in the answer and the reference answer respectively to obtain:

[0190] ;

[0191] ;

[0192] Among them, and are the semantic representation vectors of the answer and the reference answer;

[0193] Calculate their cosine similarity as the semantic similarity score:

[0194] ;

[0195] Among them, is the cosine similarity function. The closer this score is to 1, the more similar the semantics of the answer and the reference answer are.

[0196] 3. End-to-end model training, specifically including the following steps:

[0197] Based on the establishment of the subjective question paper marking model of the marking and scoring large model based on the Transformer architecture in the previous text, input the labeled dataset into the model and perform end-to-end training on the model in an autoregressive manner;

[0198] S1, Training data preparation: Divide the collected and labeled data into a training set and a validation set to prepare for the training and evaluation of the model. Divide it in a ratio of 8:2, use 80% of the data as the training set, and the remaining 20% as the validation set;

[0199] S2, Data preprocessing: Import the training dataset in batches and perform data cleaning, format unification, and anonymization processing on the input data of each batch to improve the generalization ability of the model;

[0200] S3, Fine-tuning the large model: Design an appropriate loss function and compare the scored after the model mapping with the weighted average score manually labeled to construct the loss function :

[0201] ;

[0202] Among them, is the training batch size, and are the scores predicted by the model and the weighted average score labeled in the th training batch;

[0203] S4, Loss Calculation and Gradient Update: The Adam optimization algorithm is adopted to perform gradient update on the parameters of the large-scale marking and scoring model based on the Transformer architecture according to the loss function :

[0204] ;

[0205] Among them, represents the parameter to be updated, including all weight matrices and bias vectors in the large-scale marking and scoring model based on the Transformer architecture, is the learning rate, which controls the step size of each parameter update, represents the loss function with respect to the parameter . It represents the rate of change of the loss function with respect to the parameter. By multiplying by the learning rate , the parameter can be updated according to the direction and magnitude of the gradient, so that the loss function gradually decreases, the model can better adapt to the training data, and the updated parameter will be used in the next iteration of the training process;

[0206] S5, Layer Adaptive Learning Rate: Different learning rates are assigned to different layers, with the learning rate of the lower layers being smaller than that of the higher layers, so as to protect the basic features obtained by the lower layers from being destroyed;

[0207] S6, Training Termination Condition: According to the value of the loss function on the validation set, when the loss cannot be reduced for multiple consecutive epochs, the training process is terminated;

[0208] S7, Model Saving: Save the final model parameters at the end of training;

[0209] After the above model training process, a large-scale subjective question marking and scoring model with excellent performance is obtained, laying a foundation for the actual automatic marking system.

[0210] 4. Technical Implementation, specifically including the following steps:

[0211] S1, Data Collection: Obtain the subjective question answers of candidates from the self-owned examination database and store them in a unified database for subsequent processing;

[0212] S2, Data Preprocessing: Clean, unify the format and anonymize the collected original answer data, remove invalid, duplicate or inconsistently formatted answers, convert the answer text to pure text format with standard UTF-8 encoding, and remove any information that may identify personal identity;

[0213] S3, Data Input: Tokenize and segment the preprocessed answer data according to the input requirements of the large-scale marking and scoring model based on the Transformer architecture, and input the processed data into the large-scale marking and scoring model based on the Transformer architecture in batches;

[0214] S4, Model Processing: The large-scale marking and scoring model based on the Transformer architecture matches the input answer data with the reference answers, captures context semantic features based on the self-attention mechanism, and uses the encoder layer and the feed-forward network to perform deep representation learning on the answers;

[0215] S5, Semantic Matching and Score Calculation: Adopt a bidirectional semantic matching strategy, extract the importance of each Token in the answer and the reference answer to the entire sequence, sum them up with weights to obtain a semantic representation vector, and then calculate the semantic matching score through cosine similarity as the marking result;

[0216] S6, Model Output: The large-scale marking and scoring model based on the Transformer architecture outputs the calculated marking score through the softmax layer;

[0217] S7, Feedback Generation and Result Storage: Analyze and refine the marking results output by the model, generate detailed feedback including three dimensions of content accuracy, logical coherence, and language expression, and store the marking results and feedback in the database together;

[0218] S8, Display and Query: Provide an API interface and a Web-based front-end page, allowing users to query, display, and compare the marking results and give feedback;

[0219] S9, Continuous Optimization: Based on user feedback and marking results, continuously optimize and iterate the large-scale marking and scoring model based on the Transformer architecture and the marking system, and continuously improve the accuracy, consistency of marking, and the overall performance of the system.

[0220] The following example results are shown for this process:

[0221] To verify a subjective question paper marking method based on a large model proposed by the present invention, the accuracy of the marking results for subjective question answers is shown, and the results are shown in Table 1. This embodiment also provides a comparison of the results of the method of the present invention with those of currently well-known models GPT-3.5, T5, and ELECTRA, and the evaluation indicators are mAP, computational complexity (FLOPs), and the number of parameters (Parameters).

[0222] Table 1 Quantitative Analysis of the Test Dataset on Different Models

[0223] Model mAP Parameters FLOPs GPT-3.5 82.6 28.14M 78.22 T5 80.0 25.24M 74.16 ELECTRA 85.4 36.51M 103.50 The method of the present invention 90.5 3M 10.55

[0224] As shown in Table 1, the method proposed by the present invention achieves the maximum value in mAP and the minimum values in Parameters and FLOPs, indicating that this model has high accuracy, fast processing speed, and small memory occupancy.

[0225] Example 2:

[0226] As Figure 4 shown, the present invention also provides a subjective question test paper marking device based on a large model. The device includes at least one processor and at least one memory, and also includes a communication interface and an internal bus; a computer execution program is stored in the memory; a computer execution program of a subjective question test paper marking model based on a large model constructed by the construction method described in Example 1 is stored in the memory; when the processor executes the computer execution program stored in the memory, the processor can execute a subjective question test paper marking method based on a large model. The internal bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, the buses in the drawings of this application do not limit to only one bus or one type of bus. The memory may include high-speed RAM memory, and may also include non-volatile storage NVM, such as at least one disk memory, and can also be a USB flash drive, a mobile hard disk, a read-only memory, a disk, or an optical disc, etc.

[0227] The device can be provided as a terminal, a server, or other forms of devices.

[0228] Figure 4 is a block diagram of a device shown exemplarily. The device may include one or more of the following components: a processing component, a memory, a power component, a multimedia component, an audio component, an input / output (I / O) interface, a sensor component, and a communication component. The processing component generally controls the overall operation of the electronic device, such as operations associated with display, telephone call, data communication, camera operation, and recording operation. The processing component may include one or more processors to execute instructions to complete all or part of the steps of the above method. In addition, the processing component may include one or more modules to facilitate the interaction between the processing component and other components. For example, the processing component may include a multimedia module to facilitate the interaction between the multimedia component and the processing component.

[0229] The memory is configured to store various types of data to support the operation of the electronic device. Examples of such data include instructions for any application or method operating on the electronic device, contact data, phone book data, messages, pictures, videos, and the like. The memory can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.

[0230] The communication component is configured to facilitate communication between the electronic device and other devices in a wired or wireless manner. The electronic device can access a wireless network based on communication standards, such as WiFi, 2G, or 3G, or a combination thereof. In an exemplary embodiment, the communication component receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0231] In an exemplary embodiment, the electronic device can be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above-described method.

[0232] The above are only the preferred embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

[0233] Although the specific implementation manners of the present invention have been described above, they are not limitations on the protection scope of the present invention. Those skilled in the art should understand that various modifications or deformations that can be made without creative efforts based on the technical solutions of the present invention are still within the protection scope of the present invention.

Claims

1. A method for marking subjective test papers based on a large model, characterized in that: The process includes: Step 1, data set collection and production: collect data from the own test database, collect authentic, reliable and diverse subjective question answer data, covering three types of questions: short answer questions, essay questions and case analysis questions. The collected raw data is cleaned, formatted and anonymized, and expert scoring and multi-label fusion are performed to finally obtain the subjective question answer data set; Step 2, the establishment of a subjective test paper grading model based on the Transformer architecture: the Transformer architecture-based grading model is selected as the basic model, and the input layer, word embedding layer, position encoding layer, multi-head attention mechanism layer, feedforward network layer, normalization layer and prediction output layer are designed. The two-way semantic matching calculation method is adopted. The specific steps are as follows: S1, large model selection: a large model for marking and scoring based on the Transformer architecture; S2, input layer: The Transformer-based scoring model allows a maximum input sequence length of 512. The answer text data corresponding to the fixed questions Divide into A token subsequence of length 512 , for each subsequence Add the [CLS] tag at the beginning to mark the start of the sequence and the [SEP] tag at the end to mark the end of the sequence, thus obtaining the input representation: ; in, is the input token subsequence, For splicing operation; S3, word embedding layer: The Transformer-based scoring model uses WordPiece embedding to encode the input token subsequence; S4, position encoding layer: The Transformer-based scoring model adds position encoding after word embedding, encoding the absolute position of each token subsequence in the sequence, and using sine / cosine functions to encode the position: ; in, Absolute position embedding matrix, is the position subscript, is the encoding dimension, is the sequence dimension; Embed the obtained position encoding matrix into the word embedding matrix to obtain the final embedding representation Z: ; S5, multi-head attention mechanism layer: used to capture the long-range dependencies between words in the sequence; S6, feedforward network layer: The feedforward network consists of two linear layers and a ReLU activation function: ; in, is the weight matrix, is the bias vector; S7, normalization layer: normalize the output of the feedforward network layer; S8, encoder stacking: The normalized output is residually connected to the multi-head attention input, and M layers of encoder blocks are stacked. A single encoder block contains multi-head attention, feedforward network and normalization layer. The Transformer-based scoring model captures vocabulary and grammatical features from the bottom layer to extract complex semantic and contextual features from the high layer. After M layers of encoders, the Transformer-based scoring model obtains the final contextual representation. ; S9, prediction output layer: for , add a softmax layer to output the final prediction output : ; in, and They are the weight matrix and bias vector of the softmax layer, and the final output This is the total score predicted by the Transformer-based grading model for each answer, which includes three dimensions: accuracy score, logic score, and language expression score. S10, semantic matching calculation: The output of the Transformer-based scoring model is used as the semantic representation, and a two-way matching method is designed. First, the answer and the reference answer are concatenated into a sequence; the sequence is input into the Transformer-based scoring model, and the output vector of the [CLS] tag is extracted, which is recorded as , as the semantic representation of the entire sequence, and at the same time, for each token of the answer and the reference answer, extract its The attention weight is the importance of the token to the sequence semantics: ; ; in, and The answer and reference answer are The importance of each token to the sequence semantics. and The first two in the answer and the reference answer are token, the importance of the token to the final semantic representation is: ; ; Sum the importance of all tokens in the answer and reference answer, and we get: ; ; in, and Semantic representation vectors for answers and reference answers; Calculate its cosine similarity as the semantic similarity score: ; in, is the cosine similarity function. The closer the score is to 1, the more similar the answer is to the reference answer in semantics. Step 3: Model training of the Transformer-based grading model: Use the constructed dataset to fine-tune the Transformer-based grading model, using the cross entropy loss function, combined with the Adam optimization algorithm and exponential decay learning rate scheduling; Step 4: System deployment and optimization: Deploy the trained Transformer-based grading model to the cloud server, provide an online grading interface, and perform statistical analysis on the grading results and optimize the model.

2. A method for marking subjective test papers based on a large model according to claim 1, characterized in that: The step 1 is specifically The process includes: S1, original data collection: collect data from our own test database, collect authentic, reliable and diverse subjective question answer data. Specifically, the collected data includes three main types of subjective questions and answer data: 1) Short answer questions: questions that require a simple answer in one sentence; 2) Essay questions: questions that require the respondent to combine the knowledge learned and make a multi-sentence or even multi-paragraph discussion; 3) Case analysis questions: Given a real case situation, the answerer is required to conduct a comprehensive analysis and propose a solution; After the above steps, the initial data is obtained , including questions and their corresponding answers; S2, data preprocessing: collected raw data Contains invalid, repeated or inconsistent answers, and the data needs to be cleaned, formatted and anonymized: 1) Data cleaning: remove invalid and duplicate answers, reduce noise data, and improve the quality of the data set; 2) Format unification: Format standardization, unify the answer format into a standard text format, remove garbled characters and special symbols, and convert all standardized answer texts into UTF-8 encoded plain text format, and store each answer as a text file; 3) Anonymization: Any personally identifiable information, including name, student ID number and barcode, will be removed from each answer file. Each answer file will only retain the following information: question content, answer content and score information; After the above steps, we can get the preprocessed answer data ; S3, Data Labeling: Label the answer data and mark them, including the ratings of the three dimensions of content accuracy, logic, and language expression. Specifically, set the score level for each dimension: 1) Content accuracy: The answer is scored based on its accuracy, with 5 points representing complete correctness and 0 points representing complete irrelevance to the question. 2) Logic: The logic and organization of the answer will be evaluated and scored, with 3 points for rigorous logic and smooth writing, and 0 points for confusing logic and inconsistency; 3) Language expression: score based on the accuracy, simplicity and coherence of the language, 2 points for fluent language and relevant to the topic, 0 points for incomprehensible language or irrelevant to the topic; According to the three dimensions set, experts in the field are invited to independently score each pre-processed answer document. Each answer requires L raters to score from the three dimensions and give a comprehensive score. The score of each rater is recorded as ; S4, multi-label fusion: L scores are fused using weighted average: ; in, is the weighted average score, For the The weight corresponding to each rater is determined based on the rater's experience, qualifications, and consistency of historical ratings. After the above steps, the score label is obtained. ; After steps S1 to S4, the subjective question answer files are obtained, each of which contains the question content, reference answers, and candidate answer data. And the weighted average score data , as the data set input for training the large model, denoted as .

3. A method for marking subjective test papers based on a large model according to claim 1, characterized in that: The multi-head attention mechanism layer in step 2 specifically includes the following processes: First, perform self-attention calculation: ; in, They are query sequence, key sequence and value sequence respectively. is the activation function, is the dimension of the attention head, The calculation result of the attention head is expressed as: ; in, are the linear transformation weight matrices for query, key, and value respectively; Secondly, connect h attention heads to get a multi-head attention mechanism: ; in, It is the linear transformation weight matrix output by the multi-head attention mechanism.

4. A method for marking subjective test papers based on a large model according to claim 1, characterized in that: The step 3 specifically includes the following process: Model training of the large-scale grading model based on the Transformer architecture. The labeled data set is input into the model, and the model is trained end-to-end using an autoregressive approach. S1, training data preparation: the collected and labeled data are divided into training set and validation set to prepare for model training and evaluation. The ratio of 8:2 is adopted, with 80% of the data used as training set and the remaining 20% ​​as validation set; S2, data preprocessing: import the training data set in batches, and perform data cleaning, format unification and anonymization on the input data of each batch; S3, fine-tuning the large model: designing a suitable loss function and mapping the model to the score Weighted average score compared to manual annotations Compare and build the loss function : ; in, is the training batch size, and For the The weighted average score of the model predictions and annotations in the training batches; S4, loss calculation and gradient update; S5, layer adaptive learning rate: different learning rates are assigned to different layers, and the learning rate of the lower layer is smaller than that of the higher layer; S6, training termination condition: according to the loss function value of the validation set, when the loss cannot be reduced for multiple consecutive epochs, the training process is terminated; S7, model saving: save the final model parameters at the end of training.

5. A method for marking subjective test papers based on a large model according to claim 1, characterized in that: The specific calculation formula of the normalization layer described in step 2 is as follows: ; in, and are the mean and standard deviation of the output of the feedforward network FNN calculated on a single sample, and are learnable scaling and translation parameters.

6. A method for marking subjective test papers based on a large model according to claim 1, characterized in that: Step 3: Fine-tune the Transformer-based scoring model The process includes: Adopt Adam optimization algorithm, according to loss function Gradient update of the parameters of the Transformer-based grading model: ; in, Represents the parameters to be updated, including all weight matrices and bias vectors in the Transformer-based grading model. is the learning rate, which controls the step size of each parameter update. Represents the loss function Parameters The gradient of the loss function represents the rate of change of the parameter, which is multiplied by the learning rate , update the parameters according to the direction and magnitude of the gradient , so that the loss function gradually decreases, and the updated parameters It will be used in the next iteration of training process.

7. A method for marking subjective test papers based on a large model according to claim 1, characterized in that: The step 4 specifically includes the following process: S1, data collection: obtain the examinees' answers to subjective questions from the own examination database and store them in a unified database; S2, data preprocessing: clean, unify and anonymize the collected raw answer data, remove invalid, repeated or inconsistent answers, convert the answer text into a plain text format with standard UTF-8 encoding, and remove any information that may identify an individual; S3, data input: the pre-processed answer data is segmented and divided according to the input requirements of the Transformer-based scoring model, and the processed data is input into the Transformer-based scoring model in batches; S4, model processing: The Transformer-based grading model matches the input answer data with the reference answer, captures contextual semantic features based on the self-attention mechanism, and uses the encoder layer and feedforward network to learn deep representations of the answer; S5, semantic matching and score calculation: adopt a two-way semantic matching strategy, extract the importance of each token in the answer and reference answer to the entire sequence, sum them up by weight to get the semantic representation vector, and then calculate the semantic matching score by cosine similarity as the scoring result; S6, model output: The large model for grading based on the Transformer architecture outputs the calculated scores through the softmax layer; S7, feedback generation and result storage: Analyze the scoring results output by the model, generate feedback including three dimensions: content accuracy, logical coherence, and language expression, and store the scoring results and feedback in the database; S8, Display and Query: Provides an API interface and a web-based front-end page to allow users to query, display and compare scoring results and give feedback; S9, continuous optimization: Based on user feedback and scoring results, continuously optimize and iterate the Transformer-based grading model and scoring system.

8. A subjective test paper marking device based on a large model, characterized in that: The device includes at least one processor and at least one memory, the processor and the memory are coupled; the memory stores a computer execution program of the subjective test paper grading method based on a large model as described in any one of claims 1 to 7; when the processor executes the computer execution program stored in the memory, the processor executes the subjective test paper grading.

Citation Information

Patent Citations

  • Automatic student answer scoring method for English examination translation questions

    CN112085985A

  • Automatic subjective question marking neural network model with concept enhanced representation and unidirectional attention implication

    CN113011196A