Question Scoring Method, System, Electronic Device and Storage Medium

By constructing a word frequency table and a domain adaptability pre-training model, the desensitized test question data is subjected to pseudo-text conversion and joint scoring, which solves the problem of low score accuracy of test questions in the existing technology, and achieves higher score accuracy and reliability.

CN119416771BActive Publication Date: 2025-06-27SHANDONG SAHNDA OUMASOFT CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510026946.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-06-27
Estimated Expiration
2045-01-08

AI Technical Summary

Technical Problem

It is difficult for the prior art to accurately score the test data after desensitization, traditional methods are difficult to apply effectively, and common language models are difficult to achieve ideal results in the subject field.

Method used

A word frequency table is constructed, the desensitized test questions data are converted into pseudo-text through word frequency statistics, and a joint scoring model is constructed based on the pedestal model after field adaptability pre-trained, and the classification and regression tasks are scored.

Benefits of technology

It improves the accuracy and reliability of test questions, can better consider multiple scoring dimensions in a comprehensive manner, and is suitable for test questions after desensitization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119416771B_ABST
    Figure CN119416771B_ABST
Patent Text Reader

Abstract

The present invention provides a method, system, electronic device and storage medium for scoring examination questions, belonging to the field of examination evaluation. The method includes: constructing a word frequency table for an examination subject, where the word frequency table includes the frequencies of various words in the subject; performing word frequency statistics on the desensitized examination question data, and using the word frequency table to convert the desensitized examination question data into the corresponding words in the word frequency table according to the frequency to generate a pseudo-text of the examination questions; scoring the generated pseudo-text of the examination questions based on a pre-constructed examination question scoring model to obtain an examination question scoring result; wherein the examination question scoring model is a joint scoring model based on a classification task and a regression task constructed on a base model after domain adaptation pre-training. By adopting the methods of pseudo-text conversion and domain adaptation pre-training, the understanding of the base model for domain knowledge is improved; moreover, a joint scoring model is introduced to perform machine scoring cross-verification, improving the accuracy and reliability of scoring.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of examination and assessment, and particularly to a method, system, electronic device and storage medium for scoring test questions. Background Art

[0002] In the current technical field of examination and assessment, with the development of educational informatization, the automatic scoring of test questions has become increasingly important. However, traditional scoring methods often rely on manually set rules or simple statistical methods, and it is difficult to accurately reflect the true difficulty of test questions and the actual level of candidates. In addition, in order to protect the privacy of candidates and prevent cheating, test question data usually needs to be desensitized, which further increases the difficulty of automatic scoring.

[0003] Desensitized data means replacing or encrypting personal information, sensitive information, and even all answer contents in test questions to protect data privacy. However, the desensitized data often loses some original information, making it difficult to effectively apply traditional rule-based or statistical scoring methods. Therefore, how to intelligently score desensitized test question data has become an urgent problem to be solved.

[0004] However, due to the fact that test question data usually has specific subject field characteristics, it is often difficult to achieve ideal scoring results by directly using general language models. In addition, the scoring of test questions usually involves multiple dimensions, such as difficulty, knowledge point coverage, problem-solving steps, etc., and these dimensions are difficult to accurately evaluate through a single scoring model. Therefore, it is necessary to construct a scoring model that can comprehensively consider multiple dimensions to improve the accuracy and reliability of scoring. Summary of the Invention

[0005] The purpose of the embodiments of the present invention is to provide a method, system, electronic device and storage medium for scoring test questions, which are used to fully or at least partially solve the problem of low accuracy in scoring desensitized test question data existing in the above-mentioned prior art.

[0006] To achieve the above purpose, the embodiments of the present invention provide a method for scoring test questions, including:

[0007] Construct a word frequency table for the examination subject, where the word frequency table includes the frequencies of each vocabulary in the subject appearing in the subject;

[0008] Perform word frequency statistics on the desensitized test question data, and use the word frequency table to convert the desensitized test question data into the corresponding vocabulary in the word frequency table according to the frequency, and generate a test question pseudo-text;

[0009] Score the generated pseudo-test questions based on a pre-constructed test question scoring model to obtain the test question scoring results; wherein, the test question scoring model is a joint scoring model based on classification tasks and regression tasks constructed on a base model after domain adaptation pre-training.

[0010] Optionally, the desensitized test question data is characterized as data obtained by replacing or encrypting personal information, sensitive information, and all answer contents in the test questions.

[0011] The domain adaptation pre-training process includes:

[0012] Merge the test subject corpus and the pseudo-test questions to form a dataset for domain adaptation pre-training.

[0013] Select the Chinese pre-trained language model PTM as the starting model for domain adaptation pre-training, randomly select some texts in the dataset for occlusion processing, and perform domain adaptation pre-training on the Chinese pre-trained language model PTM through the task of predicting the occluded words.

[0014] Furthermore, the process of constructing a joint scoring model based on classification tasks and regression tasks on the base model after domain adaptation pre-training includes:

[0015] Utilize the feature extraction base model after domain adaptation pre-training for extract features from the input pseudo-text sequence to obtain text sequence features :

[0016]

[0017] Perform average pooling and max pooling on the text sequence features and concatenate the features after average pooling to obtain text features :

[0018]

[0019] wherein, represents average pooling, represents max pooling, represents the concatenation operation;

[0020] The classification branch uses

[0021]

[0022] to predict the score, convert the output to the probability of each category, represents the probability value of each score category, is the classification fully connected layer, and its input dimension is , where is the dimension of the model semantic embedding, and the output dimension is the number of score types ;

[0023] The regression branch uses

[0024]

[0025] to predict the score, is the regression result, is the regression fully connected layer, and its input dimension is , and the output dimension is 1.

[0026] Furthermore, before scoring the generated pseudo test questions based on the pre-constructed test question scoring model to obtain the test question scoring result, the test question scoring method further includes fine-tuning and training the joint scoring model based on the manual scoring:

[0027] Construct a training data set, where the training data set is a triple data including pseudo test questions, category data converted from manual scores, and normalized manual scoring data , where is the pseudo test question, represents the category data converted from the manual score, represents the normalized manual scoring data:

[0028]

[0029] Among them, represents the maximum score, represents the average score, represents the manual expert scoring.

[0030] Using the triple data, train the joint scoring model based on the training method of random mini-batches. Among them, the cross-entropy is used to calculate the training loss for the classification task, and the root mean square error is used to calculate the training loss for the regression task. The Adam optimizer is used to update the parameters of the joint scoring model.

[0031] Furthermore, scoring the generated pseudo test questions based on the pre-constructed test question scoring model to obtain the test question scoring result, including:

[0032] Use the fine-tuned test question scoring model to score, and the model outputs and ;

[0033] Select The score corresponding to the category with the highest probability as the scoring result of the classification branch:

[0034]

[0035] In the formula, represents the position index for obtaining the maximum value;

[0036] Perform an inverse transformation on as the scoring result of the regression branch:

[0037] ,

[0038] In the formula, represents the rounding function;

[0039] Calculate the score difference between the classification task and the regression task. If the score difference is within the specified threshold, the scoring is valid, and the final scoring result is the average of the two:

[0040]

[0041] At the same time, select the maximum probability as the confidence level of the scoring result;

[0042] If the score difference is greater than the specified threshold, the machine scoring is invalid.

[0043] In a second aspect, the present invention also provides a test question scoring system, including:

[0044] A construction unit for constructing a word frequency table for an examination subject, where the word frequency table includes the frequencies of each vocabulary in the subject appearing in the subject;

[0045] A generation unit for performing word frequency statistics on the desensitized test question data, and using the word frequency table to convert the desensitized test question data into the corresponding vocabulary in the word frequency table according to the frequency, and generating a test question pseudo-text;

[0046] A scoring unit for scoring the generated test question pseudo-text based on a pre-constructed test question scoring model to obtain a test question scoring result; wherein, the test question scoring model is a joint scoring model based on classification tasks and regression tasks constructed on a base model after domain adaptation pre-training.

[0047] In a third aspect, the present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the above-mentioned test question scoring method are implemented.

[0048] In a fourth aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned test question scoring method are implemented.

[0049] Through the above technical solution, the method improves the understanding of the base model for domain knowledge by adopting the method of pseudo-text conversion and domain-adaptive pre-training; moreover, a joint scoring model is introduced to perform machine scoring cross-verification, so as to improve the accuracy and reliability of scoring.

[0050] Other features and advantages of the embodiments of the present invention will be described in detail in the subsequent specific implementation part. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] The drawings are used to provide a further understanding of the embodiments of the present invention, and constitute a part of the specification. Together with the following specific implementation manners, they are used to explain the embodiments of the present invention, but do not constitute a limitation to the embodiments of the present invention. In the drawings:

[0052] Figure 1 is an implementation flowchart of a test question scoring method provided by an embodiment of the present invention;

[0053] Figure 2 is a structural diagram of a joint scoring model provided by an embodiment of the present invention;

[0054] Figure 3 is a schematic diagram of the process of scoring the pseudo-text of test questions by using the fine-tuned model provided by an embodiment of the present invention;

[0055] Figure 4 is a schematic structural diagram of a test question scoring system provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0056] The following will describe in detail the specific implementation manners of the embodiments of the present invention with reference to the drawings. It should be understood that the specific implementation manners described herein are only used to illustrate and explain the embodiments of the present invention, and are not used to limit the embodiments of the present invention.

[0057] Referring to Figure 1 as shown, it is an implementation flowchart of a test question scoring method provided by an embodiment of the present invention, including the following execution steps:

[0058] Step 100: Construct a word frequency table for the examination subject, where the word frequency table includes the frequencies of each vocabulary in the subject appearing in the subject.

[0059] Step 101: Perform word frequency statistics on the desensitized test question data, and use the word frequency table to convert the desensitized test question data into the corresponding vocabulary in the word frequency table according to the frequency, and generate pseudo-text of the test questions.

[0060] It should be understood that the desensitized test question data is characterized by the data obtained by replacing or encrypting the personal information, sensitive information, and all answer contents in the test questions.

[0061] Exemplarily, for a short-answer question in a certain professional qualification exam, the full score of this question is 20 points. There are a total of 11,651 candidate answers and the corresponding scores given by professional graders. The scores of the candidate answers are shown in Table 1 below. The content is only for illustrative purposes and has no actual meaning.

[0062] Table 1: Desensitized data

[0063]

[0064] Step 102: Score the generated pseudo-text of the test questions based on a pre-constructed test question scoring model to obtain the test question scoring result; wherein, the test question scoring model is a joint scoring model based on classification tasks and regression tasks constructed on a base model after domain adaptation pre-training.

[0065] Among them, the joint scoring model includes a classification branch and a regression branch. The classification branch is used to predict the score category, and the regression branch is used to predict the specific score value, and the joint scoring model is fine-tuned based on the manual scoring data.

[0066] Specifically, the domain adaptation pre-training process includes: merging the test subject corpus and the pseudo-text of the test questions to form a dataset for domain adaptation pre-training; selecting the Chinese pre-trained language model PTM as the starting model for domain adaptation pre-training, randomly selecting some texts in the dataset for occlusion processing, and performing domain adaptation pre-training on the Chinese pre-trained language model PTM through the task of predicting the occluded words.

[0067] In some embodiments, the process of constructing a joint scoring model based on classification tasks and regression tasks on the base model after domain adaptation pre-training includes:

[0068] S1: Use the Chinese pre-trained language model PTM after domain adaptation pre-training to extract features from the input pseudo-text sequence of the test questions to obtain the text sequence features.

[0069] S2: Perform average pooling and max pooling on the text sequence features, and concatenate the features after average pooling and max pooling to obtain the text features.

[0070] S3: For the text features, use the classification fully connected layer to output the probability of each category, and use the regression fully connected layer to output the regression result.

[0071] In a specific embodiment, a joint scoring model based on classification tasks and regression tasks is constructed on the BERT base model after domain adaptation pre-training. The model is as Figure 2 shown. Use the feature extraction base model after domain adaptation pre-training for the input Feature extraction is performed on the pseudo-text sequence to obtain text sequence features :

[0072]

[0073] For the text sequence features Average pooling and max pooling are performed, and the features after pooling are concatenated to obtain text features :

[0074]

[0075] In the formula, represents average pooling, represents max pooling, represents the concatenation operation;

[0076] The classification branch uses

[0077]

[0078] to perform score prediction, The output is converted into the probability of each category, represents the probability value of each score category, is the classification fully connected layer, and its input dimension is where is the dimension of the model semantic embedding, and the output dimension is the number of score types ;

[0079] The regression branch uses

[0080]

[0081] to perform score prediction, is the regression result, is the regression fully connected layer, and its input dimension is and the output dimension is 1.

[0082] In some embodiments, before performing step 102, the following steps are further performed: fine-tuning and training the joint scoring model based on manual scoring: constructing a training data set, where the training data set is a triple data including test question pseudo-text, category data converted from manual scores, and normalized manual scoring data; using the triple data, performing joint scoring model training based on the training method of random mini-batches, where the classification task uses cross-entropy to calculate the training loss, the regression task uses root mean square error to calculate the training loss, and the Adam optimizer is used to update the parameters of the joint scoring model.

[0083] Exemplarily, a training data set is constructed. The data set is a triple , where is the pseudo-text of the test questions, represents the category data for manual score conversion, represents the normalized manual scoring data:

[0084]

[0085] Among them, represents the maximum score, represents the average score, represents the manual expert scoring.

[0086] Model fine-tuning training. Using the triple data, the model is trained based on the training method of random mini-batches. The cross-entropy is used to calculate the training loss for the classification task, and the root mean square error is used to calculate the loss for the regression task. The Adam optimizer is used to update the model parameters.

[0087] In a specific embodiment, when performing step 102, the following steps can be specifically executed:

[0088] S1020: Select the score corresponding to the category with the highest probability among the probabilities of each category output by the classification fully connected layer as the scoring result of the classification branch.

[0089] Specifically, the pseudo-text of the test questions is scored using the fine-tuned model, and the scoring result is processed to obtain the final machine score. Refer to Figure 3 as shown.

[0090] Score using the fine-tuned model, and the model outputs and ;

[0091] Select the score corresponding to the category with the highest probability as the scoring result of the classification branch:

[0092]

[0093] Among them, represents the position index for obtaining the maximum value.

[0094] S1021: After inverse transformation of the regression result output by the regression fully connected layer, it is used as the scoring result of the regression branch.

[0095] Specifically, the regression result output by the regression fully connected layer is inverse-transformed according to the following formula and used as the scoring result of the regression branch:

[0096]

[0097] In the formula, represents the rounding function, Indicates the regression result, Indicates the maximum score, Indicates the average score.

[0098] S1022: Calculate the score difference between the scoring result of the classification branch and the scoring result of the regression branch. If the score difference is within the preset threshold range, the scoring is valid, and the average value of the scoring results of the classification branch and the regression branch is used as the scoring result.

[0099] Specifically, calculate the score difference between the classification task and the regression task. If the score difference is within the specified threshold, the scoring is valid, and the final scoring result Is the average of the two:

[0100]

[0101] At the same time, select The maximum probability as the confidence level of the scoring result. If the score difference is greater than the specified threshold, the machine scoring is invalid.

[0102] In this embodiment, the scoring threshold for the two branches in this embodiment is 1 point, that is, if the classification score and the regression score are less than or equal to 1 point, the machine scoring is valid, and if it is greater than 1 point, the machine scoring is invalid. The scoring result includes: classification score, regression score, whether it is valid, machine scoring, and confidence level. An example of the scoring result is shown in Table 2:

[0103] Table 2: Scoring result

[0104]

[0105] According to the above implementation plan, complete the grading of all candidates' answers to be graded. Table 3 is the experimental calculation result obtained by using the method of this application.

[0106] Table 3: Experimental calculation result

[0107]

[0108] The above 2-point consistency rate is calculated by first counting the number of differences between machine scoring and human scoring that are less than or equal to 2 points, and then dividing by the number of valid ones.

[0109] By constructing a subject word frequency table, converting the desensitized test question data into pseudo-text, performing domain adaptation pre-training, constructing a joint scoring model, etc., intelligent scoring of desensitized data is realized, and the accuracy and reliability of scoring are improved.

[0110] Refer to Figure 4 As shown, it is a structural schematic diagram of a test question scoring system provided by an embodiment of the present invention, including:

[0111] A building unit 40 for building a word frequency table for an exam subject, where the word frequency table includes the frequencies of each vocabulary in the subject in that subject;

[0112] A generation unit 41 for performing word frequency statistics on the desensitized test question data, and using the word frequency table to convert the desensitized test question data into the corresponding vocabulary in the word frequency table according to the frequency, and generating a pseudo text of the test question;

[0113] A scoring unit 42 for scoring the generated pseudo text of the test question based on a pre-constructed test question scoring model to obtain a test question scoring result; wherein, the test question scoring model is a joint scoring model based on a classification task and a regression task constructed on a base model after domain adaptation pre-training.

[0114] The test question scoring system includes a processor and a memory. The above-mentioned building unit, generation unit, scoring unit, etc. are all stored in the memory as program units, and the corresponding functions are realized by the processor executing the above program units stored in the memory.

[0115] The processor contains a kernel, and the kernel retrieves the corresponding program unit from the memory. One or more kernels can be set, and the intelligent scoring of the desensitized data is realized by adjusting the kernel parameters, improving the accuracy and reliability of the scoring.

[0116] The memory may include non-permanent memory in a computer-readable medium, forms such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM), and the memory includes at least one storage chip.

[0117] An embodiment of the present invention provides a storage medium, on which a program is stored, and when the program is executed by a processor, the test question scoring method is realized.

[0118] An embodiment of the present invention provides a processor, and the processor is used to run a program, wherein when the program runs, the test question scoring method is executed.

[0119] An embodiment of the present invention provides a device, which includes a processor, a memory, and a program stored on the memory and executable on the processor. When the processor executes the program, the following steps are implemented: constructing a word frequency table for an examination subject, where the word frequency table includes the frequencies of each vocabulary in the subject; performing word frequency statistics on the desensitized test question data, and using the word frequency table, converting the symbols in the desensitized test question data to the corresponding vocabulary in the word frequency table according to the frequency to generate a test question pseudo-text; scoring the generated test question pseudo-text based on a pre-constructed test question scoring model to obtain a test question scoring result; where the test question scoring model is a joint scoring model based on a classification task and a regression task constructed on a base model after domain adaptation pre-training. The device in this article can be a server, a PC, a PAD, a mobile phone, etc.

[0120] The present application also provides a computer program product, which is suitable for executing a program initialized with the following method steps when executed on a data processing device: constructing a word frequency table for an examination subject, where the word frequency table includes the frequencies of each vocabulary in the subject; performing word frequency statistics on the desensitized test question data, and using the word frequency table, converting the symbols in the desensitized test question data to the corresponding vocabulary in the word frequency table according to the frequency to generate a test question pseudo-text; scoring the generated test question pseudo-text based on a pre-constructed test question scoring model to obtain a test question scoring result; where the test question scoring model is a joint scoring model based on a classification task and a regression task constructed on a base model after domain adaptation pre-training.

[0121] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0122] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the specified functions in one Figure 1 one flow or multiple flows and / or Figure 1 one block or multiple blocks.

[0123] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including an instruction device that implements the functions specified in one or more of the processes and / or blocks Figure 1 one or more of the processes and / or blocks Figure 1 specified in one or more of the blocks or blocks.

[0124] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, so that the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one or more of the processes and / or blocks Figure 1 one or more of the processes and / or blocks Figure 1 specified in one or more of the blocks or blocks.

[0125] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0126] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM) and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.

[0127] Computer-readable media includes both permanent and non-permanent, removable and non-removable media implemented by any method or technology for storing information. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media such as modulated data signals and carrier waves.

[0128] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, commodity or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising the element.

[0129] The above are only embodiments of the present application and are not used to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. A test question scoring method, characterized in that: include: Constructing a word frequency table for each test subject, wherein the word frequency table includes the frequency of each word in the subject appearing in the subject; Perform word frequency statistics on the desensitized test data, and use the word frequency table to convert the desensitized test data into corresponding words in the word frequency table according to the frequency, so as to generate test pseudo text; Scoring the generated test question pseudo text based on a pre-built test question scoring model to obtain a test question scoring result; wherein the test question scoring model is a joint scoring model based on a classification task and a regression task constructed on a base model after domain adaptability pre-training; On the domain-adaptive pre-trained base model, the process of building a joint scoring model based on classification and regression tasks includes: Leverage domain-adaptive pre-trained feature extraction base model For input Extract features from pseudo text sequences to obtain text sequence features : Text sequence features Perform average pooling and maximum pooling, and concatenate the pooled features to obtain text features : In the formula, represents average pooling, represents the maximum pooling, Represents a splicing operation; Classification branch utilization Make score predictions, Convert the output into the probability of each category, Represents the probability value of each score category, is a classification fully connected layer, whose input dimension is ,in is the dimension of the model semantic embedding, and the output dimension is the number of types of scores ; Regression branch exploitation Make score predictions, is the regression result, is a regression fully connected layer, whose input dimension is , the output dimension is 1; The generated test question pseudo text is scored based on the pre-built test question scoring model to obtain the test question scoring results, including: Use the fine-tuned test scoring model to score, and the model output and ; choose The score corresponding to the category with the highest probability is used as the scoring result of the classification branch: In the formula, Indicates the position index for obtaining the maximum value; right Perform the inverse transformation as the scoring result of the regression branch: , In the formula, represents the rounding function; Calculate the score difference between the classification task and the regression task. If the score difference is within the specified threshold, the score is valid. The final score result The average of the two: Select both The maximum probability is used as the confidence level of the scoring result; If the score difference is greater than the specified threshold, the machine score is invalid.

2. The test question scoring method according to claim 1, characterized in that: The desensitized test data refers to data in which personal information, sensitive information and all answer contents in the test questions are replaced or encrypted.

3. The test question scoring method according to claim 1, characterized in that: The domain adaptation pre-training process includes: Combine the test subject corpus and test question pseudo text to form a domain-adaptive pre-training dataset; A Chinese pre-trained language model PTM is selected as a starting model for domain adaptive pre-training, some texts in a data set are randomly selected for occlusion processing, and the Chinese pre-trained language model PTM is subjected to domain adaptive pre-training through a task of predicting occluded words.

4. The test question scoring method according to claim 1, characterized in that: Before scoring the generated test question pseudo text based on the pre-built test question scoring model to obtain the test question scoring result, the test question scoring method further includes fine-tuning the joint scoring model based on manual scoring: Construct a training data set, wherein the training data set is a triple data including test question pseudo text, category data converted from manual scores, and normalized manual scoring data ,in, is the pseudo text of the test question, represents categorical data converted from artificial scores, Represents normalized manual scoring data: In the formula, Indicates the maximum score, represents the average score, represents manual expert rating; The triplet data is used to train the joint scoring model based on a random small batch training method, wherein the classification task uses cross entropy to calculate the training loss, the regression task uses root mean square error to calculate the training loss, and the Adam optimizer is used to update the parameters of the joint scoring model.

5. A test scoring system, characterized in that: The system is used to implement the test question scoring method according to any one of claims 1 to 4; the system comprises: A construction unit, used to construct a word frequency table for an examination subject, wherein the word frequency table includes the frequency of occurrence of each word in the subject; A generating unit is used to count the word frequency of the desensitized test data, and use the word frequency table to convert the desensitized test data into corresponding words in the word frequency table according to the frequency, so as to generate a test pseudo text; The scoring unit is used to score the generated test question pseudo-text based on a pre-built test question scoring model to obtain a test question scoring result; wherein the test question scoring model is a joint scoring model based on classification tasks and regression tasks constructed on a base model after domain adaptability pre-training.

6. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the test question scoring method as described in any one of claims 1 to 4 are implemented.

7. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the test question scoring method as described in any one of claims 1 to 4 are implemented.

Citation Information

Patent Citations

  • Question generation method based on word importance weighting

    CN113128206A