Scoring method, device and equipment and computer medium
Through the comprehensive determination of feature sorting and scoring results, the problem of low scoring efficiency in the existing technology is solved, the limitations of the regression model are avoided, and the scoring efficiency is improved.
Patent Information
- Application Number
- CN202510185140.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-02-19
AI Technical Summary
In the prior art, the oral evaluation scoring scheme is less efficient, and there are problems such as linear assumption problems, overfit/underfitting problems, complexity and data imbalance.
By obtaining multiple sets of first evaluation characteristics of the test questions to be scored and multiple sets of second evaluation characteristics, reference scores and initial prediction scores of the test questions to be scored, the feature sorting results are obtained, and the scoring results of the test questions to be scored are determined based on the results, reference scores and initial prediction scores.
This solution avoids various problems when grading test questions based on regression model, and improves the efficiency of grading oral language in oral evaluation.
Smart Images

Figure CN120048292A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of intelligent scoring, and particularly relates to a scoring method, device, equipment and computer medium. Background Art
[0002] With the development of educational technology and the increasing emphasis on language ability in society, oral language ability has become one of the important components for evaluating the comprehensive quality of students. In some examinations, oral language evaluation has gradually become an indispensable link.
[0003] Compared with manual scoring, in the face of a large number of examinees, machine scoring can not only quickly process a large amount of data, shorten the scoring cycle, but also reduce the influence of subjective factors. With the gradual development of oral language evaluation technology, manual oral language evaluation is being gradually replaced by machines. How to ensure the fairness and accuracy of machine scoring has become increasingly important. To ensure the accuracy of machine scoring, a process of calibration first and then machine scoring is introduced, that is, in an oral examination, a part of the data is selected, experts give manual scores according to the scoring criteria, and the machine learns based on a part of the data with manual labels and then evaluates the entire set of data.
[0004] The existing calibration scoring is generally divided into four stages. The first stage: Select a certain amount of examination data, use a general scoring scheme to obtain general scores and output the features for calculating the general score. The second stage: Select a certain amount of data according to the general scores and send them to experts for manual scoring. The third stage: Use the features output in the first stage and the manual scores in the second stage to retrain the ranking model. The fourth stage: Replace the new ranking model, retest the entire set of data of the current examination, and output the machine score.
[0005] However, the existing technologies all rely on the way of feature engineering to perform regression on the extracted evaluation features. The conventional calibration methods basically adopt the way of feature engineering, and learn the relevant regression model based on the features of the calibration data and the manual scores. The conventional regression model may have the following problems:
[0006] Linear hypothesis problem: Many regression models assume a linear relationship between scoring and features, but in fact this relationship may be non-linear.
[0007] Overfitting / underfitting problem: If the number of training samples is insufficient or the feature selection is inappropriate, the model may have problems of overfitting (overfitting the training data and having poor generalization ability) or underfitting (failing to capture the patterns in the data).
[0008] Complexity: Oral language evaluation involves many factors, such as pronunciation, fluency, grammar, vocabulary richness, etc. A single regression model may be difficult to comprehensively capture these complex factors.
[0009] Data imbalance: The calibrated sample data may have an imbalance problem. The number of samples in some intervals is relatively large, and it is easier to cause the data model with more samples to fit better during model training, while the data with fewer samples has poor fitting.
[0010] In summary, in the related art, the related oral evaluation scoring scheme has low scoring efficiency. Summary of the Invention
[0011] The embodiments of the present application provide an implementation solution different from the prior art to solve the technical problem that the related oral evaluation scoring scheme in the related art has low scoring efficiency.
[0012] In a first aspect, the present application provides a scoring method, including:
[0013] Obtain multiple groups of first evaluation features of multiple test questions to be scored;
[0014] Obtain multiple groups of second evaluation features of multiple calibration test questions, multiple reference scores of the multiple calibration test questions, and multiple initial predicted scores of the multiple calibration test questions;
[0015] Sort the multiple groups of first evaluation features and the multiple groups of second evaluation features to obtain a target sorting result;
[0016] Determine the scoring results of the multiple test questions to be scored based on the target sorting result, the multiple reference scores, and the multiple initial predicted scores.
[0017] In a second aspect, the present application provides a scoring device, including:
[0018] An obtaining unit, configured to obtain multiple groups of first evaluation features of multiple test questions to be scored;
[0019] The obtaining unit is configured to obtain multiple groups of second evaluation features of multiple calibration test questions, multiple reference scores of the multiple calibration test questions, and multiple initial predicted scores of the multiple calibration test questions;
[0020] A sorting unit, configured to sort the multiple groups of first evaluation features and the multiple groups of second evaluation features to obtain a target sorting result;
[0021] A determining unit, configured to determine the scoring results of the multiple test questions to be scored based on the target sorting result, the multiple reference scores, and the multiple initial predicted scores.
[0022] In a third aspect, the present application provides an electronic device, including:
[0023] A processor; and
[0024] A memory, configured to store executable instructions of the processor;
[0025] Wherein, the processor is configured to execute any method in the first aspect or any possible implementation manner of the first aspect by executing the executable instructions.
[0026] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, any method in the first aspect or any possible implementation manner of the first aspect is implemented.
[0027] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program. When the computer program is executed by a processor, any method described in the first aspect or any possible implementation manner of the first aspect is implemented.
[0028] The present application provides a solution for obtaining multiple sets of first evaluation features of multiple test questions to be scored; obtaining multiple sets of second evaluation features of multiple calibration test questions, multiple reference scores of the multiple calibration test questions, and multiple initial predicted scores of the multiple calibration test questions; sorting the multiple sets of first evaluation features and the multiple sets of second evaluation features to obtain a target sorting result; and determining the scoring results of the multiple test questions to be scored based on the target sorting result, the multiple reference scores, and the multiple initial predicted scores. Scoring test questions based on feature sorting can avoid a series of problems caused by scoring test questions based on a regression model, and improve the efficiency of scoring oral tests. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. In the drawings:
[0030] Figure 1 It is a schematic flowchart of a scoring method provided by an embodiment of the present application;
[0031] Figure 2a It is a schematic flowchart of a scoring method provided by an embodiment of the present application;
[0032] Figure 2b It is a schematic diagram of the process of determining a calibration test question set provided by an embodiment of the present application;
[0033] Figure 3 It is a schematic structural diagram of a scoring device provided by an embodiment of the present application;
[0034] Figure 4A schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0035] The embodiments of the present application will be described in detail below. The examples of the embodiments are shown in the accompanying drawings. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present application, and should not be construed as a limitation to the present application.
[0036] Terms such as "first" and "second" in the description, claims and drawings of the embodiments of the present application are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0037] First, some terms in the embodiments of the present application will be explained below to facilitate the understanding of those skilled in the art.
[0038] In machine learning and statistics, overfitting and underfitting are two common concepts that describe the performance of a model on training data and unseen data. Overfitting means that the model performs too well on the training data, so that it captures the noise and random fluctuations in the training data rather than the underlying patterns in the data. Therefore, when the model faces new and unseen data, its performance is often poor. Underfitting means that the model performs poorly on the training data and cannot capture the underlying patterns in the data. Therefore, when the model faces new and unseen data, its performance is also poor.
[0039] Textual features and BERT-like features each have their own uniqueness and importance in the field of natural language processing (NLP).
[0040] Textual features refer to the basic units or attributes used to represent text data, and these features play a crucial role in text analysis and processing. Common textual features include:
[0041] Bag of Words: Regarding the text as an unordered set of words, without considering the order and context of the words, only counting the frequency of word occurrences.
[0042] TF-IDF (Term Frequency-Inverse Document Frequency): Based on the bag-of-words model, it takes into account the frequency of a word in a document and the inverse document frequency in a document collection, and is used to evaluate the importance of a word in a document.
[0043] Word Embedding: Represents words as vectors in a high-dimensional space, and these vectors can capture the semantic and syntactic relationships between words. Common word embedding models include Word2Vec, GloVe, etc.
[0044] N-gram: Considers sequences of N consecutive words in text and is used to capture local context information in the text.
[0045] Syntactic features: Such as Part-of-Speech Tagging (POS Tagging), Named Entity Recognition (NER), etc. These features provide information about the grammatical and semantic roles of words in the text.
[0046] BERT (Bidirectional Encoder Representations from Transformers) is a pre-trained language model based on the Transformer architecture and has significant features and advantages in the field of natural language processing. BERT-like features are mainly reflected in the following aspects:
[0047] Bidirectional encoding ability: BERT uses a bidirectional Transformer encoder to encode text, can utilize context information simultaneously, and capture richer semantic relationships. Deep language representation: Through pre-training on large-scale text data, BERT learns deep language representation knowledge, and these representations can capture different levels of language information from shallow syntactic features to deep semantic features.
[0048] Pre-training tasks: BERT adopts two pre-training tasks, namely Masked Language Model (MLM) and Next Sentence Prediction (NSP). These tasks help the model understand the complex semantic relationships of words in a sentence and sentence-level information.
[0049] Wide adaptability: The BERT model has wide adaptability and can be applied to various NLP tasks, such as text classification, sequence labeling, question answering systems, named entity recognition, etc., through fine-tuning.
[0050] Parameter Sharing and Efficiency: During the fine-tuning process, most of the parameters of the BERT model are shared, and only the output layer parameters related to specific tasks need to be trained. This approach not only saves computing resources but also improves the model's generalization ability. At the same time, since BERT has learned rich language representation knowledge during the pre-training stage, it can converge faster and achieve better performance during the fine-tuning stage.
[0051] RankSVM is a ranking learning algorithm based on Support Vector Machine (SVM), mainly used to solve ranking problems, especially performing well in applications in the field of information retrieval. RankSVM determines a ranking model by learning a series of object pairs, aiming to minimize ranking errors. It uses the linear programming method in the dual space to optimize the objective function, making it possible to process large-scale data sets. This algorithm takes into account the partial order information between different object pairs and endeavors to maintain these relative relationships to improve the accuracy of ranking. RankSVM treats the ranking problem as an ordered multi-class problem. In terms of the core idea, the goal of RankSVM is to find a hyperplane to distinguish different ranked objects. These objects are represented by feature vectors, and the hyperplane is composed of coefficient weights. In specific implementation, RankSVM adjusts the hyperplane by minimizing the cost function, where the cost function imposes a penalty on the occurrence of ranking errors.
[0052] With the development of educational technology and the increasing social emphasis on language ability, oral language ability has become one of the important components in evaluating students' comprehensive qualities. In some exams, oral language assessment has gradually become an indispensable part.
[0053] Compared with manual scoring, in the face of a large number of examinees, machine scoring can not only process a large amount of data quickly, shorten the scoring cycle, but also reduce the influence of subjective factors. With the gradual development of oral language assessment technology, manual oral language assessment is being gradually replaced by machines. How to ensure the fairness and accuracy of machine scoring has become increasingly important. To ensure the accuracy of machine scoring, a process of pre-calibration followed by machine scoring is introduced, that is, in an oral exam, a part of the data is selected, and experts give manual scores according to the scoring criteria. After the machine learns based on a part of the data with manual labels, it then evaluates the entire set of data.
[0054] The existing calibration scoring generally consists of four stages. The first stage: Select a certain amount of exam data, use a general scoring scheme to obtain general scores and output the features for calculating the general score. The second stage: Select a certain amount of data based on the general scores and send them to experts for manual scoring. The third stage: Use the features output in the first stage and the manual scores in the second stage to retrain the ranking model. The fourth stage: Replace the new ranking model, retest the entire set of data for the current exam, and output the machine scores.
[0055] However, existing technologies all rely on the method of feature engineering to perform regression on the extracted evaluation features. Conventional calibration methods basically adopt the method of feature engineering, and learn relevant regression models based on the features of the calibration data and the manual scores. The following problems may exist in conventional regression models:
[0056] Linear hypothesis problem: Many regression models assume a linear relationship between the scores and the features, but in fact this relationship may be non-linear.
[0057] Overfitting / underfitting problem: If the number of training samples is insufficient or the feature selection is inappropriate, the model may have problems of overfitting (overfitting the training data and having poor generalization ability) or underfitting (failing to capture the patterns in the data).
[0058] Complexity: Oral evaluation involves many factors, such as pronunciation, fluency, grammar, vocabulary richness, etc. A single regression model may be difficult to comprehensively capture these complex factors.
[0059] Data imbalance: The calibrated sample data may have an imbalance problem. The number of samples in some intervals is relatively large, and it is easier to cause the data model with more samples to be better fitted during model training, while the data with fewer samples is poorly fitted.
[0060] In summary, in the related technologies, the related oral evaluation scoring scheme has low scoring efficiency.
[0061] The following uses specific embodiments to elaborate in detail on the technical solution of the present application and how the technical solution of the present application solves the above technical problems. These several specific embodiments below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.
[0062] Figure 1 FIG. is a schematic flowchart of a scoring method provided for an exemplary embodiment of the present application. The scoring method may include: obtaining a plurality of original test questions; selecting a part of the original test questions of a certain question type from the plurality of original test questions as a plurality of initial test questions; inputting the plurality of initial test questions into an evaluation unit to obtain an evaluation result, where the evaluation result includes the evaluation features and initial predicted scores of each initial test question. Determining a calibration question set and a question set to be scored from the plurality of initial test questions according to the initial predicted scores; sending the calibration question set to a target device to enable relevant personnel to perform manual scoring on the calibration question set; training an initial ranking model based on the manual scores of the calibration question set and the evaluation results of each calibration question in the calibration question set to obtain the target ranking model; the target ranking model is used to score each question to be scored in the question set to be scored. Further, the total score of the question set to be scored can also be calculated.
[0063] Optionally, the original examination questions in this application refer to the examination questions themselves and the answers given by candidates in response to the examination questions.
[0064] Optionally, the examination questions in this application refer to the examination questions for oral evaluation.
[0065] Figure 2a FIG. 2 is a schematic flow chart of a scoring method provided for an exemplary embodiment of this application. The execution subject of this method can be any electronic device, and this method at least includes the following steps S201-S204:
[0066] S201. Obtain multiple groups of first evaluation features of multiple examination questions to be scored;
[0067] S202. Obtain multiple groups of second evaluation features of multiple calibration examination questions, multiple reference scores of the multiple calibration examination questions, and multiple initial predicted scores of the multiple calibration examination questions;
[0068] Optionally, the multiple examination questions to be scored and the multiple calibration examination questions can be the initial examination questions in the same test paper.
[0069] Optionally, the multiple examination questions to be scored and the multiple calibration examination questions can be the initial examination questions of a single question type in the same test paper.
[0070] Optionally, the multiple examination questions to be scored and the multiple calibration examination questions can be the initial examination questions in multiple test papers of the same subject.
[0071] Optionally, the multiple examination questions to be scored and the multiple calibration examination questions can be the initial examination questions of a single question type in multiple test papers of the same subject.
[0072] Optionally, the question types can be divided into reading of fixed text content and non-reading of non-fixed text.
[0073] Specifically, the specific classification rules of the question types are not limited in this application.
[0074] Optionally, the reading of fixed text content can be further divided into reading of words, sentences, passages, etc.; the non-reading of non-fixed text can be further divided into open, semi-open and closed types, and there are also differences in length, and many question types can be combined.
[0075] In some alternative embodiments of this application, the multiple examination questions to be scored and the multiple calibration examination questions are all included in multiple initial examination questions.
[0076] Optionally, the method further includes: inputting the multiple initial examination questions into an evaluation unit to obtain multiple groups of first evaluation features of the multiple to-be-graded examination questions, multiple groups of second evaluation features of the multiple calibration examination questions, the initial predicted scores of each to-be-graded examination question among the multiple to-be-graded examination questions, and the initial predicted scores of each calibration examination question among the multiple calibration examination questions.
[0077] Optionally, regarding the evaluation features in this application, for the reading question type of fixed text, they can be relevant pronunciation features calculated after forced alignment; for question types of non-fixed text, they can be text features or bert features.
[0078] In some alternative embodiments of the present application, the method further includes the following steps S01-04:
[0079] S01. Obtain multiple initial examination questions;
[0080] Optionally, the multiple initial examination questions can be multiple examination questions in the same test paper and multiple answers given by the examinee for the multiple examination questions.
[0081] Optionally, the multiple initial examination questions can be multiple examination questions in multiple test papers of the same subject and multiple answers given by the examinee for the multiple examination questions.
[0082] S02. Determine a preset number of partial examination questions from the multiple initial examination questions as the multiple calibration examination questions;
[0083] In some alternative embodiments of the present application, in S02, determining a preset number of partial examination questions from the multiple initial examination questions as the multiple calibration examination questions includes the following steps S021-S024:
[0084] S021. Obtain a preset total score;
[0085] Optionally, the preset total score can be 100 points.
[0086] S022. Obtain the total number of initial examination questions among the multiple initial examination questions;
[0087] S023. Divide the total score into multiple score segments to obtain multiple score segments;
[0088] Optionally, the multiple score segments can be: [0,10],(10,20],(20,30]...(90,100].
[0089] S024. For each of the multiple score segments, count the number of initial questions among the multiple initial questions whose initial predicted scores fall within the score segment; determine the target ratio of the number to the total number; select, from the initial questions whose initial predicted scores fall within the score segment, the target initial questions with the number whose ratio to the preset number is the target ratio as the calibration questions corresponding to the score segment.
[0090] Among them, the target initial questions with the number whose ratio to the preset number is the target ratio among the initial questions whose initial predicted scores fall within the score segment can be arbitrarily selected from the initial questions whose initial predicted scores fall within the score segment.
[0091] Among them, the total number of calibration questions corresponding to each score segment is the aforementioned preset number. All the calibration questions corresponding to each score segment are the multiple calibration questions with the preset number determined from the multiple initial questions. These multiple calibration questions form a calibration question set. Specifically, for the process of determining the calibration question set, reference can be made to Figure 2b as shown.
[0092] S03. Use at least some of the multiple initial questions other than the multiple calibration questions as the multiple questions to be scored.
[0093] Optionally, all the multiple initial questions other than the multiple calibration questions can be used as the multiple questions to be scored.
[0094] S04. Send the multiple calibration questions to a first device so that relevant personnel can score the multiple calibration questions to obtain multiple reference scores for the multiple calibration questions.
[0095] Among them, the reference score is a manual score.
[0096] S203. Sort the multiple groups of first evaluation features and the multiple groups of second evaluation features to obtain a target sorting result.
[0097] In some alternative embodiments of the present application, in S203 above, the sorting of the multiple groups of first evaluation features and the multiple groups of second evaluation features to obtain a target sorting result includes: inputting the multiple groups of first evaluation features and the multiple groups of second evaluation features into a target sorting model to obtain a target sorting result; where the target sorting model is used to sort the multiple groups of first evaluation features and the multiple groups of second evaluation features.
[0098] In some alternative embodiments of the present application, in S203 above, the sorting of the multiple groups of first evaluation features and the multiple groups of second evaluation features to obtain a target sorting result includes:
[0099] For each group of the first evaluation features among the multiple groups of the first evaluation features, construct feature pairs between the first evaluation feature and each group of the second evaluation features among the multiple groups of the second evaluation features, to obtain multiple feature pairs corresponding to the first evaluation feature; input the multiple feature pairs into a target sorting model to obtain multiple sorting results corresponding to the multiple feature pairs, and further obtain a target sorting result;
[0100] Among them, in the target sorting result, it includes multiple sorting results corresponding to multiple feature pairs corresponding to each group of the first evaluation features.
[0101] Optionally, each first evaluation feature corresponds to multiple feature pairs and multiple sorting results. The multiple sorting results corresponding to each first evaluation feature can be regarded as a group of sorting results; the multiple groups of sorting results corresponding to multiple first evaluation features are regarded as the target sorting result.
[0102] Optionally, each first evaluation feature corresponds to multiple feature pairs and multiple sorting results. The multiple sorting results corresponding to each first evaluation feature can be regarded as a group of sorting results; the total sorting result determined based on the multiple groups of sorting results corresponding to multiple first evaluation features is regarded as the target sorting result.
[0103] S204. Determine the scoring results of the multiple questions to be scored based on the target sorting result, the multiple reference scores, and the multiple initial predicted scores.
[0104] In some optional embodiments of the present application, in the foregoing S204, the determining the scoring results of the multiple questions to be scored based on the target sorting result, the multiple reference scores, and the multiple initial predicted scores includes: for each question to be scored among the multiple questions to be scored,
[0105] If the target sorting result indicates that the first evaluation feature of the question to be scored is between two groups of the second evaluation features, then determine the scoring result of the question to be scored based on a first preset algorithm, the initial predicted scores and reference scores of two calibration questions corresponding to the two groups of the second evaluation features;
[0106] If the target sorting result indicates that the first evaluation feature of the question to be scored is before the multiple groups of the second evaluation features, then determine the scoring result of the question to be scored based on a second preset algorithm, the initial predicted scores and reference scores of two calibration questions corresponding to two second evaluation features that are sorted after the first evaluation feature and are the closest to the first evaluation feature among the multiple groups of the second evaluation features.
[0107] Optionally, determining the scoring result of the test question to be scored based on the first preset algorithm, the initial predicted scores and reference scores of the two calibration test questions corresponding to the two sets of second evaluation features can be implemented through the following formula:
[0108]
[0109] where x test refers to the first evaluation feature of the test question to be scored, x train1 refers to the second evaluation feature before the first evaluation feature, x train2 refers to the second evaluation feature after the first evaluation feature, f(x train1 ) is the initial predicted score of the calibration test question corresponding to x train1 , f(x train2 ) is the initial predicted score of the calibration test question corresponding to x train2 , h(x train1 ) is the reference score of the calibration test question corresponding to x train1 , h(x train2 ) is the reference score of the calibration test question corresponding to x train2 . refers to the scoring result of the test question to be scored.
[0110] Optionally, determining the scoring result of the test question to be scored based on the second preset algorithm, the initial predicted scores and reference scores of the two calibration test questions corresponding to the two second evaluation features that are sorted after the first evaluation feature and are the closest to the first evaluation feature can be implemented through the following formula:
[0111]
[0112] where x test refers to the first evaluation feature of the test question to be scored, x train11 refers to the second evaluation feature that is sorted after the first evaluation feature and is the closest to the first evaluation feature, x train12 refers to the second evaluation feature that is sorted after the first evaluation feature and is the second closest to the first evaluation feature, f(x train11 ) is the initial predicted score of the calibration test question corresponding to x train11 , f(x train12 ) is the initial predicted score of the calibration test question corresponding to x train12 , h(x train11 ) is the reference score of the calibration test question corresponding to x train11 , h(x train12 ) is the reference score of the calibration test question corresponding to x train12 . refers to the scoring result of the test question to be scored.
[0113] In some alternative embodiments of the present application, the method further includes: training an initial ranking model based on the multiple sets of second evaluation features, the multiple reference scores of the multiple calibration questions, and the multiple initial predicted scores of the multiple calibration questions to obtain the target ranking model.
[0114] Optionally, the initial ranking model in the present application may be a RankSVM model.
[0115] In some alternative embodiments of the present application, the training of the initial ranking model based on the multiple sets of second evaluation features, the multiple reference scores of the multiple calibration questions, and the multiple initial predicted scores of the multiple calibration questions to obtain the target ranking model includes the following steps S1 - S5:
[0116] S1. Take out two sets of second evaluation features from the multiple sets of second evaluation features to form a sample feature pair;
[0117] Optionally, in S1, taking out two sets of second evaluation features from the multiple sets of second evaluation features to form a sample feature pair includes: arbitrarily taking out two sets of second evaluation features from the multiple sets of second evaluation features to form a sample feature pair.
[0118] Optionally, in S1, taking out two sets of second evaluation features from the multiple sets of second evaluation features to form a sample feature pair includes: taking out any two sets of second evaluation features that have not been taken out from the multiple sets of second evaluation features to form a sample feature pair.
[0119] S2. Obtain the true ranking result corresponding to the sample feature pair;
[0120] Optionally, each calibration question corresponds to a set of second evaluation features, and each set of second evaluation features can be represented in the following manner: X 1 , X 2 ,..., X m ;
[0121] X 1 = {x 11 , x 12 ,..., x 1n}; where x 11 , x 12 ,..., x 1n are the multi - dimensional evaluation features in X1. X 2 = {x 21 , x 22 ,..., x 2n};...
[0122] X m = {x m1 , xm2 ,..., x mn};
[0123] X 1 The reference score of the corresponding calibration test question is Y 1 ;
[0124] X 2 The reference score of the corresponding calibration test question is Y 2 ;...
[0125] X m The reference score of the corresponding calibration test question is Y m ;
[0126] When the true ranking of X 1 is before X 2 , Y 1 is greater than Y 2 , and the corresponding label can be +1;
[0127] When the true ranking of X 1 is after X 2 , Y 1 is less than Y 2 , and the corresponding label can be -1;
[0128] The feature of the sample feature pair can refer to the difference between the two second evaluation features in the sample feature pair, and the label can refer to their relative superiority and inferiority relationship.
[0129] S3. Input the sample feature pair into the initial ranking model to obtain a predicted ranking result;
[0130] Optionally, before inputting the sample feature pair into the initial ranking model, the sample feature pair can also be processed as follows: calculate the difference eigenvalue between the two second evaluation features in the sample feature pair; use the difference eigenvalue and the corresponding label as the sample data for inputting into the initial ranking model.
[0131] S4. Determine the corresponding loss information based on the true ranking result and the predicted ranking result;
[0132] Optionally, in the aforementioned S4, determining the corresponding loss information based on the true ranking result and the predicted ranking result can be achieved through the following formula:
[0133]
[0134] where X i refers to the i-th second evaluation feature, X j refers to the i-th second evaluation feature, f(X i , X j ) = w T(X i -X j ) + b, where w and b are adjustable parameters in the initial ranking model. C is a hyperparameter, ||w|| 2 is the regularization term, max(0, 1 - y ij f(X i , X j )) is the hinge loss of the ranking pair, which controls the trade-off between regularization and loss.
[0135] S5. When the loss information meets the preset conditions, the most recently determined initial ranking model is used as the target ranking model. When the loss information does not meet the preset conditions, the parameters of the initial ranking model are adjusted, and the step of taking out two sets of second evaluation features from the multiple sets of second evaluation features to form a sample feature pair is returned until the target ranking model is determined.
[0136] Optionally, when the loss information is less than the preset threshold, it is considered that the loss information meets the preset conditions.
[0137] This application proposes a pairwise-based ranking oral evaluation scheme, which no longer calculates the total score in a regression manner, removing the possible non-linear relationship between the regression model and features. It avoids the problems that may occur when the number of training samples is insufficient or the feature selection is inappropriate, such as overfitting (overfitting the training data and having poor generalization ability) or underfitting (failing to capture the patterns in the data) of the model. It also avoids the difficulty of a single regression model in comprehensively capturing complex phonemes involved in oral evaluation, such as pronunciation, fluency, grammar, and lexical richness. The scoring scheme no longer trains a regression model, solving the problem of the model effect caused by data imbalance. Moreover, this application can also select different ranking calculation schemes according to the differences in device computing capabilities.
[0138] The present application provides a scheme for obtaining multiple groups of first evaluation features of multiple test questions to be scored; obtaining multiple groups of second evaluation features of multiple calibration test questions, the multiple reference scores of the multiple calibration test questions, and the multiple initial predicted scores of the multiple calibration test questions; ranking the multiple groups of first evaluation features and the multiple groups of second evaluation features to obtain a target ranking result; and determining the scoring results of the multiple test questions to be scored based on the target ranking result, the multiple reference scores, and the multiple initial predicted scores. Scoring the test questions based on feature ranking can avoid a series of problems caused by scoring the test questions based on a regression model, improving the efficiency of scoring oral language in oral evaluation.
[0139] Figure 3 is a schematic structural diagram of a scoring device provided by an exemplary embodiment of the present application; wherein, the device includes:
[0140] An acquisition unit 31, configured to acquire multiple groups of first evaluation features of multiple questions to be scored;
[0141] The acquisition unit 31 is further configured to acquire multiple groups of second evaluation features of multiple calibration questions, multiple reference scores of the multiple calibration questions, and multiple initial predicted scores of the multiple calibration questions;
[0142] A sorting unit 32, configured to sort the multiple groups of first evaluation features and the multiple groups of second evaluation features to obtain a target sorting result;
[0143] A determination unit 33, configured to determine scoring results of the multiple questions to be scored based on the target sorting result, the multiple reference scores, and the multiple initial predicted scores.
[0144] In some alternative embodiments of the present application, when the foregoing device is used to sort the multiple groups of first evaluation features and the multiple groups of second evaluation features to obtain a target sorting result, it specifically is configured to:
[0145] Input the multiple groups of first evaluation features and the multiple groups of second evaluation features into a target sorting model to obtain a target sorting result;
[0146] Wherein, the target sorting model is used to sort the multiple groups of first evaluation features and the multiple groups of second evaluation features.
[0147] In some alternative embodiments of the present application, when the foregoing device is used to sort the multiple groups of first evaluation features and the multiple groups of second evaluation features to obtain a target sorting result, it specifically is configured to:
[0148] For each group of first evaluation features in the multiple groups of first evaluation features, construct feature pairs between the first evaluation features and each group of second evaluation features in the multiple groups of second evaluation features to obtain multiple feature pairs corresponding to the first evaluation features; input the multiple feature pairs into a target sorting model to obtain multiple sorting results corresponding to the multiple feature pairs, and further obtain a target sorting result;
[0149] Wherein, in the target sorting result, it includes multiple sorting results corresponding to multiple feature pairs corresponding to each group of first evaluation features.
[0150] In some alternative embodiments of the present application, when the foregoing device is used to determine scoring results of the multiple questions to be scored based on the target sorting result, the multiple reference scores, and the multiple initial predicted scores, it specifically is configured to:
[0151] For each question to be scored among the multiple questions to be scored,
[0152] If the target sorting result indicates that the first evaluation feature of the test question to be scored is between two sets of second evaluation features, determine the scoring result of the test question to be scored based on the first preset algorithm, the initial predicted scores and reference scores of the two calibration test questions corresponding to the two sets of second evaluation features;
[0153] If the target sorting result indicates that the first evaluation feature of the test question to be scored is before the multiple sets of second evaluation features, determine the scoring result of the test question to be scored based on the second preset algorithm, the initial predicted scores and reference scores of the two calibration test questions corresponding to the two sets of second evaluation features that are sorted after the first evaluation feature and are the closest to the first evaluation feature among the multiple sets of second evaluation features.
[0154] In some alternative embodiments of the present application, the foregoing device is further configured to:
[0155] Obtain a plurality of initial test questions;
[0156] Determine a preset number of partial test questions from the plurality of initial test questions as the plurality of calibration test questions;
[0157] Use at least some of the plurality of initial test questions other than the plurality of calibration test questions as the plurality of test questions to be scored;
[0158] Send the plurality of calibration test questions to a first device, so that relevant personnel score the plurality of calibration test questions to obtain a plurality of reference scores for the plurality of calibration test questions.
[0159] In some alternative embodiments of the present application, the foregoing device is further configured to:
[0160] Train an initial sorting model based on the multiple sets of second evaluation features, the multiple reference scores of the multiple calibration test questions, and the multiple initial predicted scores of the multiple calibration test questions to obtain the target sorting model.
[0161] In some alternative embodiments of the present application, when the foregoing device is used to train an initial sorting model based on the multiple sets of second evaluation features, the multiple reference scores of the multiple calibration test questions, and the multiple initial predicted scores of the multiple calibration test questions to obtain the target sorting model, it is specifically configured to:
[0162] Take out two sets of second evaluation features from the multiple sets of second evaluation features to form a sample feature pair;
[0163] Obtain the true sorting result corresponding to the sample feature pair;
[0164] Input the sample feature pair into the initial sorting model to obtain a predicted sorting result;
[0165] Determine corresponding loss information based on the true sorting result and the predicted sorting result;
[0166] When the loss information meets the preset conditions, use the most recently determined initial sorting model as the target sorting model. When the loss information does not meet the preset conditions, adjust the parameters of the initial sorting model, and return to execute the step of taking out two sets of second evaluation features from the multiple sets of second evaluation features to form a sample feature pair until a target sorting model is determined.
[0167] It should be understood that the device embodiments and the method embodiments can correspond to each other, and similar descriptions can refer to the method embodiments. To avoid repetition, they will not be elaborated here. Specifically, the device can execute the above method embodiments, and the foregoing and other operations and / or functions of each module in the device respectively correspond to the corresponding processes in each method in the above method embodiments. For the sake of brevity, they will not be elaborated here.
[0168] The device of the embodiments of the present application has been described above from the perspective of functional modules in combination with the drawings. It should be understood that the functional modules can be implemented in the form of hardware, or in the form of instructions in software, or in a combination of hardware and software modules. Specifically, the steps of the method embodiments in the present application can be completed by the integrated logic circuit in hardware in the processor and / or instructions in software form. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by a combination of hardware and software modules in the decoding processor. Optionally, the software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps in the above method embodiments.
[0169] Figure 4 is a schematic block diagram of an electronic device provided by an embodiment of the present application. The electronic device may include:
[0170] A memory 301 and a processor 302. The memory 301 is used to store a computer program and transmit the program code to the processor 302. In other words, the processor 302 can call and run the computer program from the memory 301 to implement the method in the embodiments of the present application.
[0171] For example, the processor 302 can be used to execute the above method embodiments according to the instructions in the computer program.
[0172] In some embodiments of the present application, the processor 302 may include, but is not limited to:
[0173] General-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and the like.
[0174] In some embodiments of the present application, the memory 301 includes, but is not limited to:
[0175] Volatile memory and / or non-volatile memory. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synch link DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0176] In some embodiments of the present application, the computer program may be divided into one or more modules, and the one or more modules are stored in the memory 301 and executed by the processor 302 to complete the method provided by the present application. The one or more modules may be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program in the electronic device.
[0177] As Figure 4 shown, the electronic device may further include:
[0178] A transceiver 303, which can be connected to the processor 302 or the memory 301.
[0179] Among them, the processor 302 can control the transceiver 303 to communicate with other devices. Specifically, it can send information or data to other devices, or receive information or data sent by other devices. The transceiver 303 can include a transmitter and a receiver. The transceiver 303 can further include an antenna, and the number of antennas can be one or more.
[0180] It should be understood that each component in the electronic device is connected through a bus system. Among them, the bus system includes, in addition to the data bus, a power bus, a control bus, and a status signal bus.
[0181] This application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a computer, the computer can execute the method of the above method embodiment. Or rather, the embodiment of this application also provides a computer program product containing instructions. When the instructions are executed by a computer, the computer executes the method of the above method embodiment.
[0182] When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of this application are generated in whole or in part. The computer can be a general-purpose computer, a dedicated computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or data center containing one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a digital video disc (DVD)), or a semiconductor medium (such as a solid state disk (SSD)), etc.
[0183] Those of ordinary skill in the art can realize that the modules and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.
[0184] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or module can be in an electrical, mechanical, or other form.
[0185] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical modules, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. For example, in each embodiment of this application, the various functional modules can be integrated in a processing module, or each module can exist physically alone, or two or more modules can be integrated in one module.
[0186] The above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, which should all be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
Claims
1. A scoring method, characterized in that: include: Obtain multiple groups of first evaluation features of multiple test questions to be scored; Acquire multiple groups of second evaluation features of multiple calibration test questions, multiple reference scores of the multiple calibration test questions, and multiple initial predicted scores of the multiple calibration test questions; Sorting the multiple groups of first evaluation features and the multiple groups of second evaluation features to obtain a target sorting result; The scoring results of the multiple test questions to be scored are determined based on the target ranking result, the multiple reference scores, and the multiple initial predicted scores.
2. The method according to claim 1, characterized in that The step of sorting the plurality of groups of first evaluation features and the plurality of groups of second evaluation features to obtain a target sorting result includes: Inputting the plurality of groups of first evaluation features and the plurality of groups of second evaluation features into a target ranking model to obtain a target ranking result; The target ranking model is used to rank the multiple groups of first evaluation features and the multiple groups of second evaluation features.
3. The method according to claim 1, characterized in that The step of sorting the plurality of groups of first evaluation features and the plurality of groups of second evaluation features to obtain a target sorting result includes: For each group of first evaluation features in the multiple groups of first evaluation features, construct a feature pair of the first evaluation feature and each group of second evaluation features in the multiple groups of second evaluation features to obtain a plurality of feature pairs corresponding to the first evaluation feature; input the plurality of feature pairs into a target ranking model to obtain a plurality of ranking results corresponding to the plurality of feature pairs, and then obtain a target ranking result; The target ranking result includes multiple ranking results corresponding to multiple feature pairs corresponding to each group of first evaluation features.
4. The method according to claim 1, characterized in that The step of determining the scoring results of the plurality of test questions to be scored based on the target ranking result, the plurality of reference scores, and the plurality of initial predicted scores includes: For each of the plurality of test questions to be graded, If the target ranking result indicates that the first evaluation feature of the test question to be scored is between the two groups of second evaluation features, then determining the scoring result of the test question to be scored based on the first preset algorithm, the initial predicted scores of the two calibration test questions corresponding to the two groups of second evaluation features, and the reference scores; If the target ranking result indicates that the first evaluation feature of the test question to be scored is before the multiple groups of second evaluation features, then the scoring result of the test question to be scored is determined based on a second preset algorithm, the initial predicted scores and reference scores of two calibrated test questions corresponding to two second evaluation features in the multiple groups of second evaluation features that are ranked after the first evaluation feature and are most adjacent to the first evaluation feature.
5. The method according to claim 1, characterized in that The method further comprises: Get multiple initial test questions; Determining a preset number of partial test questions from the plurality of initial test questions as the plurality of calibration test questions; using at least part of the plurality of initial test questions except the plurality of calibrated test questions as the plurality of test questions to be graded; The plurality of calibration test questions are sent to the first device, so that relevant personnel can score the plurality of calibration test questions to obtain a plurality of reference scores for the plurality of calibration test questions.
6. The method according to claim 2, characterized in that The method further comprises: Based on the multiple groups of second evaluation features, the multiple reference scores of the multiple calibration test questions, and the multiple initial predicted scores of the multiple calibration test questions, the initial ranking model is trained to obtain the target ranking model.
7. The method according to claim 6, characterized in that The training of the initial ranking model based on the multiple groups of second evaluation features, the multiple reference scores of the multiple calibration test questions, and the multiple initial predicted scores of the multiple calibration test questions to obtain the target ranking model includes: Taking out two groups of second evaluation features from the multiple groups of second evaluation features to form sample feature pairs; Obtaining the actual sorting results corresponding to the sample feature pairs; Inputting the sample feature pairs into an initial sorting model to obtain a predicted sorting result; Determine corresponding loss information based on the actual sorting result and the predicted sorting result; When the loss information meets the preset conditions, the most recently determined initial sorting model is used as the target sorting model. When the loss information does not meet the preset conditions, the parameters of the initial sorting model are adjusted, and the step of extracting two groups of second evaluation features from the multiple groups of second evaluation features to form a sample feature pair is returned to execute until the target sorting model is determined.
8. A scoring device, characterized in that: include: An acquisition unit, used for acquiring multiple groups of first evaluation features of multiple test questions to be scored; The acquisition unit is used to acquire multiple groups of second evaluation features of multiple calibration test questions, multiple reference scores of the multiple calibration test questions, and multiple initial predicted scores of the multiple calibration test questions; A sorting unit, used to sort the multiple groups of first evaluation features and the multiple groups of second evaluation features to obtain a target sorting result; A determination unit is used to determine the scoring results of the multiple test questions to be scored based on the target ranking result, the multiple reference scores, and the multiple initial predicted scores.
9. An electronic device, characterized in that: include: processor; as well as A memory, configured to store executable instructions of the processor; The processor is configured to perform the method of any one of claims 1 to 7 by executing the executable instructions.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Machine intelligent evaluation method and system for translation test questions
CN111767743A
Language model fusion method and device, medium and computer program product
CN113140221A
Automatic scoring method and system for subjective questions
CN114333461A
Program, device and method automatically grading from dictation voice of learner
JP2018045062A
Sample assessment
WO2022015404A1