Scoring method, device, equipment, storage medium and program product for oral examination
By using the pre-training scoring model trained in the oral exam and carrying out adaptive training, the problem of low scoring efficiency in the oral exam is solved, and automated scoring and efficiency improvement are achieved.
Patent Information
- Application Number
- CN202111405039.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-24
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2041-11-24
AI Technical Summary
The oral test score is inefficient and requires a lot of labor and time costs.
The pre-trained scoring model is trained using meta-learning method, and it is trained adaptively based on a small number of training samples to obtain the target scoring model, which is used for automated scoring.
It reduces the dependence on manual scoring, improves the scoring efficiency of oral examinations, and realizes automated scoring.
Smart Images

Figure CN114333787B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of artificial intelligence, and in particular to a scoring method, device, equipment, storage medium and program product for an oral test. Background Art
[0002] The oral test is a form of examination that tests oral ability. The types of questions used include picture description, quick response, topic description, opinion expression, etc.
[0003] During the oral test, after the examinee completes the answer, the scorer will score the answer from the perspectives of pronunciation, grammar, and accuracy of the answer to the question to obtain the test score.
[0004] Since scoring is usually done by experienced teachers or experts, it takes a lot of manpower and time, resulting in low efficiency in oral exam scoring. Summary of the invention
[0005] The present application provides a scoring method, device, equipment, storage medium and program product for an oral test, which can improve the scoring efficiency of an oral test. The technical solution is as follows:
[0006] On the one hand, an embodiment of the present application provides a scoring method for an oral test, the method comprising:
[0007] Acquire a training sample, wherein the training sample includes a sample reference answer, a sample answer audio, and a sample score for the sample answer audio of a target sample oral test question, wherein the target sample oral test question belongs to a target question type;
[0008] Training a pre-trained scoring model based on the training samples to obtain a target scoring model corresponding to the target question type, wherein the pre-trained scoring model is trained by meta-learning;
[0009] The target answer audio is scored by the target scoring model to obtain a target score for the target answer audio, wherein the target answer audio is a response to an oral test question belonging to the target question type.
[0010] On the other hand, an embodiment of the present application provides a scoring device for an oral test, the device comprising:
[0011] A first acquisition module is used to acquire a training sample, wherein the training sample includes a sample reference answer, a sample answer audio, and a sample score for the sample answer audio of a target sample oral test question, wherein the target sample oral test question belongs to a target question type;
[0012] A first training module is used to train a pre-trained scoring model based on the training samples to obtain a target scoring model corresponding to the target question type, wherein the pre-trained scoring model is obtained by training in a meta-learning manner;
[0013] A scoring module is used to score the target answer audio through the target scoring model to obtain a target score for the target answer audio, wherein the target answer audio is a response to an oral test question belonging to the target question type.
[0014] On the other hand, an embodiment of the present application provides a computer device, which includes a processor and a memory, wherein the memory stores at least one instruction, and the at least one instruction is loaded and executed by the processor to implement the scoring method for the oral test as described in the above aspects.
[0015] On the other hand, an embodiment of the present application provides a computer-readable storage medium, in which at least one instruction is stored. The at least one instruction is loaded and executed by a processor to implement the scoring method for the oral test as described in the above aspects.
[0016] On the other hand, the embodiment of the present application provides a computer program product or a computer program, which includes a computer instruction stored in a computer-readable storage medium. The processor of the computer device reads the computer instruction from the computer-readable storage medium, and the processor executes the computer instruction, so that the computer device performs the speech synthesis method described in the above aspect, or implements the scoring method for the oral test described in the above aspect.
[0017] In the embodiment of the present application, a pre-trained scoring model is obtained by pre-training using a meta-learning method. When it is necessary to score an oral test using a target question type, the pre-trained scoring model is further trained based on the training samples of the target question type to obtain a target scoring model corresponding to the target question type, thereby using the target scoring model to score the answers to the target oral test questions. Since the pre-trained scoring model is trained using a meta-learning method, that is, the pre-trained scoring model has learned prior scoring knowledge in advance, only a small amount of training samples is needed to train the target scoring model, thereby reducing the reliance on manual scoring, and after the training is completed, the target scoring model is used to realize automatic scoring of the oral test, thereby improving the scoring efficiency of the oral test. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0019] Figure 1 A schematic diagram showing an implementation environment provided by an exemplary embodiment of the present application is shown;
[0020] Figure 2 A flowchart showing a scoring method for an oral test provided by an exemplary embodiment of the present application is shown;
[0021] Figure 3 A flowchart showing a scoring method for an oral test provided by another exemplary embodiment of the present application is shown;
[0022] Figure 4 is a schematic diagram of the structure of a scoring model shown in an exemplary embodiment of the present application;
[0023] Figure 5 is a flowchart of a target scoring model scoring process shown in an exemplary embodiment of the present application;
[0024] Figure 6 is a flowchart of a meta-learning process shown in an exemplary embodiment of the present application;
[0025] Figure 7 and Figure 8 This is a comparison chart of experimental data on the task adaptation effect when using different schemes;
[0026] Fig. 9 is a flowchart of an oral test scoring process shown in an exemplary embodiment of the present application;
[0027] Fig.10 A structural block diagram of a scoring device for an oral test provided by an exemplary embodiment of the present application is shown;
[0028] Fig.11 A schematic diagram of the structure of a computer device provided by an exemplary embodiment of the present application is shown. DETAILED DESCRIPTION
[0029] In order to make the objectives, technical solutions and advantages of the present application clearer, the implementation methods of the present application will be further described in detail below with reference to the accompanying drawings.
[0030] Artificial Intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that machines have the functions of perception, reasoning and decision-making.
[0031] Artificial intelligence technology is a comprehensive discipline that covers a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, mechatronics and other technologies. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0032] The key technologies of speech technology include automatic speech recognition technology (ASR), text to speech (TTS) and voiceprint recognition technology. Enabling computers to listen, see, speak and feel is the future development direction of human-computer interaction, among which speech becomes one of the promising human-computer interaction methods. The embodiment of the present application is the application of speech technology in the oral test scenario, which is used to automatically score the audio answers to oral test questions with the help of the trained scoring model.
[0033] Since oral test questions are rich and varied, if the scoring model is trained directly for oral tests of different question types, it is necessary to rely on a large number of manually labeled training samples, resulting in a high sample preparation cost in the early stage of training. In order to reduce the reliance on manually labeled training samples while ensuring the accuracy of scoring, the embodiment of the present application proposes a scheme of using meta-learning to train a pre-trained scoring model (i.e., obtaining a unified initialization parameter for different oral tasks), and using a small amount of training samples to perform rapid adaptive training on the pre-trained scoring model. Among them, meta-learning is also called "learning to learn", that is, using previous knowledge and experience to guide the learning of new tasks, so that the model has the ability to learn to learn, so as to quickly learn new tasks based on existing knowledge.
[0034] Figure 1A schematic diagram of an implementation environment provided by an exemplary embodiment of the present application is shown. The implementation environment includes a scoring terminal 110, a server 120, and an examination terminal 130. The examination terminal 130 and the server 120 perform data communication via a communication network, and the scoring terminal 110 and the server 120 perform data communication via a communication network. Optionally, the communication network may be a wired network or a wireless network, and the communication network may be at least one of a local area network, a metropolitan area network, and a wide area network.
[0035] The scoring terminal 110 is a terminal for manual scoring, which includes but is not limited to mobile phones, computers, intelligent voice interaction devices, smart home appliances, car terminals, etc., which are not limited in the present embodiment. In some embodiments, the scoring terminal 110 is a terminal used by a scorer, who can be a teacher or a professional.
[0036] In a possible implementation, when it is necessary to train a scoring model for automatic scoring of a specific question type, the server 120 provides the scoring terminal 110 with training samples to be annotated, the training samples including sample oral test questions (of a specific question type), sample reference answers, and sample answer audio. The scoring terminal 110 plays the sample answer audio and obtains the sample score input by the scorer, thereby feeding back the sample score to the server 120.
[0037] Server 120 is a device for providing oral test scoring services. It can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.
[0038] In the embodiment of the present application, a pre-trained scoring model trained by meta-learning is provided in the server 120. When it is necessary to provide scoring services for oral exams of specific question types, the server 120 provides training samples to be annotated to the scoring terminal 110, and obtains the sample scores fed back by the scoring terminal 110, thereby adaptively training the pre-trained scoring model based on the manually annotated training samples to obtain a target scoring model corresponding to the specific question type.
[0039] The test terminal 130 is a terminal used by oral test takers, which includes but is not limited to mobile phones, computers, intelligent voice interaction devices, smart home appliances, car terminals, etc., which are not limited in the embodiments of the present application.
[0040] During the oral test, the test terminal 130 displays the oral test questions and collects audio through the audio component, thereby uploading the collected answer audio to the server 120. The server 120 scores the answer audio using the trained target scoring model and feeds the scored score back to the test terminal 130.
[0041] Indicatively, Figure 1 As shown, when the question type "topic description" needs to be automatically scored, the server 120 sends the training samples to be annotated to the scoring terminal 110. The scoring terminal 110 displays the sample oral test questions and plays the sample answer audio when receiving a click operation on the audio playback control 111. After the scorer enters the sample score in the scoring box 112 based on the sample answer audio and the sample reference answer, the scoring terminal 110 feeds back the sample score to the server 120. Based on the manually annotated training samples and the pre-trained scoring model, the server 120 trains the target scoring model corresponding to the question type "topic description".
[0042] During the oral test, the test terminal 130 displays an oral test question of the "topic description" type, and upon receiving a click operation on the recording control 131, records the test taker's answer audio, and uploads the answer audio to the server 120. The server 120 scores the answer audio based on the answer audio corresponding to the oral test question and the reference answer using the target scoring model, and feeds the obtained score back to the test terminal 130 for display by the test terminal 130.
[0043] It should be noted that in the above embodiments, the pre-trained scoring model and the target scoring model are trained by the server 120, and the scoring process is performed by the server 120. In other possible implementations, the above models can be trained by the test terminal 130 or the scoring terminal 110, and the models can be deployed on the test terminal 130 side, and the test terminal 130 scores the answer audio locally. This embodiment does not limit this. And for the convenience of description, in the following embodiments, the scoring method of the oral test is performed by a computer device as an example.
[0044] Please refer to Figure 2 , which shows a flow chart of a scoring method for an oral test provided by an exemplary embodiment of the present application.
[0045] Step 201, obtaining a training sample, wherein the training sample includes a sample reference answer, a sample answer audio, and a sample score for the sample answer audio of a target sample oral test question, and the target sample oral test question belongs to a target question type.
[0046] The target question type is a question type with automatic scoring requirements, and the sample scores in the training samples are obtained by manual annotation. Optionally, the sample scores can be based on a 1-point system, a 5-point system, a 10-point system, or a 100-point system, etc., which is not limited in this embodiment.
[0047] In a possible implementation, when receiving an automatic scoring instruction, the computer device obtains a target sample oral test question belonging to the target question type from a database based on the target question type contained in the automatic scoring instruction, and obtains a sample reference answer and a sample answer audio (the audio collected when answering the target sample oral test question) corresponding to the target sample oral test question. If the sample answer audio has not been manually annotated, it is further manually graded to obtain a sample score.
[0048] Since the target scoring model is subsequently trained on the basis of the pre-trained scoring model, compared with zero-based model training, the computer device in the embodiment of the present application only needs to obtain a small number of training samples, which helps to reduce the cost of manual annotation before training. In an illustrative example, when it is necessary to automatically score the question type of "seeing a picture and describing", the computer device obtains a sample oral test question belonging to "seeing a picture and describing", and obtains the sample reference answer of the sample oral test question, 50 sample answer audios, and the sample score of each sample audio.
[0049] Step 202, training the pre-trained scoring model based on the training samples to obtain a target scoring model corresponding to the target question type, and the pre-trained scoring model is obtained by training through meta-learning.
[0050] Optionally, the pre-trained scoring model is pre-trained and deployed by a computer device, or the pre-trained scoring model is trained by other devices and deployed in a computer device, which is not limited in this embodiment.
[0051] In some embodiments, the pre-trained scoring model is trained by meta-learning in a task-based manner. The purpose of meta-learning is to give the model good initialization parameters (i.e., the model has learned prior knowledge during the pre-training process). The initialization parameters may not perform well on the training task, but with the initialization parameters as the starting point, it can quickly adapt to new tasks and improve the model's adaptability to new tasks.
[0052] In the process of training the pre-trained scoring model, the question type corresponding to the adopted task may include the target question type, or may not include the target question type. For example, the pre-trained scoring model is obtained based on the task training corresponding to the three question types of "describe the picture", "quick answer" and "topic description", and the target question type corresponding to the training sample is "describe the picture", or the target question type corresponding to the training sample is "explanation of opinion".
[0053] In some embodiments, the process of training a pre-trained scoring model based on training samples can be called fine tuning, and the computer device adjusts the model parameters of the pre-trained scoring model based on the sample scores in the training samples as supervision, so that the target scoring model obtained after training can quickly adapt to the target question type.
[0054] Step 203, scoring the target answer audio through the target scoring model to obtain a target score for the target answer audio, where the target answer audio is a response to an oral test question of the target question type.
[0055] In one possible implementation, after completing the training of the target scoring model based on the training samples, the computer device verifies the scoring accuracy of the target scoring model using the verification samples, and further uses the target scoring model to score the actual answer audio after the scoring accuracy is verified.
[0056] Optionally, the computer device scores the target answer audio based on the target answer audio and the target reference answer through a target scoring model to obtain a target score.
[0057] To summarize, in the embodiments of the present application, a pre-trained scoring model is pre-trained using a meta-learning method. When it is necessary to score an oral test using a target question type, the pre-trained scoring model is further trained based on the training samples of the target question type to obtain a target scoring model corresponding to the target question type, thereby using the target scoring model to score the answers to the target oral test questions. Since the pre-trained scoring model is trained using a meta-learning method, that is, the pre-trained scoring model has learned prior scoring knowledge in advance, only a small amount of training samples is needed to train the target scoring model, thereby reducing the reliance on manual scoring. After the training is completed, the target scoring model is used to realize automatic scoring of the oral test, thereby improving the scoring efficiency of the oral test.
[0058] When manually scoring, the scorer will comprehensively evaluate the answer from multiple perspectives such as voice, content, and coherence. Therefore, in the embodiment of the present application, before using the scoring model for automatic scoring, it is first necessary to extract multi-dimensional features of the answer audio, and then score based on the extracted features. The specific scoring process is described below using an illustrative embodiment.
[0059] Please refer to Figure 3 , which shows a flow chart of a scoring method for an oral test provided by another exemplary embodiment of the present application.
[0060] Step 301, obtaining a training sample, wherein the training sample includes a sample reference answer, a sample answer audio, and a sample score for the sample answer audio of a target sample oral test question, and the target sample oral test question belongs to a target question type.
[0061] The implementation of this step may refer to step 201, and this embodiment will not be described in detail here.
[0062] Step 302: extract sample text features and sample acoustic features of the sample answer audio.
[0063] In the embodiment of the present application, the computer device extracts features from the answer audio from two dimensions: acoustics and text. In a possible implementation, since it is impossible to directly extract features from the answer audio in audio form, the computer first needs to perform speech recognition on the sample answer audio to obtain a sample answer text, wherein the computer device can use Automatic Speech Recognition (ASR) technology to convert the sample answer audio into a sample answer text.
[0064] Furthermore, the computer device extracts acoustic features from the sample answer audio to obtain sample acoustic features, and extracts text features from the sample answer text to obtain sample text features.
[0065] In some embodiments, the extraction of text features and acoustic features can be performed by a separate feature extraction model (independent of the scoring model) or by a feature extraction module in the scoring model (such as a neural network structure for feature extraction), which is not limited to this embodiment of the present application.
[0066] Optionally, the extracted sample acoustic features include at least one of sample pronunciation accuracy, sample pronunciation fluency, and sample pronunciation rhythm; the extracted sample text features include at least one of sample semantic features, sample keyword features, sample pragmatic features, and sample text fluency features. In the following embodiments, the specific extraction methods of acoustic features and text features will be described in detail.
[0067] Step 303, based on the sample text features and the sample acoustic features, the sample answer audio is scored by a pre-trained scoring model to obtain a predicted score for the sample answer audio.
[0068] In a possible implementation, the computer device performs feature concatenation on the extracted sample text features and sample acoustic features, and inputs the concatenated features into a pre-trained scoring model to obtain a prediction score output by the pre-trained scoring model.
[0069] Since there are differences in the scoring criteria for different question types, in order to better adapt to different scoring criteria, in one possible implementation, the pre-trained scoring model is composed of a deep neural network (DNN) and a rule vector matrix indicating the scoring criteria, wherein the rule vectors in the rule vector matrix support dynamic storage and adjustment.
[0070] The process of scoring the sample answer audio using the pre-trained scoring model is the process of scoring the sample text features and the sample acoustic features according to the scoring criteria indicated by the rule vector in the rule vector matrix.
[0071] Step 304 , training the pre-trained scoring model based on the scoring loss between the predicted score and the sample score to obtain a target scoring model.
[0072] Since the rule vector matrix in the pre-trained scoring model is not adapted to a specific task during the training process, there will be a difference between the predicted score and the sample score when the sample answer audio is scored using the pre-trained scoring model. During the training process, the computer device determines the scoring loss based on the predicted score and the sample score, and adjusts the model parameters of the pre-trained scoring model based on the scoring loss to obtain a target scoring model. Optionally, the model parameters adjusted by the computer device during the training process include the network parameters of the deep neural network and the rule vectors in the rule vector matrix.
[0073] In one possible implementation, when the scoring loss satisfies the convergence condition, or the number of training rounds reaches a preset number of rounds, the computer device determines that the model training is completed and obtains the target scoring model, which is composed of a deep neural network and a target rule vector matrix.
[0074] Step 305: extract target text features and target acoustic features of the target answer audio.
[0075] After completing the model training for the target question type through the above steps 301 to 304, when the target answer audio to be scored is obtained, the computer device first extracts features of the target answer audio to obtain target text features and target acoustic features.
[0076] In one possible implementation, a computer device performs speech recognition on a target answer audio to obtain a target answer text, and then performs text feature extraction on the target answer audio based on the target answer text and the target reference answer to obtain target text features.
[0077] Optionally, the target text feature includes at least one of a target semantic feature, a target keyword feature, a target pragmatic feature, and a target text fluency feature. The extraction process of various features is described below.
[0078] 1. Target semantic features
[0079] In a possible implementation, the computer device extracts semantic features from the target answer text to obtain target semantic features, which may include topic features, term frequency-inverse document frequency (TF-IDF) features, etc., which are not limited in the present application embodiment.
[0080] 2. Target keyword characteristics
[0081] Since the accuracy of the answer content is usually related to keywords, in one possible implementation, a computer device performs keyword extraction on the target answer text and the target reference answer respectively to obtain the first keyword in the target answer text and the second keyword in the target reference answer, thereby determining the target keyword feature of the target answer text based on the matching degree between the first keyword and the second keyword.
[0082] Optionally, the target keyword feature includes at least one of a keyword accuracy rate and a keyword recall rate. The keyword accuracy rate is determined based on the number of recalled keywords (the recalled keywords are keywords that match the first keyword and the second keyword) and the number of first keywords, and the keyword recall rate is determined based on the number of recalled keywords and the number of second keywords. For example, when the number of first keywords extracted is 5, the number of second keywords extracted is 8, and the number of recalled keywords is 4, the computer device determines that the keyword accuracy rate is 0.8 and the keyword recall rate is 0.5.
[0083] 3. Target pragmatic features
[0084] In an oral test, in addition to examining the accuracy of the expression content, it is also necessary to examine the richness and accuracy of the vocabulary, sentence patterns, and grammar used. Therefore, in a possible implementation, the computer device extracts pragmatic features from the target answer text to obtain target pragmatic features, which include at least one of vocabulary diversity, sentence pattern diversity, and grammatical accuracy.
[0085] Optionally, the computer device deduplicates the words used in the target answer text and counts them to obtain the vocabulary usage, thereby determining the vocabulary diversity based on the vocabulary usage and the total vocabulary in the target answer text; the computer device performs sentence recognition on the target answer text and counts the sentence types, thereby determining the sentence diversity based on the number of sentence types; the computer device inputs the target answer text into a pre-trained language analysis model (such as a Tensorflow syntax analysis model), and the language analysis model performs syntax analysis to obtain grammatical accuracy.
[0086] 4. Target text fluency characteristics
[0087] In one possible implementation, a computer device identifies continuously repeated content in a target answer text, such as determining the same words that appear continuously in the same sentence as continuously repeated content, determining repeated sentences that appear adjacent to each other as continuously repeated content, and so on, thereby determining the target text fluency features of the target answer text based on the proportion of the continuously repeated content in the target answer text.
[0088] It should be noted that the embodiments of the present application are only illustrative of the example in which the target text features include the above-mentioned features. In other possible implementations, other features that can characterize the accuracy, completeness, and richness of the text can also be used as target text features to improve the diversity of feature dimensions. This embodiment does not constitute a limitation to this.
[0089] In a possible implementation, the computer device extracts acoustic features from the target answer audio to obtain target acoustic features. Optionally, the target acoustic features include at least one of target pronunciation accuracy, target pronunciation fluency, and target pronunciation rhythm. The extraction process of various features is described below.
[0090] 1. Target pronunciation accuracy
[0091] In a possible implementation, the computer device performs speech recognition on the target answer audio, thereby determining the pronunciation accuracy of the target answer audio based on the Goodness Of Pronunciation (GOP) of the speech recognition result.
[0092] In some embodiments, the computer device performs at least one level of accuracy evaluation on the target answer audio from at least one granularity to obtain the target pronunciation accuracy. Wherein, when the granularity includes phoneme granularity, word granularity and sentence granularity, the at least one level of accuracy evaluation includes at least one of phoneme-level accuracy evaluation, word-level accuracy evaluation and sentence-level accuracy evaluation.
[0093] 2. Target pronunciation fluency
[0094] In a possible implementation, the computer device performs a fluency assessment on the target answer audio to obtain the target pronunciation fluency.
[0095] Since pronunciation fluency is related to speech speed and pause duration, in some embodiments, the computer device determines the target pronunciation fluency based on the average speech speed of the target answer audio, the average pronunciation duration of the pronunciation segments, and the average pause duration between the pronunciation segments. The average speech speed is determined based on the audio duration of the target answer audio and the number of words obtained by speech recognition, and the target pronunciation fluency is positively correlated with the average speech speed, negatively correlated with the average pronunciation duration, and positively correlated with the average pause duration.
[0096] 3. Target pronunciation rhythm
[0097] In a possible implementation, the computer device performs a prosody evaluation on the target answer audio to obtain the target pronunciation prosody.
[0098] In some embodiments, the computer device determines the pronunciation rhythm of the target answer audio, evaluates the correctness of word stress in the sentences in the target answer audio (i.e., determines whether the words that need to be stressed in the sentence are stressed), and evaluates the sentence boundary tone of the sentences in the target answer audio (i.e., determines whether the sentence boundary is reflected by the tone), thereby determining the target pronunciation rhythm based on the evaluation results.
[0099] It should be noted that the embodiment of the present application only uses the example that the target acoustic features include the above-mentioned features for schematic illustration. In other possible implementations, other speech features can also be used as target acoustic features to improve the diversity of feature dimensions. This embodiment does not constitute a limitation to this.
[0100] Step 306, based on the target text features and the target acoustic features, the target answer audio is scored by a target scoring model to obtain a target score for the target answer audio.
[0101] In order to better adapt to different scoring standards, in an illustrative example, Figure 4 As shown, the target scoring model obtained by training is composed of a deep neural network 41 and a target rule vector matrix 42, and an attention mechanism is integrated into the scoring process. The deep neural network 41 is composed of several hidden layers 411 and a fully connected layer 412, and the target rule vector matrix 42 is composed of target rule vectors 421 corresponding to different scoring criteria.
[0102] like Figure 5 As shown, the process of scoring using the target scoring model may include the following steps:
[0103] Step 306A, performing feature concatenation on the target text feature and the target acoustic feature to obtain the target feature.
[0104] For the extracted target text features and target acoustic features, the computer device first concatenates the two to obtain a target feature as a model input, wherein the target feature can be in the form of a feature vector.
[0105] Step 306B: input the target feature into a deep neural network to obtain a first deep feature vector and a second deep feature vector, wherein the depth of the second deep feature vector is greater than the depth of the first deep feature vector.
[0106] Furthermore, the computer device inputs the target feature into a deep neural network, and extracts the deep feature of the target feature from a hidden layer in the deep neural network to obtain a first deep feature vector and a second deep feature vector, wherein the deeper the depth of the deep feature vector, the more abstract the feature represented by the deep feature vector. In the embodiment of the present application, the first deep feature vector can be output by a shallow hidden layer, and the second deep feature vector can be output by a deep hidden layer. The embodiment of the present application does not limit the depth of the deep feature vector.
[0107] Indicatively, Figure 4 As shown, the deep neural network 41 includes a first hidden layer 4111, a second hidden layer 4112 and a third hidden layer 4113. After the computer device inputs the target feature into the deep neural network 41, it obtains the first deep feature vector output by the second hidden layer 4112 and the second deep feature vector output by the third hidden layer 4113.
[0108] Step 306C: Generate a weighted rule vector based on the first depth feature vector and the target rule vector matrix.
[0109] Since the importance of the evaluation criteria indicated by different target rule vectors in the target rule vector matrix is different, when scoring, the computer device needs to determine the rule weights corresponding to each target rule vector, and then determine the weighted rule vector of the fused rule weights.
[0110] In a possible implementation, the computer device first performs attention calculation on the first deep feature vector and the target rule vector matrix based on the attention mechanism to obtain the rule weights corresponding to each target rule vector.
[0111] Among them, the process of determining the rule weight based on the attention mechanism can be expressed by the following formula:
[0112] P = Softmax(f T M)
[0113] Among them, fT represents the transpose of the first deep feature vector, and M is the target regular vector matrix.
[0114] Indicatively, Figure 4 As shown, the size of f is 1×d, and M is composed of k target rule vectors 421 (M=[m 1 ,m 2 ,…,m k ]), the size of each target rule vector 421 is 1×d, and P=[p 1 ,p 2 ,…,p k ], where p 1 +p 2 +…+p k =1.
[0115] Furthermore, the computer device performs weighted summation on the rule weight and the target rule vector (the rule weight and the target rule vector correspond one to one) to obtain a weighted rule vector.
[0116] The process of determining the weighted rule vector based on the rule vector and the target rule vector can be expressed by the following formula:
[0117]
[0118] Among them, O represents the weighted rule vector, m i represents the i-th target rule vector, p i represents the rule weight corresponding to the i-th target rule vector, and k is the number of target rule vectors.
[0119] Indicatively, Figure 4 As shown, the computer device calculates the weighted rule vector O based on each target rule vector 421 in the target rule vector matrix 42 and its corresponding rule weight.
[0120] Step 306D, determining a target score based on the weighted rule vector and the second deep feature vector.
[0121] Furthermore, the computer device uses a fully connected layer of a deep neural network to perform processing (non-linear transformation) based on the weighted rule vector and the second deep feature vector to obtain a target score for the target answer audio.
[0122] In a possible implementation, the computer device performs vector concatenation (concat) on the weighted rule vector and the second deep feature vector to obtain a concatenated vector, and inputs the concatenated vector into a fully connected layer of a deep neural network to obtain a target score output by the fully connected layer. The deep neural network may include at least one fully connected layer, and the embodiment of the present application does not limit the number of fully connected layers.
[0123] Indicatively, Figure 4 As shown, the computer device concatenates the weighted rule vector O and the second deep feature vector output by the third hidden layer 4113, performs nonlinear transformation on the concatenated vector through the fully connected layer 412, and finally outputs the target score.
[0124] In this embodiment, a scoring model structure of a deep neural network + rule vector matrix is adopted, and the rule vector matrix is used to dynamically store and update the scoring rules to adapt to different scoring standards; and in the scoring process, the rule weight of each rule vector is determined based on the attention mechanism, and then a weighted rule vector is obtained by weighted calculation, which helps to improve the accuracy of subsequent scoring.
[0125] In a possible implementation, before training the target scoring model, the computer device first uses a meta-learning method to train a pre-trained scoring model. The training process of the pre-trained scoring model is described below.
[0126] Since the meta-learning process is trained in tasks, the computer device first needs to obtain a set of meta-learning tasks. For oral exam scenarios, the computer device can use a specific type of oral test question, a reference answer to the oral test question, several answer audios, and the annotated scores corresponding to the answer audios as a meta-learning task.
[0127] In an illustrative example, a computer device uses three question types for meta-learning, namely, picture description, quick response, and topic description. Each question type contains 4 oral questions, and each oral question contains 200 answer audios, resulting in a meta-learning task set containing 12 meta-learning tasks.
[0128] After obtaining the meta-learning task preparation, the computer device trains the pre-trained scoring model based on the meta-learning task set.
[0129] In a possible implementation, each meta-learning task is further divided into a training task and a valid task or testing task. Figure 6 As shown, the meta-learning process can include the following steps.
[0130] Step 601: Select a candidate meta-learning task from a set of meta-learning tasks.
[0131] Optionally, during each round of meta-learning, the computer device randomly selects several candidate meta-learning tasks from the meta-learning task set for this round of training.
[0132] Step 602: For each candidate meta-learning task, optimize the global model parameters of the scoring model based on the training tasks in the candidate meta-learning task to obtain the task model parameters corresponding to the candidate meta-learning task.
[0133] Optionally, for each candidate meta-learning task in the current training round, the computer device scores each answer audio in the training task through the scoring model to obtain a predicted score, and based on the loss between the predicted score and the labeled score, the gradient descent algorithm is used to optimize the global model parameters of the scoring model to obtain the task model parameters for the current candidate meta-learning task, that is, the scoring model using the task model parameters is better adapted to the current candidate meta-learning task. The loss of the candidate meta-learning task can be expressed as:
[0134]
[0135] Among them, k is the number of answer audios in the candidate meta-learning task, p i is the predicted score of the scoring model for the i-th answer audio, y i is the annotation score of the i-th answer audio.
[0136] Step 603: Determine the verification loss of the verification task in the candidate meta-learning task based on the scoring model using the task model parameters.
[0137] Optionally, the computer device scores each answer audio in the verification task by using a scoring model with task model parameters to obtain a prediction score, and determines the loss between the prediction score and the labeled score as the verification loss of the current candidate meta-learning task. The calculation process of the verification loss can refer to the above formula.
[0138] Step 604: Optimize the global model parameters based on the verification loss of each candidate meta-learning task to obtain optimized global model parameters.
[0139] Optionally, for each candidate meta-learning task in the current training round, the computer device obtains the verification loss corresponding to each candidate meta-learning task by executing the above steps 603 and 604, and sums the verification losses of different candidate meta-learning tasks, thereby optimizing the global model parameters using gradient descent according to the sum of the verification losses, thereby obtaining the optimized global model parameters.
[0140] Step 605 , when the verification loss converges, the scoring model using the optimized global model parameters is determined as the pre-trained scoring model.
[0141] During the meta-learning process, the computer device detects whether the verification loss converges. If not, the above steps 601 to 604 are repeated (based on the global model parameters optimized in the previous round); if converged, the computer device determines the scoring model using the optimized global model parameters as the pre-trained scoring model.
[0142] In a possible implementation, the computer device may use MAML (Model-Agnostic Meta-Learning) to perform meta-learning to obtain a pre-trained scoring model. The pseudo code of the process is as follows:
[0143]
[0144]
[0145] In order to verify the solution provided by the embodiment of the present application, as shown in Table 1, three types of questions are used for pre-training of meta-learning, namely, picture description, quick response and topic description. Each question type contains 4 questions, and each question contains 50 training data and 150 verification data. After the pre-training scoring model is obtained based on meta-learning training, two test sets are used for rapid adaptation of new tasks, one is the question type "picture description" included in the meta-learning training, and the other is the question type "opinion explanation" not included in the meta-learning training, to test the model's adaptability to new question types.
[0146] Table 1
[0147]
[0148] Based on the above test task data, FT-SVR, FT-BLSTM, MTL-finetune and the solution of this application were used to quickly adapt the task. The adaptation results of the "Picture Description" question type are as follows: Figure 7 As shown in the figure, the adaptation results of the “Opinion Explanation” question type are as follows Figure 8 As shown. Among them, the adaptation effect is represented by three indicators, namely, the ratio of the model prediction score to the actual score ≤0.5, the ratio ≤1, and the Pearson correlation coefficient (PCC). It can be seen that the application scheme can improve the rapid adaptation capability of tasks for both known tasks and new tasks.
[0149] In one possible application scenario, the scoring process of the oral test is as follows: Fig. 9 As shown, the steps are as follows:
[0150] 1) The teacher opens the oral test app, and the scoring terminal displays the oral test questions and plays the students' answer audio;
[0151] 2) The teacher scores the audio responses;
[0152] 3) The oral test APP sends the marked number to the server;
[0153] 4) The server sends the answer audio, reference answer, annotation score and other information to the task fast adaptation module;
[0154] 5) The task rapid adaptation module fine-tunes the pre-trained scoring model to obtain a target scoring model that is adapted to the current question type;
[0155] 6) The student opens the oral test APP, the test terminal displays the oral test questions, and obtains the student's answer;
[0156] 7) The oral test APP sends the answer audio and oral test questions to the server;
[0157] 8) The server stores the answer audio in the database;
[0158] 9) The server reads the answer audio, reference answer and question type from the database and inputs them into the target scoring model corresponding to the question type;
[0159] 10) The target scoring model scores the answer audio;
[0160] 11) The target scoring model returns the score to the server;
[0161] 12) The server returns the scores to the oral test app for students to view.
[0162] Please refer to Fig.10 , which shows a structural block diagram of a scoring device for an oral test provided by an exemplary embodiment of the present application, the device comprising:
[0163] The first acquisition module 1001 is used to acquire a training sample, wherein the training sample includes a sample reference answer, a sample answer audio, and a sample score for the sample answer audio of a target sample oral test question, wherein the target sample oral test question belongs to a target question type;
[0164] A first training module 1002 is used to train a pre-trained scoring model based on the training samples to obtain a target scoring model corresponding to the target question type, wherein the pre-trained scoring model is obtained by training in a meta-learning manner;
[0165] The scoring module 1003 is used to score the target answer audio through the target scoring model to obtain the target score of the target answer audio, where the target answer audio is the answer to the oral test question belonging to the target question type.
[0166] Optionally, the scoring module 1003 includes:
[0167] A first feature extraction unit, used to extract target text features and target acoustic features of the target answer audio;
[0168] The first scoring unit is used to score the target answer audio based on the target text feature and the target acoustic feature through the target scoring model to obtain the target score of the target answer audio.
[0169] Optionally, the target scoring model is composed of a deep neural network and a target rule vector matrix, and the target rule vector matrix is composed of target rule vectors corresponding to different scoring criteria;
[0170] The first scoring unit is used to:
[0171] Performing feature concatenation on the target text feature and the target acoustic feature to obtain a target feature;
[0172] Inputting the target feature into the deep neural network to obtain a first deep feature vector and a second deep feature vector, wherein the depth of the second deep feature vector is greater than the depth of the first deep feature vector;
[0173] Generate a weighted rule vector based on the first depth feature vector and the target rule vector matrix;
[0174] The object score is determined based on the weighted rule vector and the second deep feature vector.
[0175] Optionally, when generating a weighted rule vector based on the first deep feature vector and the target rule vector matrix, the first scoring unit is configured to:
[0176] Performing attention calculation on the first deep feature vector and the target rule vector matrix to obtain rule weights corresponding to each of the target rule vectors;
[0177] The rule weight and the target rule vector are weighted and summed to obtain the weighted rule vector.
[0178] Optionally, when determining the target score based on the weighted rule vector and the second deep feature vector, the first scoring unit is configured to:
[0179] Performing vector concatenation on the weighted rule vector and the second depth feature vector to obtain a concatenated vector;
[0180] The concatenated vector is input into a fully connected layer of the deep neural network to obtain the target score output by the fully connected layer.
[0181] Optionally, the first feature extraction unit is used to:
[0182] Extracting acoustic features of the target answer audio to obtain the target acoustic features;
[0183] Perform speech recognition on the target answer audio to obtain a target answer text; and perform text feature extraction on the target answer audio based on the target answer text and the target reference answer to obtain the target text feature.
[0184] Optionally, when performing acoustic feature extraction on the target answer audio to obtain the target acoustic feature, the first feature extraction unit is used to:
[0185] Performing at least one level of accuracy assessment on the target answer audio to obtain a target pronunciation accuracy, wherein the at least one level of accuracy assessment includes at least one of a phoneme level accuracy assessment, a word level accuracy assessment, and a sentence level accuracy assessment;
[0186] Performing a fluency assessment on the target answer audio to obtain a target pronunciation fluency;
[0187] Performing a prosody evaluation on the target answer audio to obtain a target pronunciation prosody;
[0188] At least one of the target pronunciation accuracy, the target pronunciation fluency, and the target pronunciation prosody is determined as the target acoustic feature.
[0189] Optionally, when performing text feature extraction on the target answer audio based on the target answer text and the target reference answer to obtain the target text feature, the first feature extraction unit is used to:
[0190] Extracting semantic features from the target answer text to obtain target semantic features;
[0191] Extracting a first keyword from the target answer text and a second keyword from the target reference answer; determining a target keyword feature based on a matching degree between the first keyword and the second keyword;
[0192] Extracting pragmatic features from the target answer text to obtain target pragmatic features, wherein the target pragmatic features include at least one of lexical diversity, sentence diversity, and grammatical accuracy;
[0193] Extracting text fluency features of the target answer text to obtain target text fluency features;
[0194] At least one of the target semantic feature, the target keyword feature, the target pragmatic feature, and the target text fluency feature is determined as the target text feature.
[0195] Optionally, the first training module 1002 includes:
[0196] A second feature extraction unit, used to extract sample text features and sample acoustic features of the sample answer audio;
[0197] A second scoring unit is used to score the sample answer audio based on the sample text feature and the sample acoustic feature through the pre-trained scoring model to obtain a predicted score of the sample answer audio;
[0198] A training unit is used to train the pre-trained scoring model based on the scoring loss between the predicted score and the sample score to obtain the target scoring model.
[0199] Optionally, the device further comprises:
[0200] A second acquisition module is used to acquire a meta-learning task set, wherein the meta-learning task set is composed of different meta-learning tasks, each of which contains a reference answer corresponding to the same oral test question, multiple answer audios, and multiple annotated scores;
[0201] The second training module is used to train the pre-trained scoring model based on the meta-learning task set.
[0202] Optionally, each of the meta-learning tasks consists of a training task and a verification task;
[0203] The second training module comprises:
[0204] A task selection unit, configured to select a candidate meta-learning task from the meta-learning task set;
[0205] A first optimization unit is configured to optimize the global model parameters of the scoring model for each of the candidate meta-learning tasks based on the training tasks in the candidate meta-learning tasks to obtain the task model parameters corresponding to the candidate meta-learning tasks;
[0206] a loss determination unit, configured to determine a verification loss of the verification task in the candidate meta-learning task based on a scoring model using the task model parameters;
[0207] A second optimization unit, configured to optimize the global model parameters based on the verification loss of each of the candidate meta-learning tasks to obtain the optimized global model parameters;
[0208] A determination unit is used to determine the scoring model using the optimized global model parameters as the pre-trained scoring model when the verification loss converges.
[0209] To summarize, in the embodiments of the present application, a pre-trained scoring model is pre-trained using a meta-learning method. When it is necessary to score an oral test using a target question type, the pre-trained scoring model is further trained based on the training samples of the target question type to obtain a target scoring model corresponding to the target question type, thereby using the target scoring model to score the answers to the target oral test questions. Since the pre-trained scoring model is trained using a meta-learning method, that is, the pre-trained scoring model has learned prior scoring knowledge in advance, only a small amount of training samples is needed to train the target scoring model, thereby reducing the reliance on manual scoring. After the training is completed, the target scoring model is used to realize automatic scoring of the oral test, thereby improving the scoring efficiency of the oral test.
[0210] Please refer to Fig.11 , which shows a schematic diagram of the structure of a computer device provided by an exemplary embodiment of the present application. Specifically, the computer device 1100 includes a central processing unit (CPU) 1101, a system memory 1104 including a random access memory 1102 and a read-only memory 1103, and a system bus 1105 connecting the system memory 1104 and the central processing unit 1101. The computer device 1100 may also include a basic input / output system (I / O system) 1106 that helps transmit information between various devices in the computer, and a large-capacity storage device 1107 for storing an operating system 1113, application programs 1114 and other program modules 1115.
[0211] In some embodiments, the basic input / output system 1106 may include a display 1108 for displaying information and an input device 1109 such as a mouse and a keyboard for user inputting information. The display 1108 and the input device 1109 are connected to the central processing unit 1101 through an input / output controller 1110 connected to the system bus 1105. The basic input / output system 1106 may also include an input / output controller 1110 for receiving and processing inputs from a plurality of other devices such as a keyboard, a mouse, or an electronic stylus. Similarly, the input / output controller 1110 also provides output to a display screen, a printer, or other types of output devices.
[0212] The mass storage device 1107 is connected to the central processing unit 1101 through a mass storage controller (not shown) connected to the system bus 1105. The mass storage device 1107 and its associated computer readable media provide non-volatile storage for the computer device 1100. That is, the mass storage device 1107 may include a computer readable medium (not shown) such as a hard disk or drive.
[0213] Without loss of generality, the computer-readable medium may include computer storage media and communication media. Computer storage media include volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules or other data. Computer storage media include random access memory (RAM), read-only memory (ROM), flash memory or other solid-state storage technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, cassette, magnetic tape, disk storage or other magnetic storage devices. Of course, those skilled in the art will know that the computer storage medium is not limited to the above. The above-mentioned system memory 1104 and mass storage device 1107 can be collectively referred to as memory.
[0214] The memory stores one or more programs, and the one or more programs are configured to be executed by one or more central processing units 1101. The one or more programs contain instructions for implementing the above-mentioned methods. The central processing unit 1101 executes the one or more programs to implement the methods provided by the above-mentioned various method embodiments.
[0215] According to various embodiments of the present application, the computer device 1100 can also be connected to a remote computer on the network through a network such as the Internet. That is, the computer device 1100 can be connected to the network 1112 through the network interface unit 1111 connected to the system bus 1105, or the network interface unit 1111 can be used to connect to other types of networks or remote computer systems (not shown).
[0216] The memory also includes one or more programs, which are stored in the memory and include steps executed by a computer device in the method provided in the embodiment of the present application.
[0217] An embodiment of the present application also provides a computer-readable storage medium, in which at least one instruction is stored. The at least one instruction is loaded and executed by a processor to implement the scoring method for the oral test described in any of the above embodiments.
[0218] The embodiment of the present application provides a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the scoring method for the oral test described in the above embodiment.
[0219] A person skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware or by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, and the above-mentioned storage medium may be a read-only memory, a disk or an optical disk, etc.
[0220] The above description is only an optional embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. A scoring method for an oral test, characterized in that: The method comprises: Acquire a training sample, wherein the training sample includes a sample reference answer, a sample answer audio, and a sample score for the sample answer audio of a target sample oral test question, wherein the target sample oral test question belongs to a target question type; Based on the training samples, the pre-trained scoring model is trained to obtain a target scoring model corresponding to the target question type, wherein the pre-trained scoring model is trained by meta-learning, and the target scoring model is composed of a deep neural network and a target rule vector matrix, and the target rule vector matrix is composed of target rule vectors corresponding to different scoring criteria; extracting target text features and target acoustic features of a target answer audio, wherein the target answer audio is a response to an oral test question belonging to the target question type; Performing feature concatenation on the target text feature and the target acoustic feature to obtain a target feature; Inputting the target feature into the deep neural network to obtain a first deep feature vector and a second deep feature vector, wherein the depth of the second deep feature vector is greater than the depth of the first deep feature vector; Generate a weighted rule vector based on the first depth feature vector and the target rule vector matrix; Based on the weighted rule vector and the second deep feature vector, an object score is determined.
2. The method according to claim 1, characterized in that: The step of generating a weighted rule vector based on the first depth feature vector and the target rule vector matrix includes: Performing attention calculation on the first deep feature vector and the target rule vector matrix to obtain rule weights corresponding to each of the target rule vectors; The rule weight and the target rule vector are weighted and summed to obtain the weighted rule vector.
3. The method according to claim 1, characterized in that The determining of the target score based on the weighted rule vector and the second deep feature vector comprises: Performing vector concatenation on the weighted rule vector and the second depth feature vector to obtain a concatenated vector; The concatenated vector is input into a fully connected layer of the deep neural network to obtain the target score output by the fully connected layer.
4. The method according to claim 1, characterized in that: The step of extracting target text features and target acoustic features of the target answer audio includes: Extracting acoustic features of the target answer audio to obtain the target acoustic features; Perform speech recognition on the target answer audio to obtain a target answer text; and perform text feature extraction on the target answer audio based on the target answer text and the target reference answer to obtain the target text feature.
5. The method according to claim 4, characterized in that The extracting acoustic features of the target answer audio to obtain the target acoustic features includes: Performing at least one level of accuracy assessment on the target answer audio to obtain a target pronunciation accuracy, wherein the at least one level of accuracy assessment includes at least one of a phoneme level accuracy assessment, a word level accuracy assessment, and a sentence level accuracy assessment; Performing a fluency assessment on the target answer audio to obtain a target pronunciation fluency; Performing a prosody evaluation on the target answer audio to obtain a target pronunciation prosody; At least one of the target pronunciation accuracy, the target pronunciation fluency, and the target pronunciation prosody is determined as the target acoustic feature.
6. The method according to claim 4, characterized in that The step of extracting text features from the target answer audio based on the target answer text and the target reference answer to obtain the target text features includes: Extracting semantic features from the target answer text to obtain target semantic features; Extracting a first keyword from the target answer text and a second keyword from the target reference answer; determining a target keyword feature based on a matching degree between the first keyword and the second keyword; Extracting pragmatic features from the target answer text to obtain target pragmatic features, wherein the target pragmatic features include at least one of lexical diversity, sentence diversity, and grammatical accuracy; Extracting text fluency features of the target answer text to obtain target text fluency features; At least one of the target semantic feature, the target keyword feature, the target pragmatic feature, and the target text fluency feature is determined as the target text feature.
7. The method according to any one of claims 1 to 6, characterized in that: The step of training the pre-trained scoring model based on the training sample to obtain a target scoring model corresponding to the target question type includes: Extracting sample text features and sample acoustic features of the sample answer audio; Based on the sample text features and the sample acoustic features, scoring the sample answer audio through the pre-trained scoring model to obtain a predicted score for the sample answer audio; The pre-trained scoring model is trained based on the scoring loss between the predicted score and the sample score to obtain the target scoring model.
8. The method according to any one of claims 1 to 6, characterized in that: The method further comprises: Obtain a meta-learning task set, wherein the meta-learning task set is composed of different meta-learning tasks, each of which includes a reference answer corresponding to the same oral test question, multiple answer audios, and multiple annotated scores; Based on the meta-learning task set, the pre-trained scoring model is trained.
9. The method according to claim 8, characterized in that Each of the meta-learning tasks consists of a training task and a verification task; The step of training the pre-trained scoring model based on the meta-learning task set includes: Selecting a candidate meta-learning task from the meta-learning task set; For each of the candidate meta-learning tasks, optimizing the global model parameters of the scoring model based on the training tasks in the candidate meta-learning tasks to obtain the task model parameters corresponding to the candidate meta-learning tasks; Determining a verification loss of the verification task in the candidate meta-learning task based on a scoring model using the task model parameters; Optimizing the global model parameters based on the verification loss of each of the candidate meta-learning tasks to obtain the optimized global model parameters; When the verification loss converges, the scoring model using the optimized global model parameters is determined as the pre-trained scoring model.
10. A scoring device for an oral test, characterized in that: The device comprises: A first acquisition module is used to acquire a training sample, wherein the training sample includes a sample reference answer, a sample answer audio, and a sample score for the sample answer audio of a target sample oral test question, wherein the target sample oral test question belongs to a target question type; A first training module is used to train the pre-trained scoring model based on the training samples to obtain a target scoring model corresponding to the target question type, wherein the pre-trained scoring model is obtained by training in a meta-learning manner, and the target scoring model is composed of a deep neural network and a target rule vector matrix, and the target rule vector matrix is composed of target rule vectors corresponding to different scoring criteria; A scoring module is used to extract target text features and target acoustic features of a target answer audio, wherein the target answer audio is an answer to an oral test question belonging to the target question type; perform feature splicing on the target text features and the target acoustic features to obtain target features; input the target features into the deep neural network to obtain a first deep feature vector and a second deep feature vector, wherein the depth of the second deep feature vector is greater than the depth of the first deep feature vector; generate a weighted rule vector based on the first deep feature vector and the target rule vector matrix; and determine a target score based on the weighted rule vector and the second deep feature vector.
11. A computer device, characterized in that: The computer device includes a processor and a memory, wherein the memory stores at least one instruction, and the at least one instruction is loaded and executed by the processor to implement the oral test scoring method according to any one of claims 1 to 9.
12. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores at least one instruction, and the at least one instruction is loaded and executed by the processor to implement the oral test scoring method according to any one of claims 1 to 9.
13. A computer program product, characterized in that The computer program product includes computer instructions, which are stored in a computer-readable storage medium; a processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the oral test scoring method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Full-automatic oral language evaluating management and scoring system and scoring method thereof
CN103151042A