Oral Scoring Model Training Method, Scoring Method, Device and Electronic Equipment

By introducing the first loss value and the second loss value in the oral scoring model training and training the pre-trained scoring model in combination with the meta-learning method, the problem of poor adaptability of the oral scoring model to different question types in the prior art is solved, and more efficient oral scoring model training and scoring accuracy are achieved.

CN115116474BActive Publication Date: 2025-07-01TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210502414.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-09
Publication Date
2025-07-01
Estimated Expiration
2042-05-09

AI Technical Summary

Technical Problem

The existing oral scoring model has poor adaptability to different question types, resulting in inefficient scoring in oral examinations.

Method used

By obtaining training samples, including sample speaking test questions, answer audio and corresponding sample scores, the pre-trained scoring model is used for meta-learning, and the pre-trained scoring model is trained in combination with the first loss value and the second loss value to obtain an oral scoring model suitable for the target question type.

Benefits of technology

The adaptability and scoring accuracy of the oral scoring model to the target type is improved, the number of samples required for the training process is reduced, and the training efficiency is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115116474B_ABST
    Figure CN115116474B_ABST
Patent Text Reader

Abstract

The present application discloses a method for training a spoken language scoring model, a spoken language scoring method, an apparatus, an electronic device, and a storage medium. Among them, the method for training a spoken language scoring model includes: inputting a sample response audio into a pre-trained scoring model trained by a meta-learning method to obtain a predicted score; determining a first loss value according to the sample score and the predicted score; determining a second loss value according to the magnitude relationship between the sample scores corresponding to the determined target response audios; training the pre-trained scoring model according to the first loss value and the second loss value to obtain a spoken language scoring model. In the present application, while training the pre-trained scoring model through the first loss value, the second loss value is introduced to train the pre-trained scoring model, improving the adaptability of the pre-trained scoring model to the target question type. A spoken language scoring model with high scoring ability can also be obtained through fewer training samples, thereby improving the training efficiency of the spoken language scoring model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and more specifically, to a method for training a spoken language scoring model, a spoken language scoring method, an apparatus, an electronic device, and a storage medium. Background Art

[0002] A spoken language test is a test form for examining spoken language ability. The question types of the test questions used include picture description, quick response, topic description, opinion elaboration, and so on. During the spoken language test, after the test taker finishes answering, the scorer will score the answer from the perspectives of pronunciation, grammar, and the accuracy of question answering, so as to obtain the test score.

[0003] In order to improve the scoring efficiency of the spoken language test, a neural network model can be trained based on the training samples of existing question types to obtain a spoken language scoring model. Then, the spoken language scoring model is used to score the answer audio in the spoken language test. However, the spoken language scoring model trained in this way has poor adaptability to different question types. Summary of the Invention

[0004] In view of this, embodiments of the present application provide a method for training a spoken language scoring model, a spoken language scoring method, an apparatus, an electronic device, and a storage medium.

[0005] In a first aspect, an embodiment of the present application provides a method for training a spoken language scoring model. The method includes: obtaining training samples, where the training samples include sample answer audios of sample spoken language test questions and sample scores corresponding to the sample answer audios, the sample spoken language test questions belong to a target question type, and the sample scores are obtained based on a target scoring rule corresponding to the target question type; inputting the sample answer audios into a pre-trained scoring model to obtain predicted scores corresponding to the sample answer audios, where the pre-trained scoring model is trained by a meta-learning method; determining a first loss value according to the sample scores and the predicted scores, where the first loss value represents the loss between the sample scores and the predicted scores; determining target answer audios in the sample answer audios; determining a second loss value according to the magnitude relationship between the sample scores corresponding to each target answer audio, where the second loss value represents the loss between the scoring rule of the pre-trained scoring model itself and the target scoring rule; and training the pre-trained scoring model according to the first loss value and the second loss value to obtain the spoken language scoring model.

[0006] Second aspect, an embodiment of the present application provides an oral English scoring method, and the method includes: obtaining an answer audio to be scored corresponding to a test oral English question, where the test oral English question belongs to a target question type; inputting the answer audio to be scored into an oral English scoring model to obtain an oral English score of the answer audio to be scored predicted by the oral English scoring model, where the oral English scoring model is trained by the oral English scoring model training method described in the first aspect; and outputting the oral English score of the answer audio to be scored.

[0007] Third aspect, an embodiment of the present application provides an oral English scoring model training device, and the device includes: a sample acquisition module, configured to obtain training samples, where the training samples include sample answer audios of sample oral English questions and sample scores corresponding to the sample answer audios, the sample oral English questions belong to a target question type, and the sample scores are obtained based on target scoring rules corresponding to the target question type; a first scoring module, configured to input the sample answer audio into a pre-trained scoring model to obtain a predicted score corresponding to the sample answer audio, where the pre-trained scoring model is trained by a meta-learning method; a first determination module, configured to determine a first loss value according to the sample score and the predicted score, where the first loss value represents the loss between the sample score and the predicted score; a second determination module, configured to determine target answer audios in the sample answer audio; a third determination module, configured to determine a second loss value according to the magnitude relationship between the sample scores corresponding to each of the target answer audios, where the second loss value represents the loss between the scoring rules of the pre-trained scoring model itself and the target scoring rules; and a training module, configured to train the pre-trained scoring model according to the first loss value and the second loss value to obtain the oral English scoring model.

[0008] Optionally, the third determination module is further configured to determine an assignment for each of the target answer audios according to the magnitude relationship between the sample scores corresponding to each of the target answer audios; and determine a second loss value according to the predicted score and the assignment corresponding to each of the target answer audios.

[0009] Optionally, the target answer audio includes two target answer audios; the third determination module is further configured to determine the assignment of the target answer audio with a higher sample score among the two target answer audios as a first value; and determine the assignment of the target answer audio with a lower sample score among the two target answer audios as a second value, where the first value is greater than the second value.

[0010] Optionally, the pre-trained scoring model includes a deep network, a rule vector matrix, and a fully connected layer. The rule vector matrix includes rule vectors corresponding to different scoring rules. The first scoring module is further configured to determine the feature information of the sample response audio; input the feature information into the deep network to obtain the deep features corresponding to the sample response audio; based on the deep features and the rule vector matrix, obtain weighted rule vectors; perform a concatenation operation on the weighted rule vectors and the deep features to obtain a concatenated vector; input the concatenated vector into the fully connected layer to obtain the predicted score of the sample response audio output by the fully connected layer.

[0011] Optionally, the first scoring module is further configured to perform a linear transformation operation on each dimension of the deep features to obtain transformed deep features; perform an activation process on the transformed deep features through an activation function to obtain a proportionality coefficient; based on the proportionality coefficient and the deep features, obtain processed deep features; based on the processed deep features and the rule vector matrix, obtain weighted rule vectors; based on the processed deep features and the rule vector matrix, obtain weighted rule vectors; perform a concatenation operation on the processed deep features and the weighted rule vectors to obtain a concatenated vector.

[0012] Optionally, the first scoring module is further configured to perform an attention calculation on the processed deep features and the rule vector matrix to obtain the rule weights corresponding to each of the rule vectors; based on the rule weights, perform a weighted sum of the multiple rule vectors to obtain the weighted rule vectors.

[0013] Optionally, the training sample further includes the reference answer of the sample oral test question. The first scoring module is further configured to extract acoustic features from the sample response audio to obtain acoustic features; perform speech recognition on the sample response audio to obtain a response text; based on the response text and the reference answer, obtain text features; perform feature concatenation on the acoustic features and the text features to obtain the feature information of the sample response audio.

[0014] Optionally, the first scoring module is further configured to perform at least one-level accuracy evaluation on the sample response audio to obtain pronunciation accuracy, where the at least one-level accuracy evaluation includes at least one of phoneme-level accuracy evaluation, word-level accuracy evaluation, and sentence-level accuracy evaluation; perform a fluency evaluation on the sample response audio to obtain pronunciation fluency; perform a prosody evaluation on the sample response audio to obtain pronunciation prosody; determine at least one of the pronunciation accuracy, the pronunciation fluency, and the pronunciation prosody as the acoustic features.

[0015] Optionally, the first scoring module is further configured to extract semantic features from the response text to obtain semantic features; extract first keywords in the response text and second keywords in the reference answer; determine keyword features based on the matching degree between the first keywords and the second keywords; extract pragmatic features from the response text to obtain pragmatic features, where the pragmatic features include at least one of lexical diversity, syntactic diversity, and grammatical accuracy; extract text fluency features from the response text to obtain text fluency features; and determine at least one of the semantic features, the keyword features, the pragmatic features, and the text fluency features as the text features.

[0016] Optionally, the training module is further configured to calculate the product of the second loss value and a preset parameter to obtain a product result; calculate the sum of the product result and the first loss value as the final loss value; and train the pre-trained scoring model with the final loss value to obtain the spoken language scoring model.

[0017] In a fourth aspect, an embodiment of the present application provides a spoken language scoring device, including: an audio acquisition module configured to acquire a response audio to be scored corresponding to a test spoken language question, where the test spoken language question belongs to a target question type; a second scoring module configured to input the response audio to be scored into a spoken language scoring model to obtain a spoken language score of the response audio to be scored predicted by the spoken language scoring model, where the spoken language scoring model is trained by the spoken language scoring model training method described in the first aspect; and an output module configured to output the spoken language score of the response audio to be scored.

[0018] In a fifth aspect, an embodiment of the present application provides an electronic device, including a processor and a memory; one or more programs are stored in the memory and configured to be executed by the processor to implement the above method.

[0019] In a sixth aspect, an embodiment of the present application provides a computer-readable storage medium, in which program code is stored, where the above method is executed when the program code is run by a processor.

[0020] In a seventh aspect, an embodiment of the present application provides a computer program product or a computer program, the computer program product or the computer program including computer instructions, where the computer instructions are stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the electronic device executes the above method.

[0021] A method for training a spoken language scoring model, a spoken language scoring method, an apparatus, an electronic device, and a storage medium provided by an embodiment of the present application, while training a pre-trained scoring model through a first loss value representing the loss between a sample score and a predicted score, introduces a second loss value representing the loss between the scoring rule of the pre-trained scoring model itself and the target scoring rule to train the pre-trained scoring model, improves the adaptability of the pre-trained scoring model to the target question type, so that a spoken language scoring model with high scoring ability can be obtained with fewer training samples, reduces the number of samples required in the training process, and improves the training efficiency of the spoken language scoring model. At the same time, training the pre-trained scoring model by combining the first loss value and the second loss value also improves the scoring accuracy and rationality of the pre-trained scoring model for the target question type, thereby improving the scoring ability of the spoken language scoring model. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0023] Figure 1 is a schematic diagram of an application scenario shown according to an embodiment of the present application;

[0024] Figure 2 shows a schematic diagram of a scoring interface in a scoring terminal according to an embodiment of the present application;

[0025] Figure 3 shows a schematic diagram of an exam interface in an exam terminal according to an embodiment of the present application;

[0026] Figure 4 shows a flowchart of a method for training a spoken language scoring model provided by an embodiment of the present application;

[0027] Figure 5 shows Figure 4 a flowchart of an implementation manner of step S150 in

[0028] Figure 6 shows Figure 4 a flowchart of an implementation manner of step S120 in

[0029] Figure 7 shows Figure 4 a flowchart of another implementation manner of step S120 in

[0030] Figure 8 shows Figure 6Flowchart of an implementation of step S310;

[0031] Figure 9 Schematic diagram showing the training process of the pre-training scoring model in the embodiments of the present application;

[0032] Figure 10 Flowchart of a spoken language scoring method proposed in an embodiment of the present application;

[0033] Figure 11 Schematic diagram showing the scoring process of a spoken language test in the embodiments of the present application;

[0034] Figure 12 Block diagram showing a spoken language scoring model training device proposed in an embodiment of the present application;

[0035] Figure 13 Block diagram showing a spoken language scoring device proposed in an embodiment of the present application;

[0036] Figure 14 Block diagram showing the structure of an electronic device for executing the spoken language scoring model training method according to the embodiments of the present application. Detailed implementation

[0037] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. According to the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0038] In the following description, the terms "first / second" involved are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first / second" can be interchanged with a specific order or sequence when permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0039] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0040] Artificial Intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making.

[0041] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0042] The key technologies of speech technology include Automatic Speech Recognition (ASR), Text To Speech (TTS), and voiceprint recognition technology. Enabling computers to listen, see, speak, and feel is the future development direction of human-computer interaction, and speech has become one of the promising human-computer interaction methods in the future. The embodiments of this application are the application of speech technology in the oral examination scenario, which is used to train an oral scoring model and, with the help of the trained oral scoring model, automatically score the oral response audio of oral test questions.

[0043] Figure 1 The figure shows a schematic diagram of the implementation environment provided by an exemplary embodiment of this application. This implementation environment includes a scoring terminal 110, a server 120, and an examination terminal 130. Among them, data communication is carried out between the examination terminal 130 and the server 120 through a communication network, and data communication is carried out between the scoring terminal 110 and the server 120 through a communication network. Optionally, the communication network can be a wired network or a wireless network, and this communication network can be at least one of a local area network, a metropolitan area network, and a wide area network.

[0044] The scoring terminal 110 is a terminal used for manual scoring, and this terminal includes, but is not limited to, mobile phones, computers, intelligent voice interaction devices, intelligent home appliances, vehicle-mounted terminals, etc., and the embodiments of this application do not limit this. In some embodiments, this scoring terminal 110 is the terminal used by scorers, and the scorers can be teachers or professionals.

[0045] In a possible implementation, when training a spoken language scoring model for automatically scoring specific question types, the server 120 provides the scoring terminal 110 with training samples to be annotated. The training samples include sample spoken language test questions (belonging to specific question types), sample reference answers, and sample answer audio. The scoring terminal 110 plays the sample answer audio and obtains the sample scores input by the scorers, and then feeds back the sample scores to the server 120.

[0046] For example, when automatically scoring the question type of "topic description", the server 120 sends the training samples to be annotated to the scoring terminal 110, and the scoring terminal outputs a scoring interface. As Figure 2 shown, the scoring interface in the scoring terminal 110 includes the question type ( Figure 2 the question type in is topic description), the spoken language test question, a play answer audio control 201, a score input control 202, and a confirmation control 203. The scoring terminal 110 can play the answer audio when receiving a click operation on the play answer audio control 201. After the scoring terminal 110 receives an operation of inputting a specific score in the score input control 202 and then receives a click operation on the confirmation control 203, the scoring terminal 110 takes the score in the input control 202 as the sample score and sends the obtained sample score to the server 120.

[0047] The server 120 is a device for providing spoken language exam scoring services. It can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.

[0048] In the embodiments of the present application, a pre-trained scoring model trained by using a meta-learning method is set in the server 120. When providing scoring services for a spoken language exam of a specific question type, the server 120 provides the scoring terminal 110 with training samples to be annotated and obtains the scores fed back by the scoring terminal 110, and then adaptively trains the pre-trained scoring model based on the manually annotated training samples to obtain a spoken language scoring model corresponding to the specific question type.

[0049] The exam terminal 130 is a terminal used by spoken language exam takers, including but not limited to mobile phones, computers, intelligent voice interaction devices, intelligent home appliances, vehicle-mounted terminals, etc. The embodiments of the present application do not make any limitations in this regard.

[0050] During the oral exam, the exam terminal 130 displays the oral exam questions and collects audio through the audio component, thereby uploading the collected answer audio to the server 120. The server 120 uses the trained oral scoring model to score the answer audio and feedbacks the scored score to the exam terminal 130.

[0051] For example, during the oral exam, the exam terminal 130 displays the oral exam questions of the "topic description" type sent by the server 120 and outputs the oral exam questions through the exam interface of the exam terminal 130. As Figure 3 shown in a, the exam interface of the exam terminal 130 may include the question type ( Figure 3 the question types in a, 3b, and 3c are topic descriptions), the oral exam questions, and the recording control 301. Before starting the recording, the recording control 301 may output a prompt message "Start recording". When the recording control 301 of the exam terminal 130 receives a click operation, the exam terminal 130 starts recording the audio; during the recording process, the exam interface is as Figure 3 shown in b, the audio control 301 may output a prompt message "Recording ended". When the recording control 301 receives a click operation again, the recording ends, and the answer audio is obtained. The exam terminal 130 sends the answer audio to the server 120, and the server 120 returns the oral score of the answer audio. After the exam terminal 130 receives the oral score, the exam interface is as Figure 3 shown in c, and the exam interface outputs the oral score and the corresponding oral exam questions.

[0052] It should be noted that in the above embodiments, it is described by taking the pre-trained scoring model and the oral scoring model being trained by the server 120 and the scoring process being executed by the server 120 as an example. In other possible implementation manners, the above models may be trained by the exam terminal 130 or the scoring terminal 110, and the models may be deployed on the side of the exam terminal 130, and the exam terminal 130 scores the answer audio locally. This embodiment does not make any limitations in this regard. And for the convenience of description, in the following various embodiments, it is described by taking the scoring method of the oral exam being executed by an electronic device as an example.

[0053] Please refer to Figure 4 , Figure 4 which shows a flowchart of a method for training an oral scoring model provided by an embodiment of the present application. The method can be used in an electronic device (such as Figure 1 the server 120 in), and the method includes:

[0054] S110. Obtain training samples, where the training samples include sample answer audio of sample oral exam questions and sample scores corresponding to the sample answer audio. The sample oral exam questions belong to the target question type, and the sample scores are obtained based on the target scoring rules corresponding to the target question type.

[0055] The training samples include oral test questions for training an oral English scoring model, reference answers to the oral test questions, and response audio for the oral test questions. Among them, the oral test questions for training the oral English scoring model can be used as sample oral test questions, the reference answers to the sample oral test questions can be used as sample reference answers, and the response audio for the sample oral test questions can be used as sample response audio. The sample oral test questions can be English questions, Chinese questions, Russian questions, etc. This application does not limit the language of the sample oral test questions.

[0056] The target question type is a question type with an automatic scoring requirement, and the sample scores in the training samples are obtained by manual annotation of the sample response audio by scorers according to the scoring rules. The scoring rules relied on by the scorers are used as the target scoring rules, and the target scoring rules can be rules formulated by the scorers. Usually, one target question type corresponds to one target scoring rule. For example, when the target question type is a quick response, the corresponding target scoring rule is the quick response scoring rule. Another example is that when the target question type is picture description, the corresponding target scoring rule is the picture description scoring rule.

[0057] Among them, the manually annotated sample scores can adopt a 1-point system, a 5-point system, a 10-point system, a percentage system, etc. This embodiment does not limit this.

[0058] In a possible implementation manner, when an automatic scoring instruction is received, the electronic device obtains the sample oral test questions belonging to the target question type from the database based on the target question type included in the automatic scoring instruction, and obtains the corresponding sample reference answers and sample response audio (audio collected when answering the sample oral test questions) of the sample oral test questions. If the sample response audio has not been manually annotated, it is further handed over to the scorer to score the sample response audio to obtain a sample score.

[0059] S120: Input the sample response audio into a pre-trained scoring model to obtain a predicted score corresponding to the sample response audio. The pre-trained scoring model is trained by a meta-learning method.

[0060] The pre-trained scoring model is pre-trained and deployed by the electronic device, or the pre-trained scoring model is trained by other devices and deployed in the electronic device. This embodiment does not limit this.

[0061] In some embodiments, the pre-trained scoring model is trained in a meta-learning manner on a task-by-task basis. The tasks corresponding to the pre-trained scoring model usually include tasks corresponding to multiple question types. The question types corresponding to the multiple tasks used may or may not include the target question type. For example, the pre-trained scoring model is trained based on the tasks corresponding to three question types: picture description, quick response, and topic description. The target question type of the training sample is picture description, or the target question type of the training sample is opinion elaboration.

[0062] Meta Learning means learning to learn. The purpose of meta-learning is to enable the model to acquire an ability of "learning to learn", so that it can quickly learn new tasks based on the existing "knowledge", and make the model have good initial parameters (that is, the model learns prior knowledge during the pre-training process). These initial parameters may not perform well on the training tasks, but starting from these initial parameters, the model can quickly adapt to new tasks and improve its adaptability to new tasks.

[0063] After the pre-trained scoring model is obtained, the sample response audio corresponding to the sample oral test questions of the target question type is input into the pre-trained scoring model, and the score predicted by the pre-trained scoring model is obtained as the predicted score corresponding to the sample response audio.

[0064] Generally, it is necessary to extract features from the sample response audio to obtain corresponding feature information. The feature information may include acoustic features representing the pronunciation characteristics of the sample response audio and text features representing the response text corresponding to the sample response audio. The response text corresponding to the sample response audio may refer to the text information obtained by performing speech recognition on the sample response audio.

[0065] After the feature information of the sample response audio is determined, the feature information is input into the pre-trained scoring model to obtain the predicted score output by the pre-trained scoring model.

[0066] S130. Determine a first loss value according to the sample score and the predicted score. The first loss value represents the loss between the sample score and the predicted score.

[0067] The loss between the sample score and the predicted score can be determined according to the sample score annotated for the sample response audio and the predicted score of the sample response audio predicted by the pre-trained scoring model, and used as the first loss value.

[0068] Optionally, based on the sample score and the predicted score, the first loss value can be determined according to the mean squared error loss function. The method for solving the first loss value refers to Formula 1, and Formula 1 is as follows:

[0069]

[0070] Among them, L score is the first loss value, n is the number of sample response audios in the training samples, and p i is the predicted score of the i-th sample response audio, and y i is the sample score of the i-th sample response audio.

[0071] S140. Determine a target response audio in the sample response audios.

[0072] S150. Determine a second loss value according to the magnitude relationship between the sample scores corresponding to each of the target response audios, where the second loss value represents the loss between the scoring rule of the pre-trained scoring model itself and the target scoring rule.

[0073] At least two sample response audios can be determined in the sample response audios as the target response audios, and then according to the magnitude relationship between the sample scores corresponding to each of the target response audios, the loss between the scoring rule of the pre-trained scoring model itself and the target scoring rule is determined as the second loss value, and the second loss value can more accurately reflect the loss between the target scoring rule and the scoring rule of the pre-trained scoring model itself.

[0074] The pre-trained scoring model can perform scoring prediction on the sample response audios according to the scoring rule of the pre-trained scoring model itself, and the scoring rule of the pre-trained scoring model itself can refer to the scoring rule learned by the pre-trained scoring model when training the pre-trained scoring model through meta-learning.

[0075] When the pre-trained scoring model is a model obtained through meta-learning based on the tasks of one question type, the scoring rule of the pre-trained scoring model itself is applicable to the response audios corresponding to the oral test questions under this question type. When the pre-trained scoring model is a model obtained through meta-learning based on the tasks of multiple question types, the scoring rule of the pre-trained scoring model itself is applicable to the response audios corresponding to the oral test questions under these multiple question types.

[0076] For example, the pre-trained scoring model is a model obtained through meta-learning based on the task corresponding to picture description. The scoring rule of the pre-trained scoring model itself is applicable to the response audios corresponding to the oral test questions under picture description; another example is that the pre-trained scoring model is a model obtained through meta-learning based on the tasks corresponding to picture description, quick response, and topic description respectively. The scoring rule of the pre-trained scoring model itself is applicable to the response audios corresponding to the oral test questions under the three question types of picture description, quick response, and topic description.

[0077] S160. Train the pre-trained scoring model according to the first loss value and the second loss value to obtain the oral scoring model.

[0078] After obtaining the first loss value and the second loss value, the first loss value and the second loss value can be aggregated to obtain a final loss value, and then the pre-trained scoring model can be trained through the final loss value to obtain the oral scoring model, and the obtained oral scoring model is applicable to the target question type.

[0079] The second loss value can more accurately reflect the loss between the target scoring rule and the scoring rule of the pre-trained scoring model itself. Therefore, the oral scoring model trained according to the final loss value corresponding to the second loss value can be better applicable to the target question type. Therefore, even if a small number of training samples are used, an oral scoring model with good scoring effect can be trained, thus reducing the time for training the oral scoring model and improving the training efficiency of the oral scoring model.

[0080] For example, when automatic scoring of the question type of describing a picture in words is required, the electronic device obtains sample oral test questions belonging to describing a picture in words, and obtains the sample reference answers, 10 sample answer audios of the sample oral test questions, and the sample scores of each sample answer audio.

[0081] Optionally, S160 may include: calculating the product of the second loss value and a preset parameter to obtain a product result; calculating the sum of the product result and the first loss value as the final loss value; training the pre-trained scoring model through the final loss value to obtain the oral scoring model. Among them, the preset parameter may be a value set based on requirements, and the preset parameter may refer to the weight of the second loss value, which is used to balance the influence of the first loss value and the second loss value.

[0082] The calculation method of the final loss value can refer to Formula 2, and Formula 2 is as follows:

[0083] L = L score + γ × L cons (Two)

[0084] Among them, γ is the preset parameter corresponding to the second loss value, L cons is the second loss value, and L is the final loss value.

[0085] In some embodiments, the importance of the first loss value may be stronger than that of the second loss value, and the preset parameter of the second loss value usually takes a value in the interval (0, 1), such as 0.5.

[0086] It can be understood that the target question types can include multiple different target question types. According to the training samples corresponding to different target question types, the pre-trained scoring model can be trained separately to obtain the oral scoring models corresponding to different target question types respectively. For example, according to the training samples corresponding to the two question types of picture description and quick response respectively, two pre-trained scoring models are trained separately to obtain an oral scoring model applicable to picture description and an oral scoring model applicable to quick response.

[0087] This embodiment provides a method for training an oral scoring model. By obtaining training samples, the training samples include sample answer audios of sample oral test questions and sample scores corresponding to the sample answer audios. The sample oral test questions belong to the target question type, and the sample scores are obtained based on the target scoring rules corresponding to the target question type; inputting the sample answer audio into the pre-trained scoring model to obtain a predicted score corresponding to the sample answer audio, and the pre-trained scoring model is trained by the meta-learning method; determining a first loss value according to the sample score and the predicted score, and the first loss value represents the loss between the sample score and the predicted score; determining the target answer audio in the sample answer audio; determining a second loss value according to the magnitude relationship between the sample scores corresponding to each target answer audio, and the second loss value represents the loss between the scoring rules of the pre-trained scoring model itself and the target scoring rules; training the pre-trained scoring model according to the first loss value and the second loss value to obtain the oral scoring model. In this embodiment, while training the pre-trained scoring model through the first loss value representing the loss between the sample score and the predicted score, a second loss value representing the loss between the scoring rules of the pre-trained scoring model itself and the target scoring rules is introduced to train the pre-trained scoring model, improving the adaptability of the pre-trained scoring model to the target question type, enabling the model to quickly adapt to the target question type, so that a high-scoring oral scoring model can be obtained with fewer training samples, reducing the number of samples required in the training process, and improving the training efficiency of the oral scoring model. At the same time, training the pre-trained scoring model by combining the first loss value and the second loss value also improves the scoring accuracy and rationality of the pre-trained scoring model for the target question type, thereby improving the scoring ability of the oral scoring model.

[0088] Please refer to Figure 5 , Figure 5 shows Figure 4 a flowchart of an implementation manner of step S150 in Figure 1 which can be used in an electronic device (such as the server 120 in

[0089] S210. Determine the assignment for each of the target response audios according to the magnitude relationship between the respective sample scores corresponding to each of the target response audios.

[0090] The sample response audios for training the pre-trained scoring model usually include multiple ones. Two sample response audios can be randomly determined from the multiple sample response audios as the target response audios.

[0091] The sample scores of the two determined target response audios are different. Determine the assignment for each of the target response audios according to the magnitude relationship between the respective sample scores corresponding to each of the target response audios.

[0092] For example, S210 may include: determining the assignment of the target response audio with the higher sample score among the two target response audios as the first value; determining the assignment of the target response audio with the lower sample score among the two target response audios as the second value, where the first value is greater than the second value. The first value can be 1 and the second value can be 0.

[0093] For example, the two target response audios are A1 and A2 respectively. The sample score of A1 is 0.91 (a score on a one-point scale), and the sample score of A2 is 0.84. At this time, the assignment of A2 is determined as 0, and the assignment of A1 is determined as 1.

[0094] S220. Determine the second loss value according to the respective predicted scores and assignments corresponding to each of the target response audios.

[0095] After determining the assignments corresponding to each of the target response audios, obtain the predicted scores of the target response audios predicted by the pre-trained scoring model. Determine the second loss value according to the respective predicted scores and assignments corresponding to each of the target response audios.

[0096] It can be to determine the second loss value through the cross-entropy loss function according to the respective predicted scores and assignments corresponding to each of the target response audios. The method for solving the second loss value refers to Equation 3, and Equation 3 is as follows:

[0097]

[0098] Among them, is the assignment corresponding to the i-th target response audio, is the predicted score corresponding to the i-th target response audio, and n is the number of target response audios.

[0099] In some embodiments, the predicted score corresponding to the target response audio can be on a hundred-point scale. It is necessary to normalize the predicted score to obtain a predicted score with a value in the interval (0, 1). This predicted score in the interval (0, 1) is used as the predicted score of the target response audio in Formula 3.

[0100] The scoring criteria corresponding to different question types are different, but the relative quality between two sample response audios is fixed. In this embodiment, the second loss value is determined through the sample scores and assignments of the above two target response audios, so as to model the orderliness of the scores through the second loss value, so that the second loss value can represent the loss between the scoring rules of the pre-trained scoring model itself and the target scoring rules.

[0101] In some embodiments, the second loss value can be determined by using a siamese network. The siamese network is used to measure the similarity between two inputs (two target response audios). The siamese network includes two neural networks, corresponding to two inputs. The two inputs are respectively input into the two neural networks (in this application, the two neural networks can refer to two identical pre-trained scoring models). The two neural networks respectively map the inputs to a new space, obtain the outputs corresponding to the two inputs (the predicted scores corresponding to the two target response audios), and calculate the Loss (loss value) of the two outputs as the second loss value. The second loss value is used to evaluate the similarity between the two inputs.

[0102] In this embodiment, the pre-trained scoring model is trained through the second loss value, which improves the learning efficiency of the pre-trained scoring model for the target scoring rules, improves the adaptability of the pre-trained scoring model to the target question types, and thus improves the training efficiency of the spoken language scoring model.

[0103] Please refer to Figure 6 , Figure 6 which shows Figure 4 a flowchart of an implementation manner of step S120 in Figure 1 . The method can be used in an electronic device (such as

[0104] the server 120 in

[0105] . The feature information of the sample response audio may include acoustic features representing the pronunciation characteristics of the sample response audio and text features representing the response text corresponding to the sample response audio.

[0106] S320. Input the feature information into the deep network to obtain the deep features corresponding to the sample response audio.

[0107] In this embodiment, the pre-trained scoring model may include a deep network, a rule vector matrix, and a fully-connected layer. The rule vector matrix includes rule vectors corresponding to different scoring rules. The scoring model corresponding to the pre-trained scoring model (the scoring model refers to the model with initialized parameters used to obtain the pre-trained scoring model) includes an initialized deep network and an initialized rule vector matrix. The scoring model is trained by meta-learning so that the initialized deep network learns deep representation ability and the initialized rule vector matrix learns different scoring rules, thereby obtaining the pre-trained scoring model.

[0108] Input the feature information of the sample response audio into the deep network of the pre-trained scoring model to obtain the deep features corresponding to the sample response audio output by the deep network.

[0109] S330. Based on the deep features and the rule vector matrix, obtain a weighted rule vector.

[0110] After obtaining the deep features output by the deep network of the pre-trained scoring model, according to the rule vector matrix in the pre-trained scoring model and the deep features output by the deep network of the pre-trained scoring model, obtain a weighted rule vector.

[0111] It may be to determine the weights of each rule vector in the rule vector matrix in the pre-trained scoring model according to the deep features output by the deep network of the pre-trained scoring model, and perform weighted summation on each rule vector according to the weights of each rule vector to obtain a weighted rule vector.

[0112] In some embodiments, the obtaining a weighted rule vector based on the deep features and the rule vector matrix includes: performing attention calculation on the deep features and the rule vector matrix to obtain the rule weights corresponding to each of the rule vectors; performing weighted summation on the multiple rule vectors according to the rule weights to obtain the weighted rule vector.

[0113] Attention calculation can automatically learn and calculate the contribution of input data to output data through an attention mechanism. The process of performing attention calculation according to the deep features and the rule vector matrix can refer to Equation 4, and Equation 4 is as follows:

[0114]

[0115] where M is any rule vector in the rule vector matrix, P is the rule weight corresponding to the rule vector M, is the transpose of the deep feature f1.

[0116] After obtaining the rule weights corresponding to each of the rule vectors, perform weighted summation on the multiple rule vectors to obtain the weighted rule vector.

[0117] S340. Concatenate the weighted rule vector and the depth feature to obtain a concatenated vector.

[0118] S350. Input the concatenated vector into the fully connected layer to obtain the predicted score of the sample response audio output by the fully connected layer.

[0119] After obtaining the weighted rule vector, concatenate the weighted rule vector and the depth feature to obtain a concatenated vector, and then input the concatenated vector into the fully connected layer to obtain the predicted score of the sample response audio output by the fully connected layer. Among them, the activation function of the fully connected layer can be the Sigmoid activation function.

[0120] In this embodiment, the depth feature corresponding to the sample response audio and the weighted rule vector are concatenated to obtain a concatenated vector, and the concatenated vector can accurately reflect the predicted score of the pre-trained scoring model, thereby improving the accuracy of the first loss value and the second loss value.

[0121] Please refer to Figure 7 , Figure 7 shows Figure 4 Another flowchart of step S120 in Figure 1 The method can be used in an electronic device (such as

[0122] S410. Determine the feature information of the sample response audio.

[0123] S420. Input the feature information into the deep network to obtain the depth feature corresponding to the sample response audio.

[0124] Among them, the descriptions of S410 and S420 refer to the descriptions of S310 and S320 above, and will not be elaborated here.

[0125] S430. Perform a linear transformation operation on each dimension of the depth feature to obtain a transformed depth feature; perform an activation process on the transformed depth feature through an activation function to obtain a proportionality coefficient; obtain a processed depth feature according to the proportionality coefficient and the depth feature; obtain a weighted rule vector according to the processed depth feature and the rule vector matrix.

[0126] Among them, the activation function in S430 can be the Sigmoid activation function.

[0127] By performing a linear transformation operation on each dimension of the depth feature, a transformed depth feature is obtained. The transformed depth feature is then activated through an activation function to obtain the respective proportionality coefficients for each dimension. The value of the proportionality coefficient is within the interval (0, 1). Then, each dimension of the depth feature is multiplied by the corresponding proportionality coefficient to obtain the processed depth feature. Through the above processing of the depth feature, the effects of suppressing and activating the depth feature are achieved, making the obtained processed depth feature more accurate.

[0128] The calculation process for obtaining the respective proportionality coefficients for each dimension based on the depth feature can refer to Equation Five, which is as follows:

[0129] A = Sigmoid(ω × f + b) (Five)

[0130] Where f and b are the slope and intercept, respectively, of the linear transformation operation on each dimension of the depth feature, ω is the value of any dimension of the depth feature, and A is the proportionality coefficient corresponding to ω.

[0131] In some embodiments, obtaining the weighted rule vector based on the processed depth feature and the rule vector matrix includes: performing an attention calculation on the processed depth feature and the rule vector matrix to obtain the respective rule weights corresponding to each of the rule vectors; and performing a weighted sum of the multiple rule vectors according to the rule weights to obtain the weighted rule vector.

[0132] Where the process of performing the attention calculation based on the processed depth feature and the rule vector matrix can refer to Equation Six, which is as follows:

[0133]

[0134] Where M is any one of the rule vectors in the rule vector matrix, P is the rule weight corresponding to the rule vector M, is the transpose of the processed depth feature f2.

[0135] After obtaining the respective rule weights corresponding to each of the rule vectors, a weighted sum of the multiple rule vectors is performed to obtain the weighted rule vector.

[0136] S440. Perform a concatenation operation on the processed depth feature and the weighted rule vector to obtain a concatenated vector.

[0137] After obtaining the processed depth feature, the processed depth feature is concatenated with the weighted rule vector to obtain a concatenated vector. This concatenated vector is based on the processed depth feature and more accurately reflects the predicted score of the pre-trained scoring model for the sample response audio.

[0138] S450. Input the splicing vector into the fully connected layer to obtain the predicted score of the sample response audio output by the fully connected layer.

[0139] Among them, the description of S450 refers to the description of S350 and will not be elaborated here.

[0140] In this embodiment, the depth features are suppressed and activated to obtain the processed depth features, so that the splicing vector obtained according to the processed depth features and the weighted rule vector can more accurately reflect the predicted score of the pre-trained scoring model for the sample response audio, improving the accuracy of the predicted score.

[0141] Please refer to Figure 8 , Figure 8 which shows Figure 6 a flowchart of an implementation manner of step S310 in Figure 1 . The method can be used in an electronic device (such as

[0142] S510. Extract acoustic features from the sample response audio to obtain acoustic features.

[0143] Among them, the acoustic features include at least one of pronunciation accuracy, pronunciation fluency, and pronunciation prosody. The extraction processes of various features will be described separately below.

[0144] Perform at least one-level accuracy evaluation on the sample response audio to obtain pronunciation accuracy. The at least one-level accuracy evaluation includes at least one of phoneme-level accuracy evaluation, word-level accuracy evaluation, and sentence-level accuracy evaluation.

[0145] The electronic device performs speech recognition on the response audio, and thus determines the pronunciation accuracy of the response audio based on the pronunciation confidence (Goodness Of Pronunciation, GOP) of the speech recognition result. The electronic device can perform at least one-level accuracy evaluation on the response audio from at least one granularity to obtain pronunciation accuracy. Among them, when the granularity includes phoneme granularity, word granularity, and sentence granularity, the at least one-level accuracy evaluation includes at least one of phoneme-level accuracy evaluation, word-level accuracy evaluation, and sentence-level accuracy evaluation.

[0146] The electronic device performs fluency evaluation on the sample response audio to obtain pronunciation fluency.

[0147] Since pronunciation fluency is related to speech rate and pause duration, in some embodiments, the electronic device determines pronunciation fluency based on the average speech rate of the response audio, the average pronunciation duration of pronunciation segments, and the average pause duration between pronunciation segments. Among them, the average speech rate is determined based on the audio duration of the response audio and the number of words obtained by speech recognition, and there is a positive correlation between pronunciation fluency and the average speech rate, a negative correlation between pronunciation fluency and the average pronunciation duration, and a positive correlation between pronunciation fluency and the average pause duration.

[0148] The electronic device performs a prosody evaluation on the sample response audio to obtain pronunciation prosody.

[0149] The electronic device determines to perform a pronunciation rhythm evaluation on the response audio, evaluates the correctness of word stress in the sentence of the response audio (that is, determines whether the words that need to be stressed in the sentence are stressed), and evaluates the sentence boundary tone of the sentence in the response audio (that is, determines whether the sentence boundary is reflected by the tone), so as to determine the pronunciation prosody based on the evaluation results of each item.

[0150] It should be noted that the embodiments of the present application only use the above features as examples of acoustic features for illustrative purposes. In other possible embodiments, other features that can characterize acoustic accuracy, integrity, and richness can also be used as acoustic features to increase the diversity of feature dimensions, and the present embodiment does not limit this.

[0151] S520. Perform speech recognition on the sample response audio to obtain a response text.

[0152] In this embodiment, the training sample further includes the reference answer corresponding to the sample oral test question. The electronic device performs speech recognition on the response audio to obtain a response text, and then extracts text features from the response audio based on the response text and the reference answer to obtain text features.

[0153] S530. Obtain text features according to the response text and the reference answer.

[0154] Among them, the text features may include at least one of semantic features, keyword features, pragmatic features, and text fluency features. The extraction processes of various features will be described below.

[0155] The electronic device extracts semantic features from the response text to obtain semantic features. Among them, the semantic features may include topic features, term frequency-inverse document frequency (TF-IDF) features, etc., and the embodiments of the present application do not limit this.

[0156] Since the accuracy of the answer content is usually related to keywords, the electronic device can also extract the first keywords in the answer text and the second keywords in the reference answer; based on the matching degree between the first keywords and the second keywords, the keyword feature is determined.

[0157] The keyword feature includes at least one of keyword accuracy and keyword recall. Among them, the keyword accuracy is determined based on the number of recalled keywords (the recalled keywords are the matching keywords in the first keywords and the second keywords) and the number of the first keywords, and the keyword recall is determined based on the number of recalled keywords and the number of the second keywords. For example, when the number of the first keywords extracted is 5, the number of the second keywords extracted is 8, and the number of recalled keywords is 4, the electronic device determines that the keyword accuracy is 0.8 and the keyword recall is 0.5.

[0158] In a spoken language exam, in addition to examining the accuracy of the expressed content, it is also necessary to examine the richness and accuracy of the vocabulary, sentence patterns, and grammar used. Therefore, the electronic device can also extract the pragmatic features from the answer text to obtain the pragmatic features, and the pragmatic features include at least one of vocabulary diversity, sentence pattern diversity, and grammar accuracy.

[0159] The electronic device performs deduplication statistics on the vocabulary used in the answer text to obtain the vocabulary usage amount, and thus determines the vocabulary diversity based on the vocabulary usage amount and the total amount of vocabulary in the answer text; the electronic device identifies the sentence patterns in the answer text and counts the types of sentence patterns, and thus determines the sentence pattern diversity based on the number of sentence pattern types; the electronic device inputs the answer text into a pre-trained language analysis model (such as the TensorFlow grammar analysis model), and the language analysis model performs grammar analysis to obtain the grammar accuracy.

[0160] The electronic device can also extract the text fluency features from the answer text to obtain the text fluency features. The electronic device can identify the continuously repeated content in the answer text, for example, determining the continuously occurring same vocabulary in the same sentence as the continuously repeated content, determining the adjacent repeated sentences as the continuously repeated content, etc., and thus determines the text fluency features of the answer text based on the proportion of the continuously repeated content in the answer text.

[0161] It should be noted that the embodiments of the present application only use the above features as an example for illustrative purposes. In other possible implementation manners, other features that can characterize the text accuracy, integrity, and richness can also be used as text features to increase the diversity of the feature dimensions, and the present embodiment does not limit this.

[0162] S540. Perform feature splicing on the acoustic features and the text features to obtain the feature information of the sample response audio.

[0163] For the extracted text features and acoustic features, first splice the two to obtain the feature information used as the input of the pre-training scoring model. Among them, the feature information can be in the form of a feature vector.

[0164] In this embodiment, the feature information of the sample response audio includes acoustic features and text features. The acoustic features and text features respectively include multiple features. The feature information of the sample response audio can more accurately and comprehensively reflect the specific features of the sample response audio, making the predicted score corresponding to the sample response audio more accurate and reliable.

[0165] To understand this solution more conveniently, the training method of the oral scoring model in the embodiments of the present application will be explained below in combination with a specific scenario.

[0166] Please refer to Figure 9 , Figure 9 which shows a schematic diagram of the training process of the pre-training scoring model in the embodiments of the present application.

[0167] Among them, the pre-training scoring model can include a task-related feature module and a scoring rule module. The task-related feature module includes a deep network and a fully connected layer. The scoring rule module includes a rule vector matrix. The rule vector matrix includes Z rule vectors, and Z is a non-zero integer.

[0168] After obtaining the training samples, determine the feature information of the sample response audio in the training samples, input the feature information of the sample response audio into the deep network to obtain deep features; according to the deep features, determine the proportionality coefficient, and according to the proportionality coefficient and the deep features, obtain the processed deep features.

[0169] According to the processed deep features and Z rule vectors, perform attention calculation to obtain the rule weights of each of the Z rule vectors, and perform weighted summation on the Z rule vectors according to the rule weights of each of the Z rule vectors to obtain a weighted rule vector.

[0170] Splice the weighted rule vector and the processed deep features to obtain a spliced vector, then input the spliced vector into the fully connected layer to obtain the predicted score of the sample response audio, and determine the first loss value according to the predicted score of the sample response audio and the sample score.

[0171] Two target response audios can be determined in the sample response audio, and the feature information of the two target response audios is respectively input into two neural networks in the siamese network to obtain the predicted scores output by each of the two neural networks. The two neural networks in the siamese network can be the same network model as the pre-training scoring model.

[0172] Determine the assignment values for the two target response audios according to the magnitude relationship between the sample scores of the two target response audios respectively, and determine the second loss value according to the predicted scores and the corresponding assignment values of the two target response audios.

[0173] Calculate the final loss value through the first loss value and the second loss value, and train the pre-trained scoring model with the final loss value to obtain the spoken language scoring model. Training the pre-trained scoring model with the final loss value may refer to adjusting the rule vector matrix and the parameters of the deep network in the pre-trained scoring model.

[0174] The process of training the pre-trained scoring model based on the training samples can be referred to as fine tuning, and the electronic device adjusts the model parameters of the pre-trained scoring model with the final loss value corresponding to the training samples, so that the obtained spoken language scoring model can quickly adapt to the target question type.

[0175] Please refer to Figure 10 , Figure 10 which shows a flowchart of a spoken language scoring method proposed in an embodiment of the present application. The method can be used in an electronic device (the electronic device can be Figure 1 the server 120 in

[0176] S610. Obtain the response audio to be scored corresponding to the test spoken language question, where the test spoken language question belongs to the target question type.

[0177] The test spoken language question may refer to a spoken language question used for a spoken language test. When the electronic device is a server, the test spoken language question may be sent to the examination terminal by the server. The examination terminal outputs the test spoken language question, and the candidate records the response audio for the test spoken language question through the examination terminal. This response audio is used as the response audio to be scored, and the response audio to be scored is sent to the server through the examination terminal.

[0178] Since the trained spoken language scoring model is based on the training samples of the target question type, the obtained test spoken language question can belong to the target question type, so as to ensure a relatively high accuracy of the spoken language score of the response audio to be scored corresponding to the test spoken language question predicted by the spoken language scoring model.

[0179] S620. Input the response audio to be scored into the spoken language scoring model to obtain the spoken language score of the response audio to be scored predicted by the spoken language scoring model, where the spoken language scoring model is trained by the spoken language scoring model training method described in any of the above embodiments.

[0180] The spoken language scoring model can be trained by the spoken language scoring model training method described in any of the above embodiments, which will not be elaborated here.

[0181] Since the pre-trained scoring model includes a deep network, a fully connected layer, and a rule vector matrix, the trained spoken language scoring model also includes a deep network, a fully connected layer, and a rule vector matrix. However, the parameters of the rule vector matrix and the deep network of the spoken language scoring model are different from those of the pre-trained scoring model.

[0182] The input of the spoken language scoring model is a feature vector. Therefore, it is necessary to determine the features of the audio of the response to be scored to obtain the feature information of the audio of the response to be scored. The feature information of the audio of the response to be scored may include acoustic features representing the pronunciation characteristics of the audio of the response to be scored and text features corresponding to the audio of the response to be scored.

[0183] Among them, the method for determining the feature information of the audio of the response to be scored refers to the method for determining the audio of the sample response above and will not be elaborated here.

[0184] Input the feature information of the audio of the response to be scored into the deep network of the spoken language scoring model to obtain the deep features to be scored, which serve as the deep features to be scored; determine a new proportionality coefficient according to the deep features to be scored, and obtain the processed deep features to be scored according to the new proportionality coefficient and the deep features to be scored.

[0185] Perform attention calculation based on the processed deep features to be scored and each rule vector in the rule vector matrix of the spoken language scoring model to obtain the respective weights of each rule vector, and perform weighted summation on each rule vector according to the respective weights of each rule vector to obtain a new weighted rule vector.

[0186] Concatenate the new weighted rule vector and the processed deep features to be scored to obtain a new concatenated vector, and then input the new concatenated vector into the fully connected layer of the spoken language scoring model to obtain the spoken language score of the audio of the response to be scored.

[0187] S630. Output the spoken language score of the audio of the response to be scored.

[0188] After obtaining the spoken language score of the audio of the response to be scored, the electronic device can output the spoken language score of the audio of the response to be scored.

[0189] In some embodiments, when the electronic device is a server, the server receives the audio of the response to be scored sent by the examination terminal, the server scores the audio of the response to be scored through the spoken language scoring model to obtain the corresponding spoken language score, and the server returns the spoken language score of the audio of the response to be scored to the examination terminal, and the examination terminal outputs the spoken language score of the audio of the response to be scored.

[0190] It can be understood that the oral score of the response audio to be scored output by the oral scoring model can be a normalized score (a score with a value in the range (0, 1)), and this oral score can be processed to obtain the corresponding actual score, and the actual score can be in a percentage system or a ten-point system.

[0191] In a possible implementation manner, before training the oral scoring model, the electronic device first trains a pre-training scoring model by using a meta-learning method. The training process of the pre-training scoring model is described below.

[0192] Since the meta-learning process is trained in units of tasks, the electronic device first needs to obtain a meta-learning task set. For the oral examination scenario, the electronic device can take a specific type of oral test question, the reference answer to the oral test question, several response audios, and the sample score corresponding to the response audio as a meta-learning task.

[0193] In a schematic example, the electronic device uses three types of questions for meta-learning, namely picture description, quick response, and topic description. Each type of question contains 4 oral test questions, and each oral test question contains 200 response audios, obtaining a meta-learning task set containing 12 meta-learning tasks.

[0194] After obtaining the meta-learning tasks, the electronic device trains the pre-training scoring model based on the meta-learning task set.

[0195] In a possible implementation manner, each meta-learning task is further divided into a training task and a validation task (valid task or testing task). The meta-learning process may include: selecting candidate meta-learning tasks from the meta-learning task set; for each candidate meta-learning task, optimizing the global model parameters of the scoring model based on the training tasks in the candidate meta-learning task to obtain the task model parameters corresponding to the candidate meta-learning task; determining the validation loss of the validation task in the candidate meta-learning task based on the scoring model using the task model parameters; optimizing the global model parameters based on the validation losses of each candidate meta-learning task to obtain the optimized global model parameters; and when the validation loss converges, determining the scoring model using the optimized global model parameters as the pre-training scoring model.

[0196] In each round of the meta-learning process, the electronic device randomly selects several candidate meta-learning tasks from the meta-learning task set for training in this round.

[0197] For each candidate meta-learning task in the current training round, the electronic device scores each response audio in the training task through a scoring model to obtain a predicted score, and based on the loss between the predicted score and the corresponding sample score, uses the gradient descent algorithm to optimize the global model parameters of the scoring model, obtaining the task model parameters for the current candidate meta-learning task, that is, the scoring model using these task model parameters is better adapted to the current candidate meta-learning task.

[0198] The solution algorithm for the loss value of the candidate meta-learning task can be the mean squared error loss function. The solution of the loss value of the candidate meta-learning task refers to Formula Seven, and Formula Seven is as follows:

[0199]

[0200] Where, L h is the loss value of the candidate meta-learning task, k is the number of response audios in the candidate meta-learning task, is the predicted score of the scoring model for the i-th response audio, is the sample score (manually labeled score) of the i-th response audio.

[0201] The electronic device scores each response audio in the validation task through the scoring model using the task model parameters to obtain a predicted score, and determines the loss between the predicted score and the sample score as the validation loss of the current candidate meta-learning task. Among them, the calculation process of the validation loss can refer to the above Formula Seven.

[0202] For each candidate meta-learning task in the current training round, the electronic device obtains the validation loss corresponding to each candidate meta-learning task by executing the above method, and sums the validation losses of different candidate meta-learning tasks, and then optimizes the global model parameters using gradient descent according to the sum of the validation losses, thereby obtaining the optimized global model parameters.

[0203] During the meta-learning process, the electronic device detects whether the validation loss converges. If it does not converge, the above training steps are repeated (based on the globally optimized model parameters in the previous round); if it converges, the electronic device determines the scoring model using the optimized global model parameters as the pre-trained scoring model.

[0204] In a possible implementation, the electronic device can use MAML (Model-Agnostic Meta-Learning) for meta-learning to obtain the pre-trained scoring model. The pseudo-code for this process is as follows:

[0205]

[0206] To verify the solution provided by the embodiments of the present application, as shown in Table 1, pre-training of meta-learning is performed using 3 types of questions, namely picture description, quick response, and topic description. Each type of question contains 4 questions, and each question contains 50 pieces of training data and 150 pieces of validation data. When training the oral scoring model based on the pre-training scoring model obtained through meta-learning, two test sets are used. One is the picture description question type included in the meta-learning training, and the other is the opinion elaboration question type not included in the meta-learning training, to test the adaptation ability of the oral scoring model to new question types. Table 1 is as follows:

[0207] Table 1

[0208]

[0209] Based on the above test task data, rapid adaptation training for new tasks is respectively performed based on SVR (Support Vector Regression, Support Vector Regression), BLSTM (Bidirectional Long Short-Term Memory), MTL pre-train, MAML, and the oral scoring model. Among them, the oral scoring model refers to the oral scoring model trained according to the oral scoring model training method described in any of the above embodiments, and MTL pre-train is the pre-training scoring model obtained through training by meta-learning.

[0210] The test results of the picture description question type are shown in Chart 2. Table 2 is as follows:

[0211] Table 2

[0212] Model Difference ≤ 0.5 Difference ≤ 1 PCC (%) FT-SVR 60.5 87.3 50.8 FT-BLSTM 63.2 87.5 51.5 MTL pre-train 64.1 89.6 52.3 MAML 66.5 90.7 54.5 Oral scoring model 70.8 93.4 58.2

[0213] The test results of the opinion elaboration question type are shown in Table 3. Table 3 is as follows:

[0214] Table 3

[0215] Model Difference ≤ 0.5 Difference ≤ 1 PCC (%) FT-SVR 55.1 85.3 49.8 FT-BLSTM 58.4 87.6 51.6 MTL pre-train 60.2 88.3 53.1 MAML 63.6 89.5 56.8 Oral scoring model 67.3 92.9 59.2

[0216] Among them, FT-SVR is the model obtained by performing rapid adaptation training for new tasks based on SVR, and FT-BLSTM is the model obtained by performing rapid adaptation training for new tasks based on BLSTM.

[0217] The test results are represented by three metrics, namely the proportion of score differences ≤ 0.5 levels, the proportion of score differences ≤ 1 level, and the Pearson correlation coefficient (PCC, which is used to measure the linear correlation between two variables X and Y, and its value ranges from -1 to 1). Among them, the score difference refers to the difference between the predicted score of the scoring model and the sample score of the actual manual scoring. It can be seen that by using the oral scoring model of the present application, the task can be quickly adapted with high ability for both known tasks and new tasks.

[0218] When the number of training samples is different, the Pearson correlation coefficients of different models are shown in Table 4 as follows:

[0219] Table 4

[0220] Model 0 10 20 50 FT-SVR 49.8 51.1 53.5 61.7 FT-BLSTM 51.6 52.7 55.4 64.2 MTL pre-train 53.1 56.4 60.6 73.8 MAML 56.8 61.5 63.7 78.2 Oral scoring model 59.2 63.1 65.2 79.0

[0221] It can be seen that under the condition of a certain number of samples, the oral scoring model in the present application has better effects. At the same time, under the condition of extremely few samples (such as 10), the oral scoring model of the present application has good performance.

[0222] In a possible application scenario, the scoring process of the oral examination is as Figure 11 shown, and the steps are as follows:

[0223] 1) The teacher opens the oral examination APP, and the scoring terminal displays the oral examination questions and plays the audio of the student's answer.

[0224] 2) The teacher scores the answer audio.

[0225] 3) The oral examination APP sends the marked score (the sample score corresponding to the answer audio) to the server.

[0226] 4) The server sends information such as the answer audio, reference answer, and marked score to the task quick adaptation module.

[0227] 5) The task quick adaptation module fine-tunes the pre-trained scoring model to obtain an oral scoring model adapted to the current question type.

[0228] 6) The student opens the oral examination APP, and the examination terminal displays the oral examination questions and obtains the student's answer.

[0229] 7) The oral examination APP sends the answer audio and the oral examination questions to the server.

[0230] 8) The server stores the answer audio in the database.

[0231] 9) The server reads the answer audio, reference answer, and question type from the database and inputs them into the oral scoring model corresponding to the question type.

[0232] 10) The oral language scoring model scores the response audio;

[0233] 11) The oral language scoring model returns the score (the oral language score predicted by the oral language scoring model) to the server;

[0234] 12) The server returns the score to the oral exam APP for the student to view.

[0235] Please refer to Figure 12 , Figure 12 FIG. shows a block diagram of an oral language scoring model training device proposed in an embodiment of the present application. The device 700 includes:

[0236] A sample acquisition module 710, configured to acquire training samples, where the training samples include sample response audio of sample oral questions and sample scores corresponding to the sample response audio. The sample oral questions belong to a target question type, and the sample scores are obtained based on a target scoring rule corresponding to the target question type;

[0237] A first scoring module 720, configured to input the sample response audio into a pre-trained scoring model to obtain a predicted score corresponding to the sample response audio. The pre-trained scoring model is trained by a meta-learning method;

[0238] A first determination module 730, configured to determine a first loss value according to the sample score and the predicted score. The first loss value represents the loss between the sample score and the predicted score;

[0239] A second determination module 740, configured to determine target response audio in the sample response audio;

[0240] A third determination module 750, configured to determine a second loss value according to the magnitude relationship between the sample scores corresponding to each target response audio. The second loss value represents the loss between the scoring rule of the pre-trained scoring model itself and the target scoring rule;

[0241] A training module 760, configured to train the pre-trained scoring model according to the first loss value and the second loss value to obtain the oral language scoring model.

[0242] Optionally, the third determination module 750 is further configured to determine an assignment for each target response audio according to the magnitude relationship between the sample scores corresponding to each target response audio; and determine a second loss value according to the predicted score and the assignment corresponding to each target response audio.

[0243] Optionally, the target response audio includes two target response audios; the third determination module 750 is further configured to determine the assignment value of the target response audio with a higher sample score among the two target response audios as a first value; and determine the assignment value of the target response audio with a lower sample score among the two target response audios as a second value, where the first value is greater than the second value.

[0244] Optionally, the pre-trained scoring model includes a deep network, a rule vector matrix, and a fully connected layer. The rule vector matrix includes rule vectors corresponding to different scoring rules. The first scoring module 730 is further configured to determine the feature information of the sample response audio; input the feature information into the deep network to obtain the deep features corresponding to the sample response audio; based on the deep features and the rule vector matrix, obtain weighted rule vectors; perform a concatenation operation on the weighted rule vectors and the deep features to obtain a concatenated vector; and input the concatenated vector into the fully connected layer to obtain the predicted score of the sample response audio output by the fully connected layer.

[0245] Optionally, the first scoring module 730 is further configured to perform a linear transformation operation on each dimension of the deep features to obtain transformed deep features; perform an activation process on the transformed deep features through an activation function to obtain a proportionality coefficient; based on the proportionality coefficient and the deep features, obtain processed deep features; based on the processed deep features and the rule vector matrix, obtain weighted rule vectors; based on the processed deep features and the rule vector matrix, obtain weighted rule vectors; perform a concatenation operation on the processed deep features and the weighted rule vectors to obtain a concatenated vector.

[0246] Optionally, the first scoring module 730 is further configured to perform an attention calculation on the processed deep features and the rule vector matrix to obtain the rule weights corresponding to the respective rule vectors; and perform a weighted sum of the multiple rule vectors according to the rule weights to obtain the weighted rule vectors.

[0247] Optionally, the training sample further includes the reference answer to the sample oral test question. The first scoring module 730 is further configured to extract acoustic features from the sample response audio to obtain acoustic features; perform speech recognition on the sample response audio to obtain a response text; based on the response text and the reference answer, obtain text features; and perform feature concatenation on the acoustic features and the text features to obtain the feature information of the sample response audio.

[0248] Optionally, the first scoring module 730 is further configured to perform at least one-level accuracy evaluation on the sample answer audio to obtain pronunciation accuracy, where the at least one-level accuracy evaluation includes at least one of phoneme-level accuracy evaluation, word-level accuracy evaluation, and sentence-level accuracy evaluation; perform fluency evaluation on the sample answer audio to obtain pronunciation fluency; perform prosody evaluation on the sample answer audio to obtain pronunciation prosody; and determine at least one of the pronunciation accuracy, the pronunciation fluency, and the pronunciation prosody as the acoustic feature.

[0249] Optionally, the first scoring module 730 is further configured to extract semantic features from the answer text to obtain semantic features; extract the first keywords in the answer text and the second keywords in the reference answer; determine keyword features based on the matching degree between the first keywords and the second keywords; extract pragmatic features from the answer text to obtain pragmatic features, where the pragmatic features include at least one of lexical diversity, syntactic diversity, and grammatical accuracy; extract text fluency features from the answer text to obtain text fluency features; and determine at least one of the semantic features, the keyword features, the pragmatic features, and the text fluency features as the text features.

[0250] Optionally, the training module 760 is further configured to calculate the product of the second loss value and a preset parameter to obtain a product result; calculate the sum of the product result and the first loss value as the final loss value; and train the pre-trained scoring model with the final loss value to obtain the oral scoring model.

[0251] Please refer to Figure 13 , Figure 13 , which shows a block diagram of an oral scoring device proposed in an embodiment of the present application. The device 800 includes:

[0252] An audio acquisition module 810, configured to acquire an answer audio to be scored corresponding to a test oral question, where the test oral question belongs to a target question type;

[0253] A second scoring module 820, configured to input the answer audio to be scored into the oral scoring model to obtain an oral score of the answer audio to be scored predicted by the oral scoring model, where the oral scoring model is trained by the oral scoring model training method described in any of the above embodiments;

[0254] An output module 830, configured to output the oral score of the answer audio to be scored.

[0255] It should be noted that the device embodiments in this application correspond to the foregoing method embodiments. For the specific principles in the device embodiments, reference may be made to the content in the foregoing method embodiments, which will not be elaborated here.

[0256] Figure 14 FIG. shows a structural block diagram of an electronic device for performing the method for training a spoken language scoring model according to an embodiment of the present application. The electronic device may be Figure 1 a server, etc. It should be noted that Figure 14 the computer system 1200 of the electronic device shown is only an example, and should not impose any limitations on the functions and usage scope of the embodiments of the present application.

[0257] As Figure 14 shown, the computer system 1200 includes a central processing unit (CPU) 1201, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1202 or a program loaded from a storage section 1208 into a random access memory (RAM) 1203, such as executing the method in the foregoing embodiments. In the RAM 1203, various programs and data required for system operations are also stored. The CPU 1201, ROM 1202, and RAM 1203 are connected to each other via a bus 1204. An input / output (I / O) interface 1205 is also connected to the bus 1204.

[0258] The following components are connected to the I / O interface 1205: an input section 1206 including a keyboard, a mouse, etc.; an output section 1207 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 1208 including a hard disk, etc.; and a communication section 1209 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 1209 performs communication processing via a network such as the Internet. A drive 1210 is also connected to the I / O interface 1205 as required. A removable medium 1211, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1210 as required, so that a computer program read from it can be installed into the storage section 1208 as required.

[0259] In particular, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product that includes a computer program carried on a computer-readable medium, and the computer program includes program code for performing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through the communication section 1209 and / or installed from the removable medium 1211. When the computer program is executed by the central processing unit (CPU) 1201, various functions defined in the system of the present application are performed.

[0260] It should be noted that the computer-readable medium shown in the embodiments of the present application can be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. And in the present application, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any suitable medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0261] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. Among them, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above-mentioned module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, as well as the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0262] The units involved in the embodiments described in the present application can be implemented in software or in hardware, and the described units can also be provided in a processor. Among them, the names of these units do not, in some cases, constitute a limitation on the unit itself.

[0263] As another aspect, the present application also provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiments; or may exist separately without being assembled into the electronic device. The above computer-readable storage medium carries computer-readable instructions, and when the computer-readable storage instructions are executed by a processor, the method in any of the above embodiments is implemented.

[0264] According to one aspect of the embodiments of the present application, there is provided a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of the electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the electronic device executes the method in any of the above embodiments.

[0265] It should be noted that although several modules or units of the devices for performing actions are mentioned in the above detailed description, such a division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0266] From the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (such as a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the embodiments of the present application.

[0267] After considering the specification and practicing the embodiments disclosed herein, those skilled in the art will readily conceive of other embodiments of the present application. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include known common knowledge or conventional technical means in the technical field not disclosed in the present application. It should be understood that the present application is not limited to the precise structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.

[0268] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for training an oral language scoring model, characterized in that, The method includes: Obtaining training samples, where the training samples include sample response audios of sample oral test questions and sample scores corresponding to the sample response audios. The sample oral test questions belong to a target question type, and the sample scores are obtained based on a target scoring rule corresponding to the target question type; Inputting the sample response audio into a pre-trained scoring model to obtain a predicted score corresponding to the sample response audio. The pre-trained scoring model is trained through a meta-learning method; Determining a first loss value according to the sample score and the predicted score. The first loss value represents the loss between the sample score and the predicted score; Determining a target response audio in the sample response audio; Determining a second loss value according to the magnitude relationship between the sample scores corresponding to each target response audio. The second loss value represents the loss between the scoring rule of the pre-trained scoring model itself and the target scoring rule; Training the pre-trained scoring model according to the first loss value and the second loss value to obtain the oral scoring model.

2. The method according to claim 1, wherein The determining the second loss value according to the magnitude relationship between the sample scores corresponding to each target response audio includes: Determining the assignment of each target response audio according to the magnitude relationship between the sample scores corresponding to each target response audio; Determining the second loss value according to the predicted score and the assignment corresponding to each target response audio.

3. The method according to claim 2, wherein The target response audio includes two target response audios. The determining the assignment of each target response audio according to the magnitude relationship between the sample scores corresponding to each target response audio includes: Determining the assignment of the target response audio with a higher sample score among the two target response audios as a first value; Determining the assignment of the target response audio with a lower sample score among the two target response audios as a second value, where the first value is greater than the second value.

4. The method according to claim 1, characterized in that The pre-trained scoring model includes a deep network, a rule vector matrix, and a fully connected layer. The rule vector matrix includes rule vectors corresponding to different scoring rules. The inputting the sample response audio into the pre-trained scoring model to obtain a predicted score corresponding to the sample response audio includes: Determining the feature information of the sample response audio; Inputting the feature information into the deep network to obtain a deep feature corresponding to the sample response audio; Obtaining a weighted rule vector based on the deep feature and the rule vector matrix; Performing a concatenation operation on the weighted rule vector and the deep feature to obtain a concatenated vector; Inputting the concatenated vector into the fully connected layer to obtain the predicted score of the sample response audio output by the fully connected layer.

5. The method according to claim 4, characterized in that, The obtaining a weighted rule vector based on the deep feature and the rule vector matrix includes: Performing a linear transformation operation on each dimension of the deep feature to obtain a transformed deep feature; Performing an activation process on the transformed deep feature through an activation function to obtain a proportionality coefficient; Obtaining a processed deep feature according to the proportionality coefficient and the deep feature; Based on the processed depth features and the rule vector matrix, a weighted rule vector is obtained; The concatenation operation on the weighted rule vector and the depth features to obtain a concatenated vector includes: Performing a concatenation operation on the processed depth features and the weighted rule vector to obtain a concatenated vector.

6. The method according to claim 5, characterized in that, The obtaining of the weighted rule vector based on the processed depth features and the rule vector matrix includes: Performing an attention calculation on the processed depth features and the rule vector matrix to obtain the rule weights corresponding to each of the rule vectors; According to the rule weights, performing a weighted sum on multiple rule vectors to obtain the weighted rule vector.

7. The method according to claim 4, wherein The training sample further includes the reference answer to the sample oral test question; the determining of the feature information of the sample answer audio includes: Performing acoustic feature extraction on the sample answer audio to obtain acoustic features; Performing speech recognition on the sample answer audio to obtain an answer text; According to the answer text and the reference answer, obtaining text features; Performing feature concatenation on the acoustic features and the text features to obtain the feature information of the sample answer audio.

8. The method according to claim 7, characterized in that, The performing of acoustic feature extraction on the sample answer audio to obtain acoustic features includes: Performing at least one-level accuracy evaluation on the sample answer audio to obtain pronunciation accuracy, where the at least one-level accuracy evaluation includes at least one of phoneme-level accuracy evaluation, word-level accuracy evaluation, and sentence-level accuracy evaluation; Performing fluency evaluation on the sample answer audio to obtain pronunciation fluency; Performing prosody evaluation on the sample answer audio to obtain pronunciation prosody; Determining at least one of the pronunciation accuracy, the pronunciation fluency, and the pronunciation prosody as the acoustic features.

9. The method according to claim 8, wherein The obtaining of text features according to the answer text and the reference answer includes: Performing semantic feature extraction on the answer text to obtain semantic features; Extracting a first keyword in the answer text and a second keyword in the reference answer; Based on the matching degree between the first keyword and the second keyword, determining keyword features; Performing pragmatic feature extraction on the answer text to obtain pragmatic features, where the pragmatic features include at least one of lexical diversity, syntactic diversity, and grammatical accuracy; Performing text fluency feature extraction on the answer text to obtain text fluency features; Determining at least one of the semantic features, the keyword features, the pragmatic features, and the text fluency features as the text features.

10. The method according to claim 1, characterized in that The training of the pre-trained scoring model according to the first loss value and the second loss value to obtain the oral scoring model includes: Calculating the product of the second loss value and a preset parameter to obtain a product result; Calculating the sum of the product result and the first loss value as the final loss value; Training the pre-trained scoring model through the final loss value to obtain the oral scoring model.

11. A method for oral language scoring, characterized in that, The method includes: Obtaining a to-be-scored answer audio corresponding to a test oral test question, where the test oral test question belongs to a target question type; Input the audio of the answer to be scored into the spoken language scoring model to obtain the spoken language score of the audio of the answer to be scored predicted by the spoken language scoring model, where the spoken language scoring model is trained by any one of claims 1 to 10; Output the spoken language score of the audio of the answer to be scored.

12. An oral language scoring model training device, characterized in that, The device includes: A sample acquisition module, configured to acquire training samples, where the training samples include sample answer audios of sample spoken language test questions and sample scores corresponding to the sample answer audios, the sample spoken language test questions belong to the target question type, and the sample scores are obtained based on the target scoring rules corresponding to the target question type; A first scoring module, configured to input the sample answer audio into a pre-trained scoring model to obtain a predicted score corresponding to the sample answer audio, where the pre-trained scoring model is trained by a meta-learning method; A first determination module, configured to determine a first loss value according to the sample score and the predicted score, where the first loss value represents the loss between the sample score and the predicted score; A second determination module, configured to determine a target answer audio in the sample answer audio; A third determination module, configured to determine a second loss value according to the magnitude relationship between the sample scores corresponding to each of the target answer audios, where the second loss value represents the loss between the scoring rules of the pre-trained scoring model itself and the target scoring rules; A training module, configured to train the pre-trained scoring model according to the first loss value and the second loss value to obtain the spoken language scoring model.

13. An oral language scoring device, characterized in that, The device includes: An audio acquisition module, configured to acquire the audio of the answer to be scored corresponding to the test spoken language test question, where the test spoken language test question belongs to the target question type; A second scoring module, configured to input the audio of the answer to be scored into the spoken language scoring model to obtain the spoken language score of the audio of the answer to be scored predicted by the spoken language scoring model, where the spoken language scoring model is trained by any one of claims 1 to 10; An output module, configured to output the spoken language score of the audio of the answer to be scored.

14. An electronic device, characterized in that, Includes: One or more processors; A memory; One or more applications, where the one or more applications are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs are configured to execute the method according to any one of claims 1 - 11.

15. A computer-readable storage medium, characterized in that, Program code is stored in the computer-readable storage medium, and the program code can be called by the processor to execute the method according to any one of claims 1 - 11.

Citation Information

Patent Citations

  • Spoken language pronunciation quality evaluation method, device and equipment and storage medium

    CN112700795A

  • Voice adaptive recognition method, system and device and storage medium

    CN114187900A

  • Scoring method, device and equipment for oral test, storage medium and program product

    CN114333787A