Answer scoring method and model training method, device, storage medium and equipment
By acquiring the audio and text sequences of the test takers' answers, and combining them with phoneme sequences and key points from multiple model essays to generate a global key point representation, the problem of machine scoring systems being unable to accurately assess the semantic level of key points in oral exams has been solved, resulting in more accurate and stable scoring.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- IFLYTEK CO LTD
- Filing Date
- 2022-12-26
- Publication Date
- 2026-07-24
AI Technical Summary
Existing machine scoring systems have scoring errors in oral exams, especially in their inability to accurately assess the semantic level of the key points in the candidates' answers, resulting in inaccurate and unstable scoring.
By acquiring the audio and text sequences of the candidates' answers and combining them with phoneme sequences, the full-text audio and language level representation is determined. A global key point representation is generated using multiple model essays, and multiple perspectives are integrated for scoring, with attention paid to semantic errors at the key point level.
It improves the accuracy and stability of machine scoring, can identify semantic errors at the key point level in candidates' answers, avoids the problem of focusing only on pronunciation and content level while ignoring key point level, and enhances the comprehensiveness of scoring.
Smart Images

Figure CN115905475B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, specifically to an answer scoring method, a model training method, an apparatus, a computer-readable storage medium, and a computer device. Background Technology
[0002] In recent years, with the development and progress of artificial intelligence technology, it has been used in more and more scenarios, such as oral assessment scenarios involving oral exams. In oral assessment scenarios, such as the English oral exams for the middle and high school entrance examinations, machine scoring has largely replaced human scoring. Among them, the English oral picture description topic, which examines the candidate's comprehensive pronunciation level, is evaluated based on four aspects: the candidate's pronunciation level, the completeness of the answer points, the language level of the answer content, and the fluency of pronunciation.
[0003] While machine scoring is generally superior to human scoring, it can still produce some egregious errors. For example, test-takers may answer off-topic, or their content may only match keywords but lack key points or have low similarity to them. Conversely, test-takers may describe content with keywords similar to the images in the exam but with different meanings, leading to overly high machine scores. These issues all contribute to a decrease in the accuracy of machine scoring for picture-based descriptive questions. Summary of the Invention
[0004] This application provides an answer scoring method, a model training method, an apparatus, a computer-readable storage medium, and a computer device, which can improve the stability and accuracy of machine scoring for question-and-answer sessions.
[0005] This application provides an answer scoring method, including:
[0006] The system obtains the audio of a candidate's answer during an oral exam, as well as the text sequence of the candidate's full answer and the phoneme sequence of the candidate's full answer, determined based on the audio.
[0007] The speech level of the candidate's answer is determined based on the speech and the phoneme sequence, and the language level of the candidate's answer is determined based on the text sequence.
[0008] Obtain key points from multiple sample essays for the oral exam, and determine the global key point representation of the candidate's answer based on the multiple sample essay key points and the text sequence;
[0009] The candidate's answer is scored based on the full-text speech level representation, the full-text language level representation, and the global key point representation to obtain the candidate's answer score.
[0010] This application also provides a method for training a spoken language scoring model, including:
[0011] Obtain a training dataset and an initial oral scoring model. The training dataset includes multiple training samples from the oral exam. Each training sample includes training speech for each candidate's answer, a label score for the candidate's answer, a training text sequence of the candidate's full answer determined based on the training speech, and a training phoneme sequence of the candidate's full answer.
[0012] The training speech and the training phoneme sequence are input into the speech level modeling module of the initial oral scoring model for encoding and decoding processing to determine the full-text speech level representation of the candidate's answer; and the training text sequence is input into the language level modeling module of the initial oral scoring model for language extraction processing to determine the full-text language level representation of the candidate's answer.
[0013] The key points of multiple training sample texts for the oral exam are obtained, and the key points of multiple training sample texts and the training text sequence are input into the key point processing module of the initial oral scoring module for key point semantic processing to determine the global key point representation of the candidate's answer.
[0014] The training full-text speech level representation, the training full-text language level representation, and the training global key point representation are input into the fusion module of the initial oral scoring module for scoring processing to obtain the training score of the candidate's answer;
[0015] The initial spoken language scoring model is updated based on the training score and the label score to obtain the spoken language scoring model.
[0016] This application also provides an answer scoring device, including:
[0017] The first acquisition unit is used to acquire the voice of the candidate's answer in the oral test, as well as the text sequence of the candidate's full answer and the phoneme sequence of the candidate's full answer determined based on the voice.
[0018] A speech representation unit is used to determine the full-text speech level representation of the examinee's answer based on the speech and the phoneme sequence;
[0019] A language representation unit is used to determine the full-text language level representation of the examinee's answer based on the text sequence;
[0020] The second acquisition unit is used to acquire multiple sample essays and key points for the oral exam;
[0021] The key point representation unit is used to determine the global key point representation of the candidate's answer based on multiple sample essay key points and the text sequence;
[0022] The scoring unit is used to score the candidate's answer based on the full-text speech level representation, the full-text language level representation, and the global key point representation, so as to obtain the score of the candidate's answer.
[0023] This application also provides a spoken language scoring model training device, including:
[0024] The first training acquisition unit is used to acquire a training dataset and an initial oral scoring model. The training dataset includes multiple training samples from the oral exam. Each training sample includes training speech of each candidate's answer, the label score of the candidate's answer, and a training text sequence of the candidate's full answer and a training phoneme sequence of the candidate's full answer determined based on the training speech.
[0025] A training speech representation unit is used to input the speech and the phoneme sequence into the speech level modeling module of the initial oral scoring model for encoding and decoding processing, so as to determine the training full-text speech level representation of the candidate's answer;
[0026] A training language representation unit is used to input the training text sequence into the language proficiency modeling module of the initial oral scoring model for language extraction processing, so as to determine the language proficiency representation of the candidate's full-text training response.
[0027] The second training acquisition unit is used to acquire the key points of multiple training sample texts for the oral exam;
[0028] The training key point representation unit is used to input multiple training sample key points and the training text sequence into the key point processing module of the initial oral scoring module for key point semantic processing, so as to determine the global training key point representation of the candidate's answer;
[0029] The training scoring unit is used to input the full-text speech level representation, the full-text language level representation, and the global key point representation into the fusion module of the initial oral scoring module for scoring processing, so as to obtain the training score of the candidate's answer;
[0030] An update unit is used to update the initial spoken language scoring model based on the training score and the label score to obtain a spoken language scoring model.
[0031] This application also provides a computer-readable storage medium storing a computer program adapted for loading by a processor to perform the steps described in any of the above embodiments.
[0032] This application also provides a computer device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the steps in the method described in any of the above embodiments by calling the computer program stored in the memory.
[0033] The answer scoring method, model training method, apparatus, computer-readable storage medium, and computer device provided in this application embodiment acquire the speech of a candidate's answer in an oral exam, the text sequence of the candidate's full answer determined from the speech, and the phoneme sequence of the candidate's full answer. Based on the speech and phoneme sequence, the full-text speech level representation of the candidate's answer is determined, and based on the text sequence, the full-text language level representation of the candidate's answer is determined. Multiple sample key points from the oral exam are acquired, and a global key point representation of the candidate's answer is determined based on the multiple sample key points and the text sequence. The candidate's answer is then scored based on the full-text speech level representation, the full-text language level representation, and the global key point representation to obtain a score for the candidate's answer. This application embodiment not only obtains the full-text speech level representation and the full-text language level representation but also incorporates a global key point representation, enabling the scoring method to integrate representations from multiple different perspectives, improving accuracy. Simultaneously, the scoring method can focus on semantic errors at the key point level, avoiding focusing only on pronunciation and content levels while neglecting semantic errors at the key point level, thus improving the stability and accuracy of machine scoring in oral exams. Attached Figure Description
[0034] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0035] Figure 1 This is a flowchart illustrating the answer scoring method provided in an embodiment of this application.
[0036] Figure 2 This is a schematic diagram of the network structure of the oral scoring model provided in the embodiments of this application.
[0037] Figure 3 This is a schematic diagram of the network structure of the speech ED model provided in an embodiment of this application.
[0038] Figure 4 This is a schematic diagram of a sub-process of the answer scoring method provided in the embodiments of this application.
[0039] Figure 5 This is a schematic diagram illustrating the training process of the key point extraction model provided in the embodiments of this application.
[0040] Figure 6 This is a schematic diagram illustrating the usage process of the key point extraction model provided in the embodiments of this application.
[0041] Figure 7 This is a flowchart illustrating the oral scoring model training method provided in an embodiment of this application.
[0042] Figure 8 This is a flowchart illustrating the oral scoring model training method provided in an embodiment of this application.
[0043] Figure 9 This is a schematic diagram of the answer scoring device provided in an embodiment of this application.
[0044] Figure 10 A schematic diagram of the structure of the oral scoring model training device provided in the embodiments of this application.
[0045] Figure 11 A schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation
[0046] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0047] This application provides an answer scoring method, apparatus, computer-readable storage medium, and computer device. Specifically, the answer scoring method of this application can be executed by a computer device, and the answer scoring apparatus of this application is integrated into a computer device. This integration can occur in one or more computer devices. For example, the process of training the oral scoring model in this application is executed in one computer device, while the process of using the oral scoring model is executed in another computer device. Correspondingly, the process of training the oral scoring model is integrated into one computer device, and the process of using the oral scoring model is integrated into another computer device.
[0048] The computer equipment can be a terminal device or a server. The terminal device can be a smartphone, tablet, laptop, touchscreen, game console, personal computer (PC), smart vehicle terminal, robot, etc. The server can be a standalone physical server, a service node in a blockchain system, a server cluster consisting of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud storage, and big data and artificial intelligence platforms.
[0049] Before providing a detailed description of the solutions in the embodiments of this application, we will further analyze the current scoring scheme.
[0050] In the oral assessment scenario of oral exams, taking the English oral exam's "Looking at a Picture and Speaking" topic as an example, this section scores the answers to the questions. Currently, machine scoring can produce some absurd scoring errors, including the following error scenarios.
[0051] (1) The candidate's answer is off-topic, or the content only matches the keywords but the main points are wrong / the similarity of the main points is low; the candidate's description is similar to the keywords in the content of the exam picture, but the meaning of the description is different, resulting in the phenomenon of machine scoring too high. For example, the picture shows (the answer is in English, but for ease of explanation, it is described in Chinese) "Two colleagues are loudly discussing the contents of the book in the library, and the librarian next to them is advising them to be quiet, and there is a 'quiet' icon on the wall." For candidates who answer "two classmates are singing in the library", or only mention the keyword "library", or whose answers are all about "how inspiring it is to draw on the books in the library", which have low similarity of main points, there is a problem of scoring too high. Since the existing scoring scheme models the language level and content score of the answer based on the full text level or keyword level, it cannot pay attention to semantic errors at the main point level. Therefore, the existing scoring scheme does not consider the above expression to be wrong, so the score is too high.
[0052] (2) While the candidates' answers were relatively concise, the "key points" were described completely, leading to a lower machine score. For example, the picture (the answer was in English, but for ease of explanation, it is described in Chinese) shows a boy playing chess with his grandfather at the table, while his grandmother watches them and smiles. Birds fly by outside the window, and there are blooming plants. We should spend more time with the elderly. Some candidates only answered "The boy plays chess with his grandparents on the weekend." Although there was no further detailed description, the most crucial point was addressed. From the scoring criteria, it can be seen that answers with "key points" should receive more points. Existing solutions often overfit to the length of the answer without considering whether the "key points" are matched or modeling the importance of different "key points." They assume that every expression is equally important, resulting in a low score or 0 points.
[0053] (3) The picture descriptions are divergent, and the candidates' answers are novel. The candidates' answers did not appear in the teacher's reference sample, but the content of the candidates' answers was relevant to the meaning of the test picture or belonged to the category of advanced expression, which led to the phenomenon of low machine scoring. For example, the picture is (the answer is in English, but for ease of explanation, it is expressed in Chinese) "A boy is feeding chickens on a farm, and a girl is feeding sheep grass. This story tells us that we should love animals." Some candidates' final thoughts were "This story tells us that we should love nature" or "This story tells us that we should develop more hobbies on weekends." In fact, the teacher would not deduct points based on the meaning of the picture. The reason for deducting points in the existing scoring scheme is that it can only learn a limited number of candidates' key points from the training set. For key points that have not appeared in the training set, it is impossible to score them correctly and cannot cover divergent expressions.
[0054] Therefore, embodiments of this application provide an answer scoring method to solve at least one of the above-mentioned problems. The answer scoring method, apparatus, computer-readable storage medium, and computer device described in detail below are further described.
[0055] Figure 1 This is a flowchart illustrating the answer scoring method in an embodiment of this application. The method is applied in a computer device and includes the following steps.
[0056] 101. Obtain the audio of the candidate's answer in the oral exam, as well as the text sequence of the candidate's full answer and the phoneme sequence of the candidate's full answer determined based on the audio.
[0057] Oral exams can take place in various scenarios, such as English oral exams and Mandarin oral exams. For example, an English oral exam might include a picture-based topic assessment where candidates must speak based on a given picture. The following explanation uses an English oral exam as an example.
[0058] In the oral exam, the audio of the candidate's answer is obtained. This audio is then processed through speech recognition to obtain the corresponding text, or the text sequence of the candidate's full answer, such as the word sequence in an English oral exam. This text sequence is then transformed to obtain the phoneme sequence of the candidate's full answer. Finally, the text sequence and the factor sequence of the candidate's full answer are obtained.
[0059] Since the oral assessment model will be discussed later, it will be briefly introduced here. This oral assessment model consists of six modules: a speech level modeling module, a language level modeling module, a general semantic image-text generation module (this module may be unnecessary in some cases), a key point-level semantic correctness / incorrectness module, a key point importance module, and a fusion module. Specifically... Figure 2 As shown.
[0060] The test includes three modules: a speech level modeling module to determine the full-text speech level representation, including pronunciation accuracy and fluency; a language level modeling module to determine the full-text language level representation, including language level and content; a general semantic image-text generation module to generate a large number of general model essay points based on images from the oral exam's picture-based topic section, thus expanding the model essay point library; a point-level semantic accuracy module including a point extraction model and a point matching pre-trained model, where the point extraction model extracts the corresponding student points from multiple model essay points; a point importance module to determine the global point representation of the student's answer; and a fusion module to fuse the full-text speech level representation output by the speech level modeling module, the full-text language level representation output by the language level modeling module, and the global point representation output by the point importance module, and to achieve the final scoring.
[0061] Since the general semantic image and text generation module (which may not be needed in some cases), the key point level semantic correctness module, and the key point importance module all involve the content of key points, in some embodiments, the general semantic image and text generation module (which may not be needed in some cases), the key point level semantic correctness module, and the key point importance module can also be referred to as the key point processing module.
[0062] 102. Determine the full-text phonetic representation of the candidate's answer based on the speech and phoneme sequence.
[0063] The process involves dividing the speech into multiple speech segments, determining the corresponding phoneme segments based on these segments, and then encoding and decoding each speech segment and each phoneme segment corresponding to each speech segment to obtain the hidden layer representation of each speech segment. Finally, the obtained hidden layer representations of the multiple speech segments are processed according to the phoneme level to obtain the full-text speech level representation of the candidate's answer.
[0064] This section involves the speech proficiency modeling module of the oral assessment model. This part models the speech proficiency of the test takers, mainly including the accuracy of pronunciation of each phoneme and the fluency of the entire speech.
[0065] First, the candidate's speech is divided into multiple speech segments, such as v segments, using a speech detection algorithm like VAD (Voice Activity Detection). Then, a trained automatic speech recognition (ASR) model is used to obtain multiple ASR transcribed text segments corresponding to the speech segments. The v speech segments and their corresponding transcribed text segments can be represented as {(S1,A1),(S2,A2),…,(S…A2)…A2 ... v A v Let S be represented by )} i and A i Let i represent the i-th speech segment and its corresponding transcribed text segment, respectively. The transcribed text segment is converted into its corresponding phoneme segment, for example, using the G2P (Grapheme-to-Phoneme) tool to convert the ASR transcribed text segment A. i Convert the words in the text into their corresponding phoneme segments P i This results in multiple speech segments and corresponding factor segments.
[0066] This process involves encoding and decoding each speech segment from multiple speech segments, and each phoneme segment corresponding to each speech segment from multiple phoneme segments, to obtain the hidden layer representation of each speech segment. For example, each speech segment of the examinee's answer and its corresponding phoneme segment are input into the speech encoder-decoder model in the speech level modeling module for encoding and decoding to obtain the hidden layer representation of each speech segment. This speech encoder-decoder model can be simply referred to as the speech ED model. Figure 2 , Figure 3 As shown.
[0067] in, Figure 3 This is a simplified diagram of the network structure of a speech ED model. The speech encoder layer in the speech ED model can be a 6-layer Transformer structure with convolutional layers, such as a 6-layer VGG-Transformer or Conformer. The input is speech S. i The decoder layer can be composed of three Transformer layers, with the input being a phoneme fragment P. i , Figure 3 P in i ={w,ah,n,d,ey,ih} represents the phoneme segments of the transcribed text fragment "one day". It's important to note that the speech ED model can also be other network structures that achieve the same functionality.
[0068] Each speech segment (speech X_wav) is input into the speech ED model, where it is encoded using the speech encoder (speech-encoder layer) to obtain the speech encoding result. The speech encoding result is then input into the phoneme attention layer of the speech ED model to obtain the attention weight of each phoneme in the corresponding phoneme segment. The attention weights and each phoneme in the corresponding phoneme segment are then input into the decoder (decoder layer) of the speech ED model for decoding to obtain the hidden layer output (phoneme hidden) of each phoneme in the corresponding phoneme segment. The hidden layer output of multiple phonemes in the corresponding phoneme segment is used as the hidden layer representation of the speech segment.
[0069] The hidden representations of multiple speech segments can be obtained in the same way. Averaging these hidden representations at the phoneme level yields the full-text speech level representation h, which includes both pronunciation accuracy and fluency. speech .
[0070] For example, if the ASR transcribed text A1 is "one day", then the corresponding phoneme segment P1 = {w, ah, n, d, ey, ih}, and inputting it into the speech ED model will yield a hidden layer representation h of 6 phonemes. speech1 The dimensions are [6, 512], where 512 is the dimension of the hidden layer representation of each phoneme; the ASR transcribed text A2 is “Woman buy something.”, then the corresponding phoneme segment P2 = {w, uh, m, ax, n, b, ay, s, ah, m, th, ih, ng}, and the input speech ED model will obtain the hidden layer representation h of 13 phonemes. speech2 The dimension is [13, 512]. The above "averaging the hidden representations of multiple speech segments according to phoneme level" refers to averaging the hidden representations of the same phoneme. Suppose that in one case, English has a total of 48 phonemes. To obtain the hidden representation of the phoneme "w", starting from h... speech1 Take the hidden layer representation of "w" at position 1, from h speech2 Take the hidden layer representation of "w" at position 1, and average the two hidden layer representations to obtain the final hidden layer representation of the phoneme "w". Continue in this way to obtain the phoneme-level hidden layer representation "phonehidden" with dimensions [48, 512], where 48 is the size of the English phoneme list, thus obtaining a full-text speech level representation that includes pronunciation accuracy and fluency.
[0071] It is important to note that this application does not directly obtain the corresponding score using the speech ED model. Instead, it obtains the feature output of the hidden layer as a full-text speech level representation of the candidate's answer. Alternatively, it can be understood that the speech ED model in this application does not have a final linear connection layer / linear mapping layer such as a fully connected layer and / or a softmax layer, but directly obtains the feature output of the hidden layer.
[0072] 103. Determine the full-text language level of the candidate's answer based on the text sequence.
[0073] This section involves the language proficiency modeling module of the oral assessment model. This part models the language proficiency of the test takers' answers, mainly including two aspects: language proficiency and content based on the entire text.
[0074] The speech-level modeling module uses a pre-trained language model, such as the BERT language model. This BERT language model can be a multi-layered stack of Transformer structures, such as 12 layers. This particular BERT language model uses a language representation model that has been pre-trained on massive amounts of publicly available text data, such as Wikipedia, for a cloze test task. Other pre-trained language models can also perform the same function.
[0075] The process involves inputting the text sequence of the candidate's full response (e.g., a word sequence) along with its start identifier into the language BERT model for full-text semantic extraction, resulting in the full-text hidden output. Finally, the hidden output at the position corresponding to the start identifier in the text sequence is used as the full-text language proficiency representation h of the candidate's response. lm .
[0076] For example, if a test taker's full answer is the text sequence "Woman buy something.", then the content input into the BERT language model will include the start identifier of the text sequence, such as " <cls>"Woman buys something." (This sentence appears to be part of a larger sentence.) <cls>This is the starting identifier of the text sequence. Finally, after processing by the language BERT model, it will... <cls>The hidden layer output corresponding to the position is used as h lm ,Right now Figure 2 The word "language hidden" appears in the text.
[0077] It's also important to note that this module doesn't provide a final score; instead, it directly outputs the hidden layer. Alternatively, this can be understood as the language pre-training model in this module lacking a final linear connection / linear mapping layer, such as a fully connected layer and / or a softmax layer; it only requires the corresponding hidden layer output.
[0078] 104. Obtain key points from multiple sample essays for the oral exam, and determine the overall key points of the candidate's answer based on the key points of multiple sample essays and the text sequence.
[0079] The key points of multiple sample essays for the oral exam can be obtained in advance and stored in the sample essay key point database. Among them, the sample essay key points in the teacher's reference answer are the most basic sample essay key points. These sample essay key points cannot cover some novel expressions. Therefore, in order to expand the possible key expressions in the pictures in the picture description topic, multiple sample essay key points can be generated through intelligent generation to expand the sample essay key point database.
[0080] In one scenario, the steps for obtaining multiple sample essay points for an oral exam include: obtaining the exam image for the picture-based speaking topic corresponding to the oral exam; performing image encoding processing on the exam image to obtain the image encoding result; and generating multiple sample essay points for the exam image based on the image encoding result. This embodiment obtains sample essay points automatically, expanding the sample essay point library to include more sample essay points, covering more divergent expressions. This avoids situations where the machine score is too low due to the divergent image description and the candidate's novel expression, or where the candidate's expression is not found in the teacher's reference sample, but the candidate's content fits the meaning of the exam image or belongs to the category of advanced expression, resulting in an excessively low machine score.
[0081] The part that generates multiple sample text key points involves a general semantic image and text generation module. This module includes image and text generation models, such as the CLIPCap image and text generation model. Figure 2 As shown, the image-to-text generation model takes an image from an oral exam's "Look at the picture and describe the topic" section as input and outputs a large number of sample essay points related to the image content, essentially generating a sample essay point expression library.
[0082] Specifically, the image encoding module in the image-text generation model, such as the CLIP (Contrastive Language-Image Pre-Training) model in the CLIPCap model, is used to encode the exam images to obtain the image encoding results. Then, the image decoding module in the image-text generation model, such as the GPT2 (Generative Pre-trained Transformer 2) model in the CLIPCap model, is used to decode the image encoding results to generate multiple sample key points of the exam images. In other words, the decoding module, such as the GPT2 model, is used to perform the task of generating the text descriptions corresponding to the images.
[0083] The CLIP model is a pre-trained general-purpose model that uses massive amounts of images and text for comparative learning, enabling fine-grained encoding of images and deep semantic representation. For example, inputting image information for the "Image-to-Text Generation CLIPCap Model" topic, after multiple decoding iterations, yields a library of sample key points containing k results, including k sample key points {F1, F2, ..., F...}. k For example, F1 = The boy is playing chess with the olders, F2 = The old man is sitting on a chair, and F3 = The old woman is standing by them. It's important to note that the image-text generation model in the general semantic image-text generation module, such as the CLIPCap image-text generation model, is also a pre-trained model.
[0084] After obtaining the key points of multiple sample essays, the global key point representation of the candidate's answer is determined based on the key points of the multiple sample essays and the text sequence. Determining the global key point representation of the candidate's answer involves a key point level semantic correctness module and a key point importance module.
[0085] This method involves extracting key points from each of the multiple sample essays and the corresponding text sequence to determine the key points for the candidate's answer. Then, for each sample essay key point and its corresponding candidate answer key point, attention-level semantic extraction is performed to obtain a global representation of the candidate's answer. In this embodiment, after extracting the candidate's answer key points corresponding to each sample essay key point, attention-level semantic extraction is performed to focus on semantic information at the key point level. Because it focuses on semantic information at the key point level, it can solve the problem of high machine scores caused by candidates' off-topic answers or keyword matching but incorrect or low key point similarity. Furthermore, the inclusion of an attention mechanism focuses on the importance of key points, thus addressing the issue of low machine scores in scenarios where candidates' answers are concise but the "key points" are fully described.
[0086] like Figure 4 As shown, the steps for determining the global key points of a candidate's answer based on multiple sample essays and text sequences include the following steps 201 to 203.
[0087] 201. The key points of each sample essay and the text sequence in multiple sample essays are extracted to determine the key points for the candidate's answer corresponding to each sample essay key point.
[0088] This process involves concatenating each key point from multiple model essays with a separate text sequence to obtain multiple concatenated text sequences. Each of these concatenated text sequences is then input into the key point extraction model within the key point-level semantic correctness module for key point extraction. This process yields the candidate's answer key points corresponding to each model essay key point, as well as the first similarity (also known as key point similarity) between each model essay key point and the candidate's answer key points. The first similarity plays a role in training the oral scoring model, which will be discussed later.
[0089] like Figure 5 The diagram shown illustrates the training process of the key point extraction model. The key point extraction model can include a multi-layer Transformer structure, such as a 6-layer Transformer structure, or other structures that achieve the same functionality.
[0090] Each concatenated text sequence includes a sample essay's key points and a text sequence of the candidate's full answer. For the oral exam's picture-based description topic, there are three sample essays and a text sequence of the candidate's full answer. One of the concatenated text sequences can be... <cls>Oldwoman is standing by them. <sep>Tom is playing chess with his grandpa. The grandpa is sitting with a fan. His grandma is watching them; Among them, <cls>It is the starting identifier for the key points of the sample essay. <sep>It is the start identifier of the text sequence of the candidate's full answer.
[0091] Each concatenated text sequence is input into the key point extraction model. The model extracts key points from the concatenated text sequences to mark the start position (Pos_start) and end position (Pos_end) of each model text key point in the student's answer text sequence, indicating the position of the first word and the last word. It also marks the key point similarity between the model text key points and the extracted student answer key points. simi keypoint simi ∈[0,1]. That is, the input of the key point extraction model is the text sequence of each sample essay key point and the candidate's full answer. The output of the key point extraction model has 3 items, namely the key point similarity Keypoint_Sim, the start position and the end position of the extracted candidate's answer key points.
[0092] like Figure 5 As shown, since there are three concatenated text sequences, the output of the key point extraction model is the key point similarity corresponding to the three concatenated text sequences, the start position and end position of the extracted candidate's answer key points, for example, Keypoint_Sim=1, Pos_start=0, Pos_end=6; Keypoint_Sim=1, Pos_start=8, Pos_end=14; Keypoint_Sim=0.5, Pos_start=16, Pos_end=20 respectively.
[0093] Specifically, the text corresponding to the start and end positions of each sample essay point in the candidate's answer text sequence is taken as the candidate's answer point for each sample essay point. For example, for sample essay point 1 with Pos_start=0 and Pos_end=6, the corresponding candidate's answer point text is "Tom is playing chess with his grandpa."
[0094] The key points for each sample essay corresponding to the candidate's answer are represented as {P1, P2, ..., P...} k It should be noted that some sample essays do not have corresponding student response guidelines.
[0095] like Figure 6 The diagram shown illustrates the usage flow of the key point extraction model provided in this application embodiment. The key points of the sample essay and the text sequence of the examinee's answer are input into the key point extraction model, which outputs the key points of the sample essay F. i Similarity score between the test taker's answer and the test taker's answer. And the key points of the sample essay F i Matched candidate answer points P i Let i ∈ [1, k]. The text sequence of the candidate's full answer is "He is playing chess. His grandpa is happy." There are 3 sample essay points, k = 3. We can obtain the 3 candidate answer points and the similarity between each sample essay point and the candidate's answer point. For example, P1 = "He is playing chess," with a similarity of 0.8; P2 = "His grandpa is happy," with a similarity of 0.4; P3 = "", meaning P3 is empty, with a similarity of 0.1.
[0096] The key point extraction model is obtained through pre-training, as shown in the following example. Figure 5 As shown. For each picture-based essay topic, prepare sample essays and key points for each topic, along with text sequences of full answers from multiple different test-takers. Then, train the key point extraction model. Based on a large amount of historical data from picture-based essay topics, a general key point extraction model can be pre-trained.
[0097] Therefore, although the spoken language scoring model includes the text generation model in the general semantic text generation module and the key point extraction model, the text generation model and the key point extraction model are obtained through pre-training. These two models do not participate in the modification of network parameters during the training process of the spoken language scoring model, and they are directly used during the training process.
[0098] 202. For each sample essay key point and the corresponding candidate answer key point, perform key point-level semantic extraction processing to obtain the full text key point semantic representation of the candidate answer and the hidden semantic representation of the contextual similarity between each sample essay key point and the corresponding candidate answer key point.
[0099] This step involves the point-matching part of the point-level semantic right / wrong module. The point-matching part involves a pre-trained point-matching model, such as the BERT point-matching model. Figure 2 As shown.
[0100] Specifically, the key points of each sample essay and the corresponding key points of the candidate's answer can be concatenated to obtain each key point pair text, where each key point pair text includes a start identifier. The global key point representation identifier and multiple key point pairs texts are input into the key point matching pre-trained model for key point level semantic extraction processing to obtain the full text key point semantic representation of the candidate's answer and the hidden semantic representation corresponding to multiple key point pairs texts. The hidden semantic representation at the position corresponding to the sample essay start identifier of each key point pair text is used as the hidden semantic representation of the contextual similarity between each key point pair text.
[0101] The process involves concatenating the key points of each sample essay with the corresponding key points of the candidate's answer. Assuming there are three sample essay key points, this results in three key point pairs: {(F1,P1),(F2,P2),(F3,P3)}. Each key point pair includes one sample essay key point and its corresponding candidate's answer key point. At least one start identifier is added to each key point pair. For example, a sample essay start identifier is added before each sample essay key point in each key point pair, and a candidate's answer start identifier is added before each candidate's answer key point in each key point pair to distinguish between the sample essay key points and the candidate's answer key points.
[0102] A global point identifier can be added before all point pairs. For example, the global point identifier and the three point pairs as a whole can be represented as "[MASK]". <cls>The boy is playing chess with theolders <sep>He is playing chess <cls>The old man is sitting on a chair <sep>Hisgrandpa is happy <cls>Old woman is standing by them <sep>"." Here, [MASK] is the global key identifier. <cls>This indicates the starting identifier for each key point in the sample text. <sep>This indicates the start identifier for each point in the text that corresponds to the candidate's answer.
[0103] The global key point representation identifier and multiple key point pairs (including the sample essay start identifier and the candidate's answer start identifier) are input into the key point matching BERT model. The key point matching BERT model is used to perform key point-level semantic extraction processing to obtain the hidden layer representation of each input text, including the full text key point semantic representation of the candidate's answer and the hidden layer semantic representation corresponding to multiple key point pairs.
[0104] For example, the key point matches the input "[MASK]" of the BERT model. <cls>The boy is playing chess withthe olders <sep>He is playing chess <cls>The old man is sitting on a chair <sep>His grandpa is happy <cls>Old woman is standing by them <sep>The input text consists of 37 elements, including words and symbols. After semantic extraction at the point level by the point-matching BERT model, the 37 input texts are processed to obtain hidden semantic representations for each input text. Assuming the hidden semantic representation is 512-dimensional, the point-matching BERT model will output a 37*512-dimensional hidden semantic representation.
[0105] Extracting multiple key points: For each key point in the text, extract the start identifier of the key point in the text, such as the start identifier of the key point in the sample text. <cls>The hidden semantic representation at the corresponding position serves as the hidden semantic representation of the contextual similarity between each key point and the text. This means that the final step of the key point matching BERT model only requires extracting the hidden semantic representation h at the position corresponding to the global key point representation identifier. mask And each key point begins with a marker in the text. <cls>The hidden semantic representation h at the corresponding position cls_1 ,h cls_2 ,…,h cls_k .
[0106] For example, [MASK] <cls> <cls> <cls>The hidden semantic representations at these four locations correspond to a 4*512 dimension. Among them, [MASK] <cls> <cls> <cls>From front to back, they are the global key point identifier, the first sample key point start identifier, the second sample key point start identifier, and the third sample key point start identifier. Figure 2 To distinguish them, a corresponding number is added after cls to indicate which key point in the sample text begins, making it easier to understand.
[0107] [MASK] <cls> <cls> <cls>The latent semantic representations at these four locations represent the semantic representation of the candidate's full-text answer, the latent semantic representation of the contextual similarity between the first key point and the text, the latent semantic representation of the contextual similarity between the second key point and the text, and the latent semantic representation of the contextual similarity between the third key point and the text, respectively. The extracted latent semantic representations include contextual information such as similarity / matching information of the contextual key points.
[0108] This key point matches the BERT model from input to the final extracted output, and we can see the following:
[0109] a. The input to the BERT model for key point matching includes global key point identifiers and multiple key point texts such as word sequences. It is important to note that the input includes multiple key point text pairs, rather than a single key point text pair, in order to prepare for extracting contextual information such as contextual key point information. If it were a single key point text pair, there would be no contextual information.
[0110] b. The model is limited to a pre-trained point-matching model, such as the BERT point-matching model. Therefore, it can extract contextual information from the input. Thus, the contextual similarity includes not only the similarity / matching information between the sample text and the candidate's answer text in each point pair, but also the contextual point information. This contextual point information helps determine the similarity / matching information. Since there are relationships between the contextual texts, using the preceding and following point information to assist in determining the similarity / matching information can improve accuracy.
[0111] c. The input of the BERT model for key point matching includes a lot of information, such as 37 pieces of information such as word sequences in the key point text. However, after extraction, only 4 corresponding hidden semantic representations are obtained. This achieves the integration of other information into these 4 corresponding hidden semantic representations, which to some extent reduces dimensionality and provides a basis for subsequent scoring. If a lot of information is included, it will be inconvenient to perform the final scoring.
[0112] d. The hidden semantic representation of the start identifier of each key point in the text, such as the position corresponding to the start identifier of the key point in the sample text, is used as the hidden semantic representation of the contextual similarity between each key point in the text. This effect can be achieved because the hidden semantic representation of the position corresponding to the start identifier of each key point in the text is processed by a linear connection layer / linear mapping layer during training, and the processing result is used to calculate the loss value with the first similarity in the key point extraction model. The training section will describe this in detail later.
[0113] 203. Attention mechanism processing is applied to the full text key semantic representation and the hidden layer semantic representation to obtain the global key representation after attention weighting of the hidden layer semantic representation.
[0114] The output extracted from the key point matching BERT model, i.e., the semantic representation of the key points in the whole text h mask The latent semantic representation h of the contextual similarity between each key point and the text cls_1 ,h cls_2 ,…,h cls_k We employ attention mechanisms to model the importance of different key points, thus addressing the issue where test takers' answers are concise but lack a complete description of "key points," resulting in lower machine scores.
[0115] This section involves the key point importance module of the oral assessment model, which includes a key point importance model. Specifically, the full-text key point semantic representation and the hidden semantic representation are concatenated to obtain a concatenated semantic representation. This concatenated semantic representation is then input into the key point importance model for attention weight processing, resulting in the full-text key point semantic representation and the attention weights corresponding to multiple hidden semantic representations. Finally, a weighted average is performed based on the attention weights of the multiple hidden semantic representations and their corresponding hidden semantic representations to determine the global key point representation after the attention weights are applied to the hidden semantic representations.
[0116] like Figure 2 In the process, the extracted key points are matched with the output h of the BERT model. mask h cls_1 ,h cls_2 ,…,h cls_k Together, they serve as input to the key importance model. For example, the key importance model currently includes 4*512 dimensions of input.
[0117] Among them, the key importance model is a self-attention mechanism module, and the formula for the attention weight of the self-attention mechanism module is shown in formula (1).
[0118]
[0119] Where Q = FC1(H), K = FC2(H), V = FC3(H), d k It is the dimension of key point representation, that is, the dimension of hidden semantic representation, such as d. k =512,FC i All are linear mapping layers / linear connection layers, H = {h} mask ,h cls1 ,h cls2 ,…,h clsk },here This is to prevent the inner product calculation from being too large and to avoid the softmax result from having a spike, causing one or a few values to be much larger than other values.
[0120] The importance of different points is reflected by the attention weights calculated by the softmax() function in formula (1). The output of the softmax() function is k+1 scores, where the first score represents h. mask The importance of itself, the hidden semantic representation h of the k key point pairs corresponding to the k scores. cls_i The importance of each point corresponds to the attention weights of the semantic representation of the main points in the full text and the semantic representations of multiple hidden layers, respectively. The importance of a specific point is independent of its similarity or order; it is learned from the training dataset.
[0121] After obtaining the semantic representation of the full text and the attention weights corresponding to multiple hidden semantic representations, the attention weights and corresponding hidden semantic representations of the multiple hidden semantic representations are weighted to determine the global key representation after the attention weights of the hidden semantic representations, as well as the weighted multiple hidden semantic representations.
[0122] The input to the key importance model is {h}. mask ,h cls_1 ,h cls_2 ,…,h cls_k The output is h′. mask ,h′ cls_1 ,…,h′ cls_k The output dimension and size are consistent with the input, for example, both are 4*512. The hidden semantic representation h of the k key point pairs... cls_i The corresponding attention weights are then used for weighting to obtain the weighted hidden semantic representation h′. cls_i Then, the global key point representation is determined based on the k weighted hidden semantic representations, such as by adding the k weighted hidden semantic representations together to obtain the global key point representation, etc. The global key point representation h KPcls =h′ mask .
[0123] The output of the importance module is the global importance representation h′. mask The final output is 1*512, while the k weighted hidden semantic representations h′ are discarded. cls_i .
[0124] From the input to the final output of the importance module, we can see the following two points:
[0125] a. The input of the key point importance module includes the semantic representation of the full text key points and multiple hidden layer semantic representations. That is, the input includes multi-dimensional information, such as 4*512 dimensions, while the output only takes one output, namely the global key point representation, which is 1*512 dimensions. This is equivalent to using the key point importance module to achieve dimensionality reduction. In this way, the solution in this application can be applied to any application scenario. Finally, only one global key point representation needs to be obtained, which is convenient for subsequent fusion processing.
[0126] b. The final output of the key point importance module is a global key point representation. This global key point representation is obtained by weighting the implicit semantic representation of the contextual similarity between the key points of each sample essay and the key points of the candidate's answer with the corresponding attention weights. This global key point representation integrates the importance information of different key points in the candidate's full answer, which solves the problem of the machine scoring being low when the candidate's answer is relatively concise but the "key points" are described in a complete manner.
[0127] At this point, we have obtained the full-text speech level representation indicating pronunciation accuracy and fluency, the full-text language level representation indicating language proficiency and content, and the global key point representation indicating the importance of key points. In other words, we have obtained representations of the candidate's full answer from multiple different perspectives.
[0128] 105. Based on the overall speech level, overall language level, and overall key points of the text, the examinee's answer is scored to obtain the examinee's score.
[0129] This step involves the fusion module in the spoken language scoring model, which includes the fusion model.
[0130] The full-text speech level representation, full-text language level representation, and global key point representation can be fused to obtain a fused hidden layer representation that incorporates the key point representation. The fused hidden layer representation is then input into the fusion model in the fusion module for nonlinear scoring processing to obtain the score of the examinee's answer.
[0131] Among them, the full-text speech level representation h speech The overall language level is expressed as h lm and global key point representation h KPcls spliced together to form a fused hidden layer representation h concat =(h speech ,h lm ,h KPcls The hidden layer representation is input into the fusion model for nonlinear scoring processing, so as to integrate the scoring scales from multiple different perspectives of the examinee in the current exam and obtain the score of the examinee's answer.
[0132] The formula for the fusion model is shown in formula (2) below.
[0133] Pred fuse =sigmoid(W1h) concat +b1) (2)
[0134] Among them, Pred fuse The final score for the candidate's answer is a value in the interval [0,1]. The sigmoid function represents a non-linear function, which is used for non-linear scoring.
[0135] The above embodiments not only obtain full-text speech level representation and full-text language level representation, but also add global key point representation, which makes the scoring method integrate multiple representations from different perspectives, improves accuracy, and allows the scoring method to pay attention to semantic errors at the key point level, as well as the importance of different key points, thereby improving the stability and accuracy of machine scoring in oral exams.
[0136] Figure 7 This is a flowchart illustrating the oral scoring model training method provided in this application embodiment. The method is mainly used to train an oral scoring model, which can be applied to the answer scoring method in any of the above embodiments. The method includes the following steps.
[0137] 301. Obtain the training dataset and the initial oral scoring model. The training dataset includes multiple training samples from the oral exam. Each training sample includes the training speech of each candidate's answer, the label score of the candidate's answer, and the training text sequence and training phoneme sequence of the candidate's full answer determined based on the training speech.
[0138] Among them, the label score can be a relatively accurate score of the candidate's answer in advance. For example, the candidate's answer can be scored manually to obtain the label score.
[0139] The only difference between training speech, training text sequence, and training phoneme sequence and the speech, text sequence, and phoneme sequence mentioned above is that they all have the word "training" in front of them, but their essential meaning is the same. Many other terms in the following text are also the same, and will not be explained again.
[0140] The initial oral scoring model is the model that requires network parameter updates, and its modules are the same as those in the oral scoring model.
[0141] 302. Input the training speech and training phoneme sequence into the speech level modeling module of the initial oral scoring model for encoding and decoding processing to determine the speech level representation of the candidate's full-text training response, and input the training text sequence into the language level modeling module of the initial oral scoring model for language extraction processing to determine the language level representation of the candidate's full-text training response.
[0142] The training speech can be divided into multiple training speech segments, and multiple training phoneme segments corresponding to these segments can be identified. Each training speech segment and each corresponding training phoneme segment in the speech level modeling module is then input into the speech ED model for encoding and decoding to obtain the training hidden layer representation of each training speech segment. The obtained training hidden layer representations of the multiple training speech segments are then processed at the phoneme level to obtain the training full-text speech level representation of the examinee's response. This training full-text speech level representation includes information on the accuracy and fluency of the examinee's pronunciation.
[0143] The training text sequence can be input into the language BERT model in the language proficiency modeling module for language extraction processing to obtain a training full-text language proficiency representation of the test taker's answers. This training full-text language proficiency representation includes both the language proficiency and content information of the test taker's answers.
[0144] 303. Obtain the key points of multiple training sample essays for the oral exam, and input the key points of multiple training sample essays and training text sequences into the key point processing module of the initial oral scoring model for key point semantic processing to determine the global key point representation of the candidate's answer.
[0145] The steps for obtaining multiple training sample key points for the oral exam include: obtaining training exam images for the picture-based speaking topics corresponding to the oral exam; performing image encoding processing on the training exam images to obtain training image encoding results, for example, inputting the training exam images into the image-text generation model in the general semantic image-text generation module, and using the image encoding module of the image-text generation model to perform image encoding processing on the training exam images; and generating multiple training sample key points for the training exam images based on the training image encoding results, for example, using the image decoding module of the image-text generation model to decode the training image encoding results to generate multiple training sample key points for the training exam images.
[0146] The step of inputting multiple training sample key points and training text sequences into the key point processing module of the initial oral scoring model for key point semantic processing to determine the training global key point representation of the candidate's answer includes: using the key point extraction model in the key point processing module to extract key points from each training sample key point and training text sequence to determine the candidate's answer training key points corresponding to each training sample key point, and the first similarity between the candidate's answer training key points and the training sample key points; using the key point matching pre-training model and key point importance model in the key point processing module to perform key point-level attention semantic extraction processing on each training sample key point and the corresponding candidate's answer training key points to obtain the training global key point representation of the candidate's answer.
[0147] Specifically, multiple training sample essay key points and corresponding candidate answer training key points can be input into the key point matching pre-training model of the key point level semantic correctness module for key point level semantic extraction processing, so as to obtain the semantic representation of the full text training key points of the candidate's answer, and the training hidden layer semantic representation of the contextual similarity between each training sample essay key point and the corresponding candidate answer training key point; the training full text key point semantic representation and the training hidden layer semantic representation are input into the key point importance model of the key point importance module for attention mechanism processing, so as to obtain the training global key point representation after attention weighting of the training hidden layer semantic representation.
[0148] The aforementioned step of using the key point extraction model in the key point processing module to extract key points from each training model essay key point and training text sequence in multiple training model essay key points, in order to determine the corresponding test taker answer training key points for each training model essay key point, and the first similarity between the test taker answer training key points and the training model essay key points, includes: concatenating each training model essay key point and training text sequence in multiple training model essay key points to obtain multiple training concatenated text sequences; inputting each training concatenated text sequence in the multiple training concatenated text sequences into the key point extraction model for key point extraction processing, in order to obtain the corresponding test taker answer training key points for each training model essay key point, and the first similarity between the test taker answer training key points and the training model essay key points.
[0149] The steps described above, which involve inputting the key points of each training sample text and the corresponding key points of the candidate's answer into the key point matching pre-training model of the key point level semantic correctness module for key point level semantic extraction processing to obtain the semantic representation of the full text key points of the candidate's answer and the training latent semantic representation of the contextual similarity between each training sample text key point and the corresponding candidate's answer key points, include: concatenating each training sample text key point and the corresponding candidate's answer key point into a text pair to obtain each text pair of training key points, wherein each text pair of training key points includes a start identifier; inputting the global key point representation identifier and multiple text pairs of training key points into the key point matching pre-training model of the key point level semantic correctness module for key point level semantic extraction processing to obtain the semantic representation of the full text key points of the candidate's answer and the training latent semantic representation of the text pairs of training key points; and extracting the training latent semantic representation at the position corresponding to the start identifier of each text pair of training key points as the training latent semantic representation of the contextual similarity between each text pair of training key points.
[0150] The steps described above, which involve inputting the training full-text key point semantic representation and the training hidden layer semantic representation into the key point importance model of the key point importance module for attention mechanism processing to obtain the training global key point representation after attention weighting of the training hidden layer semantic representation, include: concatenating the training full-text key point semantic representation and the training hidden layer semantic representation to obtain the training concatenated semantic representation; inputting the training concatenated semantic representation into the key point importance model of the key point importance module for attention weight processing to obtain the attention weights corresponding to the training full-text key point semantic representation and multiple training hidden layer semantic representations; and weighting the attention weights corresponding to multiple training hidden layer semantic representations and the corresponding training hidden layer semantic representations to determine the training global key point representation after attention weighting of the training hidden layer semantic representation.
[0151] 304. The training full-text speech level representation, the training full-text language level representation, and the training global key point representation are input into the fusion module of the initial oral scoring model for scoring processing to obtain the training score of the candidate's answer.
[0152] Specifically, the training full-text speech level representation, the training full-text language level representation, and the training global key point representation can be fused to obtain a training fused hidden layer representation that incorporates the key point representation; the training fused hidden layer representation is then input into the fusion model in the fusion module for nonlinear scoring processing to obtain the training score of the examinee's answer.
[0153] 305. Update the initial spoken language scoring model based on the training score and the label score to obtain the spoken language scoring model.
[0154] The loss value is calculated based on the training score and the label score. The loss value is used to update the network parameters of the initial spoken language score until the training stopping condition is met, such as the loss value converges or the number of training rounds reaches the preset number of rounds. Then the training stops and the spoken language score model is obtained.
[0155] In this embodiment, the oral scoring model is trained end-to-end. During training, optimizable network parameters include the speech proficiency modeling module, the language proficiency modeling module, the key point matching pre-trained model in the key point-level semantic correctness module, the key point importance module, and the fusion module. Figure 2 The part represented by the solid line is updated when updating network parameters. The model in the virtual representation part is a pre-trained model and does not need to be updated.
[0156] In training the oral assessment model, not only were training full-text speech level representations, including pronunciation accuracy and fluency, and training full-text language level representations, including language proficiency and content, considered, but global key point representations were also added. This allows the oral assessment model's scoring method to integrate representations from multiple different perspectives, improving accuracy. At the same time, the scoring method can focus on semantic errors at the key point level, as well as the importance of different key points, thus improving the stability and accuracy of machine scoring in oral exams.
[0157] In one embodiment, such as Figure 8 As shown, another flowchart of a spoken language scoring model training method is provided. This method is mainly used to train a spoken language scoring model, which can be applied to the answer scoring method in any of the above embodiments. The method includes the following steps.
[0158] 401. Obtain the training dataset and the initial oral scoring model. The training dataset includes multiple training samples from the oral exam. Each training sample includes the training speech of each candidate's answer, the label score of the candidate's answer, and the training text sequence and training phoneme sequence of the candidate's full answer determined based on the training speech.
[0159] 402. Input the training speech and training phoneme sequence into the speech level modeling module of the initial oral scoring model for encoding and decoding processing to determine the speech level representation of the candidate's full-text training response, and input the training text sequence into the language level modeling module of the initial oral scoring model for language extraction processing to determine the language level representation of the candidate's full-text training response.
[0160] 403. Obtain the key points of multiple training sample texts for the oral exam, and use the key point extraction model of the initial oral scoring model to extract key points from each training sample text and the training text sequence to determine the corresponding test taker's answer training key points for each training sample text key point, as well as the first similarity between the test taker's answer training key points and the training sample text key points.
[0161] 404. The key points of multiple training sample essays and the corresponding key points of the test takers' answers are input into the key point matching pre-training model of the initial oral scoring model for key point-level semantic extraction processing to obtain the semantic representation of the key points of the full training text of the test takers' answers, as well as the training hidden layer semantic representation of the contextual similarity between the key points of each training sample essay and the corresponding key points of the test takers' answers. The training hidden layer semantic representation is then linearly mapped using a linear connection layer to obtain the second similarity corresponding to each training hidden layer semantic representation.
[0162] Continuing with the example above, the output of the key point matching pre-trained model has a dimension of 4*512, which includes the semantic representation of the key points of the full training text of the candidate's answer (1*512), and the semantic representation of the contextual similarity between the key points of each training sample text and the corresponding key points of the candidate's answer (3 x 1*512).
[0163] During training, the key point matching pre-training model also includes a newly added linear connection layer / linear mapping layer, such as a fully connected layer or a softmax layer. This layer takes the previously output training full-text key point semantic representation, the training hidden layer semantic representation of the contextual similarity between each training model key point and the corresponding candidate's answer training key point, and inputs them into the newly added linear connection layer / linear mapping layer. The linear connection layer / linear mapping layer performs linear mapping processing to obtain a 4*1 dimensional output result. The output result of each training hidden layer semantic representation is used as the second similarity corresponding to that training hidden layer semantic representation. The second similarity has three data points, each corresponding to the similarity of each training hidden layer semantic representation. During training, the second similarity is trained with the corresponding first similarity as the target.
[0164] 405. Determine the second loss value based on the first similarity and the second similarity.
[0165] During the training process, an optimized loss function was set for the key point matching pre-trained model, as shown in formula (3).
[0166]
[0167] Among them, FC kp This is a newly added linear connection layer / linear mapping layer, where k is the number of training sample essays for this question. This represents the first similarity between the key points of the i-th training sample and the corresponding key points of the candidate's answer. L represents the training hidden semantic representation corresponding to the key points of the i-th training sample text. keypoint This represents the determined second loss value.
[0168] It's important to note that the loss function of the key-point matching pre-trained model cannot be simply understood. How can we ensure that the key-point matching pre-trained model... <cls>The hidden output at the corresponding position can include similarity information (the hidden representation of similarity) for each training point pair (a training essay point and a corresponding test taker's answer training point), rather than keyword information or other information. This is achieved through a loss function, which enables the point matching of the pre-trained model. <cls>The hidden layer output at the corresponding position can learn similarity information.
[0169] Furthermore, instead of directly calculating the similarity information between each training point pair, a point matching pre-trained model is used here. On the one hand, this is to optimize the network parameters of the speech proficiency modeling module, language proficiency modeling module, point-level semantic correctness module, point matching pre-trained model, point importance module, and fusion module in the end-to-end oral scoring model. If the similarity between each training point pair is directly calculated, the similarity is just a specific numerical value with nothing to optimize, and there is no way to backpropagate the parameters during training. On the other hand, using a point matching pre-trained model ensures that the result, in addition to the similarity information of each training point pair, also utilizes contextual information, using contextual point information to assist in evaluating similarity, so that the similarity also includes contextual information such as contextual point information.
[0170] Because it utilizes contextual information, the input to the key point matching pre-trained model includes multiple different training key point pairs of text. For example, training key point 1 of the sample essay and training key point 1 of the candidate's answer, training key point 2 of the sample essay and training key point 2 of the candidate's answer, training key point 3 of the sample essay and training key point 3 of the candidate's answer. Otherwise, the training latent semantic representation of the training key point pair can be obtained directly from training key point 1 of the sample essay and training key point 1 of the candidate's answer, without the need to input multiple different training key point pairs of text at the same time.
[0171] Please refer to the corresponding descriptions in the section above that uses the oral assessment model for understanding this part; it will not be repeated here.
[0172] 406. The training full-text key point semantic representation and the training hidden layer semantic representation are input into the key point importance model of the key point importance module for attention mechanism processing, so as to obtain the training global key point representation after attention weighting of the training hidden layer semantic representation.
[0173] 407. The training full-text speech level representation, the training full-text language level representation, and the training global key point representation are input into the fusion module of the initial oral scoring model for scoring processing to obtain the training score of the candidate's answer.
[0174] 408. The first loss value is determined based on the training score and the label score, and the initial spoken language scoring model is updated based on the first loss value and the second loss value to obtain the spoken language scoring model.
[0175] The optimization loss function of the fusion module can be shown in formula (4).
[0176]
[0177] Where n is the number of test takers who answered in all training samples, Y is the label score for each training sample, and Pred fuse L represents the training score obtained for each training sample. scoring This represents the first loss value.
[0178] After obtaining the first loss value and the second loss value, the initial spoken language scoring model is updated based on the first loss value and the second loss value. For example, a first coefficient and a second coefficient are determined, and the first loss value and the second loss value are weighted and summed based on the first coefficient and the second coefficient to obtain the overall loss value; the initial spoken language scoring model is updated based on the overall loss value, wherein the sum of the first coefficient and the second coefficient is 1, and the first coefficient is greater than the second coefficient.
[0179] The training loss function of the entire oral scoring model is shown in formula (5).
[0180] L tot =(1-α)L scoring +αL keypoint (5)
[0181] Where 1-α is the first coefficient and α is the second coefficient. Where 1-α > α. In practice, α can take the value 0.1.
[0182] The purpose of explicitly limiting the first coefficient to be greater than the second coefficient is to maintain the characteristics of the key point matching pre-trained model, that is, to use contextual information to represent the similarity information of each key point pair (hidden layer representation). However, the most important thing is to obtain the final oral score, so the first coefficient is greater than the second coefficient.
[0183] The method for training the oral scoring model in this embodiment enables the obtained oral scoring model to improve the stability and accuracy of machine scoring in oral examinations.
[0184] All of the above technical solutions can be combined in any way to form optional embodiments of this application, and will not be described in detail here.
[0185] To facilitate better implementation of the answer scoring method in this application, this application also provides an answer scoring device. Please refer to... Figure 9 , Figure 9 This is a schematic diagram of the structure of the answer scoring device provided in an embodiment of this application. The answer scoring device 500 may include a first acquisition unit 501, a speech representation unit 502, a language representation unit 503, a second acquisition unit 504, a key point representation unit 505, and a scoring unit 506.
[0186] The first acquisition unit 501 is used to acquire the voice of the candidate's answer in the oral test, as well as the text sequence of the candidate's full answer and the phoneme sequence of the candidate's full answer determined based on the voice.
[0187] The speech representation unit 502 is used to determine the full-text speech level representation of the candidate's answer based on the speech and the phoneme sequence.
[0188] The language representation unit 503 is used to determine the full-text language level representation of the candidate's answer based on the text sequence.
[0189] The second acquisition unit 504 is used to acquire multiple sample essays and key points for the oral exam.
[0190] The key point representation unit 505 is used to determine the global key point representation of the candidate's answer based on multiple sample key points and the text sequence.
[0191] Scoring unit 506 is used to score the candidate's answer based on the full-text speech level representation, the full-text language level representation, and the global key point representation, so as to obtain the score of the candidate's answer.
[0192] To facilitate better implementation of the oral language scoring model training method of this application, this application also provides an oral language scoring model training device. Please refer to... Figure 10 , Figure 10 This is a schematic diagram of the structure of the answer scoring device provided in an embodiment of this application. The answer scoring device 600 may include a first training acquisition unit 601, a training speech representation unit 602, a training language representation unit 603, a second training acquisition unit 604, a training key point representation unit 605, a training scoring unit 506, and an update unit 507.
[0193] The first training acquisition unit 601 is used to acquire a training dataset and an initial oral scoring model. The training dataset includes multiple training samples from the oral exam. Each training sample includes training speech of each candidate's answer, the label score of the candidate's answer, and a training text sequence and a training phoneme sequence of the candidate's full answer determined based on the training speech.
[0194] The training speech representation unit 602 is used to input the speech and the phoneme sequence into the speech level modeling module of the initial oral scoring model for encoding and decoding processing, so as to determine the training full-text speech level representation of the candidate's answer.
[0195] The training language representation unit 603 is used to input the training text sequence into the language proficiency modeling module of the initial oral scoring model for language extraction processing, so as to determine the language proficiency representation of the candidate's full-text training response.
[0196] The second training acquisition unit 604 is used to acquire the key points of multiple training sample texts for the oral exam.
[0197] The training key point representation unit 605 is used to input multiple training sample key points and the training text sequence into the key point processing module of the initial oral scoring module for key point semantic processing, so as to determine the global training key point representation of the candidate's answer.
[0198] The training scoring unit 606 is used to input the full-text speech level representation, the full-text language level representation, and the global key point representation into the fusion module of the initial oral scoring module for scoring processing, so as to obtain the training score of the candidate's answer.
[0199] The update unit 607 is used to update the initial spoken language scoring model based on the training score and the label score to obtain a spoken language scoring model.
[0200] All the above-mentioned technical solutions can be combined in any way to form optional embodiments of this application, and the beneficial effects that can be achieved are described above, and will not be repeated here.
[0201] Accordingly, embodiments of this application also provide a computer device, which can be a terminal or a server. For example... Figure 11 As shown, Figure 11 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. The computer device 700 includes a processor 701 with one or more processing cores, a memory 702 with one or more computer-readable storage media, and a computer program stored in the memory 702 and executable on the processor. The processor 701 is electrically connected to the memory 702.
[0202] The processor 701 is the control center of the computer device 700. It connects various parts of the computer device 700 through various interfaces and lines. By running or loading software programs (computer programs) and / or modules stored in the memory 702, and calling data stored in the memory 702, it performs various functions of the computer device 700 and processes data, thereby monitoring the computer device 700 as a whole.
[0203] In this embodiment, the processor 701 in the computer device 700 loads the instructions corresponding to the processes of one or more applications into the memory 702 according to the following steps, and the processor 701 runs the applications stored in the memory 702 to achieve the functions of any of the above method embodiments, such as the steps in any embodiment of the above answer scoring method and / or the steps in any embodiment of the above oral scoring model training method. Please refer to the above description for details.
[0204] The specific implementation and beneficial effects of each operation / step that the processor can execute can be found in the preceding method embodiments, and will not be repeated here.
[0205] Optional, such as Figure 11 As shown, the computer device 700 also includes: a touch screen display 703, a radio frequency circuit 704, an audio circuit 705, an input unit 706, and a power supply 707. The processor 701 is electrically connected to the touch screen display 703, the radio frequency circuit 704, the audio circuit 705, the input unit 706, and the power supply 707. Those skilled in the art will understand that... Figure 11 The computer device structure shown does not constitute a limitation on the computer device and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0206] The touch display screen 703 can be used to display a graphical user interface (GUI) and receive operation commands generated by the user interacting with the GUI. The touch display screen 703 may include a display panel and a touch panel. The display panel can be used to display information input by the user or information provided to the user, as well as various graphical user interfaces of the computer device. These graphical user interfaces can be composed of graphics, text, icons, video, and any combination thereof. Optionally, the display panel can be configured using a liquid crystal display (LCD), an organic light-emitting diode (OLED), or other similar devices. The touch panel can be used to collect touch operations performed by the user on or near it (such as operations performed by the user using a finger, stylus, or any suitable object or accessory on or near the touch panel), generate corresponding operation commands, and execute the corresponding program. The touch panel may cover the display panel. When the touch panel detects a touch operation on or near it, it transmits the data to the processor 701 to determine the type of touch event. Subsequently, the processor 701 provides corresponding visual output on the display panel based on the type of touch event. In this embodiment, the touch panel and display panel can be integrated into the touch display screen 703 to achieve input and output functions. However, in some embodiments, the touch panel and the touch display screen 703 can be implemented as two independent components to achieve input and output functions. That is, the touch display screen 703 can also be used as part of the input unit 706 to achieve input functions.
[0207] In this embodiment, the touch display screen 703 is used to present a graphical user interface and receive operation commands generated by the user interacting with the graphical user interface.
[0208] The radio frequency circuit 704 can be used to transmit and receive radio frequency signals to establish wireless communication with network devices or other computer devices, and to transmit and receive signals with network devices or other computer devices.
[0209] Audio circuitry 705 can be used to provide an audio interface between a user and a computer device via a speaker and a microphone. Audio circuitry 705 converts received audio data into electrical signals, transmits them to the speaker, and the speaker converts them into sound signals for output. Conversely, the microphone converts collected sound signals into electrical signals, which are then received by audio circuitry 705, converted back into audio data, and then processed by processor 701 before being transmitted via radio frequency circuitry 704 to, for example, another computer device, or output to memory 702 for further processing. Audio circuitry 705 may also include an earphone jack to facilitate communication between peripheral headphones and the computer device.
[0210] The input unit 706 can be used to receive input numbers, characters, or user characteristic information (such as fingerprints, iris, facial information, etc.), and to generate keyboard, mouse, joystick, optical, or trackball signal inputs related to user settings and function control.
[0211] Power supply 707 is used to supply power to various components of computer device 700. Optionally, power supply 707 can be logically connected to processor 701 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. Power supply 707 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.
[0212] although Figure 11 As not shown in the diagram, the computer device 700 may also include a camera, sensors, a wireless fidelity module, a Bluetooth module, etc., which will not be described in detail here.
[0213] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0214] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.
[0215] Therefore, embodiments of this application provide a computer-readable storage medium storing a plurality of computer programs that can be loaded by a processor to execute steps in any of the answer scoring methods provided in embodiments of this application. For example, the computer program can execute steps in any embodiment of the above-described answer scoring method, and / or steps in any embodiment of the above-described oral scoring model training method.
[0216] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.
[0217] The storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0218] Since the computer program stored in the storage medium can execute the steps in any of the answer scoring methods provided in the embodiments of this application, the beneficial effects that any of the answer scoring methods provided in the embodiments of this application can achieve can be realized, as detailed in the preceding embodiments, and will not be repeated here.
[0219] The foregoing has provided a detailed description of an answer scoring method, apparatus, storage medium, and computer device provided in the embodiments of this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.< / cls> < / cls> < / cls> < / cls> < / cls> < / cls> < / cls> < / cls> < / cls> < / cls> < / cls> < / cls> < / cls> < / sep> < / cls> < / sep> < / cls> < / sep> < / cls> < / sep> < / cls> < / sep> < / cls> < / sep> < / cls> < / sep> < / cls> < / sep> < / cls> < / sep> < / cls> < / cls> < / cls> < / cls>
Claims
1. A method for scoring answers, characterized in that, include: The method obtains the audio of a candidate's answer in an oral exam, as well as a text sequence of the candidate's full answer and a phoneme sequence of the candidate's full answer, determined based on the audio. The text sequence is all the text corresponding to the audio obtained after speech recognition processing, and the phoneme sequence is obtained after conversion processing of the text sequence. The speech is divided into multiple speech segments, and multiple phoneme segments corresponding to the multiple speech segments are determined; For each speech segment in multiple speech segments and each phoneme segment corresponding to each speech segment in multiple phoneme segments, encoding and decoding processes are performed to obtain the hidden layer representation of each speech segment; The hidden layer representations of the multiple speech segments obtained are processed at the phoneme level to obtain the full-text speech level representation of the candidate's answer; The text sequence of the candidate's full answer, plus the start identifier of the text sequence, is input into the language model for full-text semantic extraction processing to obtain the full-text hidden layer output; The hidden output at the position corresponding to the start identifier of the text sequence in the full-text hidden output is determined as the full-text language proficiency representation of the candidate's answer; The key points of multiple sample essays for the oral exam are obtained, and the key points of each sample essay and the text sequence are extracted to determine the key points for the candidate's answer corresponding to each sample essay. For each sample essay's key points and the corresponding key points of the candidate's answer, attention semantic extraction processing at the key point level is performed to obtain a global key point representation of the candidate's answer; The candidate's answer is scored based on the full-text speech level representation, the full-text language level representation, and the global key point representation to obtain the candidate's answer score.
2. The method according to claim 1, characterized in that, The step of performing attention semantic extraction processing at the key point level for each sample essay and the corresponding key points of the candidate's answer to obtain a global key point representation of the candidate's answer includes: For each sample essay key point and the corresponding candidate's answer key point, semantic extraction processing at the key point level is performed to obtain the full text key point semantic representation of the candidate's answer key point, as well as the hidden semantic representation of the contextual similarity between each sample essay key point and the corresponding candidate's answer key point; The full-text key semantic representation and the hidden layer semantic representation are processed by an attention mechanism to obtain a global key representation after attention weighting of the hidden layer semantic representation.
3. The method according to claim 2, characterized in that, The method further includes an oral assessment model, which includes a key-point level semantic correctness module. The step of performing key-point level semantic extraction processing on each model essay key point and the corresponding key points of the candidate's answer to obtain the full-text key point semantic representation of the candidate's answer and the implicit semantic representation of multiple model essay key points includes: The text is concatenated with the key points of each sample essay and the corresponding key points of the candidate's answer to obtain a text of each key point pair, wherein each text of the key point pair includes a start identifier; The global key point representation identifier and multiple key point pairs are input into the key point matching pre-trained model in the key point level semantic right and wrong module to perform key point level semantic extraction processing, so as to obtain the full text key point semantic representation of the candidate's answer and the hidden layer semantic representation corresponding to multiple key point pairs. The latent semantic representation at the position corresponding to the start identifier of each key point in the text is extracted as the latent semantic representation of the contextual similarity between each key point in the text.
4. The method according to claim 2, characterized in that, The method further includes a spoken language scoring model, which includes a key point importance module. The step of performing attention mechanism processing on the full-text key point semantic representation and the hidden semantic representation to obtain a global key point representation after attention weighting of the hidden semantic representation includes: The semantic representation of the full text's key points and the semantic representation of the hidden layer are concatenated to obtain the concatenated semantic representation; The concatenated semantic representation is input into the key importance model of the key importance module for attention weight processing to obtain the attention weights corresponding to the full-text key semantic representation and multiple hidden layer semantic representations. The global key point representation is determined by weighting the attention weights corresponding to the hidden semantic representations and the corresponding hidden semantic representations.
5. The method according to claim 1, characterized in that, The method further includes an oral assessment model, which includes a key point-level semantic correctness module. The step of extracting key points from each of the multiple model essays and the text sequence to determine the corresponding key points for the candidate's answer includes: Each key point from multiple sample texts is concatenated with the text sequence to obtain multiple concatenated text sequences. Each of the multiple concatenated text sequences is input into the key point extraction model in the key point level semantic right and wrong module for key point extraction processing, so as to obtain the key points for the candidate's answer corresponding to each sample essay key point.
6. The method according to claim 1, characterized in that, The steps for obtaining multiple sample essays for the oral exam include: Obtain the images from the picture-based speaking topics corresponding to the oral exam; The examination images are subjected to image encoding processing to obtain the image encoding result; Based on the image encoding results, multiple sample essay points for the exam image are generated.
7. The method according to claim 1, characterized in that, The method further includes an oral assessment model, which includes a fusion module. The step of scoring the candidate's answer based on the full-text speech level representation, the full-text language level representation, and the global key point representation to obtain the score of the candidate's answer includes: The full-text speech level representation, the full-text language level representation, and the global key point representation are fused together to obtain a fused hidden layer representation that incorporates the key point representation. The fusion hidden layer representation is input into the fusion model in the fusion module for nonlinear scoring processing to obtain the score of the candidate's answer.
8. A method for training a spoken language scoring model, characterized in that, include: Obtain a training dataset and an initial oral scoring model. The training dataset includes multiple training samples from the oral exam. Each training sample includes training speech for each examinee's answer, a label score for the examinee's answer, a training text sequence of the examinee's full answer determined based on the training speech, and a training phoneme sequence of the full answer. The training text sequence is all the text corresponding to the training speech obtained after speech recognition processing, and the training phoneme sequence is obtained after transformation processing of the training text sequence. The training speech is divided into multiple training speech segments, and multiple training phoneme segments corresponding to the multiple training speech segments are determined; Each training speech segment from multiple training speech segments and each training phoneme segment corresponding to each training speech segment from multiple training phoneme segments are input into the speech model in the speech level modeling module for encoding and decoding processing to obtain the training hidden layer representation of each training speech segment. The training hidden layer representations of the multiple training speech segments obtained are processed at the phoneme level to obtain the training full-text speech level representation of the candidate's answer. The training text sequence is input into the language model in the language proficiency modeling module of the initial oral scoring model for language extraction processing, so as to determine the language proficiency representation of the candidate's full-text training response. The key points of multiple training sample texts for the oral exam are obtained, and each key point of the multiple training sample texts and the training text sequence are input into the key point extraction model of the key point level semantic right and wrong module for key point extraction processing, so as to determine the key points of the candidate's answer corresponding to each training sample text. Each training sample key point and the corresponding candidate answer training key point are input into the key point matching pre-training model of the key point level semantic correctness module for key point level semantic extraction processing, so as to obtain the semantic representation of the full text training key points of the candidate answer, and the training hidden layer semantic representation of the contextual similarity between each training sample key point and the corresponding candidate answer training key point. The training full-text key point semantic representation and the training hidden layer semantic representation are input into the key point importance model of the key point importance module for attention mechanism processing to obtain the training global key point representation after attention weighting of the training hidden layer semantic representation. The training full-text speech level representation, the training full-text language level representation, and the training global key point representation are input into the fusion module of the initial oral scoring model for scoring processing to obtain the training score of the candidate's answer; The initial spoken language scoring model is updated based on the training score and the label score to obtain the spoken language scoring model.
9. The method according to claim 8, characterized in that, The method further includes: when performing key point extraction processing, obtaining the first similarity between the key points of each training sample and the key points of the candidate's answer training; After obtaining the trained hidden layer semantic representation, the process also includes: The training full-text key semantic representation and the training hidden layer semantic representation are input into the newly added linear mapping layer for linear mapping processing to obtain the second similarity corresponding to each training hidden layer semantic representation; The second loss value is determined based on the first similarity and the second similarity; The step of updating the initial spoken language scoring model based on the training score and the label score to obtain the spoken language scoring model includes: A first loss value is determined based on the training score and the label score; The initial spoken language scoring model is updated based on the first loss value and the second loss value to obtain the spoken language scoring model.
10. The method according to claim 9, characterized in that, The step of updating the initial spoken language scoring model based on the first loss value and the second loss value to obtain the spoken language scoring model includes: Determine the first coefficient and the second coefficient, and based on the first coefficient and the second coefficient, perform a weighted summation on the first loss value and the second loss value respectively to obtain the total loss value; The initial spoken language scoring model is updated based on the overall loss value, wherein the sum of the first coefficient and the second coefficient is 1, and the first coefficient is greater than the second coefficient.
11. An answer scoring device, characterized in that, include: The first acquisition unit is used to acquire the voice of the candidate's answer in the oral test, as well as the text sequence of the candidate's full answer and the phoneme sequence of the candidate's full answer determined based on the voice; wherein, the text sequence is all the text corresponding to the voice obtained after speech recognition processing of the voice, and the phoneme sequence is obtained after conversion processing of the text sequence; The speech representation unit is used to divide the speech into multiple speech segments and determine multiple phoneme segments corresponding to the multiple speech segments; to perform encoding and decoding processing on each speech segment in the multiple speech segments and each phoneme segment corresponding to each speech segment in the multiple phoneme segments to obtain the hidden layer representation of each speech segment; and to process the obtained hidden layer representations of the multiple speech segments according to the phoneme level to obtain the full-text speech level representation of the candidate's answer. The language representation unit is used to input the text sequence of the candidate's full answer, plus the start identifier of the text sequence, into the language model for full-text semantic extraction processing to obtain the full-text hidden layer output; the hidden layer output at the position corresponding to the start identifier of the text sequence in the full-text hidden layer output is determined as the full-text language level representation of the candidate's answer. The second acquisition unit is used to acquire multiple sample essays and key points for the oral exam; The key point representation unit is used to extract key points from each of the multiple sample essay key points and the text sequence to determine the corresponding key points for the candidate's answer; and to perform key point-level attention semantic extraction processing on each sample essay key point and the corresponding key points for the candidate's answer to obtain a global key point representation of the candidate's answer. The scoring unit is used to score the candidate's answer based on the full-text speech level representation, the full-text language level representation, and the global key point representation, so as to obtain the score of the candidate's answer.
12. A training device for an oral assessment model, characterized in that, include: The first training acquisition unit is used to acquire a training dataset and an initial oral scoring model. The training dataset includes multiple training samples from the oral exam. Each training sample includes training speech for each examinee's answer, a label score for the examinee's answer, a training text sequence of the examinee's full answer determined based on the training speech, and a training phoneme sequence of the full answer. The training text sequence is all the text corresponding to the training speech obtained after speech recognition processing of the training speech, and the training phoneme sequence is obtained after conversion processing of the training text sequence. The training speech representation unit is used to divide the training speech into multiple training speech segments and determine multiple training phoneme segments corresponding to the multiple training speech segments; each training speech segment in the multiple training speech segments and each training phoneme segment corresponding to each training speech segment in the multiple training phoneme segments are input into the speech model in the speech level modeling module for encoding and decoding processing to obtain the training hidden layer representation of each training speech segment; the obtained training hidden layer representations of the multiple training speech segments are processed according to the phoneme level to obtain the training full-text speech level representation of the candidate's answer; The training language representation unit is used to input the training text sequence into the language model in the language proficiency modeling module of the initial oral scoring model for language extraction processing, so as to determine the language proficiency representation of the candidate's full-text training response. The second training acquisition unit is used to acquire the key points of multiple training sample texts for the oral exam; The training key point representation unit is used to input each training model key point and the training text sequence from multiple training model key points into the key point extraction model of the key point level semantic correctness module for key point extraction processing to determine the corresponding candidate answer training key points for each training model key point; input each training model key point and the corresponding candidate answer training key points into the key point matching pre-training model of the key point level semantic correctness module for key point level semantic extraction processing to obtain the training full text key point semantic representation of the candidate answer, and the training hidden layer semantic representation of the contextual similarity between each training model key point and the corresponding candidate answer training key point; input the training full text key point semantic representation and the training hidden layer semantic representation into the key point importance model of the key point importance module for attention mechanism processing to obtain the training global key point representation after attention weighting of the training hidden layer semantic representation. The training scoring unit is used to input the training full-text speech level representation, the training full-text language level representation, and the training global key point representation into the fusion module of the initial oral scoring model for scoring processing, so as to obtain the training score of the candidate's answer; An update unit is used to update the initial spoken language scoring model based on the training score and the label score to obtain a spoken language scoring model.
13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program adapted for loading by a processor to perform the steps of the method as described in any one of claims 1-10.
14. A computer device, characterized in that, The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the steps of the method as described in any one of claims 1-10 by invoking the computer program stored in the memory.