Abnormal score detection method and device, equipment and computer readable storage medium
Patent Information
- Application Number
- CN202110214645.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-02-25
- Publication Date
- 2026-09-08
- Estimated Expiration
- 2041-02-25
AI Technical Summary
[0003]对于口语考试系统而言,其采用人机对话的方式,考生只需通过计算机和耳麦设备即可完成对口语试题的作答与全自动智能评分,由于口语考试中的开放题型的音频答案具有多样性,因此,这种全自动智能评分可能存在评分不准确的情况,然而,相关技术尚缺乏对评分进行异常检测的有效手段
[0072] Based on the first multimodal feature extracted from the audio answer using text content and the second multimodal feature extracted from the reference audio, a reference score for the audio answer is determined. Anomaly detection is then performed on the original score of the audio answer in conjunction with the reference score. This enables effective detection of abnormal scores, thereby effectively filtering out abnormal original scores and ultimately making the oral exam scoring as accurate as possible.
Smart Images

Figure CN113590772B_ABST
Abstract
Description
Technical Field
[0001] This application relates to artificial intelligence technology, and more particularly to a method, apparatus, device, and computer-readable storage medium for detecting anomaly scores. Background Technology
[0002] Artificial intelligence (AI) is the theory, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. An increasing number of AI products possess question-and-answer scoring capabilities, which can be applied to various voice scoring systems, such as encyclopedia question-and-answer systems, language testing systems for language education applications, and oral examination systems.
[0003] For oral examination systems, which use a human-computer dialogue method, candidates only need to use a computer and headset to answer oral test questions and receive fully automatic intelligent scoring. However, due to the diversity of audio answers for open-ended questions in oral examinations, this fully automatic intelligent scoring may result in inaccurate scoring. Nevertheless, the relevant technology currently lacks effective means to detect anomalies in the scoring. Summary of the Invention
[0004] This application provides a method, apparatus, device, and computer-readable storage medium for detecting anomaly scores, which can effectively detect anomaly scores.
[0005] The technical solution of this application embodiment is implemented as follows:
[0006] This application provides a method for detecting abnormal scores, including:
[0007] Obtain the audio answer to the target question, and the text content corresponding to the audio answer;
[0008] Based on the text content, multimodal features are extracted from the audio answer to obtain the first multimodal feature of the audio answer;
[0009] Obtain the second multimodal features of the reference audio corresponding to the target problem;
[0010] The first multimodal feature of the audio answer is matched with the second multimodal feature of the reference audio, and a reference score for the audio answer is determined based on the matching result.
[0011] Obtain the original score of the audio answer, and perform anomaly detection on the original score based on the reference score to obtain a detection result that characterizes whether the original score is abnormal.
[0012] In the above scheme, after performing anomaly detection on the original score based on the reference score to obtain a detection result characterizing whether the original score is abnormal, the method further includes:
[0013] When the detection result indicates that the original score is an abnormal score, a correction prompt message corresponding to the audio answer is sent;
[0014] The correction prompt information is used to prompt the audio answer to undergo scoring and correction processing.
[0015] This application provides a method for detecting abnormal scores, including:
[0016] A scoring and detection interface is presented, and at least one question and a corresponding scoring and detection function item are displayed in the scoring and detection interface.
[0017] In response to the triggering operation of the scoring detection function item corresponding to the target question, the information input interface corresponding to the target question is presented;
[0018] Based on the information input interface, the audio answer to the target question and the corresponding original score are received;
[0019] In response to a rating detection instruction triggered based on the audio answer and the corresponding original rating, a detection result is output to characterize whether the original rating is abnormal;
[0020] The detection result is obtained by performing anomaly detection on the original score based on the reference score of the audio answer. The reference score is determined based on the matching result between the first multimodal feature of the audio answer and the second multimodal feature of the reference audio of the target question.
[0021] This application provides an anomaly scoring detection device, including:
[0022] The first acquisition module is used to acquire the audio answer to the target question and the text content corresponding to the audio answer;
[0023] The feature extraction module is used to perform multimodal feature extraction on the audio answer based on the text content to obtain the first multimodal feature of the audio answer;
[0024] The second acquisition module is used to acquire the second multimodal features of the reference audio corresponding to the target problem;
[0025] The feature matching module is used to match the first multimodal features of the audio answer with the second multimodal features of the reference audio, and determine the reference score of the audio answer based on the matching result;
[0026] The scoring detection module is used to obtain the original score of the audio answer, and to perform anomaly detection on the original score based on the reference score, so as to obtain a detection result that characterizes whether the original score is abnormal.
[0027] In the above scheme, the feature extraction module is further used to extract features from the text content to obtain the text features of the text content;
[0028] Feature extraction is performed on the audio answer to obtain its audio features;
[0029] By fusing the text features and the audio features, the first multimodal feature of the audio answer is obtained.
[0030] In the above scheme, the feature extraction module is also used to perform word segmentation on the text content to obtain multiple words corresponding to the text content;
[0031] Each of the aforementioned words is feature-encoded to obtain the word features corresponding to each of the aforementioned words;
[0032] The text features of the text content are obtained by concatenating the word features corresponding to each of the aforementioned words.
[0033] In the above scheme, the feature extraction module is further used to perform bidirectional encoding processing on the word features of each word to obtain the context encoding features and context encoding features corresponding to each word;
[0034] The context encoding features and context encoding features of each word are concatenated to obtain the concatenated encoding features corresponding to each word.
[0035] The concatenated encoding features corresponding to each word are concatenated to obtain the text features of the text content.
[0036] In the above scheme, the feature extraction module is further used to concatenate the text features and audio features of each word to obtain the concatenated features of the word.
[0037] Obtain the weight corresponding to each of the aforementioned words;
[0038] Based on the acquired weights, the concatenation features of each word are weighted and summed to obtain the first multimodal feature of the audio answer.
[0039] In the above scheme, the feature processing module is further used to obtain multiple sample scores corresponding to the target question, and each sample score corresponds to at least one reference audio.
[0040] The first multimodal feature of the audio answer is matched with the second multimodal feature of each of the reference audios to obtain a first similarity value between the first multimodal feature and each second multimodal feature;
[0041] Based on the obtained first similarity values and corresponding sample scores, a reference score for the audio answer is determined.
[0042] In the above scheme, when each sample score corresponds to multiple reference audios, each sample score corresponds to multiple first similarity values. The feature processing module is also used to average the multiple first similarity values corresponding to each sample score to obtain a second similarity value corresponding to each sample score.
[0043] Obtain the aggregation degree measure corresponding to each sample score, and based on the aggregation degree measure corresponding to each sample score, normalize the second similarity value of the corresponding sample score to obtain the third similarity value of each sample score.
[0044] From the third similarity values of each sample score, the sample score corresponding to the largest third similarity value is selected as the reference score for the audio answer.
[0045] In the above scheme, the feature processing module is further configured to perform the following operations for each of the sample scores:
[0046] The similarity between features of the multiple second multimodal features corresponding to the sample score is matched to obtain multiple fourth similarity values corresponding to the sample score;
[0047] The aggregation degree measure corresponding to the sample score is obtained by averaging multiple fourth similarity values.
[0048] In the above scheme, the scoring detection module is also used to obtain the score difference between the reference score and the original score;
[0049] When the score difference exceeds the difference threshold and the maximum similarity value exceeds the similarity threshold, the original score is determined to be an abnormal score.
[0050] In the above scheme, the feature processing module is further used to extract features from the audio answer through the first feature extraction layer of the scoring model to obtain the audio features of the audio answer;
[0051] The text content is feature extracted through the second feature extraction layer of the scoring model to obtain the text features of the text content;
[0052] The scoring prediction layer of the scoring model predicts the score of the audio answer based on its audio features and text features, thus obtaining the original score of the audio answer.
[0053] In the above scheme, the rating prediction layer includes a first sub-prediction layer, a second sub-prediction layer, a third sub-prediction layer, and a rating fusion layer. The feature processing module is further used for...
[0054] Based on the audio features of the audio answer, the first sub-prediction layer is used to predict the pronunciation score of the audio answer to obtain the pronunciation score of the audio answer;
[0055] Based on the text features of the audio answer, the second sub-prediction layer is used to predict the grammar score of the audio answer to obtain the grammar score of the audio answer.
[0056] The third sub-prediction layer matches the text features of the audio answer with the text features of the reference audio, and determines the accuracy score of the audio answer based on the matching results.
[0057] The original score of the audio answer is obtained by fusing the pronunciation score, the grammar score, and the accuracy score through the scoring fusion layer.
[0058] In the above scheme, after performing anomaly detection on the original score based on the reference score to obtain a detection result characterizing whether the original score is abnormal, the device further includes:
[0059] The information sending module is used to send correction prompt information corresponding to the audio answer when the detection result indicates that the original score is an abnormal score;
[0060] The correction prompt information is used to prompt the audio answer to undergo scoring and correction processing.
[0061] This application provides an anomaly scoring detection device, including:
[0062] The first presentation module is used to present the scoring detection interface and to present at least one question and corresponding scoring detection function items in the scoring detection interface.
[0063] The second presentation module is used to present the information input interface corresponding to the target question in response to the trigger operation of the scoring detection function item corresponding to the target question;
[0064] The information receiving module is used to receive the audio answer and corresponding original score of the target question based on the information input interface;
[0065] The result output module is used to respond to a scoring detection instruction triggered based on the audio answer and the corresponding original score, and output a detection result to characterize whether the original score is abnormal.
[0066] The detection result is obtained by performing anomaly detection on the original score based on the reference score of the audio answer; the reference score is determined based on the matching result between the first multimodal feature of the audio answer and the second multimodal feature of the reference audio of the target question.
[0067] This application provides an electronic device, including:
[0068] Memory, used to store executable instructions;
[0069] The processor, when executing executable instructions stored in the memory, implements the anomaly scoring detection method provided in the embodiments of this application.
[0070] This application provides a computer-readable storage medium storing executable instructions for inducing a processor to execute and implement the abnormal scoring detection method provided in this application.
[0071] The embodiments of this application have the following beneficial effects:
[0072] Based on the first multimodal feature extracted from the audio answer using text content and the second multimodal feature extracted from the reference audio, a reference score for the audio answer is determined. Anomaly detection is then performed on the original score of the audio answer in conjunction with the reference score. This enables effective detection of abnormal scores, thereby effectively filtering out abnormal original scores and ultimately making the oral exam scoring as accurate as possible. Attached Figure Description
[0073] Figure 1 A schematic diagram of the architecture of the anomaly scoring detection system 100 provided in the embodiments of this application;
[0074] Figure 2 This is an optional structural schematic diagram of the electronic device 500 provided in an embodiment of this application;
[0075] Figure 3 A flowchart illustrating the anomaly scoring detection method provided in this application embodiment;
[0076] Figure 4 A schematic diagram of the architecture of the classification model provided in the embodiments of this application;
[0077] Figure 5 This is a schematic diagram of the multimodal feature acquisition process provided in the embodiments of this application;
[0078] Figure 6This is a schematic diagram of the examination interface provided in an embodiment of this application;
[0079] Figure 7 A flowchart illustrating the anomaly scoring detection method provided in this application embodiment;
[0080] Figure 8 This is a schematic diagram of the scoring model provided in the embodiments of this application;
[0081] Figure 9 This is a schematic diagram of the scoring model provided in the embodiments of this application;
[0082] Figure 10 A flowchart illustrating the anomaly scoring detection method provided in this application embodiment;
[0083] Figure 11 This is a schematic diagram of the scoring and detection interface provided in an embodiment of this application;
[0084] Figure 12 This is a schematic diagram of the scoring display interface provided in an embodiment of this application;
[0085] Figure 13 This is a schematic diagram of the architecture of the anomaly scoring detection system provided in the embodiments of this application;
[0086] Figure 14 A schematic diagram of the structure of the anomaly scoring detection device provided in the embodiments of this application;
[0087] Figure 15 A schematic diagram of the structure of the detection device for anomaly scoring provided in the embodiments of this application. Detailed Implementation
[0088] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0089] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0090] In the following description, the terms “first, second…” are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that “first, second…” may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.
[0091] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0092] Before providing a further detailed description of the embodiments of this application, the nouns and terms involved in the embodiments of this application will be explained, and the nouns and terms involved in the embodiments of this application shall be interpreted as follows.
[0093] 1) Speech recognition technology: Automatic Speech Recognition (ASR) aims to convert the lexical content of human speech into computer-readable input, such as keystrokes, binary codes, or character sequences.
[0094] 2) Multimodal: From the perspective of semantic perception, multimodal data involves information received by different sensory channels such as vision, hearing, touch, and smell. From the data level, multimodal data can be seen as a combination of multiple data types, such as images, numbers, text, symbols, audio, time series, or composite data forms composed of different data structures such as sets, trees, and graphs, and even combinations of various information resources from different databases and knowledge bases.
[0095] See Figure 1 , Figure 1 This is a schematic diagram of the architecture of the anomaly scoring detection system 100 provided in the embodiments of this application. To support an exemplary application, the terminal 400 connects to the server 200 through the network 300, which can be a wide area network or a local area network, or a combination of both.
[0096] In some embodiments, the anomaly scoring detection method provided in this application can be implemented by either terminal 400 or server 200. When implemented by terminal 400 alone, it can be installed on terminal 400 as a client, enabling the client on terminal 400 to have local anomaly scoring detection functionality. It can also be used as a plugin for related clients, downloaded to the client for local use as needed. In this deployment method, the anomaly scoring detection system can complete all detection processes locally without accessing an external network, ensuring absolute data security.
[0097] In some embodiments, the abnormal scoring detection method provided in this application can be implemented collaboratively by a terminal 400 and a server 200. For example, the terminal 400 collects the audio answer corresponding to the target question and sends the collected audio answer to the server 200; the server 200 obtains the audio answer corresponding to the target question, performs text conversion on the audio answer to obtain the text content of the audio answer, and extracts multimodal features from the audio answer based on the text content to obtain the first multimodal feature of the audio answer; the server 200 also obtains the second multimodal feature of the reference audio corresponding to the target question, matches the first multimodal feature of the audio answer with the second multimodal feature of the reference audio, and determines the score based on the matching result. The server 200 obtains a reference score for the audio answer and also acquires the original score of the audio answer. Based on the reference score, the server performs anomaly detection on the original score to obtain a detection result that indicates whether the original score is abnormal. When the detection result indicates that the original score is abnormal, the server 200 sends a correction prompt to the administrator to prompt further processing, such as manual intervention or using other scoring models, to reduce the generation of abnormal scores and ultimately make the score as accurate as possible. The server 200 then sends the final accurate score to the terminal for display on the terminal 400's display interface. When the detection result indicates that the original score is normal, the server 200 directly sends the original score to the terminal 400 for display on the terminal's display interface.
[0098] In some embodiments, server 200 may be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. Terminal 400 may be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, etc., but is not limited to these. Terminal 400 and server 200 can be directly or indirectly connected via wired or wireless communication, which is not limited in this embodiment.
[0099] The electronic device implementing the anomaly scoring detection method of the embodiments of this application will now be described. See also Figure 2 , Figure 2 This is an optional structural diagram of the electronic device 500 provided in the embodiments of this application. In practical applications, the electronic device 500 can be... Figure 1 Terminal 400 or server 200 in the middle, with electronic devices as Figure 1 Taking server 200 as an example, Figure 2The illustrated electronic device 500 includes at least one processor 510, a memory 550, at least one network interface 520, and a user interface 530. The various components in the electronic device 500 are coupled together via a bus system 540. It is understood that the bus system 540 is used to implement communication between these components. In addition to a data bus, the bus system 540 also includes a power bus, a control bus, and a status signal bus. However, for clarity, ... Figure 2 The general labeled all buses as Bus System 540.
[0100] The processor 510 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc.
[0101] User interface 530 includes one or more output devices 531 that enable the presentation of media content, including one or more speakers and / or one or more visual displays. User interface 530 also includes one or more input devices 532, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.
[0102] The memory 550 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state storage, hard disk drives, optical disk drives, etc. The memory 550 may optionally include one or more storage devices physically located away from the processor 510.
[0103] The memory 550 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), and the volatile memory may be random access memory (RAM). The memory 550 described in this application embodiment is intended to include any suitable type of memory.
[0104] In some embodiments, memory 550 is capable of storing data to support various operations, examples of which include programs, modules, and data structures or subsets or supersets thereof, as illustrated below.
[0105] Operating system 551 includes system programs for handling various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, driver layer, etc., for implementing various basic business functions and handling hardware-based tasks;
[0106] The network communication module 552 is used to reach other computing devices via one or more (wired or wireless) network interfaces 520, exemplary network interfaces 520 including: Bluetooth, WiFi, and Universal Serial Bus (USB), etc.
[0107] Presentation module 553 is configured to enable the presentation of information (e.g., a user interface for operating peripheral devices and displaying content and information) via one or more output devices 531 (e.g., a display screen, a speaker, etc.) associated with user interface 530;
[0108] The input processing module 554 is used to detect and translate one or more user inputs or interactions from one or more input devices 532.
[0109] In some embodiments, the anomaly scoring detection device provided in this application can be implemented in software. Figure 2 An anomaly detection device 555, which is stored in memory 550, is shown. It can be software in the form of programs and plug-ins, and includes the following software modules: a first acquisition module 5551, a feature extraction module 5552, a second acquisition module 5553, a feature processing module 5554, and a score detection module 5555. These modules are logically related and can therefore be arbitrarily combined or further divided according to the functions they implement. The functions of each module will be described below.
[0110] In other embodiments, the anomaly scoring detection device provided in this application can be implemented in hardware. As an example, the anomaly scoring detection device provided in this application can be a processor in the form of a hardware decoding processor, which is programmed to execute the anomaly scoring detection method provided in this application. For example, the processor in the form of a hardware decoding processor can be one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.
[0111] As an example, the abnormal scoring detection method provided in this application embodiment can be applied to various voice question-and-answer scoring scenarios, such as oral examination systems, brain teasers, various language education clients, and encyclopedia knowledge question-and-answer systems. A voice robot presents a target question to be answered, the user answers the target question, and the abnormal scoring detection method scores the user's audio answer and detects abnormalities in the score to determine whether the audio answer score is abnormal. When the score is abnormal, a correction prompt message is presented to correct the audio answer score so that the audio answer can be re-scored to obtain a normal score for presentation. When the score is normal, the corresponding score is presented.
[0112] Based on the above description of the anomaly scoring detection system and electronic device provided in the embodiments of this application, the anomaly scoring detection method provided in the embodiments of this application will be described next. See [link to documentation]. Figure 3 , Figure 3 This is a flowchart illustrating the anomaly scoring detection method provided in the embodiments of this application. Figure 1 The following description uses the server 200 in the embodiment of this application to illustrate the detection of the abnormal score.
[0113] Step 101: The server obtains the audio answer to the target question and the text content corresponding to the audio answer.
[0114] Here, the audio answer can be obtained directly from the server or by sending the audio answer through the terminal. In practical applications, the terminal can be equipped with a client for question and answer. When the user opens the client, the terminal presents a question and answer interface, which displays the target question to be answered and the corresponding answer function item. In response to the trigger operation of the answer function item, the terminal obtains the audio answer for the target question and sends the audio answer to the server. After receiving the audio answer sent by the terminal, the server performs text conversion on the audio answer to obtain the text content of the audio answer.
[0115] Step 102: Based on the text content, perform multimodal feature extraction on the audio answer to obtain the first multimodal feature of the audio answer.
[0116] In some embodiments, the server may extract multimodal features from the audio answer based on the text content in the following manner to obtain the first multimodal feature of the audio answer: extract features from the text content to obtain text features of the text content; extract features from the audio answer to obtain audio features of the audio answer; and fuse the text features and audio features to obtain the first multimodal feature of the audio answer.
[0117] In some embodiments, the server may extract features from the text content in the following manner to obtain the text features of the text content: perform word segmentation on the text content to obtain multiple words corresponding to the text content; encode the features of each word to obtain the word features corresponding to each word; and concatenate the word features corresponding to each word to obtain the text features of the text content.
[0118] In this application, "multiple" refers to two or more. After obtaining the audio features and text features of the audio answer, the multimodal features of the audio answer are obtained by combining the audio features and text features of the audio answer. The audio features include at least one of the following dimensions: fluency, prosody, completeness, and accuracy. Since the multimodal features comprehensively consider the features of various dimensions of the audio answer, they are used for subsequent detection processing, which can achieve accurate detection function.
[0119] In some embodiments, the text content includes multiple words. The server can fuse text features and audio features to obtain the first multimodal feature of the audio answer by: concatenating the text features and audio features of each word to obtain the word concatenation features; obtaining the weights corresponding to each word; and performing a weighted summation of the word concatenation features based on the obtained weights to obtain the first multimodal feature of the audio answer.
[0120] Here, after obtaining the text and audio features of each word, the audio and text features of each word can be concatenated to obtain the concatenated features of each word. Then, an attention mechanism (such as self-attention) is used to process the concatenated features of each word to obtain the attention features of the audio answer, which serve as the first multimodal feature of the audio answer. During the attention process, the weights corresponding to each word are obtained. Based on the obtained weights, the concatenated features of each word are weighted and summed to obtain the first multimodal feature of the audio answer. In this way, the attention mechanism can learn the dependencies between elements in the concatenated features, thereby uncovering important features in the audio answer for subsequent detection processing, achieving accurate detection.
[0121] In some embodiments, the server may also extract first multimodal features of the audio answer based on a neural network classification model, see [link to relevant documentation]. Figure 4 , Figure 4This is a schematic diagram of the architecture of the classification model provided in this application embodiment. The model includes an encoding layer, an attention layer, and a classification prediction layer. The encoding layer includes a speech encoder and a text encoder. The speech encoder is used to extract speech features from the audio answer, and the text encoder is used to extract text features from the text content corresponding to the audio answer. Both the speech encoder and the text encoder are deep neural network structures and can be composed of multiple modules, such as convolutional neural networks. The attention layer is used to fuse the audio features and text features obtained from the encoding layer to obtain the multimodal features of the audio answer. The classification prediction layer is used to predict the classification result based on the fused features. As can be seen, by inputting the audio answer to be processed and the corresponding text content into the trained classification model, the first multimodal features of the audio answer can be obtained through the attention layer.
[0122] Next Figure 4 The training process of the classification model is explained below. When training the classification model, training samples are first constructed. Each training sample consists of audio answers and text content pairs, i.e., the training sample is constructed as: training sample (audio answer, text content). Training samples include positive and negative samples. The text content in the positive samples is obtained by speech recognition of the audio answers in the training samples. The text content in the negative samples is randomly replaced with other words from the vocabulary list according to a certain probability. That is, negative samples are constructed where the audio answer and text content do not match. Positive samples are labeled with matching word tags (e.g., 1), and negative samples are labeled with non-matching word tags (e.g., 0).
[0123] After the training samples are constructed, they are input into the classification model. The speech encoder in the coding layer extracts (encodes) the acoustic features of the audio answers in the training samples, thus obtaining the audio features h of the training samples. audio The text encoder in the encoding layer extracts (encodes) text features from the text content of the training samples, thus obtaining the text features h of the training samples. text Among them, in extracting text features h text First, the text content is segmented into words to obtain multiple words corresponding to the text content. Then, feature encoding is performed on each word to obtain the word feature h corresponding to each word. word(i) (Characterizing the word features of the i-th word), and finally concatenating the word features corresponding to each word to obtain the text features of the text content, h. text =h word(1) h word(2) , ..., h word(i) .
[0124] After obtaining the audio and text features of the training samples through the encoding layer, the attention layer performs attention processing on the audio and text features of the training samples to obtain the concatenated features of each word with fused audio features, as shown in the following expression:
[0125] h word(i) =Attention(h word(i) h audio h audio )+h word(i) (1)
[0126] Among them, h word(i) The word features represent the i-th word, Attention() is the attention function, and h audio For audio features, the expression for Attention() is as follows:
[0127]
[0128] Where Q is the query vector, K is the key vector, and V is the value vector, with the vector dimensions of K and Q being d. k In this application, Q is h word(i) K is h audio V is h audio Based on the attention mechanism, the degree of matching between each word feature and the corresponding audio feature can be obtained.
[0129] A classification prediction layer (fully connected layer) is used to classify and predict the concatenated features of the fused audio features for each word, resulting in a classification score that characterizes whether the corresponding word is correctly matched. word(i) The expression is as follows:
[0130] score word(i) =sigmoid(W word h word(i) +b word (3)
[0131] Where sigmoid() is a nonlinear activation function, h word(i) The word features of the i-th word, W word b represents the weights of trainable word features. word These are the corresponding trainable bias parameters.
[0132] The optimization objective of the classification model is to minimize the cross-loss entropy H(t,p) between the classification result and the true label. Here, the difference between the classification result and the matching label of the training sample is obtained, and the value of the loss function of the classification model is determined based on the difference, as shown in the following expression:
[0133]
[0134] Where t(x) is the classification result of whether the true predicted word is correctly matched, and p(x) is the probability of the model predicting the word correctly, that is, the matching label of the training sample.
[0135] When the value of the loss function reaches a preset threshold, a corresponding error signal is determined based on the value of the loss function of the classification model. This error signal is then backpropagated through the classification model, updating the model parameters of each layer during the propagation process. Here, we explain the backpropagation process: training samples are input into the input layer of the neural network model, pass through the hidden layers, and finally reach the output layer to output the result. This is the forward propagation process of the neural network model. Since there is an error between the output result and the actual result, the error between the output result and the actual value is calculated and backpropagated from the output layer to the hidden layers until it reaches the input layer. During the backpropagation process, the values of the model parameters are adjusted according to the error. This process is iterated until convergence.
[0136] The classification model can be trained using the methods described above. After training, the multimodal features of the audio answers can be extracted. The acquisition of multimodal features involves utilizing the complementarity between multimodal features and eliminating redundancy between modalities to learn better feature representations. This allows for the discovery of important features in the audio answers for subsequent detection processing, achieving accurate detection.
[0137] See Figure 5 , Figure 5 This is a schematic diagram of the multimodal feature acquisition process provided in this application embodiment. The audio answer to be processed and the corresponding text content of the audio answer are input into the trained classification model. The audio answer is feature-encoded by the speech encoder of the encoding layer to obtain the audio features of the audio answer. The text content corresponding to the audio answer is feature-encoded by the audio encoder of the encoding layer to obtain the word features of each word. Then, the word features and audio features of each word are concatenated by the attention layer to obtain the concatenated vector of each word and obtain the weight of each word. Based on the obtained weight, the concatenated features of each word are weighted and summed to obtain the first multimodal feature of the audio answer.
[0138] In some embodiments, the server may also extract features from the text content in the following manner to obtain text features of the text content: perform bidirectional encoding on the word features of each word to obtain the context encoding features and context encoding features corresponding to each word; concatenate the context encoding features and context encoding features of each word to obtain the concatenated encoding features corresponding to each word; and concatenate the concatenated encoding features corresponding to each word to obtain text features of the text content.
[0139] Here, considering the contextual features of words, after obtaining the word features (word vectors) of each word, the word features of each word are input into a bidirectional encoding layer, such as a Bidirectional Long Short-Term Memory (Bi-LSTM) layer. The Bi-LSTM layer includes two LSTMs: one for the forward input sequence and one for the backward input sequence. The contextual encoding features corresponding to each word are extracted through a forward process (e.g., from left to right), and the contextual encoding features corresponding to each word are extracted through a backward process (e.g., from right to left). The contextual encoding features and contextual encoding features of each word are concatenated to obtain the concatenated encoding features of the corresponding word. The concatenated encoding features of each word are then concatenated to obtain the text features of the text content.
[0140] Step 103: Obtain the second multimodal features of the reference audio corresponding to the target problem.
[0141] Here, when the target question is a semi-open-ended question (such as listening and retelling, describing a picture, etc.) or an open-ended question, such as... Figure 6 When asked the open-ended question "What's your favorite sports?", Figure 6 The diagram shows an exam interface provided in this application embodiment. Different users may have different answers to this question, therefore, there may be multiple reference audios for this question.
[0142] In practical implementation, the multimodal features of the reference audio can be extracted according to the method described above. Specifically, the reference audio corresponding to the target question and the text content corresponding to the reference audio are input into the trained classification model. The reference audio is feature-encoded by the audio encoder of the encoding layer to obtain the audio features of the reference audio. The text content corresponding to the reference audio is feature-encoded by the audio encoder of the encoding layer to obtain the word features of each word. Then, the word features and audio features of each word are concatenated by the attention layer to obtain the concatenated vector of each word and obtain the weight of each word. Based on the obtained weights, the concatenated features of each word are weighted and summed to obtain the second multimodal features of the reference audio.
[0143] Step 104: Match the first multimodal features of the audio answer with the second multimodal features of the reference audio, and determine the reference score of the audio answer based on the matching results.
[0144] See Figure 7 , Figure 7This is a flowchart illustrating the anomaly scoring detection method provided in the embodiments of this application. In some embodiments, Figure 7 Show Figure 3 Step 104 can be achieved through steps 1041-1043:
[0145] Step 1041: Obtain multiple sample scores corresponding to the target question, with each sample score corresponding to at least one reference audio.
[0146] Step 1042: Perform similarity matching between the first multimodal features of the audio answer and the second multimodal features of each reference audio to obtain the first similarity value between the first multimodal features and each second multimodal feature;
[0147] Step 1043: Based on the obtained first similarity values and corresponding sample scores, determine the reference score for the audio answer.
[0148] Here, reference audio is used for anomaly detection in user audio answers. The server pre-stores one or more reference audios associated with sample scores. Each sample score corresponds to a score range. For example, for the same target question, there may be multiple sample scores such as 45, 70, 80, 90, and 100. Each sample score corresponds to one, two, or more reference audios. For example, for a score of 80, there may be reference audios in various expressions. The first multimodal features of the user's audio answer are matched with the second multimodal features of the reference audios for similarity matching. For example, cosine distance is used to calculate the similarity value between each pair of features, resulting in a first similarity value between the first multimodal features and each second multimodal feature. This yields multiple first similarity values. The sample score corresponding to the largest first similarity value among these multiple first similarity values is selected as the reference score for the audio answer.
[0149] For example, the first similarity values between the first multimodal features of the audio answer and the second multimodal features of reference audio 1 (45 points), reference audio 2 (70 points), reference audio 3 (80 points), reference audio 4 (90 points), and reference audio 5 (100 points) are 0.2, 0.4, 0.8, 0.3, and 0.6, respectively. Then, the sample score (80 points) corresponding to the largest first similarity value (0.8) is selected as the reference score for the audio answer.
[0150] In some embodiments, when each sample score corresponds to multiple reference audios, each sample score corresponds to multiple first similarity values. The server can determine the reference score for the audio answer based on the obtained first similarity values and the corresponding sample scores in the following manner:
[0151] The average of multiple first similarity values corresponding to each sample score is calculated to obtain the second similarity value for each sample score. The aggregation degree measure corresponding to each sample score is obtained, and the second similarity value of the corresponding sample score is normalized based on the aggregation degree measure to obtain the third similarity value for each sample score. From the third similarity values of each sample score, the sample score corresponding to the largest third similarity value is selected as the reference score for the audio answer.
[0152] In some embodiments, the server can obtain the aggregation degree measure corresponding to each sample score in the following manner:
[0153] Perform the following operations for each sample score: perform feature similarity matching on multiple second multimodal features of the sample score to obtain multiple fourth similarity values corresponding to the sample score; average the multiple fourth similarity values corresponding to the sample score to obtain the aggregation degree measure corresponding to the sample score.
[0154] The aggregation degree metric is used to characterize the degree of aggregation among multiple second multimodal features. Here, when each sample score corresponds to two or more reference audios, such as for a score of 80, there are multiple reference audios with different expressions, each sample score corresponds to multiple first similarity values. In this case, the average of the multiple first similarity values corresponding to each sample score is calculated to obtain the second similarity value sim(outer) between the first multimodal feature of the audio answer and the second multimodal feature of each sample score distribution. Since each sample score corresponds to multiple reference audio's second multimodal features, for each sample score, the multiple second multimodal features of the sample score are matched for feature similarity. For example, cosine distance is used to calculate the similarity value between pairs of features, resulting in multiple fourth similarity values corresponding to the sample score. The average of these multiple fourth similarity values is then calculated to obtain the aggregation degree measure sim(inner) corresponding to the sample score. Based on the aggregation degree measure sim(inner), the second similarity value sim(outer) of the sample score is normalized, for example, by dividing sim(outer) by sim(inner), to obtain the third similarity value corresponding to the sample score. This process is repeated to obtain the third similarity values corresponding to all other sample scores. From the third similarity values corresponding to all sample scores, the sample score corresponding to the largest third similarity value is selected as the reference score corresponding to the audio answer.
[0155] For example, assuming the reference audio corresponding to the sample score (80 points) is: Reference Audio 1, Reference Audio 2, and Reference Audio 3, then the first multimodal feature corresponding to the audio answer is matched with the second multimodal feature corresponding to the reference audio 1 to reference audio 3 corresponding to the sample score (80 points) to obtain three first similarity values sim(1), sim(2), and sim(3). At the same time, the second multimodal features of reference audio 1, reference audio 2, and reference audio 3 are matched with feature similarity to obtain three fourth similarity values: sim(12), sim(13), and sim(23). Then, for the sample score (80 points), the second similarity value sim(o) = 1 / 2. uter)=(sim(1), sim(2) and sim(3)) / 3, aggregation degree measure sim(inner)=(sim(12)+sim(13)+sim(23)) / 3, third similarity value=sim(outer) / sim(inner); and so on, to obtain the third similarity value of the corresponding other sample scores. For example, assuming that the sample scores include 5 scores: 45 points, 70 points, 80 points, 90 points, 100 points, etc., the final 5 third similarity values are 0.3, 0.5, 0.3, 0.6, 0.1 respectively. Select the sample score (90 points) corresponding to the largest third similarity value as the reference score corresponding to the audio answer.
[0156] Step 105: Obtain the original score of the audio answer, and perform anomaly detection on the original score based on the reference score to obtain the detection result used to characterize whether the original score is abnormal.
[0157] In some embodiments, the server may obtain the original score of the audio answer by: extracting features from the audio answer through the first feature extraction layer of the scoring model to obtain the audio features of the audio answer; extracting features from the text content through the second feature extraction layer of the scoring model to obtain the text features of the text content; and predicting the score of the audio answer based on the audio features and the text features of the audio answer through the score prediction layer of the scoring model to obtain the original score of the audio answer.
[0158] Here, after obtaining the audio answer and its corresponding text content, a scoring model is used to score the audio answer, resulting in its raw score. For example... Figure 8 As shown, Figure 8The diagram below illustrates the structure of the scoring model provided in this application embodiment. The scoring model includes a first feature extraction layer, a second feature extraction layer, and a score prediction layer. The first feature extraction layer is used to extract acoustic features from the audio answer to obtain the audio features of the audio answer. The second feature extraction layer is used to extract text features from the text content corresponding to the audio answer to obtain the text features of the text content. The score prediction layer is used to combine the audio features and text features of the audio answer to predict the score of the audio answer and obtain the original score of the audio answer.
[0159] In some embodiments, the scoring model can be trained as follows: A training sample set is constructed, wherein the training samples in the training sample set include native language audio samples and non-native language audio samples. Each training sample is labeled with an expert rating. The training samples are input into the scoring model. The first feature extraction layer of the scoring model extracts features from the training samples to obtain audio features. The second feature extraction layer of the scoring model extracts features from the text content of the training samples to obtain text features. The scoring prediction layer of the scoring model predicts the score of the training samples based on the audio and text features to obtain the predicted score. The difference between the predicted score and the labeled expert rating is obtained, and the value of the loss function is obtained based on the difference. When the value of the loss function reaches a preset threshold, the corresponding error signal is determined based on the value of the loss function of the scoring model. The error signal is backpropagated in the scoring model, and the model parameters of each layer of the scoring model are updated during the propagation process.
[0160] In some embodiments, see Figure 9 , Figure 9 The diagram below illustrates the structure of the scoring model provided in this application embodiment. The scoring prediction layer in the scoring model includes a first sub-prediction layer, a second sub-prediction layer, a third sub-prediction layer, and a scoring fusion layer. The server can predict the score of the audio answer based on the audio features and text features of the audio answer in the following manner to obtain the original score of the audio answer:
[0161] Based on the audio features of the audio answer, a first sub-prediction layer is used to predict the pronunciation score of the audio answer, resulting in a pronunciation score. Based on the text features of the audio answer, a second sub-prediction layer is used to predict the grammar score of the audio answer, resulting in a grammar score. A third sub-prediction layer is used to match the text features of the audio answer with the text features of the reference audio, and the correctness score of the audio answer is determined based on the matching results. Finally, a score fusion layer is used to fuse the pronunciation score, grammar score, and correctness score to obtain the original score of the audio answer.
[0162] Here, based on the audio features of the audio answer, the first sub-prediction layer predicts the pronunciation score of the audio answer, such as detecting the pronunciation quality (accuracy, completeness, fluency, rhythm, etc.) and predicting the pronunciation score based on the detection results. Based on the text features of the audio answer, the second sub-prediction layer predicts the grammar score of the audio answer, such as detecting the grammar quality (accuracy) and predicting the grammar score based on the detection results. The third sub-prediction layer matches the text features of the audio answer with the text features of the reference audio to determine whether the audio answer is relevant to the question, irrelevant, or complete, and determines the correctness score of the audio answer based on the matching results. Finally, the score fusion layer fuses the pronunciation score, grammar score, and correctness score to obtain the original score of the audio answer.
[0163] In some embodiments, when detecting the grammatical quality of audio answers, a target word feature corresponding to the word feature is predicted based on the word features in the text features. When the word feature does not match the target word feature, a grammatical error is detected. Based on the number of occurrences of the grammatical error, a grammatical score for the corresponding audio answer is determined. In some embodiments, for a certain type of error, the contextual features of each word feature in the text features can also be learned through a deep learning-based model, and then the word can be predicted using the contextual features. If the prediction result differs from the original word, the original word is marked as an error.
[0164] In some embodiments, the server may combine pronunciation scores, grammar scores, and accuracy scores to obtain the original score for the audio answer in the following manner:
[0165] The weights for pronunciation score, grammar score, and accuracy score are determined respectively. Based on the determined weights, the pronunciation score, grammar score, and accuracy score are weighted and summed to obtain the original score corresponding to the audio answer.
[0166] Here, different weights can be assigned to pronunciation scores, grammar scores, and accuracy scores according to the actual situation. Based on the different weights of each dimension, the original scores representing the overall quality of the image answer can be obtained.
[0167] In some embodiments, the server can perform anomaly detection on the original score based on the reference score in the following manner to obtain a detection result that characterizes whether the original score is abnormal: obtain the score difference between the reference score and the original score; when the score difference exceeds the difference threshold and the maximum first similarity value exceeds the similarity threshold, the original score is determined to be an abnormal score.
[0168] Here, when each sample score corresponds to one reference audio, if the score difference between the reference score and the original score exceeds the difference threshold and the maximum first similarity value exceeds the similarity threshold, the original score is determined to be an abnormal score; otherwise, the original score is determined to be a normal score. When each sample score corresponds to multiple reference audios, each sample score corresponds to multiple first similarity values. The reference score for the audio answer is further determined based on the third similarity value and the sample score. In this case, if the score difference between the reference score and the original score exceeds the difference threshold and the maximum third similarity value exceeds the similarity threshold, the original score is determined to be an abnormal score; otherwise, the original score is determined to be a normal score.
[0169] For example, based on the triplet data (P_raw, P_cluster, s), it can be determined whether the original score is abnormal, where P_raw is the original score, P_cluster is the reference score, and s is the largest first similarity value or the largest third similarity value. Assuming the triplet data is (60, 80, 0.9), the similarity threshold is 0.8, and the difference threshold is 10, then since 0.9 is greater than 0.8, and the difference between the reference score and the original score (20) exceeds the difference threshold (10), the original score can be determined to be an abnormal score.
[0170] In some embodiments, after the server performs anomaly detection on the original score based on the reference score and obtains a detection result that characterizes whether the original score is abnormal, it can also send correction prompt information in the following manner: when the detection result characterizes the original score as an abnormal score, correction prompt information for the corresponding audio answer is sent; wherein, the correction prompt information is used to prompt the audio answer to undergo score correction processing.
[0171] Here, when the detection result indicates that the original score is an abnormal score, the server can send a correction prompt to the terminal for unified processing, such as manual intervention or using other scoring models, to reduce the generation of abnormal scores and ultimately make the score as accurate as possible. The final accurate score is then sent to the terminal for display. When the detection result indicates that the original score is a normal score, the original score can be sent to the terminal for display.
[0172] The method for detecting anomaly scores provided in the embodiments of this application will be described next. See [link to relevant documentation]. Figure 10 , Figure 10 This is a flowchart illustrating an abnormal scoring detection method provided in an embodiment of this application. The method is applied to a scoring management terminal, such as a terminal on the scoring verification personnel's side, and includes:
[0173] Step 201: The terminal displays a scoring and detection interface, which presents at least one question and its corresponding scoring and detection function item.
[0174] Here, the terminal is located at the scoring correction personnel's location. The terminal is equipped with a detection client for detecting anomalies in the scoring of users' audio answers. When the scoring correction personnel need to detect anomalies in the scoring, they can open the correction client on the terminal. In response to this opening operation, the terminal presents a scoring detection interface, which displays one or more questions. When multiple questions are presented, each question can correspond to a scoring detection function item. In this case, the scoring detection function item is used to detect the scoring of the corresponding question. Alternatively, multiple questions can correspond to one scoring detection function item, in which case the scoring function item is used to perform batch detection of the scoring of multiple questions.
[0175] Step 202: In response to the triggering operation of the scoring detection function item corresponding to the target question, present the information input interface corresponding to the target question.
[0176] Here, when the scoring and correction personnel trigger the scoring detection function item corresponding to the target question, the terminal responds to the trigger operation and presents an information input interface for inputting the audio answer and the corresponding original score of the target question. The information input interface presents information input options, and the audio answer and the corresponding original score can be obtained based on the information input options.
[0177] Step 203: Based on the information input interface, receive the audio answer to the target question and the corresponding original score.
[0178] Here, when the scoring and correction personnel trigger the information input option, because in actual applications (such as examination scenarios), multiple candidates have answered a certain question, there are multiple corresponding audio answers for the target question. Here, the audio answers are associated with the corresponding original scores and the audio answers are associated with the corresponding candidates. Therefore, the candidate number to be detected can be selected based on the information input option, and the audio answer and the corresponding original score of the target question can be received.
[0179] Step 204: In response to a rating detection instruction triggered based on the audio answer and the corresponding original rating, output a detection result to characterize whether the original rating is abnormal.
[0180] Here, when a user triggers the start detection function for the received audio answer and its corresponding original score, the terminal responds to the trigger operation, receives the corresponding score detection instruction, and responds to the score detection instruction to perform anomaly detection on the original score of the audio answer, and obtains and presents the corresponding detection results.
[0181] See Figure 11 , Figure 11This is a schematic diagram of the scoring detection interface provided in this application embodiment. When the user clicks the scoring detection function item A1 in the scoring detection interface, the terminal responds to the click operation and presents the information input interface A2 corresponding to the target question. The information input interface presents information input options. When an information input option is clicked, the terminal presents multiple candidate options A3 that can be selected. When candidate 1 is selected, the terminal receives the audio answer of candidate 1 to the target question and the corresponding original score. In response to the trigger operation of the start detection function item A4, the terminal performs scoring anomaly detection on the original score of candidate 1's answer to the target question and presents the corresponding detection result A5.
[0182] In some embodiments, when the detection result indicates that the original score is an abnormal score, the original score can be uniformly processed, such as through manual intervention or by using other scoring models, to re-score the audio answer and obtain a normal score. This reduces the generation of abnormal scores, ultimately making the scoring as accurate as possible, and the final normal score is sent to the terminal for display on the terminal's interface. When the detection result indicates that the original score is a normal score, the original score can be sent to the terminal for display on the terminal's interface.
[0183] It should be noted that the above detection results are obtained by performing anomaly detection on the original score based on the reference score of the audio answer; the reference score is determined by the matching result between the first multimodal feature of the audio answer and the second multimodal feature of the reference audio of the target question. That is, the method of obtaining the detection results here is implemented through steps 101-105 in the above embodiment, which will not be repeated here.
[0184] The following describes an exemplary application of this application in a real-world scenario, taking oral exams as an example. More and more oral exams are adopting fully automated intelligent scoring by machines. For open-ended questions in oral exams, not only is the text content corresponding to the audio answers diverse, but the pronunciation quality of the audio answers also varies. Due to this diversity, scoring models may produce a small number of inaccurate predicted scores. Therefore, this application provides a method for detecting abnormal scores, enabling effective detection of abnormal scores.
[0185] This application provides a method for detecting abnormal scores, mainly involving two aspects: multimodal feature extraction and abnormal score detection. The process of acquiring multimodal features involves utilizing the complementarity between multimodal features and eliminating redundancy between modalities to learn better feature representations. The multimodal feature extraction involved in related technologies mainly includes two parts: joint representation and collaborative representation. Joint representation maps the information representations of multiple modalities to the same space, such as mapping text and images to the same space, or mapping speech and text to the same space, to analyze speech sentiment. Collaborative representation is responsible for mapping each modality in the multimodal dataset to its respective representation space, but the mapped vectors satisfy certain correlation constraints, obtaining the constraint relationships between vectors, such as addition, subtraction, multiplication, and division relationships.
[0186] For anomaly score detection, the prediction uncertainty analysis methods of related technologies mainly adopt two approaches: First, modeling the prediction uncertainty of the model. This modeling approach can be divided into two parts: 1. Direct modeling of uncertainty, typical methods include Gaussian process regression, Monte Carlo dropout, and deep mixture density networks. Gaussian process regression uses a Gaussian distribution to model the output, determining the mean and variance of each prediction result; Monte Carlo dropout uses multiple models to analyze the uncertainty of the model, assuming that for uncertain data, the output of each model has diversity; deep mixture density networks are similar to Gaussian process modeling, modeling the mean and variance of the results. 2. A more refined modeling of uncertainty by combining constructed data, such as generating constructed data (data far removed from the training data), and simultaneously modeling both the constructed data and the real data, enabling the model to explicitly learn that test data far removed from the training data often has a larger variance than general data. Second, defining uncertainty types (e.g., no reading at all, random reading, etc.), extracting some effective features, such as text features, classifying uncertainty types, filtering data with these types, and then inputting them into the scoring model.
[0187] However, the methods for obtaining multimodal representations of speech and text in the aforementioned related technologies do not model from the perspective of oral examination applications and are not suitable for oral examinations. Therefore, this application embodiment constructs a multimodal feature extraction method that can simultaneously combine pronunciation and text features, and can extract multimodal features from the audio answers of examinations and fuse audio and text features. Regarding the measurement of scoring uncertainty, the first approach to modeling uncertainty requires a basic examination evaluation model that can output the uncertainty of the prediction results; the second approach to classifying uncertainty types depends on how the uncertainty type is defined, limiting the possibilities of uncertainty types. Therefore, this application embodiment, based on a deep neural network, obtains multimodal features obtained by fusing audio and text features from the perspective of the matching degree of audio and text features. Based on the multimodal features of the audio answers, anomaly detection is performed on the original scores output by the scoring model of the examination scoring system to filter out anomaly score samples based on the detection results.
[0188] See Figure 12 , Figure 12 This is a schematic diagram of the scoring display interface provided in this application embodiment. The application scenario of the abnormal scoring detection method provided in this application embodiment is an oral examination scenario. The product is implemented in oral examinations and is mainly applied to open expression questions in oral examinations. The target question to be answered and the start recording button are presented on the examination interface of the terminal. When the candidate clicks the start recording button, he can start answering the question. When he clicks the end recording button, he ends the answering of the question. During this process, the terminal collects the audio answer of the candidate and sends the collected audio answer to the server. The server scores the audio answer and performs anomaly detection on the score, and then returns the normal score to the terminal for display on the terminal's display interface.
[0189] See Figure 13 , Figure 13This is a schematic diagram of the architecture of the abnormal scoring detection system provided in this application embodiment. The system includes a terminal and a server. The terminal displays the examination interface, and the user can start answering questions by clicking the "Start Recording" button on the examination interface and stop answering questions by clicking "Stop Recording." The terminal collects the audio answers of the examinee and sends the collected audio answers to the server. The server stores the audio answers in a database and reads the audio answers from the database, inputting them into a speech recognition module. The speech recognition module extracts acoustic features from the audio answers to obtain audio features, and performs text conversion on the audio answers to obtain the corresponding text content. Then, the audio features and text content of the audio answers are input into the scoring module, which scores the audio answers to obtain the original score. Simultaneously, the audio answers and their corresponding text content are input into a multimodal feature extraction module. Based on the text content of the audio answers, the multimodal feature extraction module extracts multimodal features from the audio answers to obtain the multimodal features of the audio answers. The original score is then... The multimodal features of the audio answers are input into the anomaly detection module. Based on these features, the module detects anomalies in the original scores, obtaining a detection result that indicates whether the original score is abnormal. This result is stored in the database. When the detection result indicates that the original score is abnormal, the server sends a correction prompt to the administrator to prompt further processing, such as manual intervention or using other scoring models, to reduce the generation of abnormal scores and ultimately make the scoring as accurate as possible. The final accurate score is then sent to the terminal for display. When the detection result indicates that the original score is normal, the original score is sent to the terminal for display.
[0190] Next, we will... Figure 13 The scoring module, multimodal feature extraction module, and anomaly scoring detection module involved are explained.
[0191] 1. Scoring Module
[0192] The oral exam scoring module mainly performs automatic evaluation of the user's audio answers. It generally includes two parts: 1. Extracting audio features and text content of the audio answer based on speech recognition technology. For example, based on the basic pronunciation features of speech recognition, the audio answer is transformed to obtain audio features of various pronunciations, or acoustic features are directly extracted based on speech to obtain the audio features of the audio answer; the audio answer is then converted into text to obtain the corresponding text content; 2. The audio features and text content of the audio answer are input into a trained scoring model to score the audio answer and obtain the raw score of the audio answer.
[0193] In practical implementation, the scoring model can be trained as follows: A training sample set is constructed, comprising native language audio samples and non-native language audio samples. Each training sample is labeled with an expert rating. The training samples are input into the scoring model. The first feature extraction layer of the scoring model extracts features from the training samples to obtain their audio features. The second feature extraction layer of the scoring model extracts features from the text content of the training samples to obtain their text features. The scoring prediction layer of the scoring model predicts the scores of the training samples based on their audio and text features to obtain their predicted scores. The difference between the predicted scores and the labeled expert ratings is obtained, and the value of the loss function is obtained based on this difference. When the value of the loss function reaches a preset threshold, the corresponding error signal is determined based on the value of the loss function of the scoring model. The error signal is backpropagated through the scoring model, and the model parameters of each layer of the scoring model are updated during the propagation process until convergence.
[0194] 2. Multimodal feature extraction module
[0195] In practical applications, servers can extract multimodal features of audio answers based on neural network classification models, such as... Figure 4 As shown, the classification model includes an encoding layer, an attention layer, and a classification prediction layer. The encoding layer includes a speech encoder and a text encoder. The speech encoder is used to extract speech features from the audio answer, and the text encoder is used to extract text features from the corresponding text content. Both the speech encoder and the text encoder are deep neural network structures, which can be composed of multiple modules, such as convolutional neural networks. The attention layer is used to fuse the audio features and text features obtained from the encoding layer to obtain the multimodal features of the audio answer. The classification prediction layer is used to predict the classification result based on the fused features. It can be seen that the multimodal feature extraction module is part of the classification model. The audio answer to be processed and the corresponding text content are input into the trained classification model, and the multimodal features of the audio answer can be obtained through the attention layer.
[0196] When training the classification model, training samples are first constructed. The training samples consist of audio answers and text content pairs, that is, the training sample is constructed as: training sample (audio answer, text content). The training samples include positive samples and negative samples. The text content in the positive samples is obtained by speech recognition of the audio answers in the training samples. The text content in the negative samples is randomly replaced with other words in the vocabulary according to a certain probability. That is, negative samples with mismatched audio answers and text content are constructed. Positive samples are labeled with matching word labels (such as 1), and negative samples are labeled with unmatched word labels (such as 0).
[0197] After the training samples are constructed, they are input into the classification model. The speech encoder in the coding layer extracts (encodes) the acoustic features of the audio answers in the training samples, thus obtaining the audio features h of the training samples. audio The text encoder in the encoding layer extracts (encodes) text features from the text content of the training samples, thus obtaining the text features h of the training samples. text Among them, in extracting text features h text First, the text content is segmented into words to obtain multiple words corresponding to the text content. Then, feature encoding is performed on each word to obtain the word feature h corresponding to each word. word(i) (Characterizing the word features of the i-th word), and finally concatenating the word features corresponding to each word to obtain the text features of the text content, h. text =h word(1) h word(2) , ..., h word(i) .
[0198] After obtaining the audio and text features of the training samples through the encoding layer, the attention layer performs attention processing on the audio and text features of the training samples to obtain the concatenated feature h of each word's fused audio features. word(i) =Attention(h word(i) h audio h audio )+h word(i) , where h word(i) h represents the word features of the i-th word. audio For audio features, Where Q is the query vector, K is the key vector, and V is the value vector, with the vector dimensions of K and Q being d. k In this application, Q is h word(i) K is h audio V is h audio Based on the attention mechanism, the degree of matching between each word feature and the corresponding audio feature can be obtained.
[0199] A classification prediction layer (fully connected layer) is used to classify and predict the concatenated features of the fused audio features for each word, resulting in a classification score that characterizes whether the corresponding word is correctly matched. word(i) =sigmoid(W word h word(i) +b word ), where sigmoid() is a nonlinear activation function, h word(i) The word features of the i-th word, W word b represents the weights of trainable word features. word These are the corresponding trainable bias parameters.
[0200] The optimization objective of the classification model is to minimize the cross-loss entropy H(t,p) between the classification result and the true label. Here, the difference between the classification result and the matching labels of the training samples is obtained, and the value of the loss function of the classification model is determined based on the difference. Where t(x) represents the classification result of whether the true predicted word is correctly matched, and p(x) represents the probability of the model correctly predicting the word, i.e., the matching label of the training sample. When the value of the loss function reaches a preset threshold, the corresponding error signal is determined based on the value of the loss function of the classification model; the error signal is backpropagated in the classification model, and the model parameters of each layer of the classification model are updated during the propagation process.
[0201] The classification model can be trained using the above method. After the classification model is trained, the audio answers from the evaluators can be input into the speech encoder, and the corresponding text content can be input into the text encoder. Finally, the word features h of each word can be extracted. word(i) , all word features h word(i) The average value is used to obtain the multimodal feature representation of the audio answer.
[0202] 3. Anomaly Scoring Detection Module
[0203] Here, for the same question, there are multiple sample scores (i.e., labels), such as 45, 70, 80, 90, 100, etc. Each sample score corresponds to multiple reference audios; for example, for a score of 80, there are reference audios with various expressions. Based on the original scores of the user's audio answers and the corresponding sample scores, the degree of data aggregation under each sample score is determined. The degree of aggregation is mainly determined by the distance under each sample score distribution of the training samples. First, based on the multimodal feature extraction method mentioned above, all features under a certain sample score in the training samples are obtained, that is, the multimodal features of multiple reference audios corresponding to each sample score are obtained. For each sample score, the multiple multimodal features of the sample score are matched for feature similarity. For example, cosine distance is used to calculate the similarity value between pairs of features, and multiple similarity values corresponding to the sample score are obtained. For example, assuming there are 50 reference audios under the sample score (80 points), 50*49 / 2 similarity values are obtained. The multiple similarity values corresponding to the sample score are averaged to obtain the aggregation degree measure sim(inner) corresponding to the sample score.
[0204] For the audio answers provided by the user (i.e., test data), based on the multimodal feature extraction method described above, the multimodal features of the audio answers are obtained. For each sample score, the following processing is performed: the multimodal features of the audio answers are compared with the multimodal features of multiple reference audios corresponding to the sample score to calculate the similarity, resulting in a set of similarity values corresponding to the sample score. The average of these similarities is then calculated to obtain the similarity sim(outer) between the audio answers and the distribution of each sample score.
[0205] After obtaining sim(inner) and sim(outer), the similarity sim(outer) is normalized based on sim(inner). sim(outer) is then divided by sim(inner) to obtain the similarity value between the audio answer and each sample score. For example, assuming the sample scores include 45, 70, 80, 90, and 100 points, the resulting 5 similarity values are 0.3, 0.5, 0.3, 0.6, and 0.1, respectively. The sample score corresponding to the highest similarity value (90 points) is selected as the reference score for the audio answer, ultimately yielding the feature pair (P_cluster, s) for the corresponding audio answer. Here, P_cluster is the reference score, and s is the highest similarity value. In the example above, P_cluster is 90 points, and s is 0.6.
[0206] Finally, the triplet data (P_raw, P_cluster, s) of the audio answers is obtained and input into the anomaly scoring detection module. The basic principle of the anomaly scoring detection module is that for audio answers with a large s, if P_raw and P_cluster differ significantly, it may be an anomaly sample. For example, assuming the triplet data is (60, 80, 0.9), the similarity threshold is 0.8, and the difference threshold is 10, then since 0.9 is greater than 0.8, and the difference between the reference score and the original score (20) exceeds the difference threshold (10), the original score can be determined to be an anomaly score.
[0207] The abnormal scoring detection method provided in this application embodiment was tested based on oral topic expression questions. A total of 1,000 test audio data (i.e., test audio answers) and corresponding expert-annotated scores were used. The audio test data were input into the abnormal scoring detection module, and abnormal samples were screened based on the detection results. The accuracy of the detection results was 74%, and the recall rate was 20%. Although the recall rate was low, the accuracy of the recalled samples was relatively high, which can effectively screen abnormal samples.
[0208] The following description continues to illustrate the exemplary structure of the anomaly scoring detection device 555 provided in the embodiments of this application as a software module. In some embodiments, such as Figure 14 As shown, Figure 14This is a schematic diagram of the structure of the anomaly scoring detection device provided in the embodiments of this application. The software modules stored in the anomaly scoring detection device 555 in the memory 550 include:
[0209] The first acquisition module 5551 is used to acquire the audio answer corresponding to the target question and the text content corresponding to the audio answer;
[0210] The feature extraction module 5552 is used to perform multimodal feature extraction on the audio answer based on the text content to obtain the first multimodal feature of the audio answer;
[0211] The second acquisition module 5553 is used to acquire the second multimodal features of the reference audio corresponding to the target problem;
[0212] The feature matching module 5554 is used to match the first multimodal features of the audio answer with the second multimodal features of the reference audio, and determine the reference score of the audio answer based on the matching result;
[0213] The scoring detection module 5555 is used to obtain the original score of the audio answer and perform anomaly detection on the original score based on the reference score to obtain a detection result that characterizes whether the original score is abnormal.
[0214] In some embodiments, the feature extraction module is further configured to extract features from the text content to obtain the text features of the text content;
[0215] Feature extraction is performed on the audio answer to obtain its audio features;
[0216] By fusing the text features and the audio features, the first multimodal feature of the audio answer is obtained.
[0217] In some embodiments, the feature extraction module is further configured to perform word segmentation on the text content to obtain multiple words corresponding to the text content;
[0218] Each of the aforementioned words is feature-encoded to obtain the word features corresponding to each of the aforementioned words;
[0219] The text features of the text content are obtained by concatenating the word features corresponding to each of the aforementioned words.
[0220] In some embodiments, the feature extraction module is further configured to perform bidirectional encoding processing on the word features of each word to obtain the context encoding features and context encoding features corresponding to each word;
[0221] The context encoding features and context encoding features of each word are concatenated to obtain the concatenated encoding features corresponding to each word.
[0222] The concatenated encoding features corresponding to each word are concatenated to obtain the text features of the text content.
[0223] In some embodiments, the feature extraction module is further configured to concatenate the text features and audio features of each word to obtain the concatenated features of the word.
[0224] Obtain the weight corresponding to each of the aforementioned words;
[0225] Based on the acquired weights, the concatenation features of each word are weighted and summed to obtain the first multimodal feature of the audio answer.
[0226] In some embodiments, the feature processing module is further configured to obtain multiple sample scores corresponding to the target question, each sample score corresponding to at least one reference audio.
[0227] The first multimodal feature of the audio answer is matched with the second multimodal feature of each of the reference audios to obtain a first similarity value between the first multimodal feature and each second multimodal feature;
[0228] Based on the obtained first similarity values and corresponding sample scores, a reference score for the audio answer is determined.
[0229] In some embodiments, when each sample score corresponds to multiple reference audios, each sample score corresponds to multiple first similarity values. The feature processing module is further configured to average the multiple first similarity values corresponding to each sample score to obtain a second similarity value corresponding to each sample score.
[0230] Obtain the aggregation degree measure corresponding to each sample score, and based on the aggregation degree measure corresponding to each sample score, normalize the second similarity value of the corresponding sample score to obtain the third similarity value of each sample score.
[0231] From the third similarity values of each sample score, the sample score corresponding to the largest third similarity value is selected as the reference score for the audio answer.
[0232] In some embodiments, the feature processing module is further configured to perform the following operations for each of the sample scores:
[0233] The similarity between features of the multiple second multimodal features corresponding to the sample score is matched to obtain multiple fourth similarity values corresponding to the sample score;
[0234] The aggregation degree measure corresponding to the sample score is obtained by averaging multiple fourth similarity values.
[0235] In some embodiments, the scoring detection module is further configured to obtain the score difference between the reference score and the original score;
[0236] When the score difference exceeds the difference threshold and the maximum similarity value exceeds the similarity threshold, the original score is determined to be an abnormal score.
[0237] In some embodiments, the feature processing module is further configured to extract features from the audio answer through the first feature extraction layer of the scoring model to obtain the audio features of the audio answer;
[0238] The text content is feature extracted through the second feature extraction layer of the scoring model to obtain the text features of the text content;
[0239] The scoring prediction layer of the scoring model predicts the score of the audio answer based on its audio features and text features, thus obtaining the original score of the audio answer.
[0240] In some embodiments, the rating prediction layer includes a first sub-prediction layer, a second sub-prediction layer, a third sub-prediction layer, and a rating fusion layer. The feature processing module is further used for...
[0241] Based on the audio features of the audio answer, the first sub-prediction layer is used to predict the pronunciation score of the audio answer to obtain the pronunciation score of the audio answer;
[0242] Based on the text features of the audio answer, the second sub-prediction layer is used to predict the grammar score of the audio answer to obtain the grammar score of the audio answer.
[0243] The third sub-prediction layer matches the text features of the audio answer with the text features of the reference audio, and determines the accuracy score of the audio answer based on the matching results.
[0244] The original score of the audio answer is obtained by fusing the pronunciation score, the grammar score, and the accuracy score through the scoring fusion layer.
[0245] In some embodiments, after performing anomaly detection on the original score based on the reference score to obtain a detection result characterizing whether the original score is abnormal, the apparatus further includes:
[0246] The information sending module is used to send correction prompt information corresponding to the audio answer when the detection result indicates that the original score is an abnormal score;
[0247] The correction prompt information is used to prompt the audio answer to undergo scoring and correction processing.
[0248] See Figure 15 , Figure 15 A schematic diagram of the structure of the detection device 150 for anomaly scoring in this application embodiment is shown, including:
[0249] The first presentation module 151 is used to present a scoring detection interface and to present at least one question and a corresponding scoring detection function item in the scoring detection interface.
[0250] The second presentation module 152 is used to present the information input interface corresponding to the target question in response to the trigger operation of the scoring detection function item corresponding to the target question;
[0251] Information receiving module 153 is used to receive the audio answer and corresponding original score of the target question based on the information input interface;
[0252] The result output module 154 is used to respond to a scoring detection instruction triggered based on the audio answer and the corresponding original score, and output a detection result to characterize whether the original score is abnormal.
[0253] The detection result is obtained by performing anomaly detection on the original score based on the reference score of the audio answer; the reference score is determined based on the matching result between the first multimodal feature of the audio answer and the second multimodal feature of the reference audio of the target question.
[0254] This application provides a computer program product or computer program that includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the anomaly scoring detection method described above in this application.
[0255] This application provides a computer-readable storage medium storing executable instructions, wherein the executable instructions, when executed by a processor, will cause the processor to execute the abnormal scoring detection method provided in this application.
[0256] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or it may be a variety of devices including one or any combination of the above-mentioned memories.
[0257] In some embodiments, executable instructions may take the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a standalone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0258] As an example, executable instructions may, but do not necessarily, correspond to files in a file system. They may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a Hyper Text Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple collaborating files (e.g., a file that stores one or more modules, subroutines, or code sections).
[0259] As an example, executable instructions can be deployed to execute on a single computing device, or on multiple computing devices located in one location, or on multiple computing devices distributed across multiple locations and interconnected via a communication network.
[0260] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, and improvements made within the spirit and scope of this application are included within the scope of protection of this application.
Claims
1. A method for detecting anomaly scores, characterized in that, The method includes: Obtain the audio answer to the target question, and the text content corresponding to the audio answer, wherein the text content includes multiple words; The text features of each word and the audio features of the audio answer are concatenated to obtain the concatenated features of the word. The weights corresponding to each word are obtained based on the attention mechanism, wherein the query vector of the attention mechanism is the word feature of the word, and the key vector and value vector are both the audio features; Based on the acquired weights, the concatenation features of each word are weighted and summed to obtain the first multimodal feature of the audio answer; Obtain the second multimodal features of the reference audio corresponding to the target problem; Obtain multiple sample scores corresponding to the target question, and each sample score corresponds to multiple reference audios; The first multimodal feature of the audio answer is matched with the second multimodal feature of each of the reference audios to obtain a first similarity value between the first multimodal feature and each second multimodal feature; The average of the multiple first similarity values corresponding to each sample score for the target question is calculated to obtain the second similarity value corresponding to each sample score. Obtain the aggregation degree measure corresponding to each sample score, and based on the aggregation degree measure corresponding to each sample score, normalize the second similarity value of the corresponding sample score to obtain the third similarity value of each sample score. From the third similarity values of each sample score, the sample score corresponding to the largest third similarity value is selected as the reference score for the audio answer; Obtain the score difference between the original score of the audio answer and the reference score; when the score difference exceeds the difference threshold and the maximum third similarity value exceeds the similarity threshold, determine that the original score is an abnormal score.
2. The method as described in claim 1, characterized in that, The method further includes: The text content is segmented to obtain multiple words corresponding to the text content; Each of the aforementioned words is feature-encoded to obtain the word features corresponding to each of the aforementioned words; The text features of the text content are obtained by concatenating the word features corresponding to each of the aforementioned words.
3. The method as described in claim 2, characterized in that, The step of concatenating the word features corresponding to each of the aforementioned words to obtain the text features of the text content includes: Each word's features are bidirectionally encoded to obtain the context encoding features and background encoding features corresponding to each word. The context encoding features and context encoding features of each word are concatenated to obtain the concatenated encoding features corresponding to each word. The concatenated encoding features corresponding to each word are concatenated to obtain the text features of the text content.
4. The method as described in claim 1, characterized in that, The step of obtaining the aggregation degree measure corresponding to each of the sample scores includes: Perform the following operations for each of the aforementioned sample scores: The multiple second multimodal features of the sample score are respectively matched for feature similarity to obtain multiple fourth similarity values corresponding to the sample score; The aggregation degree measure corresponding to the sample score is obtained by averaging multiple fourth similarity values.
5. The method as described in claim 1, characterized in that, The method further includes: The audio answer is extracted using the first feature extraction layer of the scoring model to obtain its audio features. The text content is feature extracted through the second feature extraction layer of the scoring model to obtain the text features of the text content; The scoring prediction layer of the scoring model predicts the score of the audio answer based on its audio features and text features, thus obtaining the original score of the audio answer.
6. The method as described in claim 5, characterized in that, The rating prediction layer includes a first sub-prediction layer, a second sub-prediction layer, a third sub-prediction layer, and a rating fusion layer. The step of predicting the rating of the audio answer based on its audio features and text features to obtain the original rating of the audio answer includes: Based on the audio features of the audio answer, the first sub-prediction layer is used to predict the pronunciation score of the audio answer to obtain the pronunciation score of the audio answer; Based on the text features of the audio answer, the second sub-prediction layer is used to predict the grammar score of the audio answer to obtain the grammar score of the audio answer. The third sub-prediction layer matches the text features of the audio answer with the text features of the reference audio, and determines the accuracy score of the audio answer based on the matching results. The original score of the audio answer is obtained by fusing the pronunciation score, the grammar score, and the accuracy score through the scoring fusion layer.
7. A method for detecting abnormal scores, characterized in that, The method includes: A scoring and detection interface is presented, and at least one question and a corresponding scoring and detection function item are displayed on the scoring and detection interface; In response to the triggering operation of the scoring detection function item corresponding to the target question, the information input interface corresponding to the target question is presented; Based on the information input interface, the audio answer to the target question and the corresponding original score are received; In response to a rating detection instruction triggered based on the audio answer and the corresponding original rating, a detection result is output to characterize whether the original rating is abnormal; The detection result is obtained by performing anomaly detection on the original score based on the reference score of the audio answer. When the score difference between the original score and the reference score exceeds a difference threshold, and the maximum third similarity value used to determine the reference score exceeds a similarity threshold, the original score is determined to be an anomalous score. The reference score is determined as follows: multiple sample scores corresponding to the target question are obtained, each sample score corresponding to multiple reference audios; the first multimodal feature of the audio answer is matched with the second multimodal feature of each reference audio to obtain a first similarity value between the first multimodal feature and each second multimodal feature; the multiple first similarity values corresponding to each sample score corresponding to the target question are averaged to obtain a second similarity value corresponding to each sample score; and the aggregation degree corresponding to each sample score is obtained. The second similarity value of the corresponding sample scores is normalized based on the aggregation degree metric corresponding to each sample score to obtain the third similarity value of each sample score; from the third similarity values of each sample score, the sample score corresponding to the largest third similarity value is selected as the reference score of the audio answer; the first multimodal feature is determined by the following method: obtaining the text content corresponding to the audio answer, the text content including multiple words; concatenating the text features of each word and the audio features of the audio answer to obtain the concatenated features of the words; obtaining the weights corresponding to each word based on the attention mechanism, the query vector of the attention mechanism being the word features of the words, and the key vector and value vector being the audio features; and weighting and summing the concatenated features of each word based on the obtained weights to obtain the first multimodal feature of the audio answer.
8. A detection device for anomaly scoring, characterized in that, The device includes: The first acquisition module is used to acquire the audio answer to the target question and the text content corresponding to the audio answer, wherein the text content includes multiple words; The feature extraction module is used to concatenate the text features of each word and the audio features of the audio answer to obtain the concatenated features of the word; obtain the weights corresponding to each word based on an attention mechanism, wherein the query vector of the attention mechanism is the word feature of the word, and the key vector and value vector are the audio features; and perform a weighted summation of the concatenated features of each word based on the obtained weights to obtain the first multimodal features of the audio answer. The second acquisition module is used to acquire the second multimodal features of the reference audio corresponding to the target problem; The feature processing module is used to obtain multiple sample scores corresponding to the target question, and each sample score corresponds to multiple reference audios; perform similarity matching between the first multimodal features of the audio answer and the second multimodal features of each reference audio to obtain a first similarity value between the first multimodal features and each second multimodal feature; average the multiple first similarity values corresponding to each sample score corresponding to the target question to obtain a second similarity value corresponding to each sample score; obtain the aggregation degree measure corresponding to each sample score, and normalize the second similarity value of the corresponding sample score based on the aggregation degree measure corresponding to each sample score to obtain a third similarity value corresponding to each sample score; select the sample score corresponding to the largest third similarity value from the third similarity values of the corresponding sample scores as the reference score of the audio answer; The scoring detection module is used to obtain the score difference between the original score of the audio answer and the reference score; when the score difference exceeds the difference threshold and the maximum third similarity value exceeds the similarity threshold, the original score is determined to be an abnormal score.
9. The apparatus as claimed in claim 8, characterized in that, The feature extraction module is further configured to perform word segmentation on the text content to obtain multiple words corresponding to the text content; perform feature encoding on each word to obtain word features corresponding to each word; and perform feature concatenation on the word features corresponding to each word to obtain the text features of the text content.
10. The apparatus as claimed in claim 9, characterized in that, The feature extraction module is also used to perform bidirectional encoding processing on the word features of each word to obtain the context encoding features and context encoding features corresponding to each word. The context encoding features and context encoding features of each word are concatenated to obtain the concatenated encoding features corresponding to each word; the concatenated encoding features corresponding to each word are then concatenated to obtain the text features of the text content.
11. An electronic device, characterized in that, include: Memory, used to store executable instructions; A processor, when executing executable instructions stored in the memory, implements the anomaly scoring detection method according to any one of claims 1 to 7.
12. A computer-readable storage medium, characterized in that, It stores executable instructions for use by a processor to implement the anomaly scoring detection method according to any one of claims 1 to 7.
13. A computer program product comprising computer instructions, characterized in that, When the computer instructions are executed by the processor, they implement the abnormal scoring detection method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Voice scoring method and system
CN103928023A