Evaluation Method, Device, Electronic Device and Storage Medium for Audio Data
By conducting uncertainty analysis on audio data and text data and selecting appropriate evaluation methods for final evaluation, the problem of inaccurate evaluation results in the prior art is solved and the accuracy of oral evaluation is improved.
Patent Information
- Application Number
- CN202110204456.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-02-23
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2041-02-23
AI Technical Summary
The existing oral automatic evaluation technology has problems of uncertainty and low accuracy in the evaluation results due to incomplete coverage of training data.
By obtaining audio data and text data for uncertainty analysis, the uncertainty analysis results of the evaluation model are determined, and based on this result, selecting the final evaluation model or other evaluation methods for the final evaluation, such as manual evaluation, to improve the accuracy of the evaluation results.
It effectively reduces the inaccuracy of evaluation results and improves the accuracy of audio data evaluation.
Smart Images

Figure CN113590741B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology. Specifically, this application relates to a method, device, electronic device, and storage medium for evaluating audio data. Background Art
[0002] With the development of artificial intelligence technology, artificial intelligence technology plays an important role in various fields. In the field of computer-aided teaching, the automatic oral evaluation technology plays an important role. The implementation of the automatic oral evaluation technology can effectively improve the efficiency of oral evaluation.
[0003] However, since the automatic oral evaluation technology targets a large number of people, including people of different ages and different oral levels, and at the same time, since the training scoring data for oral evaluation often needs to be manually annotated, which is not only time-consuming but also requires a high degree of professionalism from the operators for annotation. The above problems make the training data of the oral evaluation model often unable to fully cover all the characteristics of the evaluated person, resulting in uncertainty or errors in the scores output by the final oral evaluation model, that is, the accuracy is relatively low. Summary of the Invention
[0004] The technical solution provided by this application aims to solve at least one of the above technical defects, especially the technical defect of relatively low accuracy of audio data evaluation results. Among them, the technical solution is as follows:
[0005] In the first aspect of this application, a method for evaluating audio data is provided, including:
[0006] Obtain audio data and text data corresponding to the audio data;
[0007] Based on the audio data and text data, perform uncertainty analysis to determine the uncertainty analysis result of the result obtained by using an evaluation model to evaluate the audio data;
[0008] Based on the uncertainty analysis result, determine the evaluation result of using the evaluation model or other evaluation methods to evaluate the audio data as the final evaluation result.
[0009] In an embodiment, based on the audio data and text data, perform uncertainty analysis to determine the uncertainty analysis result of the result obtained by using an evaluation model to evaluate the audio data, including:
[0010] Based on the audio data and text data, perform speech recognition to determine the time information of the alignment between the speech and the text;
[0011] Based on the audio data and the time information, perform uncertainty analysis to determine the uncertainty analysis result of the result obtained by using an evaluation model to evaluate the audio data.
[0012] In another embodiment, uncertainty analysis is performed based on audio data and time information to determine the uncertainty analysis result of the result obtained by using an evaluation model to evaluate the audio data, including:
[0013] Extract the acoustic feature information in the audio data;
[0014] Based on the acoustic feature information and time information, determine the feature representation of the audio data;
[0015] Based on the feature representation of the audio data and the training data for training the evaluation model, determine the uncertainty parameter of the audio data;
[0016] Based on the uncertainty parameter, determine the uncertainty analysis result of the result obtained by using the evaluation model to evaluate the audio data.
[0017] In yet another embodiment, based on the acoustic feature information and time information, determine the feature representation of the audio data, including:
[0018] Use a pre-constructed acoustic feature extractor to determine the label information of the audio data based on the acoustic feature information;
[0019] Based on the label information and time information, determine the duration corresponding to each word, and average the features of the corresponding number of frames based on the duration to obtain the feature representation of each word;
[0020] Average the feature representations of all words to obtain the feature representation of the corresponding audio data.
[0021] In one embodiment, the steps of training the acoustic feature extractor include:
[0022] Obtain training data, where the training data includes frame-level acoustic feature information and corresponding true label information;
[0023] Use the training data to train the acoustic feature extractor so that the network parameters of the acoustic feature extractor are adjusted based on the cross-loss function; the cross-loss function is determined based on the probability of predicting the label information corresponding to each frame of acoustic feature information during training and the true label information.
[0024] In one embodiment, based on the feature representation of the audio data and the training data for training the evaluation model, determine the uncertainty parameter of the audio data, including:
[0025] Determine the training feature representations included under each training label in the training data for training the evaluation model;
[0026] Calculate the similarity between the training feature representations included under each training label, and determine the aggregation degree measure of each training label;
[0027] Calculate the similarity between the feature representation of the audio data and the training feature representation of the training data, and determine the similarity value between the audio data and the training data under each training label;
[0028] Normalize the similarity value based on the aggregation degree metric, and determine the result of the normalization process as the uncertainty parameter of the audio data.
[0029] In one embodiment, determine the uncertainty analysis result of the result obtained by evaluating the audio data using the evaluation model based on the uncertainty parameter, including any one of the following:
[0030] Sort the uncertainty parameters of all audio data in descending order, determine the uncertainty analysis result of the audio data corresponding to the lowest preset percentage after sorting as uncertain, and determine the uncertainty analysis results of other audio data as certain;
[0031] Calculate the mean and standard deviation of the uncertainty parameters of all audio data, determine a threshold based on the mean and standard deviation, determine the uncertainty analysis result of the audio data corresponding to the uncertainty parameter being lower than or equal to the threshold as uncertain, and determine the uncertainty analysis results of other audio data as certain.
[0032] In one embodiment, determine the evaluation result of evaluating the audio data using the evaluation model or other evaluation methods based on the uncertainty analysis result as the final evaluation result, including:
[0033] When the uncertainty analysis result is certain, determine the evaluation result of evaluating the audio data using the evaluation model as the final evaluation result;
[0034] When the uncertainty analysis result is uncertain, determine the evaluation result of evaluating the audio data using other evaluation methods as the final evaluation result.
[0035] In one embodiment, evaluate the audio data using the evaluation model, including:
[0036] Perform speech recognition based on the audio data and the text data to determine the speech feature information;
[0037] Use the evaluation model to determine the evaluation result of the audio data based on the speech feature information.
[0038] In one embodiment, it further includes:
[0039] Feedback the final evaluation result to the corresponding client to display the final evaluation result on the client.
[0040] In the second aspect of the present application, there is provided an evaluation device for audio data, including:
[0041] An acquisition module, configured to acquire audio data and text data corresponding to the audio data;
[0042] An analysis module, configured to perform uncertainty analysis based on the audio data and the text data, and determine an uncertainty analysis result of the result obtained by evaluating the audio data using an evaluation model;
[0043] A determination module, configured to determine, based on the uncertainty analysis result, an evaluation result obtained by evaluating the audio data using the evaluation model or other evaluation methods as the final evaluation result.
[0044] In a third aspect of the present application, an electronic device is provided. The electronic device includes:
[0045] One or more processors;
[0046] A memory;
[0047] One or more computer programs, where one or more computer programs are stored in the memory and configured to be executed by one or more processors. The one or more programs are configured to: execute the method provided in the first aspect.
[0048] In a fourth aspect of the present application, a computer-readable storage medium is provided. The computer storage medium is used to store computer instructions. When the computer instructions run on a computer, the computer can execute the method provided in the first aspect.
[0049] The beneficial effects brought by the technical solution provided in the present application are:
[0050] In the present application, uncertainty analysis is performed based on the acquired audio data and the text data corresponding to the audio data, and an uncertainty analysis result of the result obtained by evaluating the audio data using an evaluation model is determined. Furthermore, based on the uncertainty analysis result, it can be determined whether to use the evaluation model or other evaluation methods to evaluate the audio data as the final evaluation result. By performing uncertainty analysis on the acquired audio data in the present application, the uncertainty of the result obtained by evaluating the audio data using the evaluation model is determined, that is, the audio data for which the evaluation result may be inaccurate can be screened out; furthermore, based on the uncertainty analysis result, it can be determined whether to use the evaluation model or other evaluation methods to evaluate the audio data as the final evaluation result, so as to effectively reduce the situation where the evaluation score is inaccurate due to using the evaluation model to evaluate the audio data and improve the accuracy of audio data evaluation. Description of the Drawings
[0051] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments of the present application.
[0052] Figure 1 Flow chart of an audio data evaluation method provided by an embodiment of the present application;
[0053] Figure 2 Schematic diagram of the operation process of an acoustic feature extractor in an audio data evaluation method provided by an embodiment of the present application;
[0054] Figure 3 Schematic diagram of the process of calculating uncertainty parameters in an audio data evaluation method provided by an embodiment of the present application;
[0055] Figure 4 Implementation flowchart of an audio data evaluation method provided by an embodiment of the present application;
[0056] Figure 5 Schematic diagram of an interaction environment of an audio data evaluation method provided by an embodiment of the present application;
[0057] Figure 6 Schematic diagram of the framework of an evaluation system of an audio data evaluation method provided by an embodiment of the present application;
[0058] Figure 7a Schematic diagram of a corresponding display interface when applying the audio data evaluation method provided by an embodiment of the present application;
[0059] Figure 7b Schematic diagram of a corresponding display interface when applying the audio data evaluation method provided by an embodiment of the present application;
[0060] Figure 8 Schematic diagram of a corresponding display interface when applying the audio data evaluation method provided by an embodiment of the present application;
[0061] Figure 9 Schematic diagram of the structure of an audio data evaluation device provided by an embodiment of the present application;
[0062] Figure 10 Schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0063] The embodiments of the present application are described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described by referring to the accompanying drawings are exemplary and are only used to explain the present application, and should not be construed as a limitation to the present application.
[0064] Those skilled in the art can understand that, unless specifically stated otherwise, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the description of the present application means the presence of the stated features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. It should be understood that when we say an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any unit and all combinations of one or more of the associated listed items.
[0065] To make the objectives, technical solutions and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.
[0066] The following explains the technologies and terms related to the present application:
[0067] AI (Artificial Intelligence) is a theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning and decision-making. Artificial intelligence technology is an interdisciplinary subject, involving a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0068] In the present application, it may involve directions such as speech technology, machine learning / deep learning, etc.
[0069] Among them, the key technologies of speech technology include ASR (Automatic Speech Recognition), text-to-speech technology (TTS), and voiceprint recognition technology. Enabling computers to listen, see, speak, and feel is the future development direction of human-computer interaction, and among them, voice has become one of the most promising human-computer interaction methods. ASR technology is a technology that converts speech into text. In the embodiments of the present application, ASR technology can be used to construct a speech recognition model to process the acquired audio data.
[0070] ML (Machine Learning) is an interdisciplinary field that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning. In the embodiments of the present application, technologies related to machine learning can be used to construct an evaluation model, an uncertainty analysis module, etc.
[0071] With the development of artificial intelligence technology, artificial intelligence technology plays an important role in various fields. In the field of computer-assisted instruction, the automatic oral assessment technology plays an important role. The implementation of the automatic oral assessment technology can effectively improve the efficiency of oral evaluation. However, since the oral evaluation model constructed using machine learning technology has uncertainties, such as the uncertainty of accidental events (due to random noise in the data) and cognitive uncertainties, the evaluation results output by the final oral evaluation model are uncertain or incorrect, that is, the accuracy is relatively low.
[0072] In the related art, to solve the above-mentioned uncertainty problem, a solution for modeling uncertainty is provided. However, this solution has requirements for the basic oral evaluation model and requires the model itself to be able to output the uncertainty of the evaluation results, which increases the complexity of the model to a certain extent and correspondingly reduces the efficiency of model processing.
[0073] To solve at least one of the above problems, the present application provides a method, an apparatus, an electronic device, and a computer-readable storage medium for evaluating audio data; specifically, uncertainty analysis is performed on the evaluation result of an evaluation model for audio data, and then it is possible to determine whether to use the evaluation result of the evaluation model for evaluating the audio data as the final evaluation result based on the uncertainty analysis result, so as to effectively reduce the situation where the evaluation result obtained by the model for evaluating the audio data is incorrect and improve the accuracy of audio data evaluation.
[0074] The following uses specific embodiments to elaborate in detail on the technical solutions of the present application and how the technical solutions of the present application solve the above technical problems. These several specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below with reference to the accompanying drawings.
[0075] In an embodiment of the present application, a method for evaluating audio data is provided, as Figure 1 shown Figure 1 FIG. shows a flowchart of a method for evaluating audio data provided by an embodiment of the present application. Among them, this method can be executed by any electronic device, such as a user terminal or a server. The user terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc. The server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. However, the present application is not limited thereto. Specifically, this method includes the following steps S101-S103:
[0076] Step S101: Obtain audio data and text data corresponding to the audio data.
[0077] Specifically, the audio data can be voice data emitted by a person or data corresponding to the audio track in recorded multimedia data (such as video).
[0078] Optionally, in an embodiment of the present application, a batch (multiple) of audio data can be obtained for processing at the same time, or a single audio data can be processed; for example, when applied to the scenario of oral evaluation, multiple voices input by a batch of users can be obtained for processing at the same time, or a certain voice input by a certain user can be obtained for processing. The present application does not make any limitations in this regard. In the following embodiments, to better illustrate the embodiments of the present application, an example of obtaining a batch of audio data for processing at the same time is used for description.
[0079] In one embodiment, the text data is the basis for the user to input audio data. That is, the user can input audio data through the microphone of the terminal based on the text data. Therefore, the content represented by the audio data corresponds to the text data. When multiple audio data are obtained and processed at one time, the multiple audio data correspond to the same text data. For example, when performing oral evaluation using the audio data evaluation method provided by the embodiments of the present application at time 1, 500 audio data and 1 text data A corresponding to the audio data are obtained. For example, when performing oral evaluation using the audio data evaluation method provided by the embodiments of the present application at time 2, 1 audio data and 1 text data B corresponding to the audio data are obtained.
[0080] Optionally, when multiple audio data are obtained and processed simultaneously, although the multiple audio data all correspond to the same text data, since the pronunciation of each user is different (such as different speaking speeds), the duration of each audio data may be different.
[0081] Step S102: Perform uncertainty analysis based on the audio data and the text data to determine the uncertainty analysis result of the result obtained by evaluating the audio data using the evaluation model.
[0082] Specifically, in the embodiments of the present application, uncertainty analysis may refer to the estimation and research on various external factor changes and influences that cannot be controlled in advance during the process of evaluating audio data using the evaluation model. That is, analyze the learning situation of the evaluation model during the training process. When the model evaluates data that has not been learned, the confidence level of the obtained evaluation result is low, and the obtained evaluation result belongs to the category of uncertainty.
[0083] Among them, as Figure 4 shown, the embodiments of the present application can use a neural network to construct an uncertainty analysis module to perform uncertainty analysis based on the currently obtained audio data and text data. The specific process of uncertainty analysis will be described in subsequent embodiments.
[0084] Step S103: Determine the evaluation result of evaluating the audio data using the evaluation model or other evaluation methods as the final evaluation result based on the uncertainty analysis result.
[0085] Specifically, the uncertainty analysis result can include one of certainty and uncertainty. When the uncertainty analysis result corresponding to a certain audio data is uncertainty, that is, when the uncertainty of the evaluation result obtained by using the evaluation model to evaluate the audio data is relatively high, other evaluation methods are used to evaluate the audio data, and the evaluation result is output as the final evaluation result; when the uncertainty analysis result corresponding to a certain audio data is certainty, that is, when the certainty of the evaluation result obtained by using the evaluation model to evaluate the audio data is relatively high, the evaluation result is output as the final evaluation result.
[0086] Among them, other evaluation methods can include methods such as manual evaluation (manual review), etc. The audio data is transmitted to the user terminal for manual evaluation for display, and is evaluated by users with professional evaluation capabilities.
[0087] In the embodiment of the present application, based on the obtained audio data and the text data corresponding to the audio data, uncertainty analysis is performed to determine the uncertainty analysis result of the result obtained by using the evaluation model to evaluate the audio data. Furthermore, based on the uncertainty analysis result, it can be determined whether to use the evaluation model or other evaluation methods to evaluate the audio data, and the evaluation result is used as the final evaluation result. Through the implementation of the present application, by performing uncertainty analysis on the obtained audio data, the uncertainty of the result obtained by using the evaluation model to evaluate the audio data can be determined, that is, the audio data whose evaluation result may be inaccurate can be screened out; furthermore, based on the uncertainty analysis result, it can be determined to use the evaluation model or other evaluation methods to evaluate the audio data, and the evaluation result is used as the final evaluation result, so as to effectively reduce the situation where the evaluation score is inaccurate due to using the evaluation model to evaluate the audio data and improve the accuracy of audio data evaluation.
[0088] The specific process of uncertainty analysis will be described below. In uncertainty analysis, feature analysis is performed based on the audio data (data to be predicted), the extracted features are compared with the features of the training data used to train the evaluation model, and the similarity between the features is calculated; among them, the audio data corresponding to a lower similarity may be the data not covered by the evaluation model during training, so the corresponding uncertainty is relatively high. Based on the uncertainty analysis, an uncertainty parameter corresponding to each audio data can be obtained, and then it can be determined whether the uncertain analysis result corresponding to a certain audio data is uncertainty based on the uncertainty parameter.
[0089] In one embodiment, the uncertainty analysis based on the audio data and the text data in step S102 to determine the uncertainty analysis result of the result obtained by using the evaluation model to evaluate the audio data includes the following steps A1 - A2:
[0090] Step A1: Perform speech recognition based on audio data and text data to determine the time information of the alignment between speech and text.
[0091] Specifically, in the embodiments of the present application, speech recognition takes audio data as the research object, and through speech signal processing and pattern recognition processing, enables the machine to automatically recognize and understand the content dictated and input by the user; that is, through speech recognition, the machine transforms the speech signal in the audio data into corresponding text or commands through the recognition and understanding process.
[0092] Among them, since users of different ages and different professional levels have different speech rates, pitches, etc. when reading the same text, it is necessary to determine the time information of the alignment between speech and text in the audio data corresponding to each user. For example, for the text "I like apple", in the audio data input by user A, the pronunciation time corresponding to "I" is 1s - 1.5s, the pronunciation time corresponding to "like" is 1.6s - 2s, and the pronunciation time corresponding to "apple" is 2s - 3s.
[0093] Step A2: Perform uncertainty analysis based on audio data and time information to determine the uncertainty analysis result of the result obtained by using the evaluation model to evaluate the audio data.
[0094] Specifically, as Figure 4 shown, the obtained audio data and the processed time information of the alignment between speech and text can be used as the input data of the uncertainty analysis module, and the uncertainty analysis module performs uncertainty analysis operations based on the input data and outputs the uncertainty analysis result. The specific operations of the uncertainty analysis module will be described in subsequent embodiments.
[0095] In a feasible embodiment, the uncertainty analysis in step A2 based on audio data and time information to determine the uncertainty analysis result of the result obtained by using the evaluation model to evaluate the audio data includes the following steps B1 - B4:
[0096] Step B1: Extract the acoustic feature information in the audio data.
[0097] Optionally, the acoustic feature information in the audio data can be extracted through the acoustic model of DNN (Deep Neural Networks), or the ASR technology can be used to perform speech recognition processing on the audio data to obtain the acoustic feature information, such as MFCC (Mel Frequency Cepstrum Coefficient). The Mel frequency is proposed based on the auditory characteristics of the human ear, and it has a non-linear correspondence with the Hz frequency; MFCC is the Hz spectrum feature calculated using this correspondence between them. Among them, the step of extracting the acoustic feature information can be implemented in the uncertainty analysis module, or can be implemented in other modules before inputting the data into the uncertainty analysis module (for example, an independent speech recognition model can be used to extract the acoustic features, or it can be implemented using the speech recognition model as shown in Figure 4 shown).
[0098] Among them, when analyzing the audio data, the audio can be framed, that is, the audio data is cut into multiple small segments, and each small segment is called a frame. Correspondingly, the extracted acoustic feature information can be frame-level feature information.
[0099] Step B2: Based on the acoustic feature information and the time information, determine the feature representation of the audio data.
[0100] Specifically, as shown in Figure 4 and Figure 6 shown, step B2 can be operated using a trained acoustic feature extractor (which can also be called an acoustic feature extraction model and is part of the uncertainty analysis module).
[0101] In one embodiment, step B2 of determining the feature representation of the audio data based on the acoustic feature information and the time information includes the following steps C1 - C3 (the implementation of steps C1 - C3 can be understood as an operation for each audio data):
[0102] Step C1: Use a pre-constructed acoustic feature extractor to determine the label information of the audio data based on the acoustic feature information.
[0103] Specifically, as shown in Figure 2As shown, the acoustic feature extractor includes a feature extraction model based on a deep neural network, which can be various model structures, such as CNN (Convolutional Neural Networks), LSTM (Long Short-Term Memory), the stacking of multiple layers of networks, or the stacking of multiple layers of the same or different model structures, such as the stacking of 3-layer convolutional neural networks; and then, through the fully connected layer included in the acoustic feature extractor, the extracted deep features are non-linearly transformed to obtain the label information corresponding to the audio data.
[0104] Optionally, the acoustic feature information input to the acoustic feature extractor can be at the frame level. Correspondingly, for each frame of acoustic feature output by the acoustic feature extractor, the prediction result can be the probability distribution of frame-level senones (multi-phoneme composition units considering phoneme context, which can be triphones or multi-phonemes), or can also be understood as the probability of a certain senone label (label information).
[0105] Step C2: Based on the label information and time information, determine the duration corresponding to each word, and average the features of the corresponding number of frames based on this duration to obtain the feature representation of each word.
[0106] Specifically, based on the label information, it can be known which words are included in the current audio data. For example, when three labels are included, it can be correspondingly known that three words are included in the current audio data; accordingly, combined with the time information of speech and text alignment, the duration corresponding to each word can be determined. Then, through frame averaging processing based on the duration corresponding to each word, the feature representation of each word can be obtained. Among them, frame averaging processing can determine the corresponding number of frames based on the duration corresponding to each word (each frame is generally between 10ms - 25ms, and each word may correspond to multiple frames). Since each frame corresponds to a feature, therefore, the features corresponding to the duration corresponding to each word can be further determined, and on this basis, the mean value of the features is taken to obtain the feature representation of each word.
[0107] Step C3: Average the feature representations of all words to obtain the feature representation of the corresponding audio data.
[0108] Specifically, based on Step C2, the feature representation of each word in the current audio data can be determined. On this basis, feature averaging processing is performed, that is, the average value of the feature representations of all words is taken to obtain the feature representation of this audio data.
[0109] The following combines Figure 3 A specific application example is used to illustrate the above Steps C1 - C3.
[0110] Suppose the currently obtained audio data (a certain piece of speech) is "I like apple". After inputting this audio data into the uncertainty analysis module, the acoustic features at the frame level of the audio data are extracted. The process of extracting acoustic features may include: performing frame segmentation on the audio data. Generally, one frame is about 10ms - 25ms. Correspondingly, the audio data will be sliced into multiple speech frames (such as M speech frames {f1, f2,......, fm}). When calculating MFCC, each frame corresponds to an output column of DCT coefficients (assuming the dimension is N), then the corresponding MFCC can be represented as an N ×M matrix; where, Figure 3 The acoustic feature information of the input acoustic feature extractor shown as "I [[1,0.2,3],[1.2,3,0.5]]like [[0.3,2,3],[0.2,2,3]]apple [[1,1.2,3],[2,3.5,4]]" can be the representation form after taking the variance or standard deviation for all frames based on the matrix.
[0111] After the acoustic feature extractor extracts features based on the acoustic feature information, the label information of the audio data is output. Since the acoustic feature extractor processes the acoustic features at the frame level, the output label information can also be called the frame feature representation, which is shown as Figure 3 "I [[1,2,3],[2,3,4]]like [[1,1,3],[1,3,4]]apple [[4,2,3],[2,3,4]]".
[0112] After obtaining the frame feature representation, frame averaging processing can be performed in combination with the time information of speech - text alignment, which is the operation content corresponding to step C2. Among them, the time information is assumed to be represented as follows: the pronunciation time corresponding to "I" is 1s - 1.5s, the pronunciation time corresponding to "like" is 1.6s - 2s, and the pronunciation time corresponding to "apple" is 2s - 3s; on this basis, after performing frame averaging processing, the content shown as Figure 3 can be obtained: the feature representation of each word "I [1.5,2.5,3.5] like [1,2,3.5]apple [3,2.5,3.5]".
[0113] After obtaining the feature representation of each word, word averaging processing is performed, that is, taking the mean value of the feature representations of all words, then the content shown as Figure 3 can be obtained: the feature representation of the audio data (the feature representation of this piece of speech) "[1.83,2.33,3.5]".
[0114] In the above-mentioned embodiment, corresponding to the application scenario of English oral evaluation, the vocabulary is described in units of words. For example, the feature representation of each vocabulary corresponds to the feature representation of each word. Optionally, it can also be applied to the oral evaluation scenarios of various languages, such as Japanese, Korean, etc., and the embodiments of the present application do not limit this.
[0115] The following specifically describes the process of constructing an acoustic feature extractor.
[0116] In one embodiment, training the acoustic feature extractor includes the following steps D1 - D2:
[0117] Step D1: Obtain training data, which includes frame-level acoustic feature information and corresponding true label information.
[0118] Specifically, each training sample data in the training data respectively corresponds to frame-level deep valley feature information and the true label information corresponding to each feature; among them, the true label information includes a specific label, and this label can be a senone label.
[0119] Step D2: Use the training data to train the acoustic feature extractor, so as to adjust the network parameters of the acoustic feature extractor based on the cross-loss function; the cross-loss function is determined based on the probability of predicting the label information corresponding to each frame of acoustic feature information during training and the true label information.
[0120] Specifically, the cross-loss function can be expressed as shown in the following formula (1):
[0121]
[0122] ......Formula (1)
[0123] In formula (1), y is the true label of the senone corresponding to a certain frame of acoustic feature information, and p is the probability that the acoustic feature extractor predicts as the corresponding senone label.
[0124] Step B3: Determine the uncertainty parameter of the audio data based on the feature representation of the audio data and the training data of the training evaluation model.
[0125] Specifically, the training data used to train the evaluation model can characterize the performance of the evaluation model. By processing the feature representation of the audio data and the training data, the uncertainty situation when the evaluation model processes the audio data can be determined, that is, the confidence of the prediction result.
[0126] The following specifically describes the process of determining the uncertainty parameter.
[0127] In one embodiment, in step B3, determining the uncertainty parameter of the audio data based on the feature representation of the audio data and the training data of the training evaluation model includes the following steps E1 - E4:
[0128] Step E1: Determine the training feature representations included under each training label in the training data for training the evaluation model.
[0129] Specifically, the operation method for determining the feature representation of the audio data shown in steps B1 - B2 can be referred to to execute part of the content in step E1. For each training data, first extract the acoustic feature information in the training data, and then, based on the acoustic feature information and the time information of the alignment of speech and text, determine the feature representation of the training data. Thus, the training feature representations of each training data can be obtained. Furthermore, based on the training feature representations corresponding to the training labels, the training data distributed under each training label can be determined. For example: Among the training data, there are 4 speech data A, B, C, and D. The training feature representation corresponding to training label 1 can correspond to speech data A, C, and D. Then, there are 3 speech data distributed under training label 1.
[0130] Step E2: Calculate the similarity between the training feature representations included under each training label, and determine the aggregation degree measure of each training label.
[0131] Specifically, the similarity between every two training feature representations in the training data can be calculated through various distance functions (such as the cosine distance function). Based on the similarity between the training feature representations, a similarity feature set corresponding to each training label can be obtained. For example: If there are 100 training data corresponding to training label 1, a similarity feature set including 100 * 99 similarity features can be obtained; or, if there are speech data A, C, and D corresponding to training label 1, the similarity between the speech data can be calculated respectively, that is, the similarity between A and C, A and D, C and A, C and D, D and A, and D and C. A similarity feature set including 6 similarity features can be obtained (wherein, since AC and CA, AD and DA, CD and DC belong to the same similarity feature, in order to reduce the computational complexity of subsequent steps during processing, the same similarity features in the set can be deleted). Furthermore, for the similarity feature set of each training label, the result of obtaining the average or mode of the set is used as the aggregation degree measure sim(inner) of the training label.
[0132] Step E3: Calculate the similarity between the feature representation of the audio data and the training feature representations of the training data, and determine the similarity value between the audio data and the training data under each training label.
[0133] Specifically, assuming that there are 10 training data under training label 1, the similarity between the feature representation of the audio data and the training feature representation of each training data under this training label 1 can be calculated to obtain a similarity feature set including 10 similarity features. The result obtained by calculating the average or mode of this set can be used as the similarity value sim(outer) between the audio data and the training data under training label 1.
[0134] Step E4: Normalize the similarity value based on the aggregation degree metric, and determine the result of the normalization process as the uncertainty parameter of the audio data.
[0135] Specifically, the normalization process can be expressed as shown in the following formula (2):
[0136]
[0137] ......Formula (2)
[0138] As can be seen from formula (2), the normalization process is to calculate the proportion of the similarity value in the aggregation degree metric. Based on the normalization process, the similarity value between the audio data and each training label can be obtained. , the result obtained by the normalization process can be determined as the uncertainty parameter of the audio data. If there are 5 training labels, finally 5 similarities corresponding to the audio data can be obtained, that is, the corresponding uncertainty parameters: [0.3, 0.5, 0.3, 0.6, 0.1].
[0139] Optionally, since the evaluation model is a trained model in actual application, steps E1 and E2 above can also be processed offline. The processing of steps E1 - E2 can be understood as determining the data aggregation degree under each training label based on the training data corpus and labels. Correspondingly, step B3 can only include steps E3 - E4. When step E3 is implemented, the training feature representation can be directly obtained. When step E4 is implemented, the aggregation degree metric corresponding to each training label of the trained evaluation model can be directly obtained for operation.
[0140] Step B4: Determine the uncertainty analysis result of the result obtained by evaluating the audio data using the evaluation model based on the uncertainty parameter.
[0141] Specifically, if multiple audio data are currently obtained, the uncertainty parameters corresponding to each audio data obtained based on step B3 can be processed to obtain the uncertainty results corresponding to each audio data. If one audio data is currently obtained, the audio data can be compared with a preset threshold. If it is lower than or equal to the preset threshold, the uncertainty analysis result of the audio data is determined as uncertain; if it is higher than the preset threshold, the uncertainty analysis result of the audio data is determined as certain.
[0142] The following describes the specific process of processing multiple uncertainty parameters in obtaining the uncertainty analysis results.
[0143] In one embodiment, the uncertainty analysis result of the result obtained by using the evaluation model to evaluate the audio data based on the uncertainty parameter in step B4 includes any one of the following steps F1 - F2:
[0144] Step F1: Sort the uncertainty parameters of all audio data in descending order, determine the uncertainty analysis result of the audio data corresponding to the lowest preset percentage after sorting as uncertain, and determine the uncertainty analysis results of other audio data as certain.
[0145] Specifically, for example: If there are 10 audio data, there are correspondingly 10 uncertainty parameters. After sorting the uncertainty parameters in descending order, the uncertainty analysis result of the audio data corresponding to the lowest 10% range can be determined as uncertain, and the uncertainty analysis results of the remaining 9 audio data are determined as certain.
[0146] Step F2: Calculate the mean and standard deviation of the uncertainty parameters of all audio data, determine the threshold based on the mean and standard deviation, determine the uncertainty analysis result of the audio data corresponding to the uncertainty parameter lower than or equal to the threshold as uncertain, and determine the uncertainty analysis results of other audio data as certain.
[0147] Specifically, in step F2, the threshold is dynamically adjusted according to the situation of the uncertainty parameters corresponding to different audio data to improve the adaptability and accuracy of determining the uncertainty analysis result of the audio data based on the uncertainty parameter.
[0148] The following describes the specific process of determining the final evaluation result based on the uncertainty analysis result.
[0149] In one embodiment, as Figure 4 shown, the evaluation result obtained by using the evaluation model or other evaluation methods to evaluate the audio data based on the uncertainty analysis result in step S103 is used as the final evaluation result, including the following steps G1 - G2:
[0150] Step G1: When the uncertainty analysis result is certain, determine the evaluation result obtained by using the evaluation model to evaluate the audio data as the final evaluation result.
[0151] Specifically, the operation of the evaluation model for evaluating audio data can be performed after determining to adopt the corresponding evaluation result of the evaluation model, or can be performed synchronously during the uncertainty analysis; when performed synchronously, after determining to adopt the corresponding evaluation result of the evaluation model, the final evaluation result can be directly output, which can effectively improve the evaluation efficiency of audio data.
[0152] Step G2: When the uncertainty analysis result is uncertain, determine to adopt the evaluation result of other evaluation methods for evaluating the audio data as the final evaluation result.
[0153] Specifically, considering the problem of reducing the resources involved in additionally adopting other evaluation methods, the audio data can be evaluated by other evaluation methods after determining the corresponding evaluation result of other evaluation methods, so as to reduce the waste of resources and lower the evaluation cost of audio data.
[0154] The following specifically describes the processing process of the evaluation model for evaluating audio data.
[0155] In one embodiment, evaluating the audio data using the evaluation model includes the following steps H1-H2:
[0156] Step H1: Perform speech recognition based on the audio data and text data to determine the speech feature information.
[0157] Step H2: Use the evaluation model to determine the evaluation result of the audio data based on the speech feature information.
[0158] In the embodiment of the present application, the evaluation module can automatically evaluate the user's pronunciation. Generally, it includes two parts: 1. The operation corresponding to step H1: Extract the pronunciation confidence feature (i.e., speech feature information) based on speech recognition; 2. The operation corresponding to step H2: Construct an evaluation model based on the pronunciation confidence feature, so that the evaluation result obtained by the evaluation model for evaluating the audio data fits the scoring result of professional evaluators. Based on the trained evaluation model, input the speech and the corresponding reading text into the oral evaluation module, and output the evaluation score of the corresponding pronunciation.
[0159] Optionally, as Figure 4 shown, the operation of speech recognition can be processed by separately constructing a model, and then the speech feature information extracted by the speech recognition model can be input into the evaluation model for processing to obtain the evaluation result of the audio data.
[0160] The following specifically describes the visualization process of the final evaluation result.
[0161] In a feasible embodiment, the evaluation method for audio data provided in the embodiment of the present application further includes the following step I1:
[0162] Step I1: Feed back the final evaluation result to the corresponding user terminal to display the final evaluation result on the display interface of the user terminal.
[0163] Specifically, in the embodiment of the present application, the user can read aloud the text data (i.e., the text for shadowing) on the user terminal (as Figure 7a and Figure 7b shown). The user terminal can upload the collected audio data and text data to the server, and the server forwards them to the evaluation system. For the evaluation method of the audio data provided in the embodiment of the present application, after the evaluation system determines the final evaluation result, the server can feedback the final evaluation result to the user terminal to display the final evaluation result on the display interface of the user terminal (as Figure 8 shown).
[0164] Next, in combination with Figures 4 - 8 , an application example of the evaluation method of the audio data provided in the embodiment of the present application will be described.
[0165] In a possible embodiment, N users (students) use the terminal 400 to read aloud according to the given text for shadowing. The reading content is as Figure 7a shown "I know the fact, do you know?", and the user can click or long-press the detection control of "Start Reading Aloud" to enable the terminal 400 to turn on the microphone to collect the audio data (at this time, it is voice data) of the user's pronunciation; when the user finishes reading, the pronunciation can be ended by clicking or releasing the detection control of "End Reading Aloud" to make the terminal 400 stop collecting audio data.
[0166] After the reading ends, the terminal 400 uploads the collected audio data and text data to the server 200 through the network 300, and the server 200 calls the evaluation system 500 to evaluate the audio data.
[0167] Specifically, the server 200 can send the audio data to the uncertainty analysis module, and at the same time send the audio data and text data to the speech recognition model.
[0168] Furthermore, in the evaluation system 500, the speech text alignment result (the time information of the alignment between speech and text) output by the speech recognition model is sent to the uncertainty analysis module, and the speech feature information output by the speech recognition model is sent to the evaluation model. The speech recognition model that outputs the speech text alignment result and the speech feature information can be the same speech recognition model or different speech recognition models.
[0169] Among them, after obtaining the audio data, the uncertainty analysis module extracts frame-level acoustic features based on the audio data, and then inputs the acoustic features into the acoustic feature extractor. The uncertainty analysis result of the audio data is determined through the acoustic feature extractor and other network architectures in the uncertainty analysis module.
[0170] When the uncertainty result is certain, the evaluation result output by the evaluation model is called and returned to the user; when the uncertainty result is uncertain, the audio data is transmitted to the manual evaluation module, and the teacher conducts a score evaluation. The score finally returned to the user is the score evaluated by the teacher.
[0171] The evaluation system 500 returns the final evaluation result to the server 200, and the server 200 feeds back the final evaluation result to the terminals 400-1 to 400-N used by each user through the network 300 respectively.
[0172] After the terminal 400 obtains the final evaluation result, it will be displayed on the display interface (such as being displayed on the display interfaces 400-11 to 400-N1 respectively). The display effect is as Figure 8 shown. The read-along text is displayed on the display interface, and the quality of the spoken reading is expressed by 5 stars. Figure 8 In the evaluation result shown, a certain user gets 4 stars. If the full score is 100 points, this user can correspondingly get 80 points. Further, the final evaluation result not only includes the evaluation score, but can also correspondingly indicate the defects in the user's spoken reading, such as Figure 8 the word "know" pointed by the gesture in is the word with relatively poor spoken reading of this user.
[0173] In the embodiment of the present application, there is a situation where the evaluation system 500 can be a part of the server 200. On this basis, the execution subject of the evaluation method for audio data provided in the above embodiment can be the server 200; there is another situation where the evaluation system 500 can be borne by other independent computer devices (terminals or servers). On this basis, the execution subject of the evaluation method for audio data provided in the above embodiment is correspondingly the terminal or the server.
[0174] The following further illustrates the technical effects that the embodiment of the present application can achieve, and gives the corresponding experimental data situation.
[0175] The experiment of this application is based on a trained evaluation model. A total of 1500 pieces of test data (the audio data in the above-mentioned embodiments) and the corresponding expert annotation scores are used. The test data is input into the uncertainty analysis module and the evaluation model, and the uncertainty result and the evaluation result of the evaluation model are output. The samples with a large difference between the evaluation result and the actual expert score are used as all uncertain sample labels. The result output by the uncertainty analysis module is the uncertainty prediction value, and the accuracy of the uncertainty analysis module can be calculated. The experimental results show that the accuracy is 80% and the recall rate is 30%. Although the recall rate is low, the accuracy of the recalled test data is high, and the recalled test data can be returned to professionals for further correction of the evaluation result.
[0176] An embodiment of this application provides an evaluation device for audio data, as Figure 9 shown. The evaluation device 900 for audio data may include: an acquisition module 901, an analysis module 902, and a determination module 903. Among them, the acquisition module 901 is used to acquire audio data and text data corresponding to the audio data; the analysis module 902 is used to perform uncertainty analysis based on the audio data and the text data to determine the uncertainty analysis result of the result obtained by using the evaluation model to evaluate the audio data; the determination module 903 is used to determine the evaluation result of using the evaluation model or other evaluation methods to evaluate the audio data based on the uncertainty analysis result as the final evaluation result.
[0177] In an embodiment, when the analysis module 902 is used to perform the step of performing uncertainty analysis based on the audio data and the text data to determine the uncertainty analysis result of the result obtained by using the evaluation model to evaluate the audio data, it is further used to perform the following steps:
[0178] Perform speech recognition based on the audio data and the text data to determine the time information of the alignment between the speech and the text;
[0179] Perform uncertainty analysis based on the audio data and the time information to determine the uncertainty analysis result of the result obtained by using the evaluation model to evaluate the audio data.
[0180] In another embodiment, when the analysis module 902 is used to perform the step of performing uncertainty analysis based on the audio data and the time information to determine the uncertainty analysis result of the result obtained by using the evaluation model to evaluate the audio data, it is further used to perform the following steps:
[0181] Extract the acoustic feature information in the audio data;
[0182] Determine the feature representation of the audio data based on the acoustic feature information and the time information;
[0183] Determine the uncertainty parameter of the audio data based on the feature representation of the audio data and the training data of the training and evaluation model;
[0184] Determine the uncertainty analysis result of the result obtained by evaluating the audio data using the evaluation model based on the uncertainty parameter.
[0185] In another embodiment, when the analysis module 902 is used to execute the step of determining the feature representation of the audio data based on the acoustic feature information and the time information, it is further used to execute the following steps:
[0186] Use a pre-built acoustic feature extractor to determine the label information of the audio data based on the acoustic feature information;
[0187] Determine the duration corresponding to each word based on the label information and the time information, and average the features of the corresponding number of frames based on the duration to obtain the feature representation of each word;
[0188] Average the feature representations of all words to obtain the feature representation of the corresponding audio data.
[0189] In one embodiment, the device 900 further includes a training module for training the acoustic feature extractor. Specifically, the training module is further used to execute the following steps:
[0190] Obtain training data, where the training data includes frame-level acoustic feature information and corresponding true label information;
[0191] Use the training data to train the acoustic feature extractor so as to adjust the network parameters of the acoustic feature extractor based on the cross-loss function; the cross-loss function is determined based on the probability of predicting the label information corresponding to each frame of acoustic feature information during training and the true label information.
[0192] In one embodiment, when the analysis module 902 is used to execute the step of determining the uncertainty parameter of the audio data based on the feature representation of the audio data and the training data of the training and evaluation model, it is used to execute the following steps:
[0193] Determine the training feature representations included under each training label in the training data for training the evaluation model;
[0194] Calculate the similarity between the training feature representations included under each training label to determine the aggregation degree measure of each training label;
[0195] Calculate the similarity between the feature representation of the audio data and the training feature representations of the training data to determine the similarity value between the audio data and the training data under each training label;
[0196] Normalize the similarity value based on the aggregation degree metric, and determine the result of the normalization process as the uncertainty parameter of the audio data.
[0197] In one embodiment, when the analysis module 902 is used to perform the step of determining the uncertainty analysis result of the result obtained by evaluating the audio data using the evaluation model based on the uncertainty parameter, it is further used to perform any one of the following:
[0198] Sort the uncertainty parameters of all audio data in descending order, determine the uncertainty analysis result of the audio data corresponding to the lowest preset percentage after sorting as uncertain, and determine the uncertainty analysis results of other audio data as certain;
[0199] Calculate the mean and standard deviation of the uncertainty parameters of all audio data, determine a threshold based on the mean and standard deviation, determine the uncertainty analysis result of the audio data corresponding to the uncertainty parameter being lower than or equal to the threshold as uncertain, and determine the uncertainty analysis results of other audio data as certain.
[0200] In one embodiment, when the determination module 903 is used to perform the step of determining the evaluation result of evaluating the audio data using the evaluation model or other evaluation methods as the final evaluation result based on the uncertainty analysis result, it is further used to perform the following steps:
[0201] When the uncertainty analysis result is certain, determine the evaluation result of evaluating the audio data using the evaluation model as the final evaluation result;
[0202] When the uncertainty analysis result is uncertain, determine the evaluation result of evaluating the audio data using other evaluation methods as the final evaluation result.
[0203] In one embodiment, when the determination module 903 is used to perform the step of evaluating the audio data using the evaluation model, it is further used to perform the following steps:
[0204] Perform speech recognition based on the audio data and text data to determine the speech feature information;
[0205] Use the evaluation model to determine the evaluation result of the audio data based on the speech feature information.
[0206] In one embodiment, the device 900 further includes a feedback module, which is used to feedback the final evaluation result to the corresponding client to display the final evaluation result at the client.
[0207] The device according to the embodiment of the present application can execute the method provided by the embodiment of the present application, and the implementation principle is similar. The actions performed by each module in the device in each embodiment of the present application correspond to the steps in the method in each embodiment of the present application. For the detailed function description of each module of the device, reference can be specifically made to the description in the corresponding method shown above, and details are not described herein again.
[0208] An electronic device is provided in an embodiment of the present application. The electronic device includes: a memory and a processor; at least one program stored in the memory, which when executed by the processor, can achieve, compared with the prior art: performing uncertainty analysis based on the acquired audio data and the text data corresponding to the audio data in the present application, determining the uncertainty analysis result of the result obtained by evaluating the audio data using an evaluation model, and further being able to determine, based on the uncertainty analysis result, whether to use the evaluation model or other evaluation methods to evaluate the audio data, and using the evaluation result as the final evaluation result. Through the implementation of the present application, by performing uncertainty analysis on the acquired audio data, the uncertainty of the result obtained by evaluating the audio data using the evaluation model is determined, that is, the audio data whose evaluation result may be inaccurate can be screened out; and further, based on the uncertainty analysis result, the evaluation result of using the evaluation model or other evaluation methods to evaluate the audio data can be determined as the final evaluation result, so as to effectively reduce the situation where the evaluation score is inaccurate due to using the evaluation model to evaluate the audio data and improve the accuracy of audio data evaluation.
[0209] In an alternative embodiment, an electronic device is provided, as Figure 10 shown, Figure 10 The electronic device 1000 shown includes: a processor 1001 and a memory 1003. Among them, the processor 1001 and the memory 1003 are connected, such as through a bus 1002. Optionally, the electronic device 1000 may further include a transceiver 1004, and the transceiver 1004 can be used for data interaction between the electronic device and other electronic devices, such as data sending and / or data receiving, etc. It should be noted that in practical applications, the transceiver 1004 is not limited to one, and the structure of the electronic device 1000 does not constitute a limitation to the embodiment of the present application.
[0210] The processor 1001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logical blocks, modules, and circuits described in connection with the disclosure of this application. The processor 1001 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0211] The bus 1002 may include a path for transmitting information between the above components. The bus 1002 may be a PCI (Peripheral Component Interconnect) bus, an EISA (Extended Industry Standard Architecture) bus, or the like. The bus 1002 may be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 10 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.
[0212] The memory 1003 may be a ROM (Read Only Memory) or other type of static storage device that can store static information and instructions, a RAM (Random Access Memory), or other type of dynamic storage device that can store information and instructions. It may also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.
[0213] The memory 1003 is used to store the application program code (computer program) for executing the solution of this application, and is controlled by the processor 1001 for execution. The processor 1001 is used to execute the application program code stored in the memory 1003 to implement the content shown in the foregoing method embodiments.
[0214] Among them, the electronic device includes but is not limited to: smart phones, tablet computers, laptop computers, smart speakers, smart watches, vehicle-mounted devices, etc.
[0215] According to one aspect of the present application, there is provided a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the evaluation method of audio data provided in the above various optional implementation manners.
[0216] The embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When it runs on a computer, the computer can execute the corresponding content in the foregoing method embodiments.
[0217] It should be understood that although the steps in the flowchart of the accompanying drawings are shown in sequence according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or sub-steps or stages of other steps.
[0218] The above are only some implementation manners of the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. An evaluation method for audio data, characterized in that, Including: Obtaining audio data and text data corresponding to the audio data; Performing uncertainty analysis based on the audio data and the text data to determine an uncertainty analysis result of a result obtained by evaluating the audio data using an evaluation model, including: performing speech recognition based on the audio data and the text data to determine time information of speech-text alignment; extracting acoustic feature information from the audio data; determining a feature representation of the audio data based on the acoustic feature information and the time information; determining an uncertainty parameter of the audio data based on the feature representation of the audio data and training data for training the evaluation model; determining an uncertainty analysis result of a result obtained by evaluating the audio data using the evaluation model based on the uncertainty parameter; Based on the uncertainty analysis result, determining an evaluation result of evaluating the audio data using the evaluation model or other evaluation methods as the final evaluation result.
2. The method according to claim 1, characterized in that, The determining the feature representation of the audio data based on the acoustic feature information and the time information includes: Using a pre-constructed acoustic feature extractor to determine label information of the audio data based on the acoustic feature information; Based on the label information and the time information, determining a duration corresponding to each vocabulary, and averaging features of corresponding frames based on the duration to obtain a feature representation of each vocabulary; Averaging the feature representations of all vocabularies to obtain a feature representation of the corresponding audio data.
3. The method according to claim 2, wherein The steps of training the acoustic feature extractor include: Obtaining training data, where the training data includes frame-level acoustic feature information and corresponding true label information; Training the acoustic feature extractor using the training data to adjust network parameters of the acoustic feature extractor based on a cross-entropy loss function; the cross-entropy loss function is determined based on a probability of predicting label information corresponding to each frame of acoustic feature information during training and the true label information.
4. The method according to claim 1, wherein The determining the uncertainty parameter of the audio data based on the feature representation of the audio data and training data for training the evaluation model includes: Determining training feature representations included under each training label in the training data for training the evaluation model; Calculating a similarity between training feature representations included under each training label to determine a measure of aggregation degree of each training label; Calculating a similarity between the feature representation of the audio data and the training feature representations of the training data to determine a similarity value between the audio data and the training data under each training label; Normalizing the similarity value based on the measure of aggregation degree, and determining the result of the normalization process as the uncertainty parameter of the audio data.
5. The method according to claim 1, wherein The determining the uncertainty analysis result of a result obtained by evaluating the audio data using the evaluation model based on the uncertainty parameter includes any one of the following: Sorting the uncertainty parameters of all audio data in descending order, determining the uncertainty analysis result of the audio data corresponding to a preset percentage with the lowest sorting as uncertain, and determining the uncertainty analysis results of other audio data as certain; Calculate the mean and standard deviation of the uncertainty parameters of all audio data, determine a threshold based on the mean and standard deviation, determine that the uncertainty analysis result of the audio data corresponding to the uncertainty parameter being lower than or equal to the threshold is uncertain, and determine that the uncertainty analysis result of other audio data is certain.
6. The method according to claim 1, characterized in that Based on the uncertainty analysis result, determining the evaluation result of evaluating the audio data using an evaluation model or other evaluation methods as the final evaluation result, including: When the uncertainty analysis result is certain, determining the evaluation result of evaluating the audio data using an evaluation model as the final evaluation result; When the uncertainty analysis result is uncertain, determining the evaluation result of evaluating the audio data using other evaluation methods as the final evaluation result.
7. The method according to claim 1, wherein Evaluating the audio data using an evaluation model includes: Performing speech recognition based on the audio data and text data to determine speech feature information; Using the evaluation model to determine the evaluation result of the audio data based on the speech feature information.
8. The method according to claim 1, wherein It also includes: Feeding back the final evaluation result to the corresponding client to display the final evaluation result on the client.
9. An evaluation device for audio data, characterized in that Including: An acquisition module for acquiring audio data and text data corresponding to the audio data; An analysis module for performing uncertainty analysis based on the audio data and text data to determine the uncertainty analysis result of the result obtained by evaluating the audio data using an evaluation model, including: performing speech recognition based on the audio data and text data to determine the time information of speech and text alignment; extracting acoustic feature information from the audio data; determining the feature representation of the audio data based on the acoustic feature information and time information; determining the uncertainty parameter of the audio data based on the feature representation of the audio data and the training data for training the evaluation model; determining the uncertainty analysis result of the result obtained by evaluating the audio data using an evaluation model based on the uncertainty parameter; A determination module for determining the evaluation result of evaluating the audio data using an evaluation model or other evaluation methods as the final evaluation result based on the uncertainty analysis result.
10. The device according to claim 9, characterized in that, When the analysis module is used to determine the feature representation of audio data based on acoustic feature information and time information, it is also used for: Using a pre-constructed acoustic feature extractor to determine the label information of the audio data based on the acoustic feature information; Based on the label information and time information, determining the duration corresponding to each vocabulary, and averaging the features of the corresponding number of frames based on the duration to obtain the feature representation of each vocabulary; Averaging the feature representations of all vocabularies to obtain the feature representation of the corresponding audio data.
11. The device according to claim 10, characterized in that, The device also includes a training module for training the acoustic feature extractor; The training module is also used for: Acquiring training data, where the training data includes frame-level acoustic feature information and corresponding true label information; Train an acoustic feature extractor using training data, such that the network parameters of the acoustic feature extractor are adjusted based on a cross entropy loss function; the cross entropy loss function is determined based on the probability of predicting the label information corresponding to the acoustic feature information of each frame during training and the true label information.
12. The device according to claim 9, characterized in that, When the analysis module is used to execute the determination of the uncertainty parameter of the audio data based on the feature representation of the audio data and the training data for training the evaluation model, it is further used for: Determine the training feature representations included under each training label in the training data for training the evaluation model; Calculate the similarity between the training feature representations included under each training label, and determine the aggregation degree measure for each training label; Calculate the similarity between the feature representation of the audio data and the training feature representations of the training data, and determine the similarity value between the audio data and the training data under each training label; Perform a normalization process on the similarity value based on the aggregation degree measure, and determine the result of the normalization process as the uncertainty parameter of the audio data.
13. The device according to claim 9, characterized in that, When the analysis module is used to execute the uncertainty analysis result of the result obtained by evaluating the audio data using the evaluation model based on the uncertainty parameter, it is further used for any one of the following: Sort the uncertainty parameters of all audio data in descending order, determine the uncertainty analysis result of the audio data corresponding to the lowest preset percentage after sorting as uncertain, and determine the uncertainty analysis results of other audio data as certain; Calculate the mean and standard deviation of the uncertainty parameters of all audio data, determine a threshold based on the mean and standard deviation, determine the uncertainty analysis result of the audio data corresponding to the uncertainty parameter being lower than or equal to the threshold as uncertain, and determine the uncertainty analysis results of other audio data as certain.
14. The device according to claim 9, characterized in that, When the determination module is used to execute the determination of the evaluation result obtained by evaluating the audio data using the evaluation model or other evaluation methods as the final evaluation result based on the uncertainty analysis result, it is further used for: When the uncertainty analysis result is certain, determine the evaluation result obtained by evaluating the audio data using the evaluation model as the final evaluation result; When the uncertainty analysis result is uncertain, determine the evaluation result obtained by evaluating the audio data using other evaluation methods as the final evaluation result.
15. The device according to claim 9, characterized in that When the determination module is used to execute the evaluation of the audio data using the evaluation model, it is further used for: Perform speech recognition based on the audio data and text data to determine speech feature information; Use the evaluation model to determine the evaluation result of the audio data based on the speech feature information.
16. The device according to claim 9, characterized in that, The device further includes a feedback module for feeding back the final evaluation result to the corresponding client to display the final evaluation result at the client.
17. An electronic device, characterized in that, The electronic device includes: One or more processors; A memory; One or more computer programs, wherein the one or more computer programs are stored in the memory and configured to be executed by the one or more processors, and the one or more computer programs are configured to: execute the method according to any one of claims 1 to 8.
18. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store computer instructions, which, when run on a computer, enable the computer to execute the method described in any one of claims 1 to 8 above.
19. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method described in any one of claims 1 to 8.
Citation Information
Patent Citations
Testing method and system of semi-opened spoken language examination questions
CN102354495A
Language capability evaluation method, device and system, computer equipment and storage medium
CN110503941A