A speech evaluation method, device, electronic device and storage medium
By combining multimodal information evaluation methods with text and acoustic modalities, acoustic and semantic features are extracted, and the problem of insufficient accuracy caused by oral evaluation reliance on automatic speech recognition in the prior art is solved, achieving more efficient evaluation results and resource savings.
Patent Information
- Application Number
- CN202110579754.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-05-26
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2041-05-26
AI Technical Summary
The existing oral evaluation methods rely on the recognition accuracy of automatic speech recognition technology, resulting in insufficient accuracy of evaluation results and require a large amount of human resources.
By combining text and acoustic modalities, acoustic and semantic features are extracted, multimodal information is fused for evaluation, reducing dependence on automatic speech recognition technology, and using speech feature extraction and semantic feature extraction models to calculate the feature correlation degree to generate evaluation results.
It improves the accuracy of oral evaluation results, reduces dependence on automatic speech recognition technology, and saves human resources.
Smart Images

Figure CN113763929B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of speech processing, and particularly to a speech evaluation method, device, electronic device and storage medium. Background Art
[0002] With the continuous development of language learning, language learners of various languages can determine their individual oral language learning situations with the help of the evaluation results of oral language evaluation.
[0003] Currently, in oral language evaluation, in addition to manually scoring the oral language of language learners, it is also possible to evaluate the oral language of language learners based on intermediate features generated by automatic speech recognition technology, including text features, acoustic features, and so on. However, this method depends on the recognition accuracy of automatic speech recognition technology. Summary of the Invention
[0004] Embodiments of the present invention provide a speech evaluation method, device, electronic device and storage medium, which can extract acoustic and semantic features by simultaneously combining text and acoustic modalities, fuse multi-modal information for oral language evaluation, reduce the dependence on automatic speech recognition technology, improve the accuracy of oral language evaluation results, and save human resources.
[0005] Embodiments of the present invention provide a speech evaluation method, including:
[0006] Obtain the speech to be evaluated and the reference text corresponding to the speech to be evaluated;
[0007] Extract speech features from the speech to be evaluated to obtain the target speech features corresponding to the speech to be evaluated;
[0008] Extract semantic features from the reference text to obtain the target semantic features corresponding to the reference text;
[0009] Calculate the feature correlation degree according to the target speech features and the target semantic features to obtain the feature correlation degree between the target speech features and the target semantic features;
[0010] Based on the feature correlation degree, perform classification processing on the evaluation results of the speech to be evaluated to obtain the evaluation results corresponding to the speech to be evaluated.
[0011] Correspondingly, embodiments of the present invention further provide a speech evaluation device, including:
[0012] A data acquisition unit, configured to obtain the speech to be evaluated and the reference text corresponding to the speech to be evaluated;
[0013] A speech feature extraction unit, configured to extract speech features from the speech to be evaluated to obtain the target speech features corresponding to the speech to be evaluated;
[0014] A semantic feature extraction unit, configured to extract semantic features from the reference text to obtain target semantic features corresponding to the reference text;
[0015] A correlation degree calculation unit, configured to calculate a feature correlation degree between the target speech feature and the target semantic feature according to the target speech feature and the target semantic feature;
[0016] An evaluation result generation unit, configured to perform evaluation result classification processing on the speech to be evaluated based on the feature correlation degree to obtain an evaluation result corresponding to the speech to be evaluated.
[0017] Optionally, the speech feature extraction unit is configured to map the speech to be evaluated into a speech feature vector space according to the speech feature mapping parameters of a speech feature extraction model, obtain a target speech feature vector based on the mapping result, and use the target speech feature vector as the target speech feature corresponding to the speech to be evaluated.
[0018] Optionally, the speech feature extraction unit is configured to divide the speech to be evaluated into sub-speeches to obtain a sub-speech set;
[0019] Extract features from the sub-speeches in the sub-speech set through the speech feature mapping parameters of the speech feature extraction model to obtain speech feature sub-vectors corresponding to the sub-speeches;
[0020] Determine a speech initial feature vector of the speech to be evaluated in a speech feature vector space according to the speech feature sub-vectors;
[0021] Determine a speech correlation weight corresponding to the sub-speech according to the speech initial feature vector, where the speech correlation weight is used to indicate the correlation relationship between the sub-speeches in the sub-speech set;
[0022] Perform weighted calculation on the speech initial feature vector based on the speech correlation weight to obtain a target speech feature vector of the speech to be evaluated in a speech feature vector space.
[0023] Optionally, the semantic feature extraction unit is configured to map the reference text into a semantic feature vector space according to the semantic feature mapping parameters of a semantic feature extraction model, obtain a target semantic feature vector based on the mapping result, and use the target semantic feature vector as the target semantic feature corresponding to the reference text.
[0024] Optionally, the correlation degree calculation unit is configured to perform associated feature calculation on the target speech feature and the target semantic feature through a feature association network to obtain associated features corresponding to the target speech feature and the target semantic feature;
[0025] Based on the classification network, perform correlation analysis on the associated features to determine the feature correlation degree corresponding to the target voice feature and the target semantic feature.
[0026] Optionally, a voice model training unit is further included before the voice feature extraction unit, which is used to extract voice features from the first sample voice through a voice feature extraction model to be trained, and obtain a first voice feature vector corresponding to the first sample voice, where the first sample voice is labeled with a reference voice recognition text;
[0027] Perform text conversion on the first voice feature vector through a voice recognition model to obtain a voice recognition text corresponding to the first sample voice;
[0028] Based on the reference voice recognition text and the voice recognition text, calculate the loss of the voice feature extraction model to be trained;
[0029] Adjust the model parameters of the voice feature extraction model to be trained according to the loss to obtain a trained voice feature extraction model.
[0030] Optionally, a semantic model training unit is further included before the semantic feature extraction unit, which is used to extract semantic features from the first sample text through a semantic feature extraction model to be trained, and obtain a first sample semantic feature vector of the first sample text, where the first sample text contains at least one group of first sample text groups, and the first sample text group includes at least two first sample text sentences and the reference semantic relationship between the first sample text sentences;
[0031] Judge the semantic relationship between the first sample text sentences in each first sample text group according to the first sample semantic feature vector;
[0032] Calculate the loss of the semantic feature extraction model to be trained according to the semantic relationship and the reference semantic relationship;
[0033] Based on the loss, adjust the model parameters of the semantic feature extraction model to be trained to obtain a trained semantic feature extraction model.
[0034] Optionally, a joint training unit is further included before the voice feature extraction unit, which is used to obtain a sample pair, where the sample pair includes a second sample voice, a second sample text, and sample words that appear in the second sample voice in the second sample text, where the second sample text contains at least one group of second sample text groups, and the second sample text group includes two second sample text sentences and the reference semantic relationship between the second sample text sentences;
[0035] Using the speech feature mapping parameters of the speech feature extraction model to be trained, map the second sample speech into the speech feature vector space to obtain a second sample speech feature vector;
[0036] Using the semantic feature mapping parameters of the semantic feature extraction model to be trained, map the second sample text into the semantic feature vector space to obtain a second sample semantic feature vector;
[0037] Based on the second sample speech feature vector and the second sample semantic feature vector of the same sample pair, determine the training words in which the second sample text appears in the second sample speech;
[0038] Based on the second sample semantic feature vector, determine the semantic relationships between the second sample text sentences in each second sample text group;
[0039] According to the training words, the sample words, the semantic relationships and the reference semantic relationships, calculate the losses of the semantic feature extraction model and the speech feature extraction model to be trained;
[0040] Based on the losses, adjust the model parameters of the semantic feature extraction model and the speech feature extraction model to be trained to obtain the trained semantic feature extraction model and speech feature extraction model.
[0041] Optionally, a network training unit is further included before the correlation degree calculation unit, which is used to obtain the third sample speech features corresponding to the third sample speech and the third sample semantic features corresponding to the third sample text, where the third sample speech corresponds to the third sample text, and the third sample speech is marked with a reference evaluation result;
[0042] Perform correlation feature calculation on the third sample speech features and the third sample semantic features through the feature correlation network to be trained to obtain the sample correlation features corresponding to the third sample speech features and the third sample semantic features;
[0043] Through the classification network to be trained, perform correlation analysis on the sample correlation features to determine the feature correlation degree corresponding to the third sample speech features and the third sample semantic features;
[0044] Use the feature correlation degree as the sample evaluation result corresponding to the third sample speech, and based on the sample evaluation result and the reference evaluation result, calculate the losses of the feature correlation network and the classification network;
[0045] Based on the losses, adjust the parameters of the feature correlation network and the classification network to obtain the trained feature correlation network and classification network.
[0046] Optionally, a reference speech feature extraction unit is further included before the correlation degree calculation unit, and is configured to obtain a reference speech corresponding to the reference text;
[0047] Extract speech features from the reference speech to obtain reference speech features corresponding to the reference speech;
[0048] Correspondingly, the correlation degree calculation unit is configured to calculate the feature correlation degree between the target speech features and the reference speech features to obtain a speech feature correlation degree;
[0049] Calculate the feature correlation degree between the target speech features and the target semantic features to obtain a semantic feature correlation degree;
[0050] Based on the speech feature correlation degree and the semantic feature correlation degree, obtain the feature correlation degree between the target speech features and the target semantic features.
[0051] Optionally, a replacement text acquisition unit is further included before the semantic feature extraction unit, and is configured to obtain replacement texts corresponding to each word in the reference text;
[0052] Correspondingly, the semantic feature extraction unit is configured to perform semantic feature extraction based on the reference text and the replacement texts to obtain target semantic features corresponding to the reference text.
[0053] Optionally, the speech to be evaluated is a response speech input by a user for an evaluation question, the reference text is a preset reference answer text for the same evaluation question, and the correlation degree calculation unit is configured to calculate the feature correlation degree between the target speech features and the target semantic features, and the feature correlation degree indicates the correlation degree between the response speech and the reference answer text;
[0054] The evaluation result generation unit is configured to perform evaluation score mapping on the response speech based on the feature correlation degree, determine an evaluation score corresponding to the response speech, and use the evaluation score as an evaluation result corresponding to the response speech.
[0055] Correspondingly, an embodiment of the present invention further provides an electronic device, including a memory and a processor; the memory stores an application program, and the processor is configured to run the application program in the memory to execute the steps in any one of the speech evaluation methods provided by the embodiments of the present invention.
[0056] In addition, an embodiment of the present invention further provides a storage medium, the storage medium stores multiple instructions, and the instructions are suitable for being loaded by a processor to execute the steps in any one of the speech evaluation methods provided by the embodiments of the present invention.
[0057] By adopting the solution of the embodiment of the present invention, the voice to be evaluated and the reference text corresponding to the voice to be evaluated can be obtained, the voice feature extraction is performed on the voice to be evaluated to obtain the target voice feature corresponding to the voice to be evaluated, the semantic feature extraction is performed on the reference text to obtain the target semantic feature corresponding to the reference text, the feature correlation degree calculation is performed according to the target voice feature and the target semantic feature to obtain the feature correlation degree between the target voice feature and the target semantic feature, and based on the feature correlation degree, the evaluation result classification processing is performed on the voice to be evaluated to obtain the evaluation result corresponding to the voice to be evaluated; since in the embodiment of the present invention, after calculating the feature correlation degree between the target voice feature and the target semantic feature by combining the target voice feature and the target semantic feature, the voice to be evaluated is evaluated according to the feature correlation degree, it is possible to extract acoustic and semantic features by combining text and acoustic modalities at the same time, fuse multi-modal information for oral evaluation, reduce the dependence on automatic speech recognition technology, and improve the accuracy of oral evaluation results. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those skilled in the art can obtain other drawings without creative efforts based on these drawings.
[0059] Figure 1 It is a schematic diagram of the scenario of the voice evaluation method provided by the embodiment of the present invention;
[0060] Figure 2 It is a flowchart of the voice evaluation method provided by the embodiment of the present invention;
[0061] Figure 3 It is another flowchart of the voice evaluation method provided by the embodiment of the present invention;
[0062] Figure 4 It is a schematic diagram of the structure of the voice evaluation device provided by the embodiment of the present invention;
[0063] Figure 5 It is another schematic diagram of the structure of the voice evaluation device provided by the embodiment of the present invention;
[0064] Figure 6 It is a schematic diagram of the interface of the oral evaluation application provided by the embodiment of the present invention;
[0065] Figure 7 It is a schematic diagram of the training process of the voice feature extraction model provided by the embodiment of the present invention;
[0066] Figure 8 It is a schematic diagram of the trained voice feature extraction model provided by the embodiment of the present invention;
[0067] Figure 9 It is a schematic diagram of the training process of the semantic feature extraction model provided by an embodiment of the present invention;
[0068] Figure 10 It is a schematic diagram of joint training provided by an embodiment of the present invention;
[0069] Figure 11 It is a schematic diagram of training the feature association network and the classification network provided by an embodiment of the present invention;
[0070] Figure 12 It is a schematic diagram of another scenario of the speech evaluation method provided by an embodiment of the present invention;
[0071] Figure 13 It is a comparison chart of experimental results provided by an embodiment of the present invention;
[0072] Figure 14 It is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention;
[0073] Figure 15 It is an optional schematic diagram of the distributed system 110 applied to the blockchain system provided by an embodiment of the present invention;
[0074] Figure 16 It is an optional schematic diagram of the block structure provided by an embodiment of the present invention. Detailed implementation manners
[0075] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts fall within the protection scope of the present invention.
[0076] An embodiment of the present invention provides a speech evaluation method, device, electronic device, and storage medium. Specifically, an embodiment of the present invention provides a speech evaluation method applicable to a speech evaluation device, and the speech evaluation device can be integrated in an electronic device.
[0077] The electronic device can be a terminal device or the like, including but not limited to mobile terminals and fixed terminals. For example, mobile terminals include but are not limited to smart phones, smart watches, tablet computers, laptop computers, vehicle-mounted terminals, smart voice interaction devices, etc. Among them, fixed terminals include but are not limited to desktop computers, smart home appliances, etc.
[0078] The electronic device may also be a device such as a server, which may be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms, but is not limited thereto.
[0079] The voice evaluation method according to the embodiments of the present invention may be implemented by a server or jointly implemented by a terminal and a server.
[0080] Taking the joint implementation of the voice evaluation method by the terminal and the server as an example, the method will be described below.
[0081] As Figure 1 shown, the voice evaluation system provided by the embodiments of the present invention includes a terminal 10, a server 20, etc.; the terminal 10 is connected to the server 20 through a network, for example, through a wired or wireless network connection, etc. Among them, the terminal 10 may exist as a terminal for a user to send a user voice to be evaluated to the server 20, and the terminal includes but is not limited to a mobile phone, a computer, a smart voice interaction device, a smart home appliance, a vehicle-mounted terminal, etc.
[0082] Among them, the terminal 10 may be a terminal for a user to upload a voice to be evaluated, and is used to send the obtained voice to be evaluated to the server 20.
[0083] The server 20 may be used to obtain a voice to be evaluated and a reference text corresponding to the voice to be evaluated, extract voice features from the voice to be evaluated to obtain target voice features corresponding to the voice to be evaluated, extract semantic features from the reference text to obtain target semantic features corresponding to the reference text, calculate a feature correlation degree according to the target voice features and the target semantic features to obtain a feature correlation degree between the target voice features and the target semantic features, and perform evaluation result classification processing on the voice to be evaluated based on the feature correlation degree to obtain an evaluation result corresponding to the voice to be evaluated.
[0084] In some embodiments, the server 20 may send the evaluation result to the terminal 10, and the terminal 10 may display the evaluation result corresponding to the voice to the user.
[0085] The following will be described in detail respectively. It should be noted that the description order of the following embodiments does not limit the preferred order of the embodiments.
[0086] The embodiments of the present invention will be described from the perspective of a voice evaluation device, which may be specifically integrated in a server or a terminal.
[0087] A speech evaluation method provided by an embodiment of the present invention can be executed by a processor of a terminal or a server. For example, Figure 2 As shown, the specific process of the speech evaluation method in this embodiment can be as follows:
[0088] 201. Obtain the speech to be evaluated and the reference text corresponding to the speech to be evaluated.
[0089] Among them, the speech to be evaluated can be a speech file directly submitted by a user through an electronic device, or it can also be a speech generated by the electronic device after collecting external sounds.
[0090] For example, Figure 6 is a schematic diagram of an interface of the oral evaluation application provided by an embodiment of the present invention. A user can click on a control named "Start Recording" in the speech collection page shown as 601 in Figure 6 to trigger the electronic device to start collecting external sounds. When the user clicks on a control named "End Recording" in the end collection page shown as 602 in Figure 6 , it triggers the electronic device to end the collection of external sounds and generate the speech to be evaluated based on the collected sounds.
[0091] Among them, the reference text is the text content corresponding to the speech to be evaluated. For example, if the speech to be evaluated is the user's response speech to a certain oral test question, the reference text can be the pre-set response text to the same oral test question.
[0092] It can be understood that the oral content in the speech to be evaluated may not be the same as the content in the reference text. For example, in the display area named "Question Display Area" shown in Figure 6 , the oral test question displayed is "What's your favorite sport?", the content said by the user in the speech to be evaluated can be "My favorite sport is swimming.", while the content in the provided reference text can include "My favorite sport is basketball because basketball is a very confrontational sport; My favorite sport is yoga because yoga can make me relax physically and mentally." and so on.
[0093] It should be noted that the schematic diagram of the interface of the oral evaluation application provided by the embodiment of the present invention should not be construed as a limitation to the embodiment of the present invention.
[0094] 202. Extract speech features from the speech to be evaluated to obtain the target speech features corresponding to the speech to be evaluated.
[0095] In some optional examples, by extracting the spectrogram and frequency spectrum of the speech to be evaluated, and then performing processing such as framing, windowing, filtering, and Fourier transform on the extracted spectrogram and frequency spectrum, the target speech features corresponding to the speech to be evaluated can be obtained.
[0096] In some other alternative examples, speech feature extraction can be performed on the speech to be evaluated through machine learning and speech technologies. Among them, Machine Learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning generally include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.
[0097] The key technologies of Speech Technology include Automatic Speech Recognition (ASR), Text To Speech (TTS), and voiceprint recognition technology. Enabling computers to listen, see, speak, and feel is the future development direction of human-computer interaction, and among them, speech has become one of the most promising human-computer interaction methods in the future.
[0098] In some embodiments, step 202 may include: according to the speech feature mapping parameters of the speech feature extraction model, mapping the speech to be evaluated into the speech feature vector space, obtaining the target speech feature vector based on the mapping result, and using the target speech feature vector as the target speech feature corresponding to the speech to be evaluated.
[0099] Among them, the speech feature extraction model is a model that can perform feature extraction on speech. For example, the speech feature extraction model may include a convolutional neural network and a feature pyramid net, etc. That is to say, the step "according to the speech feature mapping parameters of the speech feature extraction model, mapping the speech to be evaluated into the speech feature vector space, obtaining the target speech feature vector based on the mapping result" may include: performing feature extraction on the speech to be evaluated through the convolutional neural network parameters of the convolutional neural network to obtain the feature information output by multiple convolutional layers in the speech feature vector space, respectively processing the feature information output by the multiple convolutional layers through the feature pyramid net, and generating the target speech feature of the speech to be evaluated according to the processing result of each layer in the feature pyramid net.
[0100] In some other alternative embodiments, a multi-head attention network or a self-attention network can be adopted to extract features from the speech to be evaluated, so as to obtain the speech correlation weights of the speech to be evaluated. For example, the specific extraction process can be to convert the initial speech feature vector into spatial vectors of multiple dimensions, and then use these spatial vectors of multiple dimensions as the speech correlation weights of each sub-speech in the sub-speech set.
[0101] That is, the step of "mapping the speech to be evaluated into the speech feature vector space according to the speech feature mapping parameters of the speech feature extraction model, and obtaining the target speech feature vector based on the mapping result" may include:
[0102] Dividing the speech to be evaluated into sub-speeches to obtain a sub-speech set;
[0103] Extracting features of the sub-speeches in the sub-speech set through the speech feature mapping parameters of the speech feature extraction model to obtain speech feature sub-vectors corresponding to the respective sub-speeches;
[0104] Determining an initial speech feature vector of the speech to be evaluated in the speech feature vector space according to the speech feature sub-vectors;
[0105] Determining the speech correlation weights corresponding to the sub-speeches according to the initial speech feature vector, where the speech correlation weights are used to indicate the correlation relationship between the sub-speeches in the sub-speech set;
[0106] Performing weighted calculation on the initial speech feature vector based on the speech correlation weights to obtain the target speech feature vector of the speech to be evaluated in the speech feature vector space.
[0107] Among them, the sub-speech is the speech obtained after dividing the speech to be evaluated, and the specific speech division rule can be set by those skilled in the art according to the actual application situation. For example, the speech division rule can be to divide the speech to be evaluated every 30 ms, or the speech to be evaluated can be divided into 10 sub-speeches with the same time length according to the time length, or the speech to be evaluated can be divided with an indefinite time length according to the spectrogram and spectrum of the speech to be evaluated, and so on. The embodiments of the present invention do not make any limitations in this regard.
[0108] Among them, the initial speech feature vector can be obtained by directly splicing the speech feature sub-vectors, or the initial speech feature vector can be obtained through vector processing processes such as weighted calculation of the speech feature sub-vectors.
[0109] For example, taking the conversion of the initial speech feature vector into spatial vectors of multiple dimensions as an example, the step of "determining the speech correlation weights corresponding to the sub-speeches according to the initial speech feature vector" may include:
[0110] The multi - head attention network is used to convert the initial speech feature vector into query vector (q), key vector (k) and value vector (v). For example, specifically, the self - attention network can be used to fuse the initial speech feature vector with transformation parameters in three dimensions respectively to obtain the query vector (q), key vector (k) and value vector (v), and the query vector (q), key vector (k) and value vector (v) are used as the speech association weights of each sub - speech in the sub - speech set.
[0111] For another example, the speech to be evaluated can be cut in the time dimension. For example, a sample point is determined every 10 ms, and the speech to be evaluated is divided into N sub - samples (sub - speeches). Each sub - sample has D dimensions, so each speech feature sub - vector can be a 1*D - dimensional vector. The speech feature sub - vectors are concatenated to obtain an N*D - dimensional initial speech feature vector. The initial speech feature vector is multiplied by the transposed vector of its own D*N dimensions, which is equivalent to calculating the similarity of each sub - speech with the other N - 1 sub - speeches, and a resulting N*N - dimensional vector is the speech association weight. The speech association weight is multiplied by the initial speech feature vector, and the resulting new N*D - dimensional vector is the target speech feature vector.
[0112] It can be understood that in order to improve the accuracy of the speech feature extraction model for feature extraction, as Figure 7 shown, the speech feature extraction model can be pre - trained. That is, before the step of "mapping the speech to be evaluated into the speech feature vector space according to the speech feature mapping parameters of the speech feature extraction model and obtaining the target speech feature vector based on the mapping result", it further includes:
[0113] Through the speech feature extraction model to be trained, the speech feature extraction of the first sample speech is carried out to obtain the first speech feature vector corresponding to the first sample speech, where the first sample speech is labeled with the reference speech recognition text;
[0114] Through the speech recognition model, the text conversion of the first speech feature vector is carried out to obtain the speech recognition text corresponding to the first sample speech;
[0115] Based on the reference speech recognition text and the speech recognition text, the loss of the speech feature extraction model to be trained is calculated;
[0116] According to the loss, the model parameters of the speech feature extraction model to be trained are adjusted to obtain the trained speech feature extraction model.
[0117] Among them, the first sample voice can be the voice that has been evaluated or to be evaluated generated during the collected oral evaluation process. To reduce the dependence on manual labor, it can also be a conventional ASR model training sample, such as datasets like LibriSpeech ASR corpus, THCHS-30, VoxForge, etc.
[0118] Among them, the speech recognition model can be a trained model or an untrained model. The speech recognition model can perform speech recognition based on the first speech feature vector and recognize the text corresponding to the first speech feature vector.
[0119] In some alternative examples, the loss of the speech feature extraction model can be obtained by solving through functions such as cross-entropy function and gradient descent method. The embodiments of the present invention do not make limitations in this regard.
[0120] It can be understood that as Figure 8 shown, in the trained speech feature extraction model, in addition to the attention network, it can also include a feed-forward neural network, etc. Technicians can add a word embedding network, etc. according to actual needs.
[0121] 203. Extract semantic features from the reference text to obtain the target semantic features corresponding to the reference text.
[0122] Among them, the target semantic features can be semantic features that represent the true meaning of the text content. The so-called semantic features, also known as sememes (SEME), are the constituent factors of the semantic units (MEME, equivalent to the smallest unit of the meaning of a semantic item), and are the distinctive features of the semantic unit. It can represent the combination relationship between words and other words.
[0123] In some examples, after performing word segmentation and other processing on the reference text, according to the preset mapping relationship between the text and the semantic features, the target speech features corresponding to the reference text can be determined.
[0124] In other examples, machine learning techniques and natural language processing techniques can be combined to extract semantic features from the reference text. Natural Language Processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between humans and computers in natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, the research in this field will involve natural language, that is, the language used by people in daily life, so it has a close connection with the research of linguistics. Natural language processing techniques usually include techniques such as text processing, semantic understanding, machine translation, robot question answering, and knowledge graphs.
[0125] In some embodiments, step 203 may include: according to the semantic feature mapping parameters of the semantic feature extraction model, mapping the reference text into the semantic feature vector space, obtaining the target semantic feature vector based on the mapping result, and using the target semantic feature vector as the target semantic feature corresponding to the reference text.
[0126] Among them, the semantic feature extraction model is a model that can extract semantic features from text. In some alternative embodiments, a multi-head attention network or a self-attention network may be used to extract semantic features from the reference text to obtain the word association weights of the reference text. For example, the specific extraction process may be to convert the semantic initial feature vector into spatial vectors of multiple dimensions, and then use these spatial vectors of multiple dimensions as the semantic association weights of each word in the word set.
[0127] That is, the step "according to the semantic feature mapping parameters of the semantic feature extraction model, mapping the reference text into the semantic feature vector space, and obtaining the target semantic feature vector based on the mapping result" may include:
[0128] Dividing the reference text into words to obtain a word set;
[0129] Performing feature extraction on the words in the word set through the semantic feature mapping parameters of the semantic feature extraction model to obtain semantic feature sub-vectors corresponding to the respective words;
[0130] Determining the semantic initial feature vector of the reference text in the semantic feature vector space according to the semantic feature sub-vectors;
[0131] Determining the semantic association weights corresponding to the words according to the semantic initial feature vector, where the semantic association weights are used to indicate the association relationship between the words in the word set;
[0132] Performing weighted calculation on the semantic initial feature vector based on the semantic association weights to obtain the target semantic feature vector of the reference text in the semantic feature vector space.
[0133] Among them, the words are the words obtained by dividing the reference text. The specific word division rules can be set by those skilled in the art according to the actual application situation. For example, the word division rules can be to divide the reference text every 10 characters, or to divide the reference text with variable word lengths according to the punctuation marks in the reference text, etc. The embodiments of the present invention do not limit this.
[0134] Among them, the semantic initial feature vector can be obtained by directly performing vector splicing on the semantic feature sub-vectors, or the semantic initial feature vector can be obtained through vector processing processes such as weighted calculation of the semantic feature sub-vectors.
[0135] It is understandable that, in order to improve the accuracy of the semantic feature extraction model for feature extraction, as Figure 9 shown, the semantic feature extraction model can be pre-trained, that is, before the step of "mapping the reference text into the semantic feature vector space according to the semantic feature mapping parameters of the semantic feature extraction model and obtaining the target semantic feature vector based on the mapping result", it further includes:
[0136] Performing semantic feature extraction on the first sample text through the semantic feature extraction model to be trained to obtain the first sample semantic feature vector of the first sample text, where the first sample text contains at least one group of first sample text groups, and each first sample text group includes at least two first sample text sentences and the reference semantic relationship between the first sample text sentences;
[0137] Judging the semantic relationship between the first sample text sentences in each first sample text group according to the first sample semantic feature vector;
[0138] Calculating the loss of the semantic feature extraction model to be trained according to the semantic relationship and the reference semantic relationship;
[0139] Based on the loss, adjusting the model parameters of the semantic feature extraction model to be trained to obtain the trained semantic feature extraction model.
[0140] Among them, as Figure 9 shown, when judging the semantic relationship between the first sample text sentences in each first sample text group according to the first sample semantic feature vector, the semantic relationship between the first sample text sentences can be determined through the semantic understanding model.
[0141] For example, the text in the first sample text group can be divided into the antecedent of the implication, denoted as P (premise), that is, the premise, and the consequent of the implication, denoted as H (hypothesis), that is, the hypothesis. There are three semantic relationships between the antecedent of the implication and the consequent of the implication. If P can infer H, it is an implicative relationship; if P cannot infer H, it is a neutral relationship. If P can infer the opposite conclusion of H, it is a contradictory relationship. After performing semantic feature extraction on P and H respectively by inputting them into two semantic feature extraction models, and then passing through the attention network, pooling layer and classification network in the semantic understanding model, three classification results of the semantic relationship between P and H are finally obtained.
[0142] Specifically, the loss of the semantic feature extraction model can be obtained by solving through the cross-entropy function, gradient descent method, etc. The embodiments of the present invention do not limit this, and the loss of the semantic feature extraction model can also be calculated by the following formula:
[0143]
[0144] Among them, p is the prediction result (the probability of each semantic relationship between P and H obtained through the training process), and y is the label (the reference semantic relationship between the first sample text sentences). i is the i-th sample group among m sample groups, and j is the j-th semantic relationship. For example, j can be 0, 1, or 2, representing the entailment relationship, neutral relationship, and contradiction relationship respectively. At this time, p represents the probability value of P and H belonging to each semantic relationship.
[0145] In some alternative examples, the reference text can be preprocessed first. For example, the sentences in the reference text can be segmented based on a dictionary, and the segmented words can be combined to obtain an initial text word set, and so on. Among them, there can also be multiple specific word segmentation algorithms based on the dictionary. For example, the maximum matching word segmentation algorithm, the shortest path word segmentation algorithm, and the word segmentation algorithm based on the n-gram model (a word segmentation algorithm).
[0146] Alternatively, a model-based word segmentation algorithm can also be used to segment the sentences in the reference text by characters, and the segmented characters can be combined to obtain text words, thereby obtaining an initial text word set. Among them, this model-based word segmentation algorithm can also include multiple types, such as generative model word segmentation algorithms, discriminant model word segmentation algorithms, and neural network word segmentation algorithms, and so on.
[0147] After obtaining the initial text word set, the text words in the initial text word set can be filtered to obtain a text word set. Among them, the filtering methods can also include multiple types. For example, stop word processing can be performed on the text words in the initial text word set. For example, the text words included in a preset stop word library can be screened out from the initial text word set to obtain the stop words in the initial text word set, and then these stop words can be filtered to obtain a text word set. Or, useless word filtering can also be performed on the text words in the initial text word set based on the regular expression corresponding to the preset useless words, and then the initial text word set after filtering out the useless words can be used as the text word set, and so on.
[0148] Among them, stop words refer to certain words or characters that are automatically filtered out before or after processing natural language data (or text) in information retrieval to save storage space and improve search efficiency. These words or characters are called Stop Words (stop words). These stop words are all manually input and not automatically generated. After generation, the stop words will form a stop word library (table). And useless words refer to some words that have nothing to do with information classification.
[0149] When extracting semantic features from the reference text, it can be after processing the reference text to obtain a text word set and then extracting semantic features from the text words, and so on.
[0150] It can be understood that by jointly training the speech feature extraction model and the semantic feature extraction model, the speech feature extraction model and the semantic feature extraction model can be optimized simultaneously, and the effect of connecting the speech feature vector space and the semantic feature vector space can also be achieved. That is, as Figure 10 shown, before the step of "mapping the reference text into the semantic feature vector space according to the semantic feature mapping parameters of the semantic feature extraction model and obtaining the target semantic feature vector based on the mapping result", it also includes:
[0151] Obtain sample pairs, where each sample pair includes a second sample speech, a second sample text, and sample words that appear in the second sample speech in the second sample text. Among them, the second sample text includes at least one set of second sample text groups, and each second sample text group includes two second sample text sentences and the reference semantic relationship between the second sample text sentences;
[0152] Map the second sample speech into the speech feature vector space through the speech feature mapping parameters of the speech feature extraction model to be trained, and obtain the second sample speech feature vector;
[0153] Map the second sample text into the semantic feature vector space through the semantic feature mapping parameters of the semantic feature extraction model to be trained, and obtain the second sample semantic feature vector;
[0154] Based on the second sample speech feature vector and the second sample semantic feature vector of the same sample pair, determine the training words that appear in the second sample speech in the second sample text;
[0155] Based on the second sample semantic feature vector, determine the semantic relationship between the second sample text sentences in each second sample text group;
[0156] Calculate the losses of the semantic feature extraction model and the speech feature extraction model to be trained according to the training words, sample words, semantic relationships, and reference semantic relationships;
[0157] Based on the losses, adjust the model parameters of the semantic feature extraction model and the speech feature extraction model to be trained to obtain the trained semantic feature extraction model and speech feature extraction model.
[0158] For example, during joint training, two different training tasks can be set, such as a word matching task and a natural language understanding task, and multi-task learning is performed on these two tasks. In the word matching task, each word in the transcribed text is mainly used to search in the speech to determine whether the word is included, that is, to determine the training words that appear in the second sample speech in the second sample text. This matching process is achieved through the multi-head attention mechanism. Finally, the classification result of whether each word appears in the speech is output. The natural language understanding task is to infer the semantic relationship between a pair of texts, that is, to input a premise and a hypothesis and output the relationship between the premise and the hypothesis. The process is similar to the separate training of the semantic feature extraction model, and the present invention will not elaborate on this.
[0159] Among them, the loss of the word matching task can be calculated by the following formula:
[0160]
[0161] Among them, y is the word matching label, p is the word matching prediction result. i is the i-th word in a sentence, n j is the number of words in a sentence, m is the number of all sentences, and ω represents the ω-th group of the second sample text group.
[0162] On the other hand, during joint training, the optimization of the model is carried out by multi-task learning. By combining the losses of the two tasks and optimizing them simultaneously, the overall loss of the joint training is shown in the following formula:
[0163] l = γ × l ω +(1 - γ) × l NLU
[0164] Among them, γ is the weight of the two tasks. The range of this weight value is between 0 and 1, and it is a parameter set in advance before joint training.
[0165] In the actual application process, there may be many synonyms that can replace the reference answers given in the reference text. For example, for "like" in the reference answer, there are also "love", "be fond of", "have a passion for", etc. that can be replaced. Therefore, when performing semantic feature extraction on the reference text, the semantic features can be extracted by combining the synonyms or near-synonyms of each word in the reference text. That is, before the step of "performing semantic feature extraction on the reference text to obtain the target semantic features corresponding to the reference text", the following steps can also be included:
[0166] Obtain the replacement texts corresponding to each word in the reference text;
[0167] Correspondingly, the step of "performing semantic feature extraction on the reference text to obtain the target semantic features corresponding to the reference text" can include:
[0168] Semantic feature extraction is performed based on the reference text and the replacement text to obtain the target semantic features corresponding to the reference text.
[0169] Among them, the replacement text can be synonyms or near-synonyms of each word, etc. When performing semantic feature extraction based on the reference text and the replacement text, the words in the reference text can be replaced with the replacement text respectively to obtain a semantic feature, and finally the target semantic feature is obtained through operations on all the semantic features. Or, after performing semantic feature extraction on the reference text, semantic feature extraction is performed on the replacement text, and finally the target semantic feature is obtained through operations on all the semantic features.
[0170] 204. Calculate the feature correlation degree based on the target speech feature and the target semantic feature to obtain the feature correlation degree between the target speech feature and the target semantic feature.
[0171] In some embodiments, for example, both the target speech feature and the target semantic feature are vectors. At this time, the feature correlation degree calculation can be directly calculating the vector similarity between the two vectors as the feature correlation degree.
[0172] In some other embodiments, the correlation feature that can represent the correlation relationship between the two features can be obtained first, and then the correlation analysis is performed on the correlation feature to determine the feature correlation degree between the two features. That is, step 204 may include:
[0173] Perform correlation feature calculation on the target speech feature and the target semantic feature through the feature correlation network to obtain the correlation features corresponding to the target speech feature and the target semantic feature;
[0174] Based on the classification network, perform correlation analysis on the correlation features to determine the feature correlation degree corresponding to the target speech feature and the target semantic feature.
[0175] Among them, the feature correlation network can be an attention network, and the correlation features between the target speech feature and the target semantic feature are calculated through the multi-head attention mechanism, etc.
[0176] In some optional examples, before the classification network, a pooling layer can also be added to process the correlation features, reduce the data volume of the correlation features, and further improve the accuracy of the representation of the correlation relationship between the correlation features and the target speech feature and the target semantic feature.
[0177] In some optional examples, in order to improve the accuracy of the finally obtained feature correlation degree, the feature correlation network and the classification network can be trained first. As Figure 11 shown, before the step of "performing correlation feature calculation on the target speech feature and the target semantic feature through the feature correlation network", it also includes:
[0178] Obtain the third sample speech features corresponding to the third sample speech and the third sample semantic features corresponding to the third sample text, where the third sample speech corresponds to the third sample text, and the third sample speech is labeled with the reference evaluation result;
[0179] Perform associated feature calculation on the third sample speech features and the third sample semantic features through the feature association network to be trained, and obtain the sample associated features corresponding to the third sample speech features and the third sample semantic features;
[0180] Through the classification network to be trained, perform association analysis on the sample associated features to determine the feature association degree corresponding to the third sample speech features and the third sample semantic features;
[0181] Based on the feature association degree, perform evaluation result classification processing on the third sample speech to obtain the sample evaluation result corresponding to the third sample speech, and calculate the losses of the feature association network and the classification network based on the sample evaluation result and the reference evaluation result;
[0182] Based on the losses, adjust the parameters of the feature association network and the classification network to obtain the trained feature association network and classification network.
[0183] Among them, the third sample speech is the speech obtained after artificial speech evaluation, and the third sample text is the reference answer text for a certain test question corresponding to the third sample speech respectively. The reference evaluation result is the result obtained by manually listening to the speech.
[0184] For example, the data set can be the data of a certain oral exam. Read a short passage, give questions, and students answer the questions. It contains n pieces of data, where there are 10 questions in total, n / 10 pieces of data for each question, and multiple reference answers are given for each question. Divide the data set, 40% is used for training, and 60% is used for prediction. The manual scoring is from 1 to 5 points, 1 point means completely answering wrong, and 5 points means completely answering correctly.
[0185] Among them, the losses of the feature association network and the classification network can be calculated based on the following formula,
[0186]
[0187] Among them, y scorej is the manual evaluation result of a certain third sample speech, p scorej is the evaluation result obtained by processing a certain third sample speech through the feature association network and the classification network. m represents that there are m third sample speeches in total.
[0188] In some embodiments, a reference speech identical to the reference text content can be obtained, and in terms of acoustics including pauses, tones, etc., the accuracy of speech evaluation can be further improved. That is, before the step of "calculating the feature correlation degree according to the target speech feature and the target semantic feature", the following steps may also be included:
[0189] Obtain a reference speech corresponding to the reference text;
[0190] Extract speech features from the reference speech to obtain reference speech features corresponding to the reference speech;
[0191] Correspondingly, the step of "calculating the feature correlation degree according to the target speech feature and the target semantic feature to obtain the feature correlation degree between the target speech feature and the target semantic feature" includes:
[0192] Calculate the feature correlation degree between the target speech feature and the reference speech feature to obtain a speech feature correlation degree;
[0193] Calculate the feature correlation degree between the target speech feature and the target semantic feature to obtain a semantic feature correlation degree;
[0194] Based on the speech feature correlation degree and the semantic feature correlation degree, obtain the feature correlation degree between the target speech feature and the target semantic feature.
[0195] Among them, when calculating the feature correlation degree between the target speech feature and the reference speech feature, an attention mechanism can be used for operation, which will not be elaborated in this embodiment of the present invention.
[0196] 205. Based on the feature correlation degree, perform evaluation result classification processing on the speech to be evaluated to obtain an evaluation result corresponding to the speech to be evaluated.
[0197] Among them, the evaluation result can be a specific score, such as 81 points, or an evaluation level such as excellent, good, etc.
[0198] For example, an evaluation result can be determined for the speech to be evaluated according to the feature correlation degree and a preset correlation degree threshold. For example, when the maximum correlation degree is 5, it can be set that when the correlation degree is greater than 4, the evaluation level is excellent, and so on.
[0199] Or, the feature correlation degree can be directly used as the evaluation result. For example, if the feature correlation degree is 4.91, the evaluation result is 4.91 points, and so on.
[0200] It can be understood that the speech to be evaluated can be the response speech input by the user for the evaluation question, and the reference text is the preset reference answer text for the same evaluation question;
[0201] Correspondingly, the step of "calculating the feature correlation degree based on the target speech feature and the target semantic feature to obtain the feature correlation degree between the target speech feature and the target semantic feature" may include:
[0202] Based on the target speech feature and the target semantic feature, calculate the feature correlation degree between the features, and the feature correlation degree indicates the correlation degree between the response speech and the reference answer text;
[0203] Correspondingly, the step of "performing evaluation result classification processing on the speech to be evaluated based on the feature correlation degree to obtain the evaluation result corresponding to the speech to be evaluated" may include:
[0204] Based on the feature correlation degree, perform evaluation score mapping on the response speech, determine the evaluation score corresponding to the response speech, and use the evaluation score as the evaluation result corresponding to the response speech.
[0205] Compare the speech evaluation method in the embodiments of the present invention with the traditional model constructed based on ASR features. The traditional model constructed based on ASR features includes two comparison models. One is to construct the model using the SVR model, and the other is to construct it using the BLSTM model combined with the attention mechanism. The evaluation metrics are the Pearson correlation coefficient and the agreement rate of manual scoring and machine scoring. The agreement rate represents the proportion of equal manual scoring and machine scoring. The results are as Figure 13 shown. It can be seen from the results that the embodiments of the present invention perform better than the traditional open-ended question evaluation model constructed based on ASR features.
[0206] As can be seen from the above, the solution of the embodiments of the present invention can obtain the speech to be evaluated and the reference text corresponding to the speech to be evaluated, extract the speech features of the speech to be evaluated to obtain the target speech feature corresponding to the speech to be evaluated, extract the semantic features of the reference text to obtain the target semantic feature corresponding to the reference text, calculate the feature correlation degree based on the target speech feature and the target semantic feature to obtain the feature correlation degree between the target speech feature and the target semantic feature, and perform evaluation result classification processing on the speech to be evaluated based on the feature correlation degree to obtain the evaluation result corresponding to the speech to be evaluated; Since in the embodiments of the present invention, after calculating the feature correlation degree between the target speech feature and the target semantic feature by combining the target speech feature and the target semantic feature, the speech to be evaluated is evaluated according to the feature correlation degree, it is possible to extract acoustic and semantic features by simultaneously combining the text and acoustic modalities, fuse multi-modal information for oral evaluation, reduce the dependence on automatic speech recognition technology, and improve the accuracy of oral evaluation results
[0207] In addition, the embodiments of the present invention do not require a large amount of manually evaluated data to train the model. Instead, a large amount of available ASR training data and natural language understanding data are used to pre-train the model, which can reduce the dependence on manually labeled speech evaluation data and save human resources.
[0208] According to the method described in the previous embodiments, the following will be further described in detail with examples.
[0209] In this embodiment, in combination with Figure 1 the system, and taking the speech feature extraction model to extract features from the speech to be detected, the semantic feature extraction model to extract features from the reference text, and obtaining the evaluation result through the vector association network and the classification network as an example for illustration.
[0210] As Figure 3 shown, the specific process of the speech evaluation method in this embodiment can be as follows:
[0211] 301. The server obtains the speech feature extraction model to be trained, trains the speech feature extraction model to be trained, and obtains the initially trained speech feature extraction model.
[0212] Among them, the training process of the speech feature extraction model includes:
[0213] Through the speech feature extraction model to be trained, extract speech features from the first sample speech to obtain the first speech feature vector corresponding to the first sample speech, where the first sample speech is labeled with the reference speech recognition text;
[0214] Through the speech recognition model, perform text conversion on the first speech feature vector to obtain the speech recognition text corresponding to the first sample speech;
[0215] Based on the reference speech recognition text and the speech recognition text, calculate the loss of the speech feature extraction model to be trained;
[0216] According to the loss, adjust the model parameters of the speech feature extraction model to be trained to obtain the initially trained speech feature extraction model.
[0217] Among them, the speech feature extraction model can be the encoder of the Transformer model, and the speech recognition model can be the decoder of the Transformer model. After training, the encoder of the Transformer model is used as the speech feature extraction model.
[0218] In some optional examples, an embedding layer can also be included before the speech feature extraction model to perform an embedding operation on the input data and input the data after the embedding operation into the speech feature extraction model.
[0219] 302. The server obtains a semantic feature extraction model to be trained, and trains the semantic feature extraction model to be trained to obtain a preliminarily trained semantic feature extraction model.
[0220] Among them, the training process of the semantic feature extraction model may include: through the semantic feature extraction model to be trained, extracting semantic features of the first sample text to obtain a first sample semantic feature vector of the first sample text. Among them, the first sample text contains at least one group of first sample text groups, and each first sample text group includes at least two first sample text sentences and the reference semantic relationship between the first sample text sentences;
[0221] According to the first sample semantic feature vector, judge the semantic relationship between the first sample text sentences in each first sample text group;
[0222] According to the semantic relationship and the reference semantic relationship, calculate the loss of the semantic feature extraction model to be trained;
[0223] Based on the loss, adjust the model parameters of the semantic feature extraction model to be trained to obtain a trained semantic feature extraction model.
[0224] When training the semantic feature extraction model, the Natural Language Understanding (NLU) task can be used to train the semantic feature extraction model.
[0225] Among them, natural language understanding is a new interdisciplinary subject, the content involves linguistics, psychology, logic, acoustics, mathematics and computer science, and is based on linguistics. The research of natural language understanding comprehensively applies the knowledge of modern phonetics, phonology, grammar, semantics and pragmatics.
[0226] 303. The server jointly trains the preliminarily trained speech feature extraction model and semantic feature extraction model to obtain a trained speech feature extraction model and semantic feature extraction model.
[0227] Among them, when jointly training, two parts of data can be input, one part is the manually transcribed data corresponding to the second sample speech, and the other part is the NLU task data, including premise and hypothesis pairs.
[0228] When jointly training, a word matching task and a natural language understanding task are adopted, and the two tasks perform multi-task learning. The word matching task is used to connect the pronunciation space and the corresponding text space, and the natural language understanding task is used to connect the semantic space. The two tasks share the semantic feature extraction model to achieve the effect of connecting the pronunciation space and the speech space.
[0229] 304. The server trains the feature association network and the classification network to be trained, and obtains the trained feature association network and classification network.
[0230] In an optional example, a scoring module including a feature association network and a classification network can be constructed based on a pre-trained speech feature extraction model and a semantic feature extraction model. When training the feature association network and the classification network, the third sample speech and the third sample text are respectively input into the speech feature extraction model and the semantic feature extraction model, and based on the multi-head attention mechanism, the pooling layer, and the non-linear transformation (MLP) layer (classification network), a scoring result in the range of 0 to 1 is finally output.
[0231] Among them, the training loss function of the scoring module can be the difference between the scoring result of the scoring module and the human evaluation result corresponding to the third sample speech.
[0232] 305. The terminal obtains the speech to be evaluated submitted by the user and sends the speech to be evaluated to the server.
[0233] As Figure 12 shown, the user submits the answering speech (speech to be evaluated) through the oral English evaluation application installed on the terminal, and the terminal sends the answering speech to the server corresponding to the oral English evaluation application through the interface of the oral English evaluation application for speech evaluation.
[0234] Among them, the terminals used by the user include but are not limited to mobile phones, computers, intelligent voice interaction devices, smart home appliances, vehicle-mounted terminals, etc.
[0235] 306. The server receives the speech to be evaluated and the question information of the speech to be evaluated, and determines the reference text corresponding to the speech to be evaluated according to the question information.
[0236] For example, the question information can be the identifier of the question such as "Question 1", "Question 2", etc., or the question information can be the question title, etc., such as "What’s your favourite sport?".
[0237] 307. The server maps the speech to be evaluated into the speech feature vector space through the speech feature mapping parameters of the speech feature extraction model, and obtains the target speech feature vector based on the mapping result.
[0238] Among them, the speech feature extraction model can include an attention network. Therefore, step 307 can include:
[0239] Dividing the speech to be evaluated into sub-speeches to obtain a sub-speech set;
[0240] Extract the speech features of the sub-speech in the sub-speech set through the speech feature mapping parameters of the speech feature extraction model to obtain the speech feature sub-vectors corresponding to each sub-speech;
[0241] Determine the initial speech feature vector of the speech to be evaluated in the speech feature vector space according to the speech feature sub-vectors;
[0242] Determine the speech association weights corresponding to the sub-speeches according to the initial speech feature vector, and the speech association weights are used to indicate the association relationship between the sub-speeches in the sub-speech set;
[0243] Based on the speech association weights, perform weighted calculation on the initial speech feature vector to obtain the target speech feature vector of the speech to be evaluated in the speech feature vector space.
[0244] In one example, the speech feature extraction model may further include a sequence encoding layer, which can interpret the order of sub-vectors in the input vector sequence. The sequence encoding layer can determine the position of the current speech feature sub-vector, etc.
[0245] In some embodiments, the attention network may be a multi-head attention network, which initializes not only a group of matrices of Q, K, and V, but multiple groups, etc.
[0246] 308. The server maps the reference text into the semantic feature vector space according to the semantic feature mapping parameters of the semantic feature extraction model, and obtains the target semantic feature vector based on the mapping result.
[0247] In some embodiments, a multi-head attention network (Multi-Head Attention) or a self-attention network (self-attention) may be used to extract the semantic features of the reference text to obtain the word association weights of the reference text. For example, the specific extraction process may be to convert the initial semantic feature vector into a spatial vector of multiple dimensions, and then use the spatial vector of multiple dimensions as the semantic association weight of each word in the word set. This is not elaborated in the embodiments of the present invention.
[0248] 309. The server performs associated feature calculation on the target speech feature vector and the target semantic feature vector through the feature association network to obtain the associated feature vector corresponding to the target speech feature vector and the target semantic feature vector.
[0249] Among them, the feature association network may be a multi-head attention network (Multi-Head Attention) or a self-attention network (self-attention) to perform associated feature calculation on the target speech feature vector and the target semantic feature vector.
[0250] For example, step 309 may include:
[0251] Divide the target speech feature vector into target speech feature sub - vectors, and divide the target semantic feature vector into target semantic feature sub - vectors;
[0252] According to the target speech feature sub - vectors and the target semantic feature sub - vectors, determine the correlation weights corresponding to the target speech feature vector and the target semantic feature vector. The correlation weights are used for the correlation relationship between the target speech feature vector and the target semantic feature vector;
[0253] Based on the correlation weights, perform weighted calculation on the target speech feature vector and the target semantic feature vector to obtain the correlation feature vector corresponding to the target speech feature vector and the target semantic feature vector.
[0254] 310. The server performs correlation analysis on the correlation feature vector based on the classification network to determine the vector correlation degree corresponding to the target speech feature vector and the target semantic feature vector.
[0255] Among them, the classification network can be a non - linear transformation (MLP) layer, etc. In some alternative examples, a pooling layer can also be added before the classification network, which can be used to compress the amount of data and parameters and reduce over - fitting of the final classification result.
[0256] In some alternative examples, the pooling layer can be a max - pooling layer or an average - pooling layer, etc. Those skilled in the art can set it according to the actual situation, and the embodiments of the present invention do not limit this.
[0257] 311. The server performs classification processing on the evaluation result of the speech to be evaluated based on the feature correlation degree to obtain the evaluation result corresponding to the speech to be evaluated, and sends the evaluation result to the terminal.
[0258] For example, after receiving the evaluation result, the terminal can display the evaluation result in the evaluation result display area shown in 603 in Figure 6 etc. Or, the user can trigger the control of "view score" and then display the page shown in 603 in Figure 6 etc.
[0259] As can be seen from the above, the embodiments of the present invention can extract acoustic and semantic features by simultaneously combining text and acoustic modalities, fuse multi - modal information for oral evaluation, reduce the dependence on automatic speech recognition technology, and improve the accuracy of oral evaluation results.
[0260] In addition, the embodiments of the present invention do not require a large amount of manual evaluation data to train the model, but use a large amount of available ASR training data and natural language understanding data to pre - train the model, which can reduce the dependence on manually labeled speech evaluation data and save human resources.
[0261] To better implement the above method, correspondingly, an embodiment of the present invention further provides a speech evaluation device.
[0262] Reference Figure 4 , the speech evaluation device may include:
[0263] A data acquisition unit 401, which can be used to acquire the speech to be evaluated and the reference text corresponding to the speech to be evaluated;
[0264] A speech feature extraction unit 402, which can be used to extract speech features from the speech to be evaluated to obtain the target speech features corresponding to the speech to be evaluated;
[0265] A semantic feature extraction unit 403, which can be used to extract semantic features from the reference text to obtain the target semantic features corresponding to the reference text;
[0266] A correlation degree calculation unit 404, which can be used to calculate the feature correlation degree according to the target speech features and the target semantic features to obtain the feature correlation degree between the target speech features and the target semantic features;
[0267] An evaluation result generation unit 405, which can be used to classify the evaluation results of the speech to be evaluated based on the feature correlation degree to obtain the evaluation result corresponding to the speech to be evaluated.
[0268] Optionally, the speech feature extraction unit 402 can be used to map the speech to be evaluated into the speech feature vector space according to the speech feature mapping parameters of the speech feature extraction model, and obtain the target speech feature vector based on the mapping result, and use the target speech feature vector as the target speech features corresponding to the speech to be evaluated.
[0269] Optionally, the speech feature extraction unit 402 can be used to divide the speech to be evaluated into sub-speeches to obtain a sub-speech set;
[0270] Extract features from the sub-speeches in the sub-speech set through the speech feature mapping parameters of the speech feature extraction model to obtain the speech feature sub-vectors corresponding to the sub-speeches;
[0271] Determine the initial speech feature vector of the speech to be evaluated in the speech feature vector space according to the speech feature sub-vectors;
[0272] Determine the speech correlation weight corresponding to the sub-speech according to the initial speech feature vector, and the speech correlation weight can be used to indicate the correlation relationship between the sub-speeches in the sub-speech set;
[0273] Based on the speech correlation weight, perform weighted calculation on the initial speech feature vector to obtain the target speech feature vector of the speech to be evaluated in the speech feature vector space.
[0274] Optionally, the semantic feature extraction unit 403 can be used to map a reference text into a semantic feature vector space according to the semantic feature mapping parameters of a semantic feature extraction model, obtain a target semantic feature vector based on the mapping result, and use the target semantic feature vector as the target semantic feature corresponding to the reference text.
[0275] Optionally, the correlation degree calculation unit 404 can be used to perform correlation feature calculation on the target speech feature and the target semantic feature through a feature correlation network to obtain the correlation feature corresponding to the target speech feature and the target semantic feature;
[0276] Based on a classification network, perform correlation analysis on the correlation feature to determine the feature correlation degree corresponding to the target speech feature and the target semantic feature.
[0277] Optionally, as Figure 5 shown, a speech model training unit 406 can also be included before the speech feature extraction unit 402, which can be used to extract speech features from a first sample speech through a speech feature extraction model to be trained to obtain a first speech feature vector corresponding to the first sample speech, where the first sample speech is annotated with a reference speech recognition text;
[0278] Perform text conversion on the first speech feature vector through a speech recognition model to obtain the speech recognition text corresponding to the first sample speech;
[0279] Calculate the loss of the speech feature extraction model to be trained based on the reference speech recognition text and the speech recognition text;
[0280] Adjust the model parameters of the speech feature extraction model to be trained according to the loss to obtain a trained speech feature extraction model.
[0281] Optionally, a semantic model training unit 407 can also be included before the semantic feature extraction unit 403, which can be used to extract semantic features from a first sample text through a semantic feature extraction model to be trained to obtain a first sample semantic feature vector of the first sample text, where the first sample text contains at least one group of first sample text groups, and each first sample text group can include at least two first sample text sentences and the reference semantic relationship between the first sample text sentences;
[0282] Judge the semantic relationship between the first sample text sentences in each first sample text group according to the first sample semantic feature vector;
[0283] Calculate the loss of the semantic feature extraction model to be trained according to the semantic relationship and the reference semantic relationship;
[0284] Adjust the model parameters of the semantic feature extraction model to be trained based on the loss to obtain a trained semantic feature extraction model.
[0285] Optionally, a joint training unit 408 may be further included before the speech feature extraction unit 402, which may be used to obtain sample pairs. A sample pair may include a second sample speech, a second sample text, and sample words that appear in the second sample speech in the second sample text. Among them, the second sample text may include at least one group of second sample text groups, and a second sample text group may include two second sample text sentences and the reference semantic relationship between the second sample text sentences;
[0286] Map the second sample speech to the speech feature vector space through the speech feature mapping parameters of the speech feature extraction model to be trained, and obtain the second sample speech feature vector;
[0287] Map the second sample text to the semantic feature vector space through the semantic feature mapping parameters of the semantic feature extraction model to be trained, and obtain the second sample semantic feature vector;
[0288] Based on the second sample speech feature vector and the second sample semantic feature vector of the same sample pair, determine the training words that appear in the second sample speech in the second sample text;
[0289] Based on the second sample semantic feature vector, determine the semantic relationship between the second sample text sentences in each second sample text group;
[0290] Calculate the losses of the semantic feature extraction model and the speech feature extraction model to be trained according to the training words, sample words, semantic relationship, and reference semantic relationship;
[0291] Based on the losses, adjust the model parameters of the semantic feature extraction model and the speech feature extraction model to be trained to obtain the trained semantic feature extraction model and speech feature extraction model.
[0292] Optionally, a network training unit 409 may be further included before the correlation degree calculation unit 404, which may be used to obtain the third sample speech feature corresponding to the third sample speech and the third sample semantic feature corresponding to the third sample text. Among them, the third sample speech corresponds to the third sample text, and the third sample speech is marked with a reference evaluation result;
[0293] Perform correlation feature calculation on the third sample speech feature and the third sample semantic feature through the feature correlation network to be trained, and obtain the sample correlation feature corresponding to the third sample speech feature and the third sample semantic feature;
[0294] Perform correlation analysis on the sample correlation feature through the classification network to be trained, and determine the feature correlation degree corresponding to the third sample speech feature and the third sample semantic feature;
[0295] Use the feature correlation degree as the sample evaluation result corresponding to the third sample voice, and calculate the losses of the feature correlation network and the classification network based on the sample evaluation result and the reference evaluation result;
[0296] Based on the losses, adjust the parameters of the feature correlation network and the classification network to obtain the trained feature correlation network and classification network.
[0297] Optionally, a reference speech feature extraction unit 410 may also be included before the correlation degree calculation unit 404, which can be used to obtain the reference speech corresponding to the reference text;
[0298] Extract the speech features of the reference speech to obtain the reference speech features corresponding to the reference speech;
[0299] Correspondingly, the correlation degree calculation unit 404 can be used to calculate the feature correlation degree between the target speech features and the reference speech features to obtain the speech feature correlation degree;
[0300] Calculate the feature correlation degree between the target speech features and the target semantic features to obtain the semantic feature correlation degree;
[0301] Based on the speech feature correlation degree and the semantic feature correlation degree, obtain the feature correlation degree between the target speech features and the target semantic features.
[0302] Optionally, a replacement text acquisition unit 411 may also be included before the semantic feature extraction unit 403, which can be used to obtain the replacement text corresponding to each word in the reference text;
[0303] Correspondingly, the semantic feature extraction unit 403 can be used to extract semantic features based on the reference text and the replacement text to obtain the target semantic features corresponding to the reference text.
[0304] Optionally, the speech to be evaluated is the response speech input by the user for the evaluation question, and the reference text is the preset reference answer text for the same evaluation question. The correlation degree calculation unit 404 can be used to calculate the feature correlation degree between the target speech features and the target semantic features, and the feature correlation degree indicates the correlation degree between the response speech and the reference answer text;
[0305] The evaluation result generation unit 405 can be used to map the evaluation score of the response speech based on the feature correlation degree, determine the evaluation score corresponding to the response speech, and use the evaluation score as the evaluation result corresponding to the response speech.
[0306] As can be seen from the above, through the speech evaluation device, by simultaneously combining the text and acoustic modalities, acoustic and semantic features can be extracted, and multi-modal information can be fused for oral evaluation, reducing the dependence on automatic speech recognition technology and improving the accuracy of oral evaluation results.
[0307] In addition, the embodiments of the present invention do not require a large amount of artificial evaluation data to train the model. Instead, a large amount of available ASR training data and natural language understanding data are used to pre-train the model, which can reduce the dependence on artificially labeled speech evaluation data and save human resources.
[0308] In addition, the embodiments of the present invention further provide an electronic device, which can be a terminal, a server, etc. For example, Figure 14 as shown, which shows a schematic structural diagram of the electronic device involved in the embodiments of the present invention. Specifically:
[0309] The electronic device may include a radio frequency (RF) circuit 901, a memory 902 including one or more computer-readable storage media, an input unit 903, a display unit 904, a sensor 905, an audio circuit 906, a wireless fidelity (WiFi) module 907, a processor 908 including one or more processing cores, and a power supply 909 and other components. Those skilled in the art can understand that Figure 14 the structure of the electronic device shown in [[ ]] does not limit the electronic device, and it may include more or fewer components than shown, or combine some components, or have different component arrangements.
[0310] Among them:
[0311] The RF circuit 901 can be used for receiving and transmitting information or signals during a call. In particular, after receiving the downlink information from the base station, it is handed over to one or more processors 908 for processing. Additionally, data related to the uplink is sent to the base station. Generally, the RF circuit 901 includes, but is not limited to, an antenna, at least one amplifier, a tuner, one or more oscillators, a Subscriber Identity Module (SIM) card, a transceiver, a coupler, a Low Noise Amplifier (LNA), a duplexer, etc. In addition, the RF circuit 901 can also communicate with the network and other devices via wireless communication. The wireless communication can use any communication standard or protocol, including but not limited to the Global System of Mobile communication (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, Short Messaging Service (SMS), etc.
[0312] The memory 902 can be used to store software programs and modules. The processor 908 executes various functional applications and data processing by running the software programs and modules stored in the memory 902. The memory 902 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area can store data created according to the use of the electronic device (such as audio data, a phone book, etc.), etc. In addition, the memory 902 can include high-speed random access memory and can also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. Correspondingly, the memory 902 can also include a memory controller to provide access to the memory 902 for the processor 908 and the input unit 903.
[0313] The input unit 903 can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control. Specifically, in a specific embodiment, the input unit 903 may include a touch-sensitive surface and other input devices. The touch-sensitive surface, also known as a touch display screen or a touchpad, can collect touch operations of the user thereon or nearby (such as operations of the user using a finger, a stylus or any suitable object or accessory on or near the touch-sensitive surface), and drive the corresponding connection device according to a preset program. Optionally, the touch-sensitive surface may include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the touch orientation of the user, detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into contact coordinates, and then sends it to the processor 908, and can receive and execute the commands sent by the processor 908. In addition, various types such as resistive, capacitive, infrared, and surface acoustic wave can be used to implement the touch-sensitive surface. In addition to the touch-sensitive surface, the input unit 903 may further include other input devices. Specifically, the other input devices may include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, a joystick, etc.
[0314] The display unit 904 can be used to display information input by the user or information provided to the user, as well as various graphical user interfaces of the terminal. These graphical user interfaces can be composed of graphics, text, icons, videos and any combination thereof. The display unit 904 may include a display panel. Optionally, the display panel may be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), etc. Further, the touch-sensitive surface may cover the display panel. When the touch-sensitive surface detects a touch operation thereon or nearby, it is transmitted to the processor 908 to determine the type of touch event. Subsequently, the processor 908 provides a corresponding visual output on the display panel according to the type of touch event. Although in Figure 14 it, the touch-sensitive surface and the display panel are implemented as two independent components to achieve input and input functions, but in some embodiments, the touch-sensitive surface and the display panel can be integrated to achieve input and output functions.
[0315] The electronic device may further include at least one sensor 905, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor. Among them, the ambient light sensor can adjust the brightness of the display panel according to the brightness of the ambient light, and the proximity sensor can turn off the display panel and / or the backlight when the electronic device is moved to the ear. As a kind of motion sensor, the gravity acceleration sensor can detect the magnitude of acceleration in all directions (generally three axes). When stationary, it can detect the magnitude and direction of gravity, and can be used in applications for identifying the posture of the mobile phone (such as horizontal and vertical screen switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc. As for other sensors that the electronic device can also be configured with, such as gyroscopes, barometers, hygrometers, thermometers, infrared sensors, etc., they will not be elaborated here.
[0316] The audio circuit 906, the speaker, and the microphone can provide an audio interface between the user and the electronic device. The audio circuit 906 can transmit the electrical signal converted from the received audio data to the speaker, and the speaker converts it into a sound signal for output. On the other hand, the microphone converts the collected sound signal into an electrical signal, which is received by the audio circuit 906 and then converted into audio data. After the audio data is output to the processor 908 for processing, it is sent to another terminal, for example, via the RF circuit 901, or the audio data is output to the memory 902 for further processing. The audio circuit 906 may also include an earphone jack to provide communication between the peripheral earphone and the electronic device.
[0317] WiFi belongs to short - range wireless transmission technology. The electronic device can help users send and receive emails, browse the web, and access streaming media through the WiFi module 907. It provides users with wireless broadband Internet access. Although Figure 14 the WiFi module 907 is shown, it can be understood that it does not belong to the essential components of the electronic device and can be omitted completely within the scope of not changing the essence of the invention according to needs.
[0318] The processor 908 is the control center of the electronic device. It connects various parts of the entire mobile phone using various interfaces and circuits. By running or executing the software programs and / or modules stored in the memory 902, and by calling the data stored in the memory 902, it executes various functions of the electronic device and processes data, thereby performing an overall detection of the mobile phone. Optionally, the processor 908 may include one or more processing cores; preferably, the processor 908 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, the user interface, and application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above - mentioned modem processor may not be integrated into the processor 908 either.
[0319] The electronic device further includes a power supply 909 (such as a battery) for powering each component. Preferably, the power supply can be logically connected to the processor 908 through a power management system, so as to manage functions such as charging, discharging, and power consumption management through the power management system. The power supply 909 may further include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator.
[0320] Although not shown, the electronic device may further include a camera, a Bluetooth module, etc., which will not be elaborated here. Specifically, in this embodiment, the processor 908 in the electronic device will load the executable files corresponding to the processes of one or more application programs into the memory 902 according to the following instructions, and the processor 908 will run the application programs stored in the memory 902 to implement various functions as follows:
[0321] Obtain the speech to be evaluated and the reference text corresponding to the speech to be evaluated;
[0322] Extract speech features from the speech to be evaluated to obtain the target speech features corresponding to the speech to be evaluated;
[0323] Extract semantic features from the reference text to obtain the target semantic features corresponding to the reference text;
[0324] Calculate the feature correlation degree according to the target speech features and the target semantic features to obtain the feature correlation degree between the target speech features and the target semantic features;
[0325] Based on the feature correlation degree, perform classification processing on the evaluation results of the speech to be evaluated to obtain the evaluation results corresponding to the speech to be evaluated.
[0326] The system involved in the embodiment of the present invention may be a distributed system formed by connecting a client and multiple nodes (any form of computer device in the access network, such as a server, a terminal) in a network communication manner.
[0327] Taking the distributed system as a blockchain system as an example, see Figure 15 , Figure 15FIG. 0 is an alternative structural schematic diagram of the distributed system 100 provided by an embodiment of the present invention applied to a blockchain system, which is formed by multiple nodes (any form of computing device accessing the network, such as a server or a user terminal) and a client. A peer-to-peer (P2P) network is formed among the nodes. The P2P protocol is an application layer protocol running on top of the Transmission Control Protocol (TCP). In a distributed system, any machine such as a server or a terminal can join and become a node. A node includes a hardware layer, an intermediate layer, an operating system layer, and an application layer. In this embodiment, the voice to be evaluated, the reference text, the training data, etc. can be stored in the shared ledger of the blockchain system through the nodes. A computer device (such as a terminal or a server) can obtain the voice to be evaluated based on the record data stored in the shared ledger.
[0328] See Figure 15 the functions of each node in the blockchain system shown, and the functions involved include:
[0329] 1) Routing, a basic function of a node, which is used to support communication between nodes.
[0330] In addition to the routing function, a node may also have the following functions:
[0331] 2) Application, which is used to be deployed in the blockchain, implement specific services according to actual business requirements, record the data related to the implemented functions to form record data, carry a digital signature in the record data to indicate the source of the task data, and send the record data to other nodes in the blockchain system. When other nodes verify the source and integrity of the record data successfully, the record data is added to the temporary block.
[0332] For example, the services implemented by the application include:
[0333] 2.1) Wallet, which is used to provide the function of conducting electronic currency transactions, including initiating a transaction (that is, sending the transaction record of the current transaction to other nodes in the blockchain system. After other nodes verify successfully, as a response to acknowledging the validity of the transaction, the record data of the transaction is deposited into the temporary block of the blockchain; of course, the wallet also supports querying the remaining electronic currency in the electronic currency address;
[0334] 2.2) Shared ledger, which is used to provide functions such as storage, query, and modification of account data, send the record data of the operations on the account data to other nodes in the blockchain system. After other nodes verify the validity, as a response to acknowledging the validity of the account data, the record data is deposited into the temporary block, and a confirmation can also be sent to the node that initiated the operation.
[0335] 2.3) Smart contract, a computerized protocol that can execute the terms of a contract, implemented by code deployed on a shared ledger for execution when certain conditions are met. According to actual business requirements, the code is used to complete automated transactions, such as querying the logistics status of the goods purchased by the buyer and transferring the buyer's electronic currency to the merchant's address after the buyer signs for the goods. Of course, smart contracts are not limited to executing contracts for transactions, but can also execute contracts for processing received information.
[0336] 3) Blockchain, which includes a series of blocks (Block) that are sequentially connected in the order of generation. Once a new block is added to the blockchain, it will not be removed again. The block records the record data submitted by nodes in the blockchain system.
[0337] See Figure 16 , Figure 16 is an optional schematic diagram of the block structure provided by the embodiments of the present invention. Each block includes the hash value of the transaction records stored in this block (the hash value of this block) and the hash value of the previous block. The blocks are connected through the hash values to form a blockchain. In addition, the block may also include information such as the timestamp when the block is generated. Blockchain, essentially a decentralized database, is a string of data blocks generated by using cryptographic methods. Each data block contains relevant information for verifying the validity of its information (anti-counterfeiting) and generating the next block.
[0338] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions or by controlling relevant hardware through instructions. The instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.
[0339] For this reason, the embodiments of the present invention provide a storage medium that stores multiple instructions that can be loaded by a processor to execute the steps in any one of the voice evaluation methods provided by the embodiments of the present invention. For example, the instructions can execute the following steps:
[0340] Obtain the voice to be evaluated and the reference text corresponding to the voice to be evaluated;
[0341] Extract voice features from the voice to be evaluated to obtain the target voice features corresponding to the voice to be evaluated;
[0342] Extract semantic features from the reference text to obtain the target semantic features corresponding to the reference text;
[0343] Calculate the feature correlation degree according to the target voice features and the target semantic features to obtain the feature correlation degree between the target voice features and the target semantic features;
[0344] Based on the feature correlation degree, the speech to be evaluated is classified to obtain the evaluation result corresponding to the speech to be evaluated.
[0345] For the specific implementation of each of the above operations, reference may be made to the foregoing embodiments, which will not be elaborated herein.
[0346] Among them, the storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disc, etc.
[0347] Since the instructions stored in the storage medium can execute the steps in any of the speech evaluation methods provided by the embodiments of the present invention, the beneficial effects achievable by any of the speech evaluation methods provided by the embodiments of the present invention can be realized. For details, refer to the foregoing embodiments, which will not be elaborated herein.
[0348] According to one aspect of the present application, there is also provided a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the electronic device executes the methods provided in the various alternative implementations in the above embodiments. <s
[0349] The above has introduced in detail a speech evaluation method, device, electronic device and storage medium provided by the embodiments of the present invention. Specific examples are used herein to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A speech evaluation method, characterized in that, Including: Obtain the speech to be evaluated and the reference text corresponding to the speech to be evaluated; Extract speech features from the speech to be evaluated to obtain the target speech features corresponding to the speech to be evaluated; Extract semantic features from the reference text to obtain the target semantic features corresponding to the reference text; Calculate the feature correlation degree according to the target speech features and the target semantic features to obtain the feature correlation degree between the target speech features and the target semantic features, including: performing associated feature calculation on the target speech features and the target semantic features through a feature association network to obtain the associated features corresponding to the target speech features and the target semantic features; based on a classification network, performing association analysis on the associated features to determine the feature correlation degree corresponding to the target speech features and the target semantic features; Based on the feature correlation degree, perform evaluation result classification processing on the speech to be evaluated to obtain the evaluation result corresponding to the speech to be evaluated.
2. The speech evaluation method according to claim 1, characterized in that The extracting speech features from the speech to be evaluated to obtain the target speech features corresponding to the speech to be evaluated includes: According to the speech feature mapping parameters of the speech feature extraction model, map the speech to be evaluated into the speech feature vector space, obtain the target speech feature vector based on the mapping result, and use the target speech feature vector as the target speech features corresponding to the speech to be evaluated.
3. The voice evaluation method according to claim 2, wherein The according to the speech feature mapping parameters of the speech feature extraction model, mapping the speech to be evaluated into the speech feature vector space, and obtaining the target speech feature vector based on the mapping result includes: Divide the speech to be evaluated into sub-speeches to obtain a sub-speech set; Extract features from the sub-speeches in the sub-speech set through the speech feature mapping parameters of the speech feature extraction model to obtain the speech feature sub-vectors corresponding to the sub-speeches; According to the speech feature sub-vectors, determine the speech initial feature vector of the speech to be evaluated in the speech feature vector space; Determine the speech association weight corresponding to the sub-speech according to the speech initial feature vector, and the speech association weight is used to indicate the association relationship between the sub-speeches in the sub-speech set; Based on the speech association weight, perform weighted calculation on the speech initial feature vector to obtain the target speech feature vector of the speech to be evaluated in the speech feature vector space.
4. The speech evaluation method according to claim 2, wherein The extracting semantic features from the reference text to obtain the target semantic features corresponding to the reference text includes: According to the semantic feature mapping parameters of the semantic feature extraction model, map the reference text into the semantic feature vector space, obtain the target semantic feature vector based on the mapping result, and use the target semantic feature vector as the target semantic features corresponding to the reference text.
5. The voice evaluation method according to claim 2, wherein Before the according to the speech feature mapping parameters of the speech feature extraction model, mapping the speech to be evaluated into the speech feature vector space, and obtaining the target speech feature vector based on the mapping result, further includes: Extract the speech features of the first sample speech through the speech feature extraction model to be trained, and obtain the first speech feature vector corresponding to the first sample speech, where the first sample speech is labeled with a reference speech recognition text; Perform text conversion on the first speech feature vector through a speech recognition model to obtain the speech recognition text corresponding to the first sample speech; Calculate the loss of the speech feature extraction model to be trained based on the reference speech recognition text and the speech recognition text; Adjust the model parameters of the speech feature extraction model to be trained according to the loss to obtain the trained speech feature extraction model.
6. The speech evaluation method according to claim 4, wherein Before mapping the reference text to the semantic feature vector space according to the semantic feature mapping parameters of the semantic feature extraction model and obtaining the target semantic feature vector based on the mapping result, it further includes: Extract the semantic features of the first sample text through the semantic feature extraction model to be trained, and obtain the first sample semantic feature vector of the first sample text, where the first sample text contains at least one group of first sample text groups, and each first sample text group includes at least two first sample text sentences and the reference semantic relationship between the first sample text sentences; Judge the semantic relationship between the first sample text sentences in each first sample text group according to the first sample semantic feature vector; Calculate the loss of the semantic feature extraction model to be trained according to the semantic relationship and the reference semantic relationship; Adjust the model parameters of the semantic feature extraction model to be trained based on the loss to obtain the trained semantic feature extraction model.
7. The voice evaluation method according to claim 4, wherein Before mapping the reference text to the semantic feature vector space according to the semantic feature mapping parameters of the semantic feature extraction model and obtaining the target semantic feature vector based on the mapping result, it further includes: Obtain a sample pair, where the sample pair includes a second sample speech, a second sample text, and a sample word that appears in the second sample speech in the second sample text, where the second sample text contains at least one group of second sample text groups, and each second sample text group includes two second sample text sentences and the reference semantic relationship between the second sample text sentences; Map the second sample speech to the speech feature vector space through the speech feature mapping parameters of the speech feature extraction model to be trained to obtain a second sample speech feature vector; Map the second sample text to the semantic feature vector space through the semantic feature mapping parameters of the semantic feature extraction model to be trained to obtain a second sample semantic feature vector; Determine the training words of the second sample text that appear in the second sample speech based on the second sample speech feature vector and the second sample semantic feature vector of the same sample pair; Determine the semantic relationship between the second sample text sentences in each second sample text group based on the second sample semantic feature vector; Calculate the losses of the semantic feature extraction model and the speech feature extraction model to be trained according to the training words, the sample words, the semantic relationship, and the reference semantic relationship; Based on the loss, adjust the model parameters of the semantic feature extraction model and the speech feature extraction model to be trained, and obtain the trained semantic feature extraction model and speech feature extraction model.
8. The speech evaluation method according to claim 1, characterized in that Before calculating the correlation features of the target speech feature and the target semantic feature through the feature correlation network, it further includes: Obtain the third sample speech feature corresponding to the third sample speech and the third sample semantic feature corresponding to the third sample text, where the third sample speech corresponds to the third sample text, and the third sample speech is labeled with a reference evaluation result; Calculate the correlation features of the third sample speech feature and the third sample semantic feature through the feature correlation network to be trained, and obtain the sample correlation features corresponding to the third sample speech feature and the third sample semantic feature; Through the classification network to be trained, perform correlation analysis on the sample correlation features to determine the feature correlation degree corresponding to the third sample speech feature and the third sample semantic feature; Based on the feature correlation degree, perform evaluation result classification processing on the third sample speech to obtain the sample evaluation result corresponding to the third sample speech, and calculate the losses of the feature correlation network and the classification network based on the sample evaluation result and the reference evaluation result; Based on the loss, adjust the parameters of the feature correlation network and the classification network to obtain the trained feature correlation network and classification network.
9. The speech evaluation method according to claim 1, wherein Before calculating the feature correlation degree according to the target speech feature and the target semantic feature, it further includes: Obtain the reference speech corresponding to the reference text; Extract the speech feature of the reference speech to obtain the reference speech feature corresponding to the reference speech; The calculating the feature correlation degree between the target speech feature and the target semantic feature according to the target speech feature and the target semantic feature includes: Calculate the feature correlation degree between the target speech feature and the reference speech feature to obtain the speech feature correlation degree; Calculate the feature correlation degree between the target speech feature and the target semantic feature to obtain the semantic feature correlation degree; Based on the speech feature correlation degree and the semantic feature correlation degree, obtain the feature correlation degree between the target speech feature and the target semantic feature.
10. The speech evaluation method according to claim 1, wherein Before extracting the semantic feature of the reference text to obtain the target semantic feature corresponding to the reference text, it further includes: Obtain the replacement text corresponding to each word in the reference text; The extracting the semantic feature of the reference text to obtain the target semantic feature corresponding to the reference text includes: Based on the reference text and the replacement text, extract the semantic feature to obtain the target semantic feature corresponding to the reference text.
11. The voice evaluation method according to any one of claims 1-4, characterized in that, The speech to be evaluated is the response speech input by the user for the evaluation question, and the reference text is the preset reference answer text for the same evaluation question; The calculating the feature correlation degree between the target speech feature and the target semantic feature according to the target speech feature and the target semantic feature includes: Calculate the feature correlation degree between the features according to the target voice feature and the target semantic feature, where the feature correlation degree indicates the correlation degree between the response voice and the reference answer text; Based on the feature correlation degree, perform evaluation result classification processing on the speech to be evaluated to obtain the evaluation result corresponding to the speech to be evaluated, including: Based on the feature correlation degree, perform evaluation score mapping on the response voice, determine the evaluation score corresponding to the response voice, and use the evaluation score as the evaluation result corresponding to the response voice.
12. A voice evaluation device, characterized in that, Including: A data acquisition unit for acquiring the speech to be evaluated and the reference text corresponding to the speech to be evaluated; A voice feature extraction unit for extracting voice features from the speech to be evaluated to obtain the target voice feature corresponding to the speech to be evaluated; A semantic feature extraction unit for extracting semantic features from the reference text to obtain the target semantic feature corresponding to the reference text; A correlation degree calculation unit for calculating the feature correlation degree between the target voice feature and the target semantic feature according to the target voice feature and the target semantic feature, including: performing correlation feature calculation on the target voice feature and the target semantic feature through a feature correlation network to obtain the correlation features corresponding to the target voice feature and the target semantic feature; based on a classification network, performing correlation analysis on the correlation features to determine the feature correlation degree corresponding to the target voice feature and the target semantic feature; An evaluation result generation unit for performing evaluation result classification processing on the speech to be evaluated based on the feature correlation degree to obtain the evaluation result corresponding to the speech to be evaluated.
13. An electronic device, characterized in that, Including a memory and a processor; the memory stores an application program, and the processor is configured to run the application program in the memory to execute the steps in the voice evaluation method according to any one of claims 1 to 11.
14. A storage medium, characterized in that, The storage medium stores multiple instructions, and the instructions are suitable for being loaded by a processor to execute the steps in the voice evaluation method according to any one of claims 1 to 11.
15. A computer program product, characterized in that, The computer program product includes computer instructions, and the computer instructions are stored in a computer-readable storage medium; a processor of an electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions to enable the electronic device to execute the steps in the voice evaluation method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Pronunciation evaluation method and device, electronic equipment and storage medium
CN111199750A