Speech evaluation method, device, equipment and storage medium
Through the proposed voice evaluation method, the audio signal is evaluated using the acoustic characteristics of the audio signal and the dialect phoneme dictionary, which solves the problem that the existing technology is difficult to deal with dialect voice, and realizes accurate evaluation of multiple voices, improving user experience.
Patent Information
- Application Number
- CN202210325744.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-29
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2042-03-29
AI Technical Summary
Existing voice evaluation technology is difficult to meet the various needs of users, especially when dealing with dialect voice, which makes it difficult to achieve accurate evaluation, resulting in poor user experience.
A speech evaluation method is proposed to evaluate the audio signal by obtaining the acoustic characteristics of the audio signal and combining dialect phoneme dictionary. The method includes extracting the acoustic features of the speech frame, determining the phoneme probability, performing phoneme alignment, and calculating the evaluation results based on the dialect scaling factor.
This method can effectively expand the scope of use of voice evaluation technology, meet various needs of users, improve user experience, and realize accurate evaluation of dialect voice.
Smart Images

Figure CN114627896B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of speech processing, and more particularly to a speech evaluation method, device, equipment and storage medium. Background Art
[0002] In recent years, with the continuous progress of technology, speech technology is being applied to various industries. For example, in the past one or two years, online education has been booming, and speech evaluation technology has been used by many online education platforms, etc., to score the pronunciation of users and judge whether the pronunciation is standard.
[0003] Most of the current speech evaluation technologies evaluate Mandarin Chinese speech, standard British and American English speech, etc. This greatly limits the scope of use of speech evaluation technology, is difficult to meet the various needs of users, and brings a poor user experience.
[0004] Therefore, there is an urgent need for a new speech evaluation technology to solve the above problems. Summary of the Invention
[0005] In view of the above problems, the present invention is proposed. According to one aspect of the present invention, there is provided a speech evaluation method, including: obtaining an audio signal to be evaluated; extracting acoustic features of each speech frame in the audio signal; using the acoustic features to determine the probability that the pronunciation of the speech frame of the audio signal is each phoneme in a phoneme dictionary, so as to obtain pronunciation information of the speech frame, wherein the phoneme dictionary includes dialect phonemes; obtaining standard text information corresponding to the audio signal to be evaluated; based on a dialect dictionary, determining the phonemes corresponding to the standard text information, wherein the phonemes of the characters in the dialect dictionary include standard phonemes and dialect phonemes; based on the acoustic features and pronunciation information of each speech frame in the audio signal, aligning the speech frames of the audio signal with the phonemes corresponding to the standard text information, so as to obtain alignment information of the speech frames; and determining an evaluation result of the audio signal relative to the standard text information according to the pronunciation information and alignment information of the speech frames.
[0006] Exemplarily, the method further includes: classifying the audio signal into dialects to obtain a dialect scaling factor of the audio signal, determining the dialect scaling factor as a first value for the case where the audio signal is dialect speech, and determining the dialect scaling factor as a second value for the case where the audio signal is non-dialect speech, wherein the first value is greater than the second value; wherein, determining the evaluation result of the audio signal relative to the text information also depends on the dialect scaling factor.
[0007] Exemplarily, classifying the audio signal into dialects to obtain a dialect scaling factor of the audio signal includes: using a binary classification model to determine the probabilities that the audio signal is dialect speech or non-dialect speech respectively; determining the dialect scaling factor based on the probability that the audio signal is dialect speech and / or the probability that the audio signal is non-dialect speech.
[0008] Exemplarily, determining a dialect scaling factor based on the probability that the audio signal is dialect speech and / or the probability that the audio signal is non-dialect speech includes: dividing the probability that the audio signal is dialect speech by the probability that the audio signal is non-dialect speech to obtain a quotient value; and determining the quotient value as the dialect scaling factor.
[0009] Exemplarily, determining an evaluation result of the audio signal relative to the text information includes: determining the accuracy A, fluency B, and completeness C of the audio sentence in the audio signal according to the pronunciation information and alignment information of the speech frames; calculating the score S of the audio sentence in the audio signal using the following formula, S = δ * (a * A + b * B + c * C), where δ represents the dialect scaling factor, and a, b, and c respectively represent the weights of the accuracy A, fluency B, and completeness C; and determining the evaluation result based on the score S of the audio sentence in the audio signal.
[0010] Exemplarily, a dialect identifier is set for the dialect phonemes, and the method further includes: extracting the acoustic features of each speech frame in the training audio signal; training an acoustic model using the acoustic features of the speech frames of the training audio signal, wherein the value of the loss function of the acoustic model is determined based on the calculation result of the acoustic model and the dialect identifier; and using the acoustic features to determine the probability that the pronunciation of the speech frame of the audio signal is each phoneme in the phoneme dictionary to obtain the pronunciation information of the speech frame, including: inputting the acoustic features into the acoustic model so that the acoustic model outputs the probability that the pronunciation of the speech frame of the audio signal is each phoneme in the phoneme dictionary.
[0011] Exemplarily, aligning the speech frames of the audio signal with the phonemes corresponding to the text information based on the acoustic features and pronunciation information of each speech frame in the audio signal to obtain the alignment information of the speech frames includes: generating a search space corresponding to the standard text information based on the phonemes corresponding to the standard text information, where the search space includes a dialect phoneme path formed by the dialect phonemes; and determining the phoneme corresponding to each speech frame based on the acoustic features and pronunciation information of each speech frame of the audio signal and the search space.
[0012] According to another aspect of the present invention, there is also provided a speech evaluation device, including:
[0013] A data acquisition module, configured to acquire the audio signal to be evaluated and the standard text information corresponding to the audio signal;
[0014] A feature extraction module, configured to extract the acoustic features of the audio signal;
[0015] A calculation module, configured to use the acoustic features to determine the probability that the pronunciation of the speech frame of the audio signal is each phoneme in the phoneme dictionary to obtain the pronunciation information of the speech frame, where the phoneme dictionary includes dialect phonemes;
[0016] A phoneme determination module, configured to determine phonemes corresponding to standard text information based on a dialect dictionary, where the phonemes of the characters in the dialect dictionary include standard phonemes and dialect phonemes;
[0017] An alignment module, configured to align the speech frames of the audio signal with the phonemes corresponding to the standard text information based on the acoustic features and pronunciation information of each speech frame in the audio signal, so as to obtain the alignment information of the speech frames;
[0018] An evaluation result determination module, configured to determine an evaluation result of the audio signal relative to the standard text information according to the pronunciation information and alignment information of the speech frames.
[0019] According to another aspect of the present invention, there is also provided a speech evaluation device, including a sound collection device, an input device, a processor, and a memory, where the sound collection device is configured to obtain an audio signal to be evaluated and send it to the processor; the input device is configured to input standard text information corresponding to the audio signal to be evaluated and send it to the processor; computer program instructions are stored in the memory, and when the computer program instructions are run by the processor, they are used to execute the speech evaluation method as described above.
[0020] According to still another aspect of the present invention, there is also provided a storage medium, on which program instructions are stored, and when the program instructions are run, they are used to execute the speech evaluation method as described above.
[0021] In the above technical solution, a speech evaluation can be performed on an audio signal with a dialect pronunciation. Thus, the usage range of the speech evaluation technology can be effectively expanded, and further, the various needs of users can be met, and the user experience can be improved.
[0022] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features, and advantages of the present invention more obvious and understandable, the following specifically describes the embodiments of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] By describing the embodiments of the present invention in more detail in conjunction with the drawings, the above and other purposes, features, and advantages of the present invention will become more obvious. The drawings are used to provide a further understanding of the embodiments of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation to the present invention. In the drawings, the same reference numerals generally represent the same components or steps.
[0024] Figure 1 FIG. shows a schematic flowchart of a speech evaluation method according to an embodiment of the present invention;
[0025] Figure 2Shows a schematic flowchart of training an acoustic model for forward calculation according to an embodiment of the present invention;
[0026] Figure 3 Shows a schematic flowchart of obtaining alignment information of speech frames by aligning speech frames of an audio signal with phonemes corresponding to standard text information according to an embodiment of the present invention;
[0027] Figure 4a Shows a schematic diagram of a search space according to an example of the prior art;
[0028] Figure 4b Shows a schematic diagram of a search space according to an embodiment of the present invention;
[0029] Figure 5 Shows a schematic flowchart of dialect classification of an audio signal to obtain a dialect scaling factor of the audio signal according to an embodiment of the present invention;
[0030] Figure 6 Shows a schematic flowchart of determining a dialect scaling factor based on the probability that an audio signal is dialect speech and / or the probability of the dialect speech of the audio signal according to an embodiment of the present invention;
[0031] Figure 7 Shows a schematic flowchart of a speech evaluation method according to another embodiment of the present invention;
[0032] Figure 8 Shows a schematic block diagram of a speech evaluation device according to an embodiment of the present invention; and
[0033] Figure 9 Shows a schematic block diagram of a speech evaluation device according to an embodiment of the present invention. Detailed implementation manners
[0034] In order to make the objectives, technical solutions and advantages of the present invention more obvious, the exemplary embodiments according to the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments of the present invention. It should be understood that the present invention is not limited by the exemplary embodiments described herein. Based on the embodiments of the present invention described in the present invention, all other embodiments obtained by those skilled in the art without creative efforts shall fall within the protection scope of the present invention.
[0035] Due to regional cultural differences, there will be certain differences in the pronunciation, vocabulary, or grammar of languages in different regions. This difference is particularly prominent in pronunciation, thus forming local dialects. For example, Sichuan dialect and Tianjin dialect in Chinese are only spoken in a specific region. As mentioned above, in the existing speech evaluation technology, only the standard Mandarin pronunciation of Chinese or the standard British and American pronunciations of English are evaluated for users. However, in actual application scenarios, there are often requirements for dialect evaluation. For example, for data support companies, they expect to conduct speech evaluation on dialect voices to perform quality inspection on the corpus. However, for the dialect evaluation requirements of users, the existing technology cannot be achieved. Therefore, this application proposes a new speech evaluation method to solve the above technical problems.
[0036] It can be understood that in the following text, Chinese or English is used as an example of speech for description. However, the embodiments of the present invention are not limited to being applied to these two languages, and it can be applied to any language with dialects, such as German.
[0037] Figure 1 Fig. 7 shows a schematic flowchart of a speech evaluation method 100 according to an embodiment of the present invention. As Figure 1 shown, the speech evaluation method 100 may include the following steps.
[0038] Step S110, obtain the audio signal to be evaluated.
[0039] Exemplarily, the person to be tested can issue the speech to be evaluated to the electronic device. The speech to be evaluated can be any speech suitable for speech evaluation. The speech to be evaluated can be received by using the sound collection device (such as a microphone) of the electronic device and converted through an analog / digital conversion circuit to convert the analog signal, that is, the speech to be evaluated, into a digital signal that can be recognized and processed by the electronic device, that is, the audio signal. Thus, the audio signal corresponds to the speech to be evaluated issued by the subject. Alternatively, the audio signal to be evaluated pre-stored therein can also be obtained from other devices or storage media through data transmission technology.
[0040] Step S120, extract the acoustic features of each speech frame in the audio signal obtained in step S110.
[0041] Preferably, the audio signal can be preprocessed. The preprocessing may include processing such as filtering and framing. For example, first, a filter and sampling are performed on an audio signal, thereby reducing the interference of signals with other frequencies except for human vocalization and / or the 50 Hz current frequency signal. In addition, the audio signal can also be framed. Framing refers to the operation of dividing the audio signal into multiple small segments, thereby obtaining multiple speech frames, where each small segment after division is called a speech frame. Each frame of the audio signal after framing has the characteristic of short-term stationarity. Optionally, the frame length of each speech frame can be set to a reasonable value between 20 milliseconds and 40 milliseconds, such as 25 milliseconds.
[0042] Feature extraction operations can be performed on the preprocessed audio signal. For each speech frame, the acoustic features of the speech frame are extracted. The acoustic features can be represented by a multi-dimensional vector, which includes the content information corresponding to the speech frame. The acoustic features may include Mel Frequency Cepstral Coefficient (MFCC) features, Mel Scale Filter Banks acoustic features, and Perceptual Linear Prediction Coefficient (PLP) features, etc. In this application, no limitation is imposed on the method for extracting acoustic features, and any existing or future technology that can achieve acoustic feature extraction is within the protection scope of this application. Exemplarily and non-limitingly, the code for extracting the corresponding acoustic features can be directly called from the KALDI open-source code.
[0043] Step S130, using the acoustic features to determine the probabilities that the speech frames of the audio signal are pronounced as each phoneme in the phoneme dictionary, so as to obtain the pronunciation information of the speech frames. The phoneme dictionary includes dialect phonemes.
[0044] Exemplarily, the probability P(q|o), that is, the probabilities that the speech frames of the audio signal are pronounced as each phoneme in the phoneme dictionary, can be obtained through the forward calculation of a neural network. Among them, o represents the phoneme corresponding to the acoustic features extracted from a speech frame in the audio signal, and q represents the phoneme in the phoneme dictionary. In Chinese speech evaluation, the phoneme dictionary can be obtained by modeling the initials, finals, and tones. In English speech evaluation, the phoneme dictionary can be obtained by modeling English phonemes. According to the embodiments of the present invention, in addition to the standard phonemes corresponding to, for example, standard Chinese pronunciation, the phoneme dictionary also includes dialect phonemes. Therefore, the embodiments of this application add dialect phonemes to the existing phoneme dictionary to realize the evaluation of the speech to be evaluated with a dialect pronunciation. It can be understood that compared with the standard phonemes, the dialect phonemes may be different in pronunciation.
[0045] In a specific embodiment, if an audio signal is divided into 400 speech frames after frame segmentation, and the phoneme dictionary includes 38 phonemes, then a 400 * 38 matrix can be obtained through forward calculation of the acoustic features of the speech frames. Among them, the number of rows of the matrix represents the number of speech frames in the audio signal, and the number of columns of the matrix represents the number of phonemes in the phoneme dictionary. Each element in the matrix represents the probability that the corresponding speech frame in the audio signal is pronounced as the corresponding phoneme. Thus, this matrix is used to represent the pronunciation information of the speech frames of the audio signal.
[0046] Step S140, obtain the standard text information corresponding to the audio signal to be evaluated.
[0047] Exemplarily, the standard text information is the correct text information corresponding to the audio signal to be evaluated. Optionally, the user can input the standard text information using the input device (such as a keyboard) of the electronic device or obtain the pre-stored standard text information from other electronic devices or storage devices. It can be understood that the execution order between step S140 and the foregoing steps S110, S120, and S130 can be arbitrary. For example, step S140 can be executed before step S110, after step S110, or during the execution of steps S110 to S130.
[0048] Step S150, based on the dialect dictionary, determine the phonemes corresponding to the standard text information. The phonemes of the characters in the dialect dictionary include standard phonemes and dialect phonemes.
[0049] It can be understood that the standard text information obtained in step S140 has corresponding correct phonemes. Based on the dialect dictionary, the phonemes corresponding to the standard text information can be determined. Among them, the dialect dictionary includes characters and the phonemes corresponding to the characters. The phonemes of the characters include standard phonemes and dialect phonemes. For the convenience of understanding and description, in this application, the dialect phonemes in the dialect dictionary include Tianjin dialect phonemes as an example for the following description. In one embodiment, the obtained standard text information is "camphor tree". Correspondingly, based on the standard phonemes of the characters in the Tianjin dialect dictionary, the phonemes corresponding to this standard text can be determined: zh ang_1sh u_4 and z ang_1 s u_4.
[0050] Step S160, based on the acoustic features and pronunciation information of each speech frame in the audio signal, align the speech frames of the audio signal with the phonemes corresponding to the standard text information to obtain the alignment information of the speech frames.
[0051] Exemplarily, forced alignment is used to achieve the alignment of speech frames of an audio signal and phonemes corresponding to text information. Forced alignment can be implemented using models such as Gaussian Mixture Model (GMM)-Hidden Markov Model (HMM), Long Short Term Memory (LSTM)-Connectionist Temporal Classification (CTC), Convolutional Neural Networks (CNN), and Recurrent Neural Network (RNN). For forced alignment, its known inputs include the acoustic features and pronunciation information of speech frames and the phonemes corresponding to standard text information. After forced alignment, an output vector can be obtained. The length of this output vector is the number of frames of the speech frames of the audio signal. This output vector indicates what the correct phoneme corresponding to each speech frame is, that is, the alignment information of the speech frames. Exemplarily, a dialect dictionary may include 32 standard phonemes and 20 dialect phonemes. For each phoneme in the dialect dictionary, including standard phonemes and dialect phonemes, an index number is assigned to it. Each element in the obtained output vector described above can represent the index number in the dialect dictionary of the phoneme in the standard text information that is aligned with the frame. For example, if the value of the 51st element in the output vector is 2, it means that the phoneme aligned with the 51st frame in the audio signal is the phoneme with the index number 2 in the dialect dictionary.
[0052] Exemplarily, forced alignment can be performed by performing an optimal path search on an HCLG graph. The grammar weighted finite state machine transducer (G graph) in the HCLG graph is generated based on the standard text information. Thus, regardless of what the audio signal is like, it will ultimately correspond to the path corresponding to the standard text information. Thereby, the phoneme that each speech frame in the audio signal should correspond to is determined, that is, forced alignment is achieved.
[0053] In step S170, according to the pronunciation information and alignment information of the speech frames, the evaluation result of the audio signal relative to the standard text information is determined. The more ideal the evaluation result is, the higher the matching degree between the audio signal and the standard text information.
[0054] As described above, the pronunciation information of the speech frames of the audio signal represents the probabilities of the speech frames being pronounced as various possible phonemes, and the alignment information of the speech frames represents which phoneme among the phonemes corresponding to the standard text information the speech frames are aligned with. Based on the above two aspects, the evaluation result of the audio signal can be determined. The greater the probability that the speech frame is pronounced as the phoneme it is aligned with, the more ideal the evaluation result; otherwise, vice versa. For example, the evaluation result can be represented as a specific score. The magnitude of the score represents the matching degree between the audio signal to be evaluated and the standard text information. The greater the score, the higher the matching degree between the two.
[0055] In the above technical solution, speech evaluation can be performed on the audio signal with dialect pronunciation. Thus, the usage range of the speech evaluation technology can be effectively expanded, and furthermore, the various needs of users can be met, and the user experience can be improved.
[0056] Exemplarily, a dialect identifier can be set for the dialect phonemes. For example, the "#dlt" in the phoneme "z ang_1 s u_4#dlt" is the dialect identifier, and based on this dialect identifier, it can be identified whether the phoneme is a dialect phoneme. Figure 2 Fig. shows a schematic flowchart of training an acoustic model for implementing forward calculation according to an embodiment of the present invention. For the foregoing step S130, using the acoustic features to determine the probabilities of the speech frames of the audio signal being pronounced as each phoneme in the phoneme dictionary to obtain the pronunciation information of the speech frames can be implemented using the trained acoustic model. As Figure 2 shown, training the acoustic model for implementing step S130 may include the following steps.
[0057] Step S210, extracting the acoustic features of each speech frame in the training audio signal.
[0058] Exemplarily, the training audio signal can be any audio signal suitable for acoustic feature extraction training. Similar to step S120, the training audio signal can be preprocessed before extracting the acoustic features of each speech frame in the training audio signal. The preprocessing operations have been described in step S120 above, and will not be elaborated here for the sake of brevity. Feature extraction operations can be performed on the preprocessed training audio signal. Similarly, the feature extraction operations have also been described in step S120, and will not be elaborated here for the sake of brevity.
[0059] Step S220, training the acoustic model using the acoustic features of the speech frames of the training audio signal. Among them, the value of the loss function of the acoustic model is determined based on the calculation result of the acoustic model and the dialect identifier.
[0060] Exemplarily, an acoustic model is used to determine the probabilities that the speech frames of a training audio signal are pronounced as various phonemes in a phoneme dictionary. Based on the determined probabilities and a dialect identifier, the value of the loss function of the acoustic model can be determined. For the case where the training audio signal is a dialect audio signal, if the dialect identifier appears, the value of the loss function can be reduced; otherwise, the value of the loss function is increased.
[0061] The acoustic model trained as described above is used for forward calculation. In other words, step S130 of determining the probabilities that the speech frames of an audio signal are pronounced as various phonemes in a phoneme dictionary using acoustic features may include: inputting the acoustic features into the acoustic model so that the acoustic model outputs the probabilities that the speech frames of the audio signal are pronounced as various phonemes in the phoneme dictionary.
[0062] It can be understood that, taking Chinese as an example, some Chinese characters are polyphonic. According to the above technical solution, dialect phonemes are marked using a dialect identifier. Thereby, the dialect phonemes and the phonemes of polyphonic characters are effectively distinguished, ensuring the training effect of the acoustic model, and further, more accurate dialect evaluation results can be obtained based on the trained acoustic model in the subsequent speech evaluation process.
[0063] Figure 3 FIG. shows a schematic flowchart of obtaining alignment information of speech frames by aligning the speech frames of an audio signal with the phonemes corresponding to standard text information according to an embodiment of the present invention. As Figure 3 shown, step S160 may include the following steps.
[0064] Step S161, generating a search space corresponding to the standard text information based on the phonemes corresponding to the standard text information, where the search space includes a dialect phoneme path formed by dialect phonemes.
[0065] Figure 4a FIG. shows a schematic diagram of a search space according to an example of the prior art. As Figure 4a shown, the inputs of this search space are all standard phonemes, and the output is a word. Moreover, each phoneme has only one output path. Figure 4a The text information corresponding to the shown search space is "This is a camphor tree", and its corresponding phonemes are: zh e_4 sh i_4 zh ang_1 sh u_4.
[0066] As described above, the dialect dictionary includes the standard phonemes and dialect phonemes of characters. For example, the Tianjin dialect phonemes are included in the dialect dictionary. Correspondingly, the Tianjin dialect phonemes of "This is a camphor tree" are: zh e_4 sh i_4 z ang_1 s u_4. Based on this, the phonemes of the text information and the corresponding search space can be determined according to the dialect dictionary. Figure 4b FIG. shows a schematic diagram of a search space according to an embodiment of the present invention. As Figure 4bAs shown, the input of the search space includes standard phonemes and dialect phonemes, and its output is also words. However, different from the prior art, the search space of this embodiment of the present application includes two paths: a standard phoneme path and a dialect phoneme path.
[0067] Step S162: Based on the acoustic features and pronunciation information of each speech frame of the audio signal and the search space, determine the phoneme corresponding to each speech frame.
[0068] For example, if the obtained audio signal is Tianjin dialect speech, according to the acoustic features and pronunciation information of each speech frame obtained above, the optimal path, that is, the dialect phoneme path, can be selected in the search space generated in step S161 above. Optionally, referring to Figure 4b , in the dialect phoneme path, for the dialect phonemes therein, a dialect identifier "#dlt" is marked behind them. On the contrary, if the obtained audio signal is Mandarin speech, the standard phoneme path will be selected. Based on the selected path, the phoneme corresponding to each speech frame can be determined.
[0069] Thus, through the search space including the dialect phoneme path, the correct phoneme corresponding to each speech frame can be determined, ensuring the accuracy of the dialect speech evaluation result.
[0070] Exemplarily, method 100 may further include step S180: Classify the audio signal into dialects to obtain a dialect scaling factor of the audio signal. Determine the dialect scaling factor as a first value for the case where the audio signal is dialect speech. Determine the dialect scaling factor as a second value for the case where the audio signal is non-dialect speech. Wherein the first value is greater than the second value. It can be understood that in this embodiment, step S180 is executed prior to step S170. And, in this embodiment, determining the evaluation result of the audio signal relative to the text information is based not only on the pronunciation information and alignment information of the speech frames, but also on the dialect scaling factor.
[0071] Exemplarily, a dialect classification module can be used to calculate a dialect scaling factor. The input of the dialect classification module is an audio signal, and the output is a dialect scaling factor. Still taking the evaluation result as a specific evaluation score as an example for illustration. The dialect scaling factor can be used as a scaling coefficient to act on the evaluation score to achieve further calculation of the evaluation score. Specifically, when the audio signal is a dialect voice, the dialect scaling factor output by the dialect classification calculation module is a first value. For example, this first value can be any reasonable value greater than 1. Then, the dialect scaling factor with this first value is applied to the evaluation score. Since it is greater than 1, the evaluation score can be amplified. In other words, when the audio signal is a dialect voice, the corresponding dialect scaling factor can be used to amplify the evaluation score to improve its evaluation score. On the contrary, when the audio signal is a non-dialect voice, the dialect scaling factor output by the dialect classification calculation module is a second value. The second value is less than the first value, and it can be any reasonable value less than or equal to 1. When the dialect scaling factor is the second value, the evaluation score can be reduced. That is, when the audio signal is a non-dialect voice, the corresponding dialect scaling factor can be used to reduce the evaluation score or keep it unchanged.
[0072] Thus, the use of the dialect scaling factor ensures that dialect voices and non-dialect voices obtain more reasonable evaluation scores during the speech evaluation process, effectively improving the accuracy of dialect speech evaluation.
[0073] Figure 5 FIG. shows a schematic flowchart of performing dialect classification on an audio signal according to step S180 of an embodiment of the present invention to obtain a dialect scaling factor of the audio signal. As Figure 5 shown, step S180 may include the following steps.
[0074] Step S181, using a binary classification model, determine the probabilities that the audio signal is a dialect voice or a non-dialect voice respectively.
[0075] Exemplarily, the binary classification model can be a trained LSTM-CNN neural network model. According to the above, the obtained audio signal to be evaluated can be input into the binary classification model. The binary classification model can calculate the corresponding probabilities that the audio signal is a dialect voice or a non-dialect voice. For example, P1 represents the probability that the audio signal is a dialect voice, and P2 represents the probability that the audio signal is a non-dialect voice.
[0076] Step S182, based on the probability that the audio signal is a dialect voice and / or the probability that the audio signal is a non-dialect voice, determine the dialect scaling factor.
[0077] The probability P1 that the audio signal is a dialect voice and / or the probability P2 that the audio signal is a non-dialect voice are obtained according to the above step S181. Exemplarily, the value of the dialect scaling factor can be determined according to the numerical values of P1 and / or P2. For example, a look-up table of the numerical relationship between the probability P1 that the audio signal is a dialect voice and the value of the dialect scaling factor is preset. When P1 is any value between 70% and 80%, correspondingly, the dialect scaling factor of 1.5 can be found in the look-up table. When P1 is any value between 60% and 70%, correspondingly, the dialect scaling factor of 1.2 can be found in the look-up table, and so on. It can be understood that the above correspondence is only exemplary and does not mean a limitation on the dialect scaling factor.
[0078] It can be understood that using the binary classification model to determine the dialect scaling factor greatly reduces the workload of the system, improves the calculation efficiency, and further improves the overall speed of the speech evaluation.
[0079] Figure 6 FIG. shows a schematic flowchart of step S182 according to an embodiment of the present invention for determining the dialect scaling factor based on the probability that the audio signal is a dialect voice and / or the probability that the audio signal is a dialect voice. As Figure 6 shown, step S182 may include the following steps.
[0080] Step S182a, divide the probability that the audio signal is a dialect voice by the probability that the audio signal is a non-dialect voice to obtain a quotient value.
[0081] According to the above step S181, the probability P1 that the audio signal is a dialect voice and the probability P2 that the audio signal is a non-dialect voice can be obtained. Dividing P1 by P2 can obtain a quotient value.
[0082] Step S182b, determine that the quotient value is the dialect scaling factor.
[0083] Exemplarily, the above quotient value can be directly determined as the dialect scaling factor. Wherein, if the audio signal is a dialect voice, that is, P1 > P2, the value of the dialect scaling factor is greater than 1. Otherwise, the value of the dialect scaling factor is less than or equal to 1.
[0084] The algorithm of the above technical solution is simple and easy to implement. And the calculation amount is small, reducing the calculation burden of the system. Most importantly, the technical solution comprehensively considers the probability that the audio signal is a dialect voice and the probability that the audio signal is a non-dialect voice to obtain the dialect scaling factor. Using this dialect scaling factor, the dialect speech evaluation can be carried out more accurately.
[0085] Exemplarily, step S170 for determining the evaluation result of the audio signal relative to the text information may include the following steps.
[0086] First, according to the pronunciation information and alignment information of the speech frames, determine the accuracy A, fluency B, and integrity C of the audio sentence in the audio signal.
[0087] The accuracy of the phonemes, words, and sentences corresponding to the standard text information can be determined in turn according to the probability of the corresponding phonemes of each speech frame in the pronunciation information and the correct phonemes aligned by the speech frames in the alignment information. The accuracy indicates whether the phoneme, word, or sentence in the audio signal is pronounced accurately. Specifically, the index number of the correct phoneme aligned by the speech frame can be determined according to the alignment information. For example, the audio signal is divided into 20 frames in total. The phonemes of the standard text information corresponding to this audio signal are: zh e_4 sh i_4 z ang_1 s u_4#dlt. According to the index numbers in the alignment information, the first to fifth frames are all aligned with the phoneme "zh". Correspondingly, the probability that the pronunciation of this speech frame is the phoneme "zh" can be found in the pronunciation information of these speech frames. Multiplying this probability by 100 can obtain a percentage score. Based on this score, the accuracy of the phoneme "zh" can be determined. The average value of the scores of the first to fifth frames can be taken as the accuracy of the phoneme "zh". It can be understood that the accuracies of multiple phonemes zh e_4 sh i_4 z ang_1 s u_4#dlt can be obtained according to the above method. Averaging the accuracy scores of multiple phonemes corresponding to a word can obtain the accuracy score of this word, such as for "zhangshu". Further, averaging the accuracy scores of multiple words corresponding to a sentence can obtain the accuracy score of this sentence. Among them, whether it is for phonemes, words, or sentences, a corresponding accuracy threshold is set for each. Taking the accuracy threshold as 80 points as an example, when the accuracy score exceeds 80, it can be considered that the corresponding phoneme, word, or sentence is pronounced correctly.
[0088] The fluency can be expressed as the number of words pronounced correctly per second, which reflects the fluency of the audio signal and is related to the speech rate and the number of pauses. Optionally, a fluency threshold is set for the fluency, for example, 3 words / second. When the fluency does not reach the fluency threshold, the fluency score can be appropriately reduced. In this application, the degree of reduction is not limited.
[0089] The integrity can be expressed as the number of words in a sentence whose scores exceed the threshold, which reflects the integrity of the audio signal reading the standard text information. This threshold includes the accuracy threshold and the fluency threshold. Preferably, when the scores of the words in a sentence exceed both the accuracy threshold and the fluency threshold at the same time, they are regarded as qualified words. Then, judge how many qualified words there are in a sentence, and determine the integrity score according to the ratio of the number of qualified words to the total number of words in a sentence. In this application, the relationship between the ratio and the integrity score is not limited.
[0090] Then, calculate the score S of the audio sentence in the audio signal using the following formula: S = δ * (a * A + b * B + c * C). Here, δ represents the dialect scaling factor, and a, b, and c represent the weights of accuracy A, fluency B, and completeness C respectively. Optionally, a, b, and c can be parameters obtained through machine learning or can be reasonably set according to experience. Specifically, according to the above formula, multiply the three obtained scores by their corresponding weights respectively and then sum them to get an initial score. Furthermore, multiply the initial score by the dialect scaling factor before the initial score, and scale the initial score to a corresponding degree according to the value of the dialect scaling factor. The scaled result is the score S of the audio sentence.
[0091] Finally, determine the evaluation result based on the score S of the audio sentence in the audio signal.
[0092] Optionally, the score S of the audio sentence can be directly regarded as the evaluation result. Alternatively, the score S of the audio sentence can also be classified according to grades. For example, scores between 0 - 60 belong to poor, 61 - 80 belong to good, and 81 - 100 belong to excellent. And regard the classified results, namely excellent, good, and poor, as the evaluation result.
[0093] According to the above solution, the final evaluation result can be obtained by integrating the scores in three aspects, ensuring the reliability of the evaluation result. In addition, since the consideration of the dialect scaling factor is added to the evaluation result, the accuracy of the evaluation result is further ensured.
[0094] Figure 7 shows a schematic flowchart of a speech evaluation method according to another embodiment of the present invention. As Figure 7 shown, first, the audio signal to be evaluated and the standard text information corresponding to the audio signal can be obtained simultaneously or at different times. Feature extraction is performed on the audio signal to obtain the acoustic features of each speech frame. Using the extracted acoustic features, the probabilities of the corresponding phonemes of the pronunciation of the speech frames of the audio signal in the phoneme dictionary can be determined. This probability can be represented by a matrix and is called pronunciation information. For the standard text information, its corresponding phonemes can be determined based on the dialect dictionary. Among them, the phonemes of the characters in the dialect dictionary include standard phonemes and dialect phonemes. Then, according to the acoustic features and pronunciation information of each speech frame, the phonemes of the speech frames of the audio signal are aligned with the phonemes corresponding to the standard text information. Thus, the alignment information of the speech frames is obtained. Among them, the alignment information includes the information of which phoneme in the phonemes corresponding to the standard text information the speech frame of the audio signal corresponds to. In addition, the audio signal can also be classified by dialect to obtain the dialect scaling factor. According to the pronunciation information, alignment information, and dialect scaling factor of the speech frames of the audio signal, the evaluation result of the audio signal relative to the standard text information can be determined.
[0095] According to another aspect of the present invention, there is provided a speech evaluation device. Figure 8 The schematic block diagram of a speech evaluation device 800 according to an embodiment of the present invention is shown. As Figure 8 shown, the speech evaluation device 800 may include the following modules.
[0096] A data acquisition module 810, configured to acquire an audio signal to be evaluated and standard text information corresponding to the audio signal.
[0097] A feature extraction module 820, configured to extract acoustic features of each speech frame of the audio signal.
[0098] A calculation module 830, configured to determine the probabilities that the pronunciations of the speech frames of the audio signal are each phoneme in a phoneme dictionary, so as to obtain pronunciation information of the speech frames, where the phoneme dictionary includes dialect phonemes.
[0099] A phoneme determination module 840, configured to determine phonemes corresponding to the standard text information based on a dialect dictionary, where the phonemes of the characters in the dialect dictionary include standard phonemes and dialect phonemes.
[0100] An alignment module 850, configured to align the speech frames of the audio signal and the phonemes corresponding to the standard text information based on the acoustic features and pronunciation information of each speech frame in the audio signal, so as to obtain alignment information of the speech frames.
[0101] An evaluation result determination module 860, configured to determine an evaluation result of the audio signal relative to the standard text information according to the pronunciation information and alignment information of the speech frames.
[0102] It should be noted that each component of the device should be understood as a functional module established to implement each step of the program flow or each step of the method. Each functional module is not an actual functional segmentation or separation limitation. The device defined by such a set of functional modules should be understood as mainly implementing the functional module architecture of the solution through the computer program recorded in the specification, rather than being understood as an entity device mainly implementing the solution through hardware means.
[0103] According to still another aspect of the present invention, there is also provided a speech evaluation device. Figure 9 The schematic block diagram of a speech evaluation device 900 according to an embodiment of the present invention is shown. As Figure 9As shown in the figure, the speech evaluation device 900 may include a sound collection device 910, an input device 920, a processor 930, and a memory 940. Among them, the sound collection device 910 is used to obtain the audio signal to be evaluated and send it to the processor 930. The input device 920 is used to input the standard text information corresponding to the audio signal to be evaluated and send it to the processor 930. The memory 940 stores computer program instructions, which are used to execute the speech evaluation method as described above when the computer program instructions are run by the processor.
[0104] According to another aspect of the present invention, a storage medium is also provided. Program instructions are stored on the storage medium, and when the program instructions are run by a computer or a processor, the computer or the processor is caused to execute the corresponding steps of the speech evaluation method of the embodiments of the present invention, and is used to implement the corresponding modules in the speech wake-up device and equipment according to the embodiments of the present invention. The storage medium may include, for example, the storage component of a tablet computer, the hard disk of a personal computer, a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a portable compact disc read-only memory (CD-ROM), a USB memory, or any combination of the above storage media. The computer-readable storage medium may be any combination of one or more computer-readable storage media.
[0105] It can be understood that those of ordinary skill in the art can read the above description of the speech evaluation method to understand the specific implementation manners and beneficial effects of the speech evaluation device, the speech evaluation equipment, and the storage medium. For the sake of brevity, they will not be repeated here.
[0106] Although example embodiments have been described herein with reference to the accompanying drawings, it should be understood that the above example embodiments are merely exemplary and are not intended to limit the scope of the present invention thereto. Those of ordinary skill in the art can make various changes and modifications therein without departing from the scope and spirit of the present invention. All such changes and modifications are intended to be included within the scope of the present invention as claimed in the appended claims.
[0107] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present invention.
[0108] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed.
[0109] In the specification provided herein, a large number of specific details are set forth. However, it can be understood that the embodiments of the present invention can be practiced without these specific details. In some instances, well-known methods, structures, and technologies have not been shown in detail so as not to obscure the understanding of this specification.
[0110] Similarly, it should be understood that, in order to streamline the present invention and assist in understanding one or more of the various inventive aspects, in the description of the exemplary embodiments of the present invention, the various features of the present invention are sometimes grouped together into a single embodiment, figure, or description thereof. However, the methods of the present invention should not be construed as reflecting the intention that the claimed invention requires more features than are expressly recited in each claim. Rather, as reflected by the corresponding claims, the inventive point lies in that the corresponding technical problems can be solved with features fewer than all the features of a single disclosed embodiment. Therefore, the claims following the detailed description are hereby expressly incorporated into the detailed description, where each claim itself serves as a separate embodiment of the present invention.
[0111] Those skilled in the art can understand that, except for features that are mutually exclusive, any combination can be used for all the features disclosed in this specification (including the accompanying claims, abstract, and drawings) and all the processes or units of any method or device so disclosed. Unless otherwise expressly stated, each feature disclosed in this specification (including the accompanying claims, abstract, and drawings) can be replaced by an alternative feature that provides the same, equivalent, or similar purpose.
[0112] In addition, those skilled in the art can understand that, although some of the embodiments described herein include certain features included in other embodiments rather than other features, the combination of the features of different embodiments means that it is within the scope of the present invention and forms different embodiments. For example, in the claims, any one of the claimed embodiments can be used in any combination.
[0113] Each component embodiment of the present invention can be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. Those skilled in the art should understand that a microprocessor or a digital signal processor (DSP) can be used in practice to implement some or all of the functions of some modules in the speech evaluation device according to the embodiments of the present invention. The present invention can also be implemented as a device program (for example, a computer program and a computer program product) for executing part or all of the methods described herein. Such a program implementing the present invention can be stored on a computer-readable medium, or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, or provided on a carrier signal, or in any other form.
[0114] It should be noted that the above embodiments illustrate the present invention rather than limit the present invention, and those skilled in the art can design alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses shall not be construed as limiting the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present invention can be implemented by means of hardware including several different elements and by means of a suitably programmed computer. The use of the words first, second, and third, etc. does not denote any order. These words can be interpreted as names.
[0115] As described above, it is only the specific implementation manner of the present invention or the description of the specific implementation manner. The protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, and all of them should be covered by the protection scope of the present invention. The protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A voice evaluation method, including: obtaining an audio signal to be evaluated; extracting acoustic features of each speech frame in the audio signal; using the acoustic features to determine the probability that the pronunciation of the speech frame of the audio signal is each phoneme in a phoneme dictionary, so as to obtain pronunciation information of the speech frame, wherein the phoneme dictionary includes dialect phonemes; obtaining standard text information corresponding to the audio signal to be evaluated; based on a dialect dictionary, determining the phonemes corresponding to the standard text information, wherein the phonemes of the characters in the dialect dictionary include standard phonemes and dialect phonemes; aligning the speech frames of the audio signal and the phonemes corresponding to the standard text information based on the acoustic features and pronunciation information of each speech frame in the audio signal, so as to obtain alignment information of the speech frames; and determining an evaluation result of the audio signal relative to the standard text information according to the pronunciation information and the alignment information of the speech frames; the method further includes: classifying the audio signal into a dialect to obtain a dialect scaling factor of the audio signal, determining the dialect scaling factor to be a first value for the case where the audio signal is dialect speech, and determining the dialect scaling factor to be a second value for the case where the audio signal is non-dialect speech, wherein the first value is greater than the second value; wherein, the determination of the evaluation result of the audio signal relative to the text information is also based on the dialect scaling factor.
2. The method according to claim 1, wherein, the classifying the audio signal into a dialect to obtain a dialect scaling factor of the audio signal includes: using a binary classification model to determine the probabilities that the audio signal is dialect speech or non-dialect speech respectively; determining the dialect scaling factor based on the probability that the audio signal is dialect speech and / or the probability that the audio signal is non-dialect speech.
3. The method according to claim 2, wherein, the determining the dialect scaling factor based on the probability that the audio signal is dialect speech and / or the probability that the audio signal is non-dialect speech includes: dividing the probability that the audio signal is dialect speech by the probability that the audio signal is non-dialect speech to obtain a quotient value; determining the quotient value as the dialect scaling factor.
4. The method according to any one of claims 1 to 3, wherein, the determination of the evaluation result of the audio signal relative to the text information includes: determining the accuracy A, fluency B and integrity C of the audio sentence in the audio signal according to the pronunciation information and the alignment information of the speech frames; calculating the score S of the audio sentence in the audio signal by using the following formula, S = δ * (a * A + b * B + c * C), where δ represents the dialect scaling factor, and a, b and c respectively represent the weights of the accuracy A, fluency B and integrity C; and determining the evaluation result based on the score S of the audio sentence in the audio signal.
5. The method according to claim 1, wherein, the dialect phonemes are provided with dialect identifiers, the method further includes: extracting acoustic features of each speech frame in the training audio signal; Train an acoustic model using the acoustic features of the speech frames of the training audio signal, where the value of the loss function of the acoustic model is determined based on the calculation result of the acoustic model and the dialect identifier; determining the probability that the speech frame of the audio signal is pronounced as each phoneme in the phoneme dictionary using the acoustic features to obtain pronunciation information of the speech frame, including: Input the acoustic features into the acoustic model so that the acoustic model outputs the probability that the speech frame of the audio signal is pronounced as each phoneme in the phoneme dictionary.
6. The method according to claim 1, wherein, Based on the acoustic features and pronunciation information of each speech frame in the audio signal, align the speech frames of the audio signal with the phonemes corresponding to the text information to obtain alignment information of the speech frames, including: Generate a search space corresponding to the standard text information based on the phonemes corresponding to the standard text information, where the search space includes a dialect phoneme path formed by dialect phonemes; Determine the phoneme corresponding to each speech frame based on the acoustic features and pronunciation information of each speech frame of the audio signal and the search space.
7. A speech evaluation device, including: A data acquisition module for acquiring an audio signal to be evaluated and standard text information corresponding to the audio signal; A feature extraction module for extracting the acoustic features of the audio signal; A calculation module for determining the probability that the speech frame of the audio signal is pronounced as each phoneme in the phoneme dictionary using the acoustic features to obtain pronunciation information of the speech frame, where the phoneme dictionary includes dialect phonemes; A phoneme determination module for determining the phonemes corresponding to the standard text information based on a dialect dictionary, where the phonemes of the characters in the dialect dictionary include standard phonemes and dialect phonemes; An alignment module for aligning the speech frames of the audio signal with the phonemes corresponding to the standard text information based on the acoustic features and pronunciation information of each speech frame in the audio signal to obtain alignment information of the speech frames; An evaluation result determination module for determining the evaluation result of the audio signal relative to the standard text information according to the pronunciation information and alignment information of the speech frames; The device is further configured to perform dialect classification on the audio signal to obtain a dialect scaling factor of the audio signal, determine that the dialect scaling factor is a first value for the case where the audio signal is dialect speech, and determine that the dialect scaling factor is a second value for the case where the audio signal is non-dialect speech, where the first value is greater than the second value; wherein, the evaluation result determination module determines the evaluation result of the audio signal relative to the text information also according to the dialect scaling factor.
8. A speech evaluation device, including a sound collection device, an input device, a processor, and a memory, wherein, The sound collection device is configured to acquire an audio signal to be evaluated and send it to the processor; The input device is configured to input standard text information corresponding to the audio signal to be evaluated and send it to the processor; The computer program instructions are stored in the memory and are used to execute the speech evaluation method according to any one of claims 1 to 6 when being run by the processor.
9. A storage medium, on which program instructions are stored, and the program instructions are used to execute the speech evaluation method according to any one of claims 1 to 6 when being run.
Citation Information
Patent Citations
Method for recognizing Gan dialect voice and Gan dialect point
CN109410914A
Spoken language evaluation method and device, equipment, storage medium and program product
CN114203201A