A pronunciation evaluation method, device, equipment and storage medium

By combining the acoustic model and the pause language model, a multi-dimensional pronunciation evaluation method is generated, which solves the problem of inaccurate evaluation results in the existing technology, realizes a comprehensive evaluation of pronunciation accuracy, omission errors and fluency, and improves the accuracy of the evaluation results.

CN115691554BActive Publication Date: 2025-10-14GUANGZHOU SHIYUAN ELECTRONICS CO LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202110846916.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-07-26
Publication Date
2025-10-14
Estimated Expiration
2041-07-26

AI Technical Summary

Technical Problem

Existing pronunciation evaluation methods can only evaluate pronunciation quality from a single dimension and cannot comprehensively and accurately assess pronunciation errors, resulting in inaccurate evaluation results.

Method used

The acoustic model is used to align the audio to be evaluated with the reference text, generate test text containing letters and blank characters, calculate the posterior probability of the letters to evaluate pronunciation accuracy, and evaluate pronunciation fluency through the pause language model. Combined with missed reading detection, multi-dimensional evaluation is achieved.

Benefits of technology

The accuracy of pronunciation assessment has been improved, and it can comprehensively evaluate pronunciation accuracy, omission errors and fluency, providing more accurate assessment results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115691554B_ABST
    Figure CN115691554B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a pronunciation evaluation method, device and equipment and a storage medium. The method comprises: obtaining to-be-evaluated audio and corresponding reference text, aligning the to-be-evaluated audio and the corresponding reference text through a preset acoustic model to obtain a first test text of the to-be-evaluated audio; merging continuous same letters in the first test text to obtain a second test text, calculating posterior probabilities of each letter in the second test text, and determining pronunciation accuracy of a corresponding letter according to the posterior probabilities; determining a missed letter according to letters in the second test text and letters in the corresponding reference text; deleting or replacing a blank symbol in the second test text with a pause symbol to obtain a third test text, calculating a language model perplexity of the third test text according to a preset pause language model, and determining pronunciation fluency of the to-be-evaluated audio according to the language model perplexity. The above technical means solve the problem of single evaluation dimension of the existing pronunciation evaluation method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of auxiliary learning technology, and in particular to a pronunciation evaluation method, apparatus, device, and storage medium. Background Art

[0002] Pronunciation quality assessment technology is a sub-method of computer-assisted language learning. It requires universities to accurately identify learners' pronunciation errors, provide objective letter-graded evaluations, and help them correct them. Pronunciation errors include mispronunciation, omissions, extraneous readings, and pauses.

[0003] Existing pronunciation assessment methods use CTC (Connectionist Temporal Classification) to detect instantaneous regions of nonlinear relationships between pronunciation parameters and acoustic parameters to detect mispronunciations. Alternatively, features such as phrase pauses based on pitch information are extracted to detect pronunciation fluency. However, the inventors found that these two pronunciation assessment methods only assess pronunciation quality for a specific type of pronunciation error and are unable to evaluate pronunciation in other dimensions, resulting in incomplete and inaccurate assessment results. Summary of the Invention

[0004] The embodiments of the present application provide a pronunciation evaluation method, apparatus, device and storage medium to solve the problem of a single evaluation dimension in existing pronunciation evaluation methods and improve the accuracy of the evaluation results.

[0005] In a first aspect, an embodiment of the present application provides a pronunciation evaluation method, comprising:

[0006] Acquire an audio to be evaluated and a corresponding reference text, align the audio to be evaluated and the corresponding reference text using a preset acoustic model, and obtain a first test text of the audio to be evaluated, where the first test text includes letters and blank characters in the corresponding reference text;

[0007] Combining consecutive identical letters in the first test text to obtain a second test text, calculating the posterior probability of each letter in the second test text, and determining the pronunciation accuracy of the corresponding letter based on the posterior probability;

[0008] determining missed letters based on the letters in the second test text and the letters in the corresponding reference text;

[0009] The blank character in the second test text is deleted or replaced with a pause character to obtain a third test text, the language model perplexity of the third test text is calculated according to a preset pause language model, and the pronunciation fluency of the audio to be evaluated is determined according to the language model perplexity.

[0010] In a second aspect, an embodiment of the present application provides a pronunciation evaluation device, comprising:

[0011] A test text determination module is configured to obtain to-be-evaluated audio and corresponding reference text, align the to-be-evaluated audio and the corresponding reference text through a preset acoustic model, and obtain first test text of the to-be-evaluated audio, the first test text containing letters and blank symbols in the corresponding reference text;

[0012] An accuracy evaluation module is configured to merge consecutive identical letters in the first test text to obtain second test text, calculate posterior probability of each letter in the second test text, and determine pronunciation accuracy of the corresponding letter according to the posterior probability;

[0013] An omission evaluation module is configured to determine omitted letters according to letters in the second test text and letters in the corresponding reference text;

[0014] A fluency evaluation module is configured to delete or replace blank symbols in the second test text with pause symbols to obtain third test text, calculate language model perplexity of the third test text according to a preset pause language model, and determine pronunciation fluency of the to-be-evaluated audio according to the language model perplexity.

[0015] In a third aspect, an embodiment of the present application provides a pronunciation evaluation device, comprising:

[0016] One or more processors;

[0017] A memory for storing one or more programs;

[0018] When the one or more programs are executed by the one or more processors, the one or more processors implement the pronunciation evaluation method according to the first aspect.

[0019] In a fourth aspect, an embodiment of the present application provides a computer readable storage medium having a computer program stored thereon, the program being executed by a processor to implement the pronunciation evaluation method according to the first aspect.

[0020] The above-mentioned pronunciation evaluation method, device, equipment and storage medium evaluate the pronunciation accuracy of the audio according to the good pronunciation of each letter by taking the posterior probability of each letter in the second test text as the pronunciation goodness of the corresponding letter. The audio omission errors are detected by comparing each letter in the second test text with the letters in the reference text. The occurrence probability of words and pauses in the third test text with pauses is predicted by a preset pause language model, and the language confusion of the third test text is calculated based on the occurrence probability, thereby evaluating the pronunciation fluency of the audio according to the language model confusion. Through the above-mentioned technical means of evaluating the pronunciation accuracy, omission errors and pronunciation fluency of the audio, pronunciation evaluation from multiple dimensions is achieved, and the accuracy of the evaluation results is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 This is a flowchart of a pronunciation evaluation method provided by one embodiment of the present application;

[0022] Figure 2 is a schematic diagram of a state transition table provided in an embodiment of the present application;

[0023] Figure 3 is a schematic diagram of a transfer path network provided in an embodiment of the present application;

[0024] Figure 4 is a schematic diagram of the optimal path provided by an embodiment of the present application;

[0025] Figure 5 This is a structural diagram of a pronunciation evaluation device provided by one embodiment of the present application;

[0026] Figure 6 This is a structural diagram of a pronunciation evaluation device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0027] The present application will be further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are intended to explain the present application, not to limit the present application. It should also be noted that, for ease of description, the accompanying drawings only illustrate portions relevant to the present application, not all structures.

[0028] It should be noted that, in this document, relational terms such as first and second are used solely to distinguish one entity, operation, or object from another, and do not necessarily require or imply the existence of any actual relationship or order between these entities, operations, or objects. For example, the terms "first" and "second" in the first sample set and the second sample set are used to distinguish between different sample sets.

[0029] The pronunciation evaluation method provided in the embodiments of the present application can be executed by a pronunciation evaluation device. The pronunciation evaluation device can be implemented in the form of software and / or hardware. The pronunciation evaluation device can be composed of two or more physical entities, or can be composed of one physical entity. For example, the pronunciation evaluation device can be a smart device such as a mobile phone, a tablet computer or a computer.

[0030] The pronunciation evaluation device is installed with at least one type of operating system, wherein the operating system includes but is not limited to an Android system, a Linux system and a Windows system. The pronunciation evaluation device can install at least one application program based on the operating system. The application program can be an application program provided by the operating system, or can be an application program downloaded from a third-party device or a server. In the embodiments, the pronunciation evaluation device at least has an application program that can execute the pronunciation evaluation method. Therefore, the pronunciation evaluation device can also be the application program itself.

[0031] For ease of understanding, the pronunciation evaluation device is exemplarily described as a mobile phone in the embodiments.

[0032] Figure 1 FIG. 1 is a flowchart of a pronunciation evaluation method according to an embodiment of the present application. Referring to FIG. 1, the pronunciation evaluation method includes the following steps. Figure 1

[0033] S110, obtaining a to-be-evaluated audio and a corresponding reference text, aligning the to-be-evaluated audio and the corresponding reference text through a preset acoustic model to obtain a first test text of the to-be-evaluated audio, and the first test text contains letters and blank symbols in the corresponding reference text.

[0034] The to-be-evaluated audio is audio that needs to be evaluated for pronunciation. The to-be-evaluated audio can be obtained by a microphone when a user reads text content displayed on a screen of a mobile phone. The reference text is the text content displayed on the screen of the mobile phone and read by the user. For example, the screen of the mobile phone displays the text content “I have a cat”. When the user reads “I have a cat”, the microphone built in the mobile phone collects the audio of the user reading “I have a cat”. Correspondingly, the mobile phone obtains the to-be-evaluated audio and the corresponding reference text to evaluate the pronunciation quality of the to-be-evaluated audio according to the reference text. However, the traditional pronunciation evaluation method can only evaluate the pronunciation quality from a single dimension, such as detecting the accuracy or fluency of pronunciation, which leads to an incomplete evaluation and an inaccurate evaluation result. To solve the problem, the pronunciation evaluation method provided in the embodiments aims to evaluate the pronunciation quality of the audio from multiple dimensions of pronunciation error types, so as to improve the certainty of the evaluation result.

[0035] ​Furthermore, since the audio to be evaluated is an audio file, which contains waveform information, waveform information is difficult to be used directly to evaluate the pronunciation quality. Therefore, the audio to be evaluated is generally converted into a corresponding test text, and the test text is used to characterize the pronunciation information of the audio to be evaluated. The test text is a text sequence composed of letters and blank characters of the reference text. The letters can be understood as the letters corresponding to each phoneme when the user reads the reference text aloud, and the blank characters can be understood as the uncertain pronunciations between the phonemes when reading the reference text aloud, which include uncertain phonemes, noises and pauses. The uncertain phonemes refer to letters whose pronunciation phonemes cannot be determined. Exemplarily, in order to determine the test text corresponding to the audio to be evaluated, the audio to be evaluated and the corresponding reference text can be aligned by a preset acoustic model to obtain a first test text. The preset acoustic model is a pre-trained CTC model, and the first test text is a test text with the same frame length as the audio to be evaluated. In this embodiment, the steps of aligning the audio to be evaluated and the corresponding reference text by the acoustic model specifically include S1101-S1104:

[0036] S1101. Insert a blank character before and after each letter of the reference text to obtain a fifth test text.

[0037] Take the reference text "cat" for example. Enter a blank character before and after each letter of "cat", and the fifth test text is εcεaεtε, where ε is the placeholder corresponding to the blank character.

[0038] S1102. Determine a transition path network including at least one transition path based on the frame length of the audio to be evaluated, the fifth test text, and a preset state jump condition; wherein the state jump condition includes jumping from a blank character before a letter to a blank character after the letter.

[0039] An exemplary description is given with the frame length of the audio to be evaluated being 10. A state transition table is constructed according to the frame length of the audio to be evaluated and the fifth test text. Figure 2 This is a schematic diagram of the state transition table provided in the embodiment of the present application. Figure 2 As shown, the column length of the state transition table is equal to the frame length of the audio being evaluated, t is the time frame of the audio being evaluated, and s is the possible states of the audio being evaluated. Since the audio being evaluated is the audio read aloud by the user with reference to the reference text, the letters and placeholders in the fifth test text represent the possible states of the audio being evaluated.

[0040] Further, according to the preset state transition condition, a transition path network is constructed in the state transition table. For example, the state transition condition includes: if the state s at time t is a placeholder, then the state s at time t can be transitioned from the state s at time t-1, the state s-1 at time t-1, and the state s-2 at time t-2; if the state s at time t is a letter, and the letter of the state s is the same as the letter of the state s-2, then the state s at time t can be transitioned from the state s at time t-1 and the state s-1 at time t-1; if the state s at time t is a letter, and the letter of the state s is different from the letter of the state s-2, then the state s at time t can be transitioned from the state s at time t-1, the state s-1 at time t-1, and the state s-2 at time t-2. Wherein, s is the current state, s-1 is the first previous state of s, s-2 is the second previous state of s, for example, referring to Figure 2 , if the current state s is a, then the state s-1 is ε, and the state s-2 is c. t is the current time, t-1 is the previous time of t, for example, t2 is the previous time of t3. Under the constraint of the above state condition, a transition path network containing multiple transition paths from t1 to t10 can be constructed. Figure 3 is a schematic diagram of the transition path network provided by the embodiment of the present application. As shown in Figure 3 , the transition path network contains multiple transition paths from t1 to t10. Since the user may miss reading when reading the reference text, for example, miss reading the phonemes of the letter c or t, the transition path network will also contain the placeholder ε after the letter c, the letter a, or the letter t at time t1, and the placeholder ε before the letter t, the letter a, or the letter c at time t10. However, since there are too many possible transition paths, Figure 3 only the transition paths of the missing reading of the middle letter are shown, for example, the transition path from the placeholder of the s-2 state at time t-1 to the placeholder of the s state at time t will retain the letter missing reading. It should be noted that no matter which state the transition path starts from and which state the transition path ends at, the transition path must comply with the above state transition condition.

[0041] S1103, calculate the posterior probability of the letters and blank symbols on the transition path, and determine the optimal path in the transition path network according to the posterior probability of the letters and blank symbols on the transition path.

[0042] Exemplarily, the pre-trained CTC model is used to evaluate the audio input, and the posterior probability of each time frame corresponding to each transition path is calculated by the CTC model, so as to determine the optimal path in the transition path network according to the posterior probability. The state corresponding to each time frame of the optimal path can be understood as the user reading the corresponding letter or uncertain letter at the corresponding time frame, and the pause. In this embodiment, the optimal path is searched in the transition path network according to the posterior probability of the state corresponding to each time frame in the transition path network by using the Viterbi algorithm. The Viterbi algorithm is used to search the optimal path, which can effectively reduce the processing amount of data and accelerate the search efficiency of the optimal path.

[0043] S1104, determining the character sequence corresponding to the optimal path as the first test text.

[0044] Exemplarily, the character sequence corresponding to the optimal path is obtained according to the characters or placeholders corresponding to each time frame on the optimal path in the order from left to right along the time frame. Figure 4 is a schematic diagram of the optimal path provided by the embodiment of the present application. As shown in Figure 4 , the character sequence corresponding to the optimal path is εεcεεεaaεt, that is, the first test text is εεcεεεaaεt.

[0045] S120, merging the same letters in the first test text to obtain a second test text, calculating the posterior probability of each letter in the second test text, and determining the pronunciation accuracy of the corresponding letter according to the posterior probability.

[0046] Exemplarily, the same letters appearing continuously in the first test text are the letters read by the user, and the audio to be tested includes a plurality of continuous audio frames recording the user reading the phonemes, which can be understood as the duration of the user reading the phonemes. In order to evaluate the subsequent omission error and the pronunciation accuracy of the user reading the phonemes of each letter in the reference text, the same letters appearing continuously are merged into one letter, and only one audio frame feature of the phoneme is retained. Exemplarily, the first test text is εεcεεεaaεt. In εεcεεεaaεt, a is a same letter appearing continuously, and the two a in the first test text are merged into one to obtain the second test text εεcεεεaεt. After obtaining the second test text, since the letters in the second test text correspond one by one to the letters in the reference text, the pronunciation accuracy of each letter in the second test text can be evaluated according to the posterior probability of each letter. In this embodiment, the step of determining the posterior probability of each letter in the second test text comprises:

[0047] S1201, determining the posterior probability of the letter on the optimal path as the posterior probability of the letter in the first test text.

[0048] Wherein, the posterior probability is the probability of each state in each time frame in the to-be-evaluated audio predicted by the pre-trained CTC model, and the lower the posterior probability of a letter is, the lower the probability of the user reading the letter phoneme in the time frame is, that is, the user is less likely to pronounce the letter phoneme in the time frame, and thus the posterior probability of the letter can be understood as the pronunciation goodness of the user reading the letter. The letter on the optimal path indicates that the user has pronounced the letter phoneme, and thus if the posterior probability of the letter corresponding to the optimal path is low, it indicates that the letter phoneme pronounced by the user is not accurate enough, which makes the CTC model output a lower posterior probability when predicting the probability of the letter.

[0049] For example, the first test text is the character sequence corresponding to the optimal path, and thus the posterior probability of the letter on the optimal path is determined as the posterior probability of the corresponding letter in the first test text.

[0050] S1202, determine the posterior probability of the letter appearing alone in the first test text as the posterior probability of the corresponding letter in the second test text.

[0051] For example, the second test text is obtained by merging the same letters appearing continuously in the first test text, and the letters appearing alone in the first test text are included in the second test text without change, and thus the posterior probability of the letter appearing alone in the first test text is directly determined as the posterior probability of the corresponding letter in the second test text.

[0052] S1203, calculate the average posterior probability of the letters appearing continuously in the first test text, and determine the average posterior probability as the posterior probability of the corresponding letter in the second test text.

[0053] For example, the same letters appearing continuously in the first test text are merged and included in the second test text, and the same letters appearing continuously correspond to respective posterior probabilities. Therefore, the respective posterior probabilities of the same letters appearing continuously, and the blank symbols within a certain range before and after the letters are averaged to obtain the average posterior probability of the letter, so as to determine the average posterior probability as the posterior probability of the corresponding letter in the second test text. For example, the posterior probabilities of the letter a in the first test text εεcεεεaaεεt are q1 and q2, the posterior probabilities of ε within a range of 2 before and after the letter a are q3, q4, q5 and q6, and the average value (q1+q2+q3+q4+q5+q6) / 6 of the six posterior probabilities is determined as the posterior probability of the letter a in the second test text εεcεεεaεεt.

[0054] S130, determine the omitted letter according to the letter in the second test text and the letter in the corresponding reference text.

[0055] For example, since the state jump condition used when constructing the transfer path network in this embodiment includes jumping from the blank character before the letter to the blank character after the letter, if the optimal path includes this state jump situation, the letters between the two blank character states will not be retained in the character sequence of the optimal path. By comparing the letters in the character sequence corresponding to the optimal path with the letters in the reference text, it is possible to detect letter omission errors that occur when the user reads the reference text. Furthermore, the second test text merges the same letters that appear consecutively and only retains one audio frame of the letters pronounced by the user, so that the second test text can be directly compared with the reference text.

[0056] In one embodiment, the blank character in the second test text is deleted to obtain a fourth test text, and the fourth test text is compared with the corresponding reference text to determine the missed letters. Exemplarily, when comparing the second test text and the reference text for letters, the placeholders in the second test text are interference items, so the placeholders in the second test text are deleted and only the letters in the second test text are retained. For example, ε in the second test text εεcεεεaεt is deleted to obtain a fourth test text cat, and the fourth test text cat is compared with the reference text cat to determine that the user has not made a letter missed error. In this embodiment, if the second test text is εεεεεεεaεt, ε is deleted to obtain the fourth test text at, and the fourth test text at is compared with the reference text cat to determine that the user has missed the letter c.

[0057] S140. Delete or replace the blank character in the second test text with a pause character to obtain a third test text, calculate the language model perplexity of the third test text according to a preset pause language model, and determine the pronunciation fluency of the audio to be evaluated according to the language model perplexity.

[0058] Among them, the preset pause language model is a pre-trained language model that can detect pauses. The language model perplexity can be understood as estimating the probability of occurrence of a sentence based on the probability of occurrence of each word or pause in a sentence output by the pause language model. Exemplarily, the language model perplexity of the audio to be evaluated is calculated using the preset pause language model to determine whether the pause position in the audio to be evaluated is appropriate. If the audio to be evaluated pauses at an inappropriate position, the language model perplexity will be higher.

[0059] For example, before calculating the language model perplexity of the audio to be evaluated, the pause positions in the audio to be evaluated are first determined, and pause symbols are inserted at the corresponding pause positions, so as to detect whether the pause positions of the audio to be evaluated are appropriate. In this embodiment, the steps of determining the pause positions of the audio to be evaluated include S1401-S1402:

[0060] S1401. Determine the sequence length of the blank character sequence according to the blank character sequence in the second test text.

[0061] A blank character sequence is a sequence of multiple consecutive εs, such as εεε. A long blank character sequence can be interpreted as a pause when the user is reading the reference text. Therefore, the length of the blank character sequence in the second test text is used to determine whether the blank character sequence represents a pause in reading.

[0062] S1402. Replace the blank character sequence whose sequence length meets the preset length threshold in the second test text with a pause character, and delete the blank character sequence whose sequence length does not meet the preset length threshold and the single blank character in the second test text to obtain a third test text.

[0063] Exemplarily, the preset length threshold can be regarded as the shortest pause duration when the user pauses in reading aloud. If the blank character sequence in the second test text is greater than or equal to the preset length threshold, it indicates that a pause state occurs in the time frame corresponding to the blank character sequence. If the blank character sequence in the second test text is less than the preset length threshold, it indicates that no pause state occurs in the time frame corresponding to the blank character sequence. After determining the blank character sequence in which a pause state occurs in the second test text, the position of the blank character sequence in the second test text is set as the pause position, and the blank character sequence is replaced with a pause character. However, the blank character sequence in which no pause state occurs and the blank character that appears alone in the second test text will affect the subsequent evaluation of pronunciation fluency, so the blank character sequence in which no pause state occurs and the blank character that appears alone in the second test text need to be deleted. For example, if the second test text is εIεεhεaεεvεeεεεεεεεεεεεεεεεεεεεεaεεcεεεaεεt and the preset length threshold is 20, it can be determined that a pause occurs in the time frame corresponding to the εεεεεεεεεεεεεεεεεεεε sequence in the second test text. Therefore, the εεεεεεεεεεεεεεεεεεεεεε sequence in the second test text is replaced with a pause character, and all ε in the second test text are deleted, resulting in the fifth test text I have|a cat, where "|" is a pause character.

[0064] In one embodiment, after determining the pause position in the second test text, a pause character may be inserted at the corresponding pause position in the fourth test text. For example, the second test text is εIεεhεaεεvεeεεεεεεεεεεεεεεεεεεεεεεεaεεcεεεaεεt, and the fourth test text is I have a cat. The pause character appears between the letters e and a, so the pause character is inserted between the letters e and a in the fourth test text, resulting in the fifth test text I have|a cat.

[0065] Furthermore, after determining the pause position of the audio to be evaluated, the language model perplexity is calculated based on the pre-trained pause language model. In this embodiment, the steps of calculating the language model perplexity specifically include S1403-S1404:

[0066] S1403: Determine the occurrence probabilities of words and pauses in the third test text according to a preset pause language model.

[0067] S1404: Substitute the occurrence probability into a preset perplexity calculation formula to calculate the language model perplexity of the third test text. The perplexity calculation formula is:

[0068]

[0069] Among them, N * is the total number of words and pauses in the third test text, is the word or pause of the third test text, Represents a text sequence of The probability of occurrence of words or pauses output by the pause language model.

[0070] For example, the pause language model can predict the probability of each word and pause symbol appearing at a corresponding position in the third test text, such as predicting the probability of “|” in the third test text “I have|a cat” appearing after the text sequence “I have”.

[0071] In this embodiment, the pronunciation fluency of the audio to be evaluated can be measured by the relative difference between the language model perplexity of the reference text and the language model perplexity of the third test text. Exemplarily, the steps of measuring the pronunciation fluency of the audio to be evaluated specifically include S1405-S1406:

[0072] S1405 : Calculate the language model perplexity of the reference text, and subtract the language model perplexity of the third test text from the language model perplexity of the reference text to obtain a relative difference between the third test text and the reference text.

[0073] For example, a standard pause position in a reference text is determined in advance, and a pause character is inserted at the pause position to obtain a pause reference text. The occurrence probability of letters and pause characters in the pause reference text is predicted using a pause language model, and the occurrence probability is substituted into the above-mentioned perplexity calculation formula to calculate the language model perplexity of the reference text. To improve the efficiency of pronunciation evaluation, the language model perplexity of the reference text can be calculated in advance, and the language model perplexity of the reference text can be directly called when evaluating pronunciation fluency.

[0074] Furthermore, the language model perplexity of the third test text is subtracted from the language model perplexity of the reference text to obtain the relative difference G between the third test text and the reference text, where G=PP(S * )-PP(S), PP(S) is the language model perplexity of the reference text.

[0075] S1406: Determine the pronunciation fluency of the audio to be evaluated based on the relative difference.

[0076] For example, when pausing at inappropriate places, the language model perplexity will be higher. Therefore, the pronunciation fluency evaluation index can be obtained by the relative difference in the language model perplexity of the two text paths to achieve the purpose of evaluating pronunciation fluency.

[0077] In summary, the pronunciation evaluation method provided by this embodiment uses the posterior probability of each letter in the second test text as the pronunciation goodness of the corresponding letter to evaluate the pronunciation accuracy of the audio according to the pronunciation goodness of each letter. By comparing each letter in the second test text with the letters in the reference text, omission errors in the audio are detected. The occurrence probability of words and pauses in the third test text with pauses is predicted by a preset pause language model, and the language confusion of the third test text is calculated based on the occurrence probability, thereby evaluating the pronunciation fluency of the audio according to the language model confusion. Through the above-mentioned technical means of evaluating the pronunciation accuracy, omission errors and pronunciation fluency of the audio, pronunciation evaluation from multiple dimensions is achieved, and the accuracy of the evaluation results is improved.

[0078] Figure 5 This is a structural diagram of a pronunciation evaluation device provided by an embodiment of the present application. Figure 5 The pronunciation evaluation device includes: a test text determination module 201, an accuracy evaluation module 202, an accuracy evaluation module 203 and a fluency evaluation module 204.

[0079] The test text determination module is configured to obtain the audio to be evaluated and the corresponding reference text, align the audio to be evaluated and the corresponding reference text using a preset acoustic model, and obtain a first test text of the audio to be evaluated, where the first test text contains letters and blank characters in the corresponding reference text;

[0080] The accuracy evaluation module is configured to merge consecutive identical letters in the first test text to obtain a second test text, calculate posterior probabilities of the letters in the second test text, and determine pronunciation accuracy of the corresponding letters according to the posterior probabilities.

[0081] The omission evaluation module is configured to determine the omitted letters according to the letters in the second test text and the letters in the corresponding reference text.

[0082] The fluency evaluation module is configured to delete or replace blank symbols in the second test text with pause symbols to obtain a third test text, calculate a language model perplexity of the third test text according to a preset pause language model, and determine pronunciation fluency of the audio to be evaluated according to the language model perplexity.

[0083] On the basis of the above embodiment, the test text determination module comprises: a state construction unit configured to insert a blank symbol before and after each letter in the reference text to obtain a fifth test text; a network construction unit configured to determine a transition path network comprising at least one transition path according to a frame length of the audio to be evaluated and the fifth test text, and a preset state transition condition; wherein the state transition condition comprises jumping from the blank symbol before the letter to the blank symbol after the letter; an optimal path determination unit configured to calculate posterior probabilities of the letters and the blank symbols on the transition path, and determine an optimal path in the transition path network according to the posterior probabilities of the letters and the blank symbols on the transition path; and a test text determination unit configured to determine a character sequence corresponding to the optimal path as the first test text.

[0084] On the basis of the above embodiment, the accuracy evaluation module comprises: a first posterior probability determination unit configured to determine the posterior probabilities of the letters on the optimal path as the posterior probabilities of the letters in the first test text; a second posterior probability determination unit configured to determine the posterior probabilities of the letters appearing alone in the first test text as the posterior probabilities of the corresponding letters in the second test text; and a third posterior probability determination unit configured to calculate average posterior probabilities of the letters appearing consecutively in the first test text, and determine the average posterior probabilities as the posterior probabilities of the corresponding letters in the second test text.

[0085] On the basis of the above embodiment, the omission evaluation module comprises: a deletion unit configured to delete the blank symbols in the second test text to obtain a fourth test text, compare the fourth test text with the corresponding reference text, and determine the omitted letters.

[0086] Based on the above embodiment, the fluency assessment module includes: a sequence length determination unit, configured to determine the sequence length of the blank character sequence based on the blank character sequence in the second test text; a pause position determination unit, configured to replace the blank character sequence in the second test text whose sequence length meets the preset length threshold with a pause character, and delete the blank character sequence in the second test text whose sequence length does not meet the preset length threshold and the blank character that appears alone, to obtain a third test text.

[0087] Based on the above embodiment, the fluency assessment module further includes: an occurrence probability calculation unit configured to determine the occurrence probability of words and pauses in the third test text according to a preset pause language model; a perplexity calculation unit configured to substitute the occurrence probability into a preset perplexity calculation formula to calculate the language model perplexity of the third test text, where the perplexity calculation formula is:

[0088]

[0089] Among them, N * is the total number of words and pauses in the third test text, is the word or pause of the third test text, Represents a text sequence of The probability of occurrence of words or pauses output by the pause language model.

[0090] Based on the above embodiment, the fluency evaluation module further includes: a relative difference calculation unit, configured to calculate the language model perplexity of the reference text, subtract the language model perplexity of the third test text from the language model perplexity of the reference text, and obtain the relative difference between the third test text and the reference text; and a fluency evaluation unit, configured to determine the pronunciation fluency of the audio to be evaluated based on the relative difference.

[0091] In summary, the pronunciation evaluation device provided by this embodiment uses the posterior probability of each letter in the second test text as the pronunciation goodness of the corresponding letter to evaluate the pronunciation accuracy of the audio according to the pronunciation goodness of each letter. By comparing each letter in the second test text with the letters in the reference text, omission errors in the audio are detected. The occurrence probability of words and pauses in the third test text with pauses is predicted by a preset pause language model, and the language confusion of the third test text is calculated based on the occurrence probability, thereby evaluating the pronunciation fluency of the audio according to the language model confusion. Through the above-mentioned technical means of evaluating the pronunciation accuracy, omission errors and pronunciation fluency of the audio, pronunciation evaluation from multiple dimensions is achieved, and the accuracy of the evaluation results is improved.

[0092] It is worth noting that in the above-mentioned embodiments of the pronunciation evaluation device, each unit and module included is only divided according to functional logic, but is not limited to the above-mentioned division, as long as the corresponding function can be realized; in addition, the specific name of each functional unit is only for easy mutual differentiation, and does not serve to limit the protection scope of the present application.

[0093] The pronunciation evaluation device provided by the embodiments of the present application is contained in a pronunciation evaluation device, and can be used to execute the pronunciation evaluation method provided by any of the above-mentioned embodiments, and has the corresponding functions and beneficial effects.

[0094] Figure 6 is a structural schematic diagram of a pronunciation evaluation device provided by an embodiment of the present application. As shown in Figure 6 , the pronunciation evaluation device includes a processor 30, a memory 31, an input device 32, an output device 33, and a display screen 34; the number of processors 30 in the pronunciation evaluation device can be one or more, Figure 6 , taking one processor 30 as an example; the number of display screens 34 in the pronunciation evaluation device can be one or more, Figure 6 , taking one display screen 34 as an example; the processor 30, the memory 31, the input device 32, the output device 33, and the display screen 34 in the pronunciation evaluation device can be connected through a bus or other means, Figure 6 , taking connection through a bus as an example.

[0095] The memory 31, as a kind of computer readable storage medium, can be used to store software programs, computer executable programs and modules, such as program instructions / modules corresponding to the pronunciation evaluation method in the embodiments of the present application (for example, the test text determination module 201, the accuracy evaluation module 202, the accuracy evaluation module 203 and the fluency evaluation module 204 in the pronunciation evaluation device). The processor 30 executes the software programs, instructions and modules stored in the memory 31, thereby performing various functional applications and data processing of the pronunciation evaluation device, that is, realizing the above-mentioned pronunciation evaluation method.

[0096] The memory 31 can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system and at least one application program required by a function; the data storage area can store data created according to the use of the pronunciation evaluation device, etc. In addition, the memory 31 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some examples, the memory 31 can further include a memory remotely arranged with respect to the processor 30, and these remote memories can be connected to the pronunciation evaluation device through a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.

[0097] The input device 32 can be used to receive inputted digital or letter information, and to generate key signal input related to user settings and function control of the pronunciation evaluation apparatus. The output device 33 can include an audio output device such as a speaker.

[0098] The pronunciation evaluation apparatus described above comprises a pronunciation evaluation device, which can be used to perform any pronunciation evaluation method, and has corresponding functions and beneficial effects.

[0099] The embodiment of the present application also provides a computer readable storage medium, which stores a computer program, and the program is executed by a processor to perform the pronunciation evaluation method provided by the above embodiment.

[0100] Of course, the computer executable instructions of the computer readable storage medium provided by the embodiment of the present application are not limited to the method operations described above, and can also perform the related operations in the pronunciation evaluation method provided by any embodiment of the present application.

[0101] Through the above description of the embodiments, those skilled in the art can clearly understand that the present application can be realized by means of software and necessary universal hardware, and of course can also be realized by hardware, but in many cases the former is a better embodiment. Based on such understanding, the technical solutions of the present application or the part that contributes to the prior art can be embodied in the form of a software product, which can be stored in a computer readable storage medium, such as a floppy disk, a read-only memory (ROM), a random access memory (RAM), a FLASH memory, a hard disk or an optical disk, and includes a number of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods of various embodiments of the present application.

[0102] Note that the above are only the preferred embodiments of the present application and the technical principles applied. Those skilled in the art will understand that the present application is not limited to the specific embodiments herein, and those skilled in the art can make various obvious changes, re-adjustments and substitutions without departing from the scope of the present application. Therefore, although the present application has been described in more detail through the above embodiments, the present application is not limited to the above embodiments, and can include more other equivalent embodiments without departing from the concept of the present application, and the scope of the present application is determined by the scope of the appended claims.

Claims

1. A pronunciation evaluation method, characterized in that: include: Acquire an audio to be evaluated and a corresponding reference text, align the audio to be evaluated and the corresponding reference text using a preset acoustic model, and obtain a first test text of the audio to be evaluated, where the first test text includes letters and blank characters in the corresponding reference text; Combining consecutive identical letters in the first test text to obtain a second test text, calculating the posterior probability of each letter in the second test text, and determining the pronunciation accuracy of the corresponding letter based on the posterior probability; determining missed letters based on the letters in the second test text and the letters in the corresponding reference text; Determining the length of a blank character sequence in the second test text; replacing blank character sequences in the second test text whose lengths meet a preset length threshold with a pause character, and deleting blank character sequences in the second test text whose lengths do not meet the preset length threshold and single blank characters, to obtain a third test text; The language model perplexity of the third test text is calculated according to a preset pause language model, and the pronunciation fluency of the audio to be evaluated is determined according to the language model perplexity.

2. The method according to claim 1, characterized in that The step of aligning the audio to be evaluated and the corresponding reference text using a preset acoustic model to obtain a first test text having the same length as the audio to be evaluated includes: Insert a blank symbol before and after each letter of the reference text to obtain a fifth test text; Determining a transition path network including at least one transition path based on the frame length of the audio to be evaluated, the fifth test text, and a preset state transition condition; wherein the state transition condition includes a transition from a blank character before a letter to a blank character after the letter; Calculating the posterior probabilities of letters and blank characters on the transition path, and determining an optimal path in the transition path network based on the posterior probabilities of letters and blank characters on the transition path; The character sequence corresponding to the optimal path is determined as the first test text.

3. The method according to claim 2, characterized in that Calculating the posterior probability of each letter in the second test text includes: Determining the posterior probabilities of letters on the optimal path as the posterior probabilities of letters in the first test text; Determine the posterior probability of a letter that appears alone in the first test text as the posterior probability of the corresponding letter in the second test text; Calculate the average posterior probability of letters that appear consecutively in the first test text, and determine the average posterior probability as the posterior probability of the corresponding letters in the second test text.

4. The method according to claim 1, wherein The step of determining the missed letters according to the letters in the second test text and the letters in the corresponding reference text comprises: The blank character in the second test text is deleted to obtain a fourth test text, and the fourth test text is compared with the corresponding reference text to determine the missed letters.

5. The method according to claim 1, wherein Calculating the language model perplexity of the third test text according to the preset pause language model includes: Determining the occurrence probabilities of words and pauses in the third test text according to a preset pause language model; Substitute the occurrence probability into a preset perplexity calculation formula to calculate the language model perplexity of the third test text. The perplexity calculation formula is: in, is the total number of words and pauses in the third test text, is a word or pause character of the third test text, Represents a text sequence of The probability of occurrence of the word or pause character output by the pause language model.

6. The method according to claim 5, characterized in that Determining the pronunciation fluency of the audio to be evaluated according to the language model perplexity includes: Calculating the language model perplexity of the reference text, and subtracting the language model perplexity of the third test text from the language model perplexity of the reference text to obtain a relative difference between the third test text and the reference text; The pronunciation fluency of the audio to be evaluated is determined according to the relative difference.

7. A pronunciation evaluation device, characterized in that: include: a test text determination module configured to obtain an audio to be evaluated and a corresponding reference text, align the audio to be evaluated and the corresponding reference text using a preset acoustic model, and obtain a first test text of the audio to be evaluated, where the first test text contains letters and blank characters in the corresponding reference text; an accuracy evaluation module configured to combine consecutive identical letters in the first test text to obtain a second test text, calculate a posterior probability of each letter in the second test text, and determine the pronunciation accuracy of the corresponding letter based on the posterior probability; a missed reading evaluation module configured to determine missed reading letters based on letters in the second test text and letters in the corresponding reference text; The fluency evaluation module is configured to determine the sequence length of the blank character sequence based on the blank character sequence in the second test text; replace the blank character sequence in the second test text whose sequence length meets the preset length threshold with a pause character, and delete the blank character sequence and the single blank character in the second test text whose sequence length does not meet the preset length threshold to obtain a third test text; calculate the language model perplexity of the third test text based on a preset pause language model; and determine the pronunciation fluency of the audio to be evaluated based on the language model perplexity.

8. A pronunciation evaluation device, characterized in that: include: one or more processors; a memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the pronunciation evaluation method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that A computer program is stored thereon, and when the program is executed by a processor, the pronunciation evaluation method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Method for automatic evaluation based on generalized fluent spoken language fluency

    CN101740024A

  • System and method for assessing proficiency in Putonghua

    CN102568475A

  • Spoken language pronunciation evaluation method and system for minority language, and storage medium

    CN112967711A