A pronunciation evaluation method and system, and an electronic device

By employing word-phoneme sequences and decoding networks under a unified phoneme rule in English pronunciation evaluation, and using a single acoustic model for evaluation, the problem of having to evaluate British and American pronunciation separately in existing technologies is solved, achieving a more efficient evaluation method.

CN119889345BActive Publication Date: 2025-11-28GUANGZHOU SHIYUAN ELECTRONICS CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311398949.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-10-25
Publication Date
2025-11-28
Estimated Expiration
2043-10-25

AI Technical Summary

Technical Problem

Existing English pronunciation assessment technologies require the use of separate American and British pronunciation assessment models, resulting in complex and inefficient assessment methods.

Method used

A decoding network is constructed using word-phoneme sequences under a unified phoneme rule, and an acoustic model is used for evaluation, which reduces the complexity of the evaluation method.

Benefits of technology

By using a unified phoneme rule-based evaluation method, the evaluation process is simplified, evaluation efficiency is improved, and it is compatible with both British and American pronunciation evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119889345B_ABST
    Figure CN119889345B_ABST
Patent Text Reader

Abstract

The application provides a pronunciation evaluation method and system and an electronic device. The method comprises the following steps: obtaining to-be-evaluated audio and to-be-evaluated text corresponding to the to-be-evaluated audio; obtaining a word-phoneme sequence, wherein the word-phoneme sequence comprises a plurality of words, a phoneme sequence of an American phonetic alphabet sequence corresponding to each word under a unified phoneme rule, and a phoneme sequence of a British phonetic alphabet sequence corresponding to each word under the unified phoneme rule; the unified phoneme rule comprises a one-to-one correspondence relationship between the British phonetic alphabet and the phoneme and a one-to-one correspondence relationship between the American phonetic alphabet and the phoneme; and in the unified phoneme rule, at least one phoneme corresponds to one British phonetic alphabet and one American phonetic alphabet; constructing a decoding network based on the word-phoneme sequence and the to-be-evaluated text; obtaining an acoustic model; and outputting evaluation information based on the to-be-evaluated audio, the acoustic model and the decoding network. The method does not need to use an American pronunciation evaluation model and a British pronunciation evaluation model for evaluation, thereby reducing the complexity of the evaluation mode.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of pronunciation evaluation technology, and in particular to a pronunciation evaluation method and system, and an electronic device. Background Technology

[0002] Pronunciation assessment technology is a subfield of computer-assisted language learning. It analyzes a user's audio recordings and outputs scores on indicators such as pronunciation accuracy, fluency, and completeness, allowing the user to correct their pronunciation problems based on these scores.

[0003] In English learning, given the existence of two main pronunciation guidelines—British and American English—and the user's pronunciation preference being unknown beforehand, a pronunciation assessment system needs to be compatible with both British and American English pronunciation assessment technologies. Existing technologies for assessing English pronunciation fall into two categories. The first category pre-defines pronunciation guidelines (American / British) and then uses the corresponding pronunciation model for assessment. However, because this type of technology assesses American and British English independently, it cannot address situations where the user's pronunciation preference is unknown beforehand. The second category uses both British and American pronunciation assessment models simultaneously to calculate the pronunciation probabilities under the corresponding British and American English rules for the input audio, and then selects the optimal model as the final result using a certain strategy. Currently, assessing English pronunciation requires two independent models, one for British English and one for American English, making the assessment method quite complex. Summary of the Invention

[0004] This application provides a pronunciation evaluation method, system, and electronic device, which reduces the complexity of the evaluation method.

[0005] In a first aspect, embodiments of this application provide a pronunciation evaluation method, the method comprising: acquiring an audio to be evaluated and a text to be evaluated corresponding to the audio; acquiring a word-phoneme sequence, wherein the word-phoneme sequence includes multiple words, a phoneme sequence of American phonetic symbols corresponding to each word under a unified phoneme rule, and a phoneme sequence of British phonetic symbols corresponding to each word under the unified phoneme rule, the unified phoneme rule including a one-to-one correspondence between British phonetic symbols and phonemes, and a one-to-one correspondence between American phonetic symbols and phonemes, and in the unified phoneme rule, at least one phoneme corresponds to both a British phonetic symbol and an American phonetic symbol; constructing a decoding network based on the word-phoneme sequence and the text to be evaluated; acquiring an acoustic model; and outputting evaluation information based on the audio to be evaluated, the acoustic model, and the decoding network.

[0006] In some embodiments, the method for constructing the word-phoneme sequence includes: obtaining multiple words, British phonetic transcription sequences corresponding to each word, and American phonetic transcription sequences corresponding to each word; obtaining the unified phoneme rule; obtaining a phoneme sequence of the British phonetic transcription sequence corresponding to each word under the unified phoneme rule based on the unified phoneme rule and each British phonetic transcription sequence; obtaining a phoneme sequence of the American phonetic transcription sequence corresponding to each word under the unified phoneme rule based on the unified phoneme rule and each American phonetic transcription sequence; and constructing the word-phoneme sequence based on each word, the phoneme sequence of the American phonetic transcription sequence corresponding to each word under the unified phoneme rule, and the phoneme sequence of the British phonetic transcription sequence corresponding to each word under the unified phoneme rule.

[0007] In some embodiments, obtaining the acoustic model includes: obtaining training audio and training text corresponding to the training audio; and training a preset model based on the training audio, the training text, and the word-phoneme sequence to obtain the acoustic model.

[0008] In some embodiments, the acoustic model is a deep neural network-hidden Markov model.

[0009] In some embodiments, outputting evaluation information based on the audio to be evaluated, the acoustic model, and the decoding network includes: determining the acoustic score corresponding to each audio frame in the audio to be evaluated based on the acoustic model; determining the optimal path based on each acoustic score and the decoding network; and outputting the evaluation information based on the optimal path.

[0010] In some embodiments, determining the optimal path based on each of the acoustic scores and the decoding network includes: determining the optimal path using the Viterbi algorithm based on each of the acoustic scores and the decoding network.

[0011] In some embodiments, the evaluation information includes at least one of the following: boundary information of each phoneme, posterior probability information corresponding to the phoneme output at each time moment, and pronunciation mode.

[0012] In a second aspect, embodiments of this application provide an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method described in any of the embodiments of the first aspect above.

[0013] Thirdly, embodiments of this application provide a pronunciation evaluation system, which includes a voice acquisition device and an electronic device as described in the second aspect, wherein the voice acquisition device is connected to the electronic device.

[0014] Fourthly, embodiments of this application also provide a computer-readable storage medium storing computer-executable instructions for causing a computer to perform the method described in any of the embodiments of the first aspect above.

[0015] Fifthly, embodiments of this application also provide a computer program product, the computer program product including a computer program stored on a computer-readable storage medium, the computer program including program instructions, which, when executed by a computer, cause the computer to perform the method described in any of the embodiments of the first aspect above.

[0016] Compared with the prior art, the beneficial effects of this application are as follows: Unlike the prior art, the embodiments of this application provide a pronunciation evaluation method, system, and electronic device. The method includes acquiring an audio to be evaluated and a corresponding text to be evaluated; acquiring a word-phoneme sequence; constructing a decoding network based on the word-phoneme sequence and the text to be evaluated; acquiring an acoustic model; and outputting evaluation information based on the audio to be evaluated, the acoustic model, and the decoding network. In this method, since both the British and American phonetic sequences of each word in the word-phoneme sequence are represented using phonemes from the Unified Phoneme Rules, when evaluating the audio using the decoding network constructed based on the word-phoneme sequence under the Unified Phoneme Rules and an acoustic model, only one acoustic model is needed to obtain the evaluation information. There is no need to use separate American and British pronunciation evaluation models, thus reducing the complexity of the evaluation method. Attached Figure Description

[0017] One or more embodiments are illustrated by way of example with reference to the accompanying drawings. These illustrations do not constitute a limitation on the embodiments. Elements / modules and steps with the same reference numerals in the drawings are represented as similar elements / modules and steps. Unless otherwise stated, the figures in the drawings do not constitute a limitation on scale.

[0018] Figure 1 This is a schematic diagram of the structure of a pronunciation evaluation system provided in an embodiment of this application;

[0019] Figure 2 This is a structural block diagram of an electronic device provided in an embodiment of this application;

[0020] Figure 3 This is a schematic flowchart of a pronunciation evaluation method provided in an embodiment of this application;

[0021] Figure 4 This is a schematic diagram of a decoding network provided in an embodiment of this application;

[0022] Figure 5 This is a schematic diagram of another decoding network provided in an embodiment of this application. Detailed Implementation

[0023] The present application will now be described in detail with reference to specific embodiments. These embodiments will help those skilled in the art to further understand the present application, but do not limit the present application in any way. It should be noted that those skilled in the art can make several modifications and improvements without departing from the concept of the present application. These all fall within the protection scope of the present application.

[0024] To facilitate understanding of this application, a more detailed description is provided below with reference to the accompanying drawings and specific embodiments. Unless otherwise defined, all technical and scientific terms used in this specification have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the application. The term "and / or" as used in this specification includes any and all combinations of one or more of the associated listed items.

[0025] It should be noted that, unless there is a conflict, the various features in the embodiments of this application can be combined with each other, all of which are within the protection scope of this application. Furthermore, although functional modules are divided in the device schematic diagram, in some cases, they can be divided differently from those in the device. In addition, the terms "first" and "second" used herein do not limit the data or execution order, but only distinguish between identical or similar items with essentially the same function and effect.

[0026] Currently, English pronunciation assessment algorithms typically employ two independent models—one for American pronunciation and one for British pronunciation—to evaluate audio and obtain the probabilities of American and British pronunciation. For example, a database of British and American accents can be established first, with phoneme-level annotation. Then, a multi-task acoustic model is trained, generally using a neural network model with a shared underlying layer and two branches in the output layer: one for American phoneme classification and the other for British phoneme classification. Next, based on the user's input speech information, the text is segmented into word sequences, and the British and American phoneme sequences for each word are extracted. The probability of British / American pronunciation for each word is then calculated. Finally, after normalization, the probability of British / American pronunciation for each word is converted, and the overall probability of British / American pronunciation for the entire text is obtained. It is evident that this assessment method requires separate evaluations of American and British pronunciations, resulting in complexity and low efficiency.

[0027] To address the aforementioned technical problems, embodiments of this application provide a method, system, and electronic device for pronunciation evaluation. In this method, an acoustic model is used, and a decoding network constructed from word-phone sequence under a unified phoneme rule outputs evaluation information. This eliminates the need to use separate American and British pronunciation evaluation models, reducing the complexity of the evaluation method and improving evaluation efficiency.

[0028] Firstly, see [the following] Figure 1 , Figure 1 This is a schematic diagram of a pronunciation evaluation system provided in an embodiment of this application. The system includes an electronic device 100 and a voice acquisition device 200. The electronic device 100 is connected to the voice acquisition device 200. The voice acquisition device 200 and the electronic device 100 can be directly or indirectly connected via wired or wireless communication, such as through a network. The network can be a wide area network or a local area network, or a combination of both. This application does not impose any limitations on this.

[0029] The voice acquisition device 200 can be used to acquire training audio or audio to be evaluated. The voice acquisition device 200 may include devices such as a microphone. It can record the user's audio and send the audio to the electronic device 100. The electronic device 100 acquires the training audio or audio to be evaluated, executes the pronunciation evaluation method provided in the embodiments of this application, and obtains the corresponding evaluation information.

[0030] Electronic device 100 can be a smartphone, tablet computer, laptop computer, desktop computer, etc.

[0031] In this embodiment, the electronic device 100 can be used to execute the pronunciation evaluation method provided in this application, and has the corresponding functional modules and beneficial effects for executing the method. Technical details not described in detail in this embodiment can be found in the methods provided in the embodiments of this application.

[0032] Secondly, please see Figure 2 It illustrates the hardware structure of an electronic device 100 capable of performing the pronunciation evaluation method described in this application. The electronic device 100 may be... Figure 1 The electronic device 100 shown.

[0033] Please see Figure 2 The electronic device 100 includes a processor 201 and a memory 202 connected via a communication link. Here, the communication link can be established via a bus. Figure 2 The bus connection between China and Israel is illustrated by example. It is understood that... Figure 2 The electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.

[0034] For details, please see Figure 2 It illustrates the hardware structure of an electronic device 100 capable of performing the pronunciation evaluation method described in this application. The electronic device 100 may be... Figure 1 The electronic device 100 shown.

[0035] Please see Figure 2 The electronic device 100 includes a processor 101 and a memory 102 connected via a communication link. Here, the communication link can be established via a bus. Figure 2 The bus connection between China and Israel is illustrated by example. It is understood that... Figure 2 The electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.

[0036] The processor 101 is configured to support the electronic device 100 in performing corresponding functions in the pronunciation evaluation method. The processor 101 can be a central processing unit (CPU), a network processor (NP), a hardware chip, or any combination thereof. The aforementioned hardware chip can be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The aforementioned PLD can be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0037] The memory 102, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules, such as the program instructions / modules corresponding to the pronunciation evaluation method in the embodiments of this application. The processor 101 can implement the pronunciation evaluation method in any of the following method embodiments by running the non-transitory software programs, instructions, and modules stored in the memory 102.

[0038] Memory 102 may include volatile memory (VM), such as random access memory (RAM); memory 102 may also include non-volatile memory (NVM), such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid-state drive (SSD); memory 102 may also include combinations of the above types of memory.

[0039] It is understood that the electronic device 100 also includes other supporting hardware and software. Hardware may include a keyboard, mouse, monitor, etc. Software may include an operating system, which is a program that manages and controls the hardware and software resources of the electronic device. Software may also include various applications (apps). Other parts of the improved electronic device 100 not involved in the embodiments of this application will not be described here.

[0040] As can be understood from the above, the pronunciation evaluation method provided in this application embodiment can be implemented by various types of electronic devices with processing capabilities, such as being implemented by the processor of an electronic device or by other devices with computing capabilities. Other devices with computing capabilities may be smart terminals or servers that are communicatively connected to the electronic device.

[0041] The above-described product can perform the methods provided in the embodiments of this application, and has the corresponding functional modules and beneficial effects for performing the methods. Technical details not described in detail in this embodiment can be found in the methods provided in the embodiments of this application.

[0042] The pronunciation evaluation method provided in this application is described below with reference to exemplary applications and implementations of the electronic device provided in the embodiments of this application. Please refer to... Figure 3 , Figure 3 This is a schematic flowchart of the pronunciation evaluation method provided in an embodiment of this application. It is understood that the execution entity of this training method can be one or more processors of an electronic device. The method includes:

[0043] Step S100: Obtain the audio to be evaluated and the corresponding text to be evaluated.

[0044] The text to be evaluated refers to a sequence of words. In this application, the text to be evaluated is a sequence of English words. This text to be evaluated can be acquired or pre-stored; for example, it can be acquired in real time by an interactive application, such as a language learning application, by the user reading the text aloud.

[0045] The audio to be evaluated refers to the speech data based on the text to be evaluated, which is to be used for pronunciation assessment. For example, speech data that needs to be assessed for spoken language can be used as the audio to be evaluated. In this application, the audio to be evaluated is in English, which can be British pronunciation, American pronunciation, or a mixture of both. The audio to be evaluated is acquired in real time; for example, it can be acquired in real time through interactive applications such as language learning applications, using the speech data of a user reading the text to be evaluated.

[0046] Step S200: Obtain the word-phoneme sequence, wherein the word-phoneme sequence contains multiple words, the phoneme sequence of the American phonetic symbols corresponding to each word under the unified phoneme rule, and the phoneme sequence of the British phonetic symbols corresponding to each word under the unified phoneme rule. The unified phoneme rule contains a one-to-one correspondence between British phonetic symbols and phonemes, and a one-to-one correspondence between American phonetic symbols and phonemes. In the unified phoneme rule, at least one phoneme corresponds to both a British phonetic symbol and an American phonetic symbol.

[0047] A phoneme sequence contains multiple phonemes arranged in sequence. A phoneme is the smallest unit of speech that is divided according to the natural properties of speech. Based on the analysis of the articulation action, the pronunciation of a word can be composed of one or more phonemes.

[0048] In the Unified Phoneme Rules, one British phonetic symbol corresponds to one phoneme, and one American phonetic symbol corresponds to one phoneme. There exists at least one phoneme that corresponds to both a British and an American phonetic symbol. The Unified Phoneme Rules can be a table of correspondences between phonemes and symbols, containing multiple phonemes, their correspondences with British phonetic symbols, and their correspondences with American phonetic symbols. The total number of phonemes is less than the sum of the total number of British and American phonetic symbols.

[0049] Specifically, both American and British phonetic symbols use the same notation for consonant phonemes. Therefore, under the unified phoneme rules, phonemes corresponding to consonant phonemes can be directly represented using either American or British phonetic symbols, meaning there are a total of 24 phonemes corresponding to consonant phonemes. Phonemes corresponding to vowel phonemes can be partially represented using unified symbols, as shown in Table 1 below.

[0050] Table 1. Correspondence between phonemes, American phonetic symbols, and British phonetic symbols.

[0051]

[0052]

[0053] As shown in Table 1, there are 23 phonemes corresponding to vowel phonemes. Therefore, the total number of phonemes under this unified phoneme rule is 47, which is less than the sum of the total number of British and American phonemes. After obtaining the above unified phoneme rule, the American and British phonetic sequences of words can be converted to obtain the phoneme sequence under the unified phoneme rule. In practical applications, the phoneme symbols in the unified phoneme rule can be formulated according to actual needs, and there is no need to be bound by the limitations of this embodiment.

[0054] In some embodiments, the method for constructing a word-phoneme sequence includes step S211: obtaining multiple words, the corresponding British phonetic symbols for each word, and the corresponding American phonetic symbols for each word. Step S212: obtaining a unified phoneme rule. Step S213: obtaining the phoneme sequence of the British phonetic symbols for each word under the unified phoneme rule, based on the unified phoneme rule and the corresponding British phonetic symbols. Step S214: obtaining the phoneme sequence of the American phonetic symbols for each word under the unified phoneme rule, based on the unified phoneme rule and the corresponding American phonetic symbols. Step S215: constructing a word-phoneme sequence based on each word, the phoneme sequence of the American phonetic symbols for each word under the unified phoneme rule, and the phoneme sequence of the British phonetic symbols for each word under the unified phoneme rule.

[0055] The electronic device can obtain multiple words and their corresponding American phonetic symbols by acquiring American word-phoneme sequences, and obtain multiple words and their corresponding American phonetic symbols by acquiring British word-phoneme sequences. After obtaining the American and British phonetic symbol sequences, the data can be integrated to obtain a correspondence table between each word, the American phonetic symbol sequence, and the British phonetic symbol sequence, as shown in Table 2 below.

[0056] Table 2. Correspondence between words, American phonetic symbols, and British phonetic symbols.

[0057]

[0058] Next, a unified phoneme rule is obtained. For example, after obtaining the correspondence table between phonetic symbols and phonemes, the American and British phonetic symbol sequences of each word are represented by phonemes in the unified phoneme rule to obtain the phoneme sequence corresponding to each word. In order to distinguish between American and British pronunciations, a distinguishing symbol can be added before the phoneme sequence. For example, the distinguishing symbols "0", "1", and "2" can be used to distinguish between American and British pronunciations. "0" indicates that British and American pronunciations are the same, "1" indicates that British pronunciation is unique, and "2" indicates that American pronunciation is unique. For example, Table 2 above can be converted to obtain Table 3 below, which is the word-phoneme sequence.

[0059] Table 3 Word-phoneme sequence

[0060]

[0061] As can be seen, the word-phoneme sequence is a set of phoneme sequences containing each word, the American phonetic transcription sequence corresponding to each word under the unified phoneme rule, and the British phonetic transcription sequence corresponding to each word under the unified phoneme rule. After the above steps, the electronic device can construct the word-phoneme sequence.

[0062] Step S300: Construct a decoding network based on word-phoneme sequences and the text to be evaluated.

[0063] The decoding network can be called a search graph, which represents all possible language spaces.

[0064] Specifically, if the text to be evaluated contains the word "teacher", then the unified phoneme identifier corresponding to "teacher" in the text to be evaluated can be obtained based on Table 3, and a system can be constructed as follows: Figure 4 The decoding network shown in the diagram uses an expression of the form "x:y". The character "x" to the left of the colon ":" represents the input phoneme of the node, and the character "y" to the right of the colon ":" represents the output character of the node. "sil" represents a silence phoneme, "SIL" represents a silence character, and "eps" represents a meaningless character. For example, on the path from node 4 to node 6, if the input phoneme of node 4 is "er", the output character of node 4 is "teacher_us".

[0065] Additionally, pronunciation indicator symbols can be added to the British and American pronunciation paths in the decoding network, such as... Figure 4 The text uses the character "_us" to indicate the American pronunciation path and the character "_uk" to indicate the British pronunciation path. In practical applications, these pronunciation indicator symbols can be set according to actual needs and are not limited to the limitations of this embodiment. Figure 4 As shown, it contains two British pronunciation paths and one American pronunciation path. The British pronunciation paths are 0-1-2-3-4-7-8-9 and 0-1-2-3-4-5-9, respectively, while the American pronunciation path is 0-1-2-3-4-6.

[0066] It should be noted that, Figure 4 The decoding network shown is a phoneme-level decoding diagram. In actual decoding, because the acoustic model outputs acoustic scores corresponding to phoneme states, this decoding network needs to be expanded to a state-level decoding diagram. Typically, one phoneme can correspond to at least three acoustic states; for example, one phoneme can correspond to three acoustic states, which respectively correspond to the beginning, middle, and end of the pronunciation. For example, for the path from node 4 to node 6, its state-level decoding diagram is as follows... Figure 5 As shown, where, Figure 5 Node 0 in the middle corresponds to Figure 4 Node 4 in Figure 5 Node 3 in the middle corresponds to Figure 5 Node 6 in Figure 5 In the diagram, node 0 has one candidate input state, one candidate directed path, and one candidate output character. Nodes 1 and 2 each have two candidate input states, two candidate directed paths, and two candidate output characters, with one candidate input state corresponding to one candidate directed path and one candidate output character. For example, for... Figure 5 In the example, node 1 has candidate input states er_0 and er_1, a candidate directed path of node 1-node 1, and a corresponding candidate output character of eps. Another candidate directed path is node 1-node 2, and its corresponding candidate output character is eps.

[0067] In this embodiment, by constructing a decoding network, the target state corresponding to each audio frame of the audio data can be analyzed frame by frame, and the pronunciation information in the audio data can be determined. In addition, by adding pronunciation mode indicator symbols to the British and American pronunciation paths when constructing the decoding network, it can be determined whether the user's pronunciation mode is biased towards American or British pronunciation after subsequent decoding.

[0068] Step S400: Obtain the acoustic model.

[0069] An acoustic model is an artificial intelligence model used for acoustic recognition, which can be pre-trained using machine learning. An acoustic model outputs an acoustic score corresponding to the audio being evaluated. Acoustic models can be based on Hidden Markov Models (HMMs), such as Gaussian Mixture Model-Hidden Markov Model (GMM-HMM) or Deep Neural Network-Hidden Markov Model (DNN-HMM), etc.

[0070] Acoustic models can be built on a single-phoneme basis. However, the pronunciation of a phoneme usually varies depending on the phonemes preceding and following it, exhibiting contextual relevance. Therefore, acoustic models can also be built using triphones as the modeling unit, where a triphone is composed of three phonemes. For example, a-b+c represents the pronunciation of phoneme b, where the preceding phoneme is a and the following phoneme is c. Furthermore, a phoneme should have at least three acoustic states, such as the acoustic states corresponding to the onset, pronunciation, and end phases of a phoneme. Therefore, the minimum number of acoustic state classifications is the number of phonemes * 3. Understandably, the larger the number of classifications, the more computationally complex the model generally becomes.

[0071] Preferably, a DNN-HMM can be used as the acoustic model, as its phoneme state classification accuracy is superior to that of GMM-HMM, thus improving the evaluation accuracy. This acoustic model can be pre-stored in a database, allowing electronic devices to retrieve it from the database.

[0072] Step S500: Based on the audio to be evaluated, the acoustic model, and the decoding network, output the evaluation information.

[0073] Specifically, an acoustic model can be used to determine the acoustic score corresponding to each frame of the audio to be evaluated, and then the optimal path can be determined based on the decoding network. Finally, the evaluation information can be obtained based on the optimal path.

[0074] The evaluation information may include at least one of the following: phoneme boundary information, posterior probability information corresponding to the phoneme output at each time step, and pronunciation mode. The phoneme boundary information includes the start and end frames corresponding to each phoneme. The posterior probability information corresponding to the phoneme output at each time step may include the posterior probability values ​​of each phoneme at the same time step, and the maximum value of all posterior probabilities at the same time step. The pronunciation mode is either British or American pronunciation. After obtaining the above evaluation information, it can be used as input to a subsequent scoring system to obtain a score. Alternatively, it can be evaluated based on multiple preset evaluation indicators to obtain corresponding evaluation results for display to the user. The preset evaluation indicators may include at least one of the following: pronunciation accuracy, sentence completeness, fluency, and stress position accuracy. The evaluation results may be at least one of the following: pronunciation accuracy, sentence completeness, fluency, and stress position accuracy.

[0075] In this embodiment, since both the British and American phonetic sequences of each word in the word-phoneme sequence are represented using phonemes in the Unified Phoneme Rules, the decoding network constructed using an acoustic model and based on the word-phoneme sequence under the Unified Phoneme Rules only needs to use one acoustic model to obtain the evaluation information during the evaluation of the audio to be evaluated. There is no need to use American and British phonetic evaluation models separately, which reduces the complexity of the evaluation method.

[0076] In some embodiments, obtaining the acoustic model includes:

[0077] Step S410: Obtain the training audio and the training text corresponding to the training audio.

[0078] Training text refers to a sequence of words. In this application, the training text is a sequence of English words. This training text can be acquired or pre-stored; for example, it can be acquired in real time through an interactive application, such as a language learning application, by the user reading the training text aloud.

[0079] Training audio refers to the speech data based on the training text to be used for pronunciation evaluation. For example, speech data for oral assessment can be used as training audio. In this application, the training audio is in English, which can be British pronunciation, American pronunciation, or a mixture of both. The training audio can be acquired in real time or pre-stored. For example, it can be acquired in real time through interactive applications such as language learning applications by having the user read the training text aloud, and used as training audio.

[0080] Step S420: Train the preset model based on the training audio, training text and word-phoneme sequence to obtain the acoustic model.

[0081] Specifically, a pre-defined model can be trained using training audio, training text, and word-phoneme sequences to obtain a DNN-HMM. It should be noted that the distinguishing symbols in the word-phoneme sequences do not need to be input into the pre-defined model; that is, they do not need to participate in model training. The specific training process can refer to existing techniques and is not limited here.

[0082] In this embodiment, an acoustic model can be obtained by training audio, training text, and word-phoneme sequences without pre-storing the acoustic model.

[0083] In some embodiments, step S500 includes:

[0084] Step S510: Based on the acoustic model, determine the acoustic score corresponding to each audio frame in the audio to be evaluated.

[0085] Specifically, the audio to be evaluated is input into the acoustic model, which can then obtain the acoustic score for each frame of audio based on the audio features of the audio to be evaluated.

[0086] Step S520: Determine the optimal path based on each acoustic score and the decoding network.

[0087] Next, in the constructed decoding network, starting from point 0, the Viterbi algorithm is used to search for the optimal path frame by frame until all audio frame information is consumed. Each step forward consumes the information of one audio frame. For example, at frame 20, there are two candidate nodes, namely... Figure 5Given nodes 1 and 2 in the audio, for candidate node 1, its forward radiating paths are node 1-node 1 and node 1-node 2; for candidate node 2, its forward radiating paths are node 2-node 2 and node 2-node 3. There are a total of four candidate paths. These four candidate paths are traversed, and the cumulative probability or cumulative cost for each path is calculated. The node on the path with the highest probability or lowest cumulative cost is selected as the optimal node corresponding to the 20th frame, completing one search step, which consumes the information of one audio frame. Thus, after consuming all audio frames, the optimal path is obtained by connecting the optimal nodes corresponding to each audio frame. Typically, the cumulative cost of a candidate path can be the sum of the transfer weight cost and the acoustic cost.

[0088] Step S530: Output evaluation information based on the optimal path.

[0089] Based on the optimal path, at least one of the following can be obtained: boundary information of each phoneme, posterior probability information of the phonemes output at each time, and pronunciation mode.

[0090] As can be seen, the evaluation information can be determined based on the above method. Furthermore, by employing the Viterbi algorithm to progressively calculate the optimal node corresponding to each audio frame, the exhaustive search for all possible results is avoided. This efficient calculation method can improve the efficiency of searching for the optimal path while ensuring accuracy.

[0091] This application also provides a non-volatile computer-readable storage medium storing computer-executable instructions that are executed by one or more processors, for example, to perform the pronunciation evaluation method described in the above embodiments. The computer-readable storage medium can be a storage medium such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; it can also be a device including one or any combination of the above-mentioned memories. The executable instructions can take the form of programs, software, software modules, scripts, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and can be deployed in any form, including being deployed as a standalone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0092] This application also provides a computer program product, including a computing program stored on a non-volatile computer-readable storage medium. The computer program includes program instructions that, when executed by a computer, cause the computer to perform the pronunciation evaluation method described in the above embodiments.

[0093] It should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0094] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a general-purpose hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the related technology, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for at least one computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments.

[0095] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and not to limit them; under the concept of this application, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations of different aspects of this application as described above, which are not provided in detail for the sake of brevity; although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

Claims

1. A method of evaluating a pronunciation, characterized by, The method comprises: obtaining to-be-evaluated audio and to-be-evaluated text corresponding to the to-be-evaluated audio, the to-be-evaluated text comprising a word sequence; obtaining a word-phoneme sequence, wherein the word-phoneme sequence comprises a plurality of words, a phoneme sequence of an American English phonetic alphabet sequence corresponding to each of the words under a unified phoneme rule, and a phoneme sequence of a British English phonetic alphabet sequence corresponding to each of the words under the unified phoneme rule, the unified phoneme rule comprising a one-to-one correspondence between the British English phonetic alphabet and phonemes and a one-to-one correspondence between the American English phonetic alphabet and phonemes, and at least one phoneme corresponding to one British English phonetic alphabet and one American English phonetic alphabet under the unified phoneme rule; based on the word-phoneme sequence and the to-be-evaluated text, constructing a decoding network, wherein, when the phoneme sequence of the American English phonetic alphabet sequence corresponding to a word under the unified phoneme rule is different from the phoneme sequence of the British English phonetic alphabet sequence corresponding to the word under the unified phoneme rule, the decoding network of the word comprises a British English pronunciation path and an American English pronunciation path; obtaining an acoustic model; based on the to-be-evaluated audio, the acoustic model, and the decoding network, outputting evaluation information.

2. The pronunciation evaluation method of claim 1, wherein, The method for constructing the word-phoneme sequence comprises: obtaining a plurality of words, a British English phonetic alphabet sequence corresponding to each of the words, and an American English phonetic alphabet sequence corresponding to each of the words; obtaining the unified phoneme rule; obtaining, according to the unified phoneme rule and each of the British English phonetic alphabet sequences, a phoneme sequence of each of the British English phonetic alphabet sequences under the unified phoneme rule; obtaining, according to the unified phoneme rule and each of the American English phonetic alphabet sequences, a phoneme sequence of each of the American English phonetic alphabet sequences under the unified phoneme rule; based on each of the words, the phoneme sequence of each of the American English phonetic alphabet sequences under the unified phoneme rule, and the phoneme sequence of each of the British English phonetic alphabet sequences under the unified phoneme rule, constructing the word-phoneme sequence.

3. The pronunciation evaluation method of claim 2, wherein, The method for obtaining the acoustic model comprises: obtaining training audio and training text corresponding to the training audio; based on the training audio, the training text, and the word-phoneme sequence, training a preset model to obtain the acoustic model.

4. The pronunciation evaluation method of claim 3, wherein, The acoustic model is a deep neural network-hidden Markov model.

5. The pronunciation evaluation method according to any one of claims 1 to 4, wherein, The method for outputting the evaluation information based on the to-be-evaluated audio, the acoustic model, and the decoding network comprises: based on the acoustic model, determining an acoustic score corresponding to each audio frame in the to-be-evaluated audio; based on each of the acoustic scores and the decoding network, determining an optimal path; based on the optimal path, outputting the evaluation information.

6. The pronunciation evaluation method of claim 5, wherein, The method for determining the optimal path based on each of the acoustic scores and the decoding network comprises: based on each of the acoustic scores and the decoding network, determining the optimal path by using a Viterbi algorithm.

7. The pronunciation evaluation method according to any one of claims 1 to 4, wherein, The evaluation information comprises at least one of phoneme boundary information, posterior probability information of phonemes output at each time, and a pronunciation manner.

8. An electronic device, comprising: The method comprises: at least one processor; and a memory in communication connection with the at least one processor; wherein The memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1 to 7.

9. A speech evaluation system characterized by, The electronic device comprises a voice collection device and the electronic device as claimed in claim 8. The voice collection device is connected to the electronic device.

10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer executable instructions for causing a computer to perform the method of any one of claims 1 to 7.

Citation Information

Patent Citations

  • English accent identification system

    CN109493846A

  • Pronunciation evaluation method and device, equipment and storage medium

    CN116631441A